Visual determination of sleep state

A video-based machine learning approach for sleep state analysis in rodents provides high-throughput, non-invasive differentiation of wakefulness, REM, and non-REM states, addressing the limitations of existing invasive and inaccurate methods.

JP7894894B2Active Publication Date: 2026-07-24JACKSON LAB THE +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
JACKSON LAB THE
Filing Date
2022-06-27
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Current methods for sleep state analysis in rodents, such as EEG/EMG recordings, have low throughput due to the need for invasive procedures and manual scoring, while non-invasive methods like activity assessment and imaging techniques struggle to distinguish between wakefulness, REM, and non-REM states accurately.

Method used

A computer-implemented method using high-resolution video data and machine learning models to analyze features like respiration, movement, and posture to differentiate between wakefulness, REM, and non-REM sleep states in rodents, employing segmentation, feature extraction, and classification techniques.

Benefits of technology

Enables high-throughput, non-invasive sleep state differentiation with improved accuracy, facilitating large-scale mechanistic studies for therapeutic discoveries.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007894894000010
    Figure 0007894894000010
  • Figure 0007894894000011
    Figure 0007894894000011
  • Figure 0007894894000012
    Figure 0007894894000012
Patent Text Reader

Abstract

The systems and methods described herein provide techniques for determining sleep state data by processing video data of a subject. The systems and methods may determine a plurality of features from the video data and may use the plurality of features to determine the sleep state data of the subject. In some embodiments, the sleep state data may be based on frequency domain features and / or time domain features corresponding to the plurality of features.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] (Related Applications) This application claims the benefit of U.S. Provisional Application No. 63 / 215,511, filed Jun. 27, 2021, under 35 U.S.C. § 119(e), the disclosure of which is incorporated herein by reference in its entirety.

[0002] In some aspects, the present invention relates to determining a subject's sleep state by processing video data using a machine learning model.

[0003] (Government Support) The present invention was made with government support under DA041668 (NIDA), DA048634 (NIDA), and HL094307 (NHLBI) awarded by the National Institutes of Health. The government has certain rights in the invention.

Background Art

[0004] Sleep is a complex behavior regulated by homeostatic processes, and its function is crucial for survival. Sleep disorders and circadian disorders are found in many diseases, including neuropsychiatric disorders, neurodevelopmental disorders, neurodegenerative disorders, physiological disorders, and metabolic disorders. Sleep function and circadian function have a bidirectional relationship with these diseases, and changes in sleep and circadian patterns can lead to or cause disease states. While the bidirectional relationship between sleep and many diseases is well explained, their genetic etiologies are not fully elucidated. In fact, treatment for sleep disorders is limited due to a lack of knowledge about sleep mechanisms. Rodents serve as readily available models of human sleep due to their similarities in sleep biology, and mice in particular are genetically traceable models for studying the mechanisms of sleep and potential therapeutic agents. One reason for this significant gap in treatment is the technological barrier that hinders reliable phenotyping of a large number of mice for evaluating sleep states. The gold standard for sleep analysis in rodents is the use of electroencephalography / electromyography (EEG / EMG) recordings. This method has low throughput because it requires surgery for electrode implantation and often necessitates manual scoring of records. While newer methods utilizing machine learning models are beginning to automate EEG / EMG scoring, data generation still suffers from low throughput. Additionally, the use of tethered electrodes can restrict animal movement and potentially alter animal behavior.

[0005] To overcome the low throughput limitations of some existing systems, several non-invasive approaches for sleep analysis have been investigated. These include activity assessment using beambreak systems or imaging techniques where a certain amount of inactivity is interpreted as sleep. Piezo-pressure sensors have also been used as a simpler and more sensitive way to access activity. However, these methods only assess wakefulness relative to sleep states and cannot distinguish between wakefulness, rapid eye movement (REM) states, and non-REM states. This is important because activity determination of sleep states can be inaccurate in humans and rodents with low general activity. Other methods for assessing sleep states include pulsed Doppler-based methods that access movement and respiration, as well as whole-body plethysmography that directly measures respiratory patterns. Both of these approaches require specialized equipment. Field sensors that detect respiration and other movements have also been used to assess sleep states. [Overview of the Initiative]

[0006] According to embodiments of the present invention, a computer implementation method is provided, the method comprising: receiving video data representing an image of a subject; using the video data to determine a plurality of features corresponding to the subject; and using the plurality of features to determine sleep state data about the subject. In some embodiments, the method also comprises using a machine learning model to process the video data to determine segmentation data showing a first set of pixels corresponding to the subject and a second set of pixels corresponding to the background. In some embodiments, the method also comprises processing the segmentation data to determine elliptic fitting data corresponding to the subject. In some embodiments, determining the plurality of features comprises processing the segmentation data to determine the plurality of features. In some embodiments, the plurality of features comprises a plurality of visual features for each video frame of the video data. In some embodiments, the method also comprises determining a time-domain feature for each visual feature of the plurality of visual features, wherein the plurality of features comprises the time-domain feature. In some embodiments, determining the time-domain feature comprises determining one of kurtosis data, mean data, median data, standard deviation data, maximum data, and minimum data. In some embodiments, the method also comprises determining a frequency-domain feature for each visual feature of the plurality of visual features, wherein the plurality of features comprises the frequency-domain feature. In some embodiments, determining frequency-domain features includes determining one of the following: power spectral density kurtosis, power spectral density skewness, mean power spectral density, total power spectral density, maximum data, minimum data, mean data, and standard deviation of power spectral density. In some embodiments, the method also further includes determining time-domain features for each of a plurality of features, determining frequency-domain features for each of a plurality of features, and using a machine learning classifier to process the time-domain features and the frequency-domain features to determine sleep state data. In some embodiments, the method also includes using a machine learning classifier to process a plurality of features to determine sleep states for video frames of video data, where sleep states are one of wakefulness, REM sleep, and non-REM sleep.In some embodiments, sleep state data indicates the duration of sleep states, one or more durations and / or frequency intervals of one or more of the wakefulness, REM, and non-REM states, and changes in one or more sleep states. In some embodiments, the method also includes determining a plurality of body regions of a subject using a plurality of features, wherein each of the plurality of body regions corresponds to a video frame of video data, and determining sleep state data based on changes in the plurality of body regions in the video. In some embodiments, the method also includes determining a plurality of width ratios using a plurality of features, wherein each of the plurality of width ratios corresponds to a video frame of video data, and determining sleep state data based on changes in the plurality of width ratios in the video. In some embodiments, determining sleep state data includes detecting a transition from a non-REM state to a REM state based on changes in the body region or body shape of the subject, wherein the changes in the body region or body shape are the result of muscle relaxation. In some embodiments, the method also includes determining a plurality of width ratios for a subject, wherein the width ratios among the plurality of width ratios correspond to video frames of video data; determining time-domain features using the plurality of width ratios; determining frequency-domain features using the plurality of width ratios, wherein the time-domain and frequency-domain features represent abdominal movement of the subject; and determining sleep state data using the time-domain and frequency-domain features. In some embodiments, the video captures the subject in its natural state. In some embodiments, the natural state of the subject includes the absence of invasive detection means within or on the subject. In some embodiments, the invasive detection means includes one or both of electrodes attached to the subject and electrodes inserted into the subject. In some embodiments, the video is high-resolution video. In some embodiments, the method also includes using a machine learning classifier to process a plurality of features to determine a plurality of sleep state predictions for each of a single video frame of video data; and using a transition model to process the plurality of sleep state predictions to determine the transition between a first sleep state and a second sleep state.In some embodiments, the transition model is a Hidden Markov Model. In some embodiments, the subjects are rodents, and optionally mice. In some embodiments, the subjects are genetically modified subjects.

[0007] According to another aspect of the present invention, a method is provided for determining the sleep state of a subject, the method comprising monitoring the subject's responses, the means for monitoring comprising any embodiment of the computer implementation method described above. In some embodiments, the sleep state comprises one or more of sleep stages, sleep intervals, changes in sleep stages, and non-sleep intervals. In some embodiments, the subject has a sleep disorder or condition. In some embodiments, the sleep disorder or condition comprises one or more of sleep apnea, insomnia, and narcolepsy. In some embodiments, the sleep disorder or condition is the result of brain injury, depression, mental illness, neurodegenerative disease, restless legs syndrome, Alzheimer's disease, Parkinson's disease, obesity, overweight, the effects of administered drugs and / or alcohol consumption, a neurological condition that can alter the sleep state, or a metabolic disorder or condition that can alter the sleep state. In some embodiments, the method also comprises administering a therapeutic agent to the subject before receiving video data. In some embodiments, the therapeutic agent comprises one or more of sleep stimulants, sleep inhibitors, and agents that can alter one or more sleep stages of the subject. In some embodiments, the method also includes performing behavioral therapy on a subject. In some embodiments, the behavioral therapy includes sensory therapy. In some embodiments, the sensory therapy is light exposure therapy. In some embodiments, the subject is a genetically modified subject. In some embodiments, the subject is a rodent, and optionally a mouse. In some embodiments, the mouse is a genetically modified mouse. In some embodiments, the subject is an animal model of a sleep disorder. In some embodiments, sleep state data determined for the subject is compared with control sleep state data. In some embodiments, the control sleep state data is sleep state data from a control subject determined by a computer implementation method. In some embodiments, the control subject does not have the sleep disorder or condition of the subject. In some embodiments, the control subject is not administered the therapeutic drug or behavioral therapy administered to the subject. In some embodiments, the control subject is administered a different dose of the therapeutic drug than the dose administered to the subject.

[0008] According to another aspect of the present invention, a method is provided for identifying the effectiveness of a candidate therapeutic agent and / or candidate behavioral therapy for treating a sleep disorder or condition of interest, the method comprising administering the candidate therapeutic agent and / or candidate behavioral therapy to a test subject and determining sleep state data for the test subject, the means of determination comprising any embodiment of the aforementioned computer implementation method, the determination showing a change in sleep state data in the test subject identifies the effect of the candidate therapeutic agent or candidate behavioral therapy, respectively, on the sleep disorder or condition of interest. In some embodiments, the sleep state data includes one or more data from sleep stages, sleep intervals, changes in sleep stages, and non-sleep intervals. In some embodiments, the test subject has a sleep disorder or condition. In some embodiments, the sleep disorder or condition includes one or more of sleep apnea, insomnia, and narcolepsy. In some embodiments, the sleep disorder or condition is the result of brain injury, depression, mental illness, neurodegenerative disease, restless legs syndrome, Alzheimer's disease, Parkinson's disease, obesity, overweight, the effects of administered drugs and / or alcohol consumption, a neurological condition that can alter sleep states, or a metabolic disorder or condition that can alter sleep states. In some embodiments, candidate therapeutic agents and / or candidate behavioral therapies are administered to the test subject one or more times before or during the reception of video data. In some embodiments, the candidate therapeutic agents include one or more of sleep stimulants, sleep inhibitors, and agents that can alter one or more sleep stages of the test subject. In some embodiments, the behavioral therapy includes sensory therapy. In some embodiments, the sensory therapy is light exposure therapy. In some embodiments, the subject is a genetically modified subject. In some embodiments, the test subject is a rodent, and optionally a mouse. In some embodiments, the mouse is a genetically modified mouse. In some embodiments, the test subject is an animal model of a sleep disorder. In some embodiments, the sleep state data determined for the test subject are compared to control sleep state data. In some embodiments, the control sleep state data is sleep state data from a control subject determined by a computer implementation method.In some embodiments, the control subject does not have the sleep disorder or condition of the test subject. In some embodiments, the control subject is not administered the candidate therapy administered to the test subject. In some embodiments, the control subject is administered a different dose of the candidate therapy than the dose administered to the test subject. In some embodiments, the control subject is administered a different regimen of candidate behavioral therapy than the regimen of candidate therapy administered to the test subject. In some embodiments, the behavioral therapy regimen includes one or more therapeutic features such as the length of the behavioral therapy, the intensity of the behavioral therapy, the light intensity in the behavioral therapy, and the frequency of the behavioral therapy. [Brief explanation of the drawing]

[0009] For a more complete understanding of this disclosure, please refer to the following description in conjunction with the attached drawings.

[0010] [Figure 1] Figure 1 is a conceptual diagram of a system for determining the sleep state data of a subject using video data, according to an embodiment of the present disclosure. [Figure 2A] Figure 2A is a flowchart showing the process for determining sleep state data according to an embodiment of the present disclosure. [Figure 2B] Figure 2B is a flowchart showing a process for determining sleep state data for multiple objects represented in a video, according to an embodiment of the present disclosure. [Figure 3] Figure 3 is a conceptual diagram of a system for training components that determine sleep state data according to an embodiment of the present disclosure. [Figure 4] Figure 4 is a block diagram conceptually illustrating exemplary components of a device according to an embodiment of the present disclosure. [Figure 5] Figure 5 is a block diagram conceptually illustrating exemplary components of a server according to an embodiment of the present disclosure. [Figure 6A] Figure 6A shows a schematic diagram illustrating the organization of data collection, annotation, feature generation, and classifier training according to embodiments of the present disclosure. [Figure 6B]Figure 6B shows a schematic diagram of frame-level information used for visual features, and according to embodiments of this disclosure, a trained neural network was used to generate a segmentation mask of mouse-related pixels for use in downstream classification. [Figure 6C] Figure 6C shows schematic diagrams of multiple frames of a video containing multiple objects, and according to embodiments of the present disclosure, instance segmentation techniques are used to generate segmentation masks for individual objects even when the objects are in close proximity to each other. [Figure 7A] Figure 7A shows an exemplary graph of selected signals in the time and frequency domains within one epoch, displaying m00 (region of the segmentation mask) (leftmost column), the FFT of the corresponding signal (middle column), and the autocorrelation of the signal (rightmost column) for arousal, non-REM, and REM states. [Figure 7B] Figure 7B shows an exemplary graph of a selected signal in the time and frequency domains within a single epoch, similar to Figure 6A, illustrating the time and frequency domain wl_ratio. [Figure 8A] Figures 8A and 8B show plots illustrating the extraction of respiratory signals from the video. Figure 8A shows exemplary spectral analysis plots for rem and non-rem epochs. Continuous wavelet transform spectral responses (top panel) and associated dominant signals (each in the lower left panel), as well as histograms of the dominant signals (each in the lower right panel). Non-rem epochs typically showed lower mean and standard deviations than rem epochs. [Figure 8B] Figure 8B shows a plot examining a larger timescale of the epoch, indicating that the non-REM signal remained stable until the REM development period. The dominant frequency was the typical mouse respiratory rate frequency. [Figure 9A]Figures 9A-C show graphs illustrating the validation of respiratory signals in video data for wl_ratio measurement. Figure 9A shows the mobile cutoff used to select sleep epochs in C57BL / 6J vs C3H / HeJ respiratory rate analysis. Below the 10% quantile cutoff threshold (black vertical line), epochs consisted of 90.2% non-REM (red line), 8.1% REM (green line), and 1.7% wakefulness (blue line). [Figure 9B] Figure 9B shows a comparison of dominant frequencies between strains observed during sleep epochs (blue for males, orange for females). [Figure 9C] Figure 9C shows, using the C57BL / 6J annotated epoch, that a higher standard deviation was observed at the dominant frequency in the REM state (blue line) than in the non-REM state (orange line). [Figure 9D] Figure 9D shows that the increase in standard deviation was consistent across all animals. [Figure 10A] Figures 10A–D show graphs and tables illustrating the performance metrics of the classifiers. Figure 10A shows the classifier performance compared at different stages, starting with the XgBoost classifier, adding an HMM model, increasing features to include seven Hu moments, and integrating SINDLE annotations to improve epoch quality. It was observed that the overall accuracy improved with each of these steps. [Figure 10B] Figure 10B shows the top 20 most important features for a classifier. [Figure 10C] Figure 10C shows the confusion matrix obtained from 10-fold cross-validation. [Figure 10D] Figure 10D shows the precision-recall table. [Figure 11A] Figures 11A-D show graphs illustrating the validation of visual scoring. Figure 11A shows hypnograms of visual scoring and EEG / EMG scoring. [Figure 11B] Figure 11B shows plots of visually scored sleep stages (top) and predicted stages (bottom) for a mouse (B6J_7) over 24 hours. [Figure 11C] Figures 11C - D show the comparison of human - based scoring and visual scoring across all C57BL / 6J mice, indicating a high agreement between the two methods. The data were plotted in 1 - hour bins over 24 hours (Figure 11C), [Figure 11D] and plotted over a 24 - hour or 12 - hour period (Figure 11D). [Figure 12] Figure 12 shows a bar graph presenting the results of additional data augmentation for the classifier model. **DETAILED DESCRIPTION OF THE INVENTION**

[0011] The present disclosure relates to determining a subject's sleep state by processing video data of the subject using one or more machine - learning models. The subject's respiration, movement, or posture are useful in themselves for differentiating sleep states. In some embodiments of the present disclosure, a combination of features of respiration, movement, and posture is used to determine the subject's sleep state. By using such a combination of features, the accuracy rate of predicting the sleep state increases. The term "sleep state" is used with respect to the rapid - eye - movement (REM) sleep state and the non - rapid - eye - movement (non - REM) sleep state. The methods and systems of the present invention can be used to evaluate and distinguish the subject's REM sleep state, non - REM sleep state, and wakefulness (non - sleep) state.

[0012] To distinguish between wakefulness, non-REM, and REM states, some embodiments use video-based methods with high-resolution video, based on the determination that information about sleep states is encoded in the video data. When transitioning from non-REM to REM states, subtle changes in the region and shape of the subject are observed, possibly due to the relaxation of REM states. Over the past few years, significant improvements have been made in the field of computer vision, primarily due to advances in the fields of machine learning, especially deep learning. Some embodiments use advanced machine vision methods to significantly improve visual sleep state classification. Some embodiments involve extracting features from video data related to the subject's respiration, movement, and / or posture. Some embodiments combine these features to determine sleep states in subjects such as mice. Embodiments of this disclosure involve non-invasive video-based methods that can be implemented with low hardware investment and yield high-quality sleep state data. The ability to reliably, non-invasively, and in a high-throughput manner access sleep states enables large-scale mechanistic studies necessary for therapeutic discoveries.

[0013] Figure 1 conceptually illustrates a system 100 (e.g., an automated sleep state system 100) for determining target sleep state data using video data. The automated sleep state system 100 may operate using various components shown in Figure 1. The automated sleep state system 100 may include an image capture device 101, a device 102, and one or more systems 105, connected via one or more networks 199. The image capture device 101 may be part of another device (e.g., device 400 shown in Figure 4), or contained within such a device, or connected to such a device, and may also be a camera, or a high-speed video camera, or other type of device capable of capturing images or video. In addition to the image capture device, device 101 may include motion detection sensors, infrared sensors, temperature sensors, ambient condition detection sensors, and other sensors configured to detect various characteristics / environmental conditions. Device 102 may be a laptop, desktop, tablet, smartphone, or other type of computing device capable of displaying data, and may also include one or more components described in relation to device 400 below.

[0014] The image capture device 101 may capture video (or one or more images) of a target and transmit video data 104 representing the video to the system 105 for processing as described herein. The video may be of a target in an open field arena. In some cases, the video data 104 may correspond to images (image data) captured by the device 101 at specific time intervals, so that the images capture the target over a period of time. In some embodiments, the video data 104 may be high-resolution video of the target.

[0015] System 105 may include one or more components shown in Figure 1 and may be configured to process video data 104 to determine the sleep state data of a subject. System 105 may also generate sleep state data 152 corresponding to the subject, which may represent one or more sleep states of the subject observed in the video (e.g., awake / non-sleep state, non-REM state, and REM state). System 105 may send the sleep state data 152 to device 102 for output to the user in order to observe the results of processing the video data 104.

[0016] In some embodiments, the video data 104 may include two or more target videos, and the system 105 may process the video data 104 to determine the sleep state data for each target represented in the video data 104.

[0017] System 105 may be configured to determine various data from video data 104 about the subject. To determine this data and to determine sleep state data 152, System 105 may include several different components. As shown in Figure 1, System 105 may include a segmentation component 110, a feature extraction component 120, a spectral analysis component 130, a sleep state classification component 140, and a post-classification component 150. System 105 may include fewer or more components than those shown in Figure 1. In some embodiments, these various components may be located on the same physical system 105. In other embodiments, one or more of the various components may be located on different / separate physical systems 105. Communication between the various components may occur directly or via the network 199. Communication between device 101, System 105, and device 102 may occur directly or via the network 199.

[0018] In some embodiments, one or more components shown as part of system 105 may be located in device 102 or a computing device (e.g., device 400) connected to image acquisition device 101.

[0019] At a high level, the system 105 may be configured to process the video data 104 to determine multiple features corresponding to the subject, and to use these features to determine the subject's sleep state data 152.

[0020] Figure 2A is a flowchart of a process 200 for determining target sleep state data 152 according to an embodiment of the present disclosure. One or more steps of process 200 may be performed in a different order / sequence than that shown in Figure 2A. One or more steps of process 200 may be performed by components of system 105 shown in Figure 1.

[0021] In step 202 of process 200 shown in Figure 2A, system 105 may receive image data 104 representing an image of the object. In some embodiments, the image data 104 may be received by a segmentation component 110 or provided to the segmentation component 110 by system 105 for processing. In some embodiments, the image data 104 may be an image capturing the object in its natural state. The object may be in its natural state when no invasive methods are applied to the object (e.g., electrodes are not inserted into or attached to the object, dye / color markings are not applied to the object, surgical procedures are not performed on the object, invasive detection means are not in or on the object, etc.). The image data 104 may be a high-resolution image of the object.

[0022] In step 204 of process 200 shown in Figure 2A, the segmentation component 110 may perform a segmentation process using the video data 104 to determine the ellipse data 112 (shown in Figure 1). The segmentation component 110 may process the video data 104 to generate a segmentation mask that identifies objects in the video data 104, and then employ techniques to generate ellipse fit / representation for the objects. The segmentation component 110 may use one or more techniques (e.g., one or more ML models) for object tracking of video / image data and may be configured to identify objects. The segmentation component 110 may generate a segmentation mask for each video frame of the video data 104. The segmentation mask may indicate which pixels in the video frame correspond to objects and / or which pixels in the video frame correspond to background / non-objects. The segmentation component 110 may use a machine learning model to process the video data 104 and determine segmentation data that shows a first set of pixels corresponding to the subject and a second set of pixels corresponding to the background.

[0023] A video frame, as used herein, may be a portion of video data 104. Video data 104 may be divided into multiple parts / frames of the same length / time. For example, a video frame may be 1 millisecond of video data 104. When determining data such as a segmentation mask for a video frame of video data 104, a component of system 105, such as a segmentation component 110, may process a set of video frames (a window of video frames). For example, to determine a segmentation mask for an instant video frame, the segmentation component 110 may process (i) a set of video frames that occur before the instant video frame (with respect to time) (e.g., the three video frames before the instant video frame), (ii) the instant video frame, and (iii) a set of video frames that occur after the instant video frame (with respect to time) (e.g., the three video frames after the instant video frame). Thus, in this embodiment, the segmentation component 110 may process seven video frames to determine a segmentation mask for one video frame. This type of processing may be referred to in this specification as window-based processing of video frames.

[0024] Using the segmentation mask of the video data 104, the segmentation component 110 may determine the ellipse data 112. The ellipse data 112 may be an ellipse fit for the object (an ellipse drawn around the object's body). For different types of objects, the system 105 may be configured to determine different shape fits / representations (e.g., circular fit, rectangular fit, square fit, etc.). The segmentation component 110 may determine the ellipse data 112 as a subset of pixels in the segmentation mask corresponding to the object. The ellipse data 112 may include a subset of pixels. The segmentation component 110 may determine the ellipse fit for the object for each video frame of the video data 104. The segmentation component 110 may determine the ellipse fit for the video frame using the window-based processing of the video frames described above. The ellipse data 112 may be a vector or a matrix of pixels representing the ellipse fit for all video frames of the video data 104. The segmentation component 110 may process the segmentation data to determine the elliptic fit data 112 corresponding to the target.

[0025] In some embodiments, the ellipse fit 112 for an object may define some parameters of the object. For example, the ellipse fit may correspond to the position of the object and may include coordinates (e.g., x and y) representing the pixel position of the object in the video frame of the video data 104 (e.g., the center of the ellipse). The ellipse fit may correspond to the length of the major axis and the length of the minor axis of the object. The ellipse fit may include the sine and cosine of the vector angle of the major axis. The angle may be defined with respect to the direction of the major axis. The major axis may extend from the tip of the head or nose of the object to an end of the object's body, such as the base of the tail. The ellipse fit may also correspond to the ratio between the length of the major axis and the length of the minor axis of the object. In some embodiments, the ellipse data 112 may include the aforementioned measurements for all video frames of the video data 104.

[0026] In some embodiments, the segmentation component 110 may use one or more neural networks to process the video data 104 to determine the segmentation mask and / or elliptic data 112. In other embodiments, the segmentation component 110 may use other ML models, such as encoder-decoder architectures, to determine the segmentation mask and / or elliptic data 112.

[0027] The elliptic data 112 may also include confidence scores for the segmentation components 110 when determining the elliptic fit for the video frame. Alternatively, the elliptic data 112 may include probabilities or likelihoods of the elliptic fit corresponding to the subject.

[0028] In embodiments in which the video data 104 captures two or more objects, the segmentation component 110 may identify each of the captured objects and determine ellipse data 112 for each of the captured objects. The ellipse data 112 for each object may be provided separately (in parallel or sequentially) to the feature extraction component 120 for processing.

[0029] In step 206 of process 200 shown in Figure 2A, the feature extraction component 120 may determine multiple features using elliptic data 112. The feature extraction component 120 may determine multiple features for each video frame of video data 104. In some exemplary embodiments, the feature extraction component 120 may determine 16 features for each video frame of video data 104. The determined features may be stored as frame feature data 122 shown in Figure 1. The frame feature data 122 may be a vector or matrix containing the values ​​of multiple features corresponding to each video frame of video data 104. The feature extraction component 120 may determine multiple features by processing segmentation data (determined by segmentation component 110) and / or elliptic data 112.

[0030] The feature extraction 120 may determine multiple features so as to include multiple target visual features for each video frame of the video data 104. The following are example features that may be determined by the feature extraction component 120 and included in the frame feature data 122.

[0031] The feature extraction component 120 may process pixel information contained in the ellipse data 112. In some embodiments, the feature extraction component 120 may determine the major axis length, minor axis length, and the ratio of the major axis length to the minor axis length for each video frame of the video data 104. These features may already be contained in the ellipse data 112, or the feature extraction component 120 may determine these features using the pixel information contained in the ellipse data 112. The feature extraction component 120 may also determine the area of ​​the object (e.g., surface area) using the ellipse fit information contained in the ellipse data 112. The feature extraction component 120 may determine the position of the object, which is represented as the center pixel of the ellipse fit. The feature extraction component 120 may also determine the change in the position of the object based on the change in the center pixel of the ellipse fit from one video frame of the video data 104 to another (later occurring) video frame. The feature extraction component 120 may also determine the perimeter (e.g., circumference) of the ellipse fit.

[0032] The feature extraction component 120 may determine one or more (e.g., seven) Hu moments. Hu moments (also known as Hu moment invariants) may be a set of seven numbers calculated using the central moments of the image / video frame, which are invariant under image transformations. The first six moments have been proven invariant under translational motion, scaling, rotation, and inversion, while the sign of the seventh moment changes under image inversion. In image processing, computer vision, and related fields, image moments are a specific weighted average (moment) of the intensity of image pixels, or a function of such moments, usually chosen to have some attractive properties or interpretations. Image moments are useful for describing a subject after segmentation. The feature extraction component 120 may determine the Hu image moments, which are a numerical description of the segmentation mask of the subject, through integration and linear combination of central image moments.

[0033] In step 208 of process 200 shown in Figure 2A, the spectral analysis component 130 may perform spectral analysis using multiple features to determine frequency-domain features 132 and time-domain features 134. The spectral analysis component 130 may use signal processing techniques to determine frequency-domain features 132 and time-domain features 134 from frame feature data 122. In some embodiments, the spectral analysis component 130 may determine a set of time-domain features and a set of frequency-domain features for each feature (from feature data 122) of each video frame of video data 104 in an epoch. In an exemplary embodiment, the spectral analysis component 130 may determine six time-domain features for each feature of each video frame in an epoch. In some embodiments, the spectral analysis component 130 may determine 14 frequency-domain features for each feature of each video frame in an epoch. The epoch may be the duration of the video data 104, for example, 10 seconds, 5 seconds, etc. The frequency domain features 132 may be vectors or matrices representing frequency domain features determined for each feature of the feature data 122 and for each epoch of the video frame. The time domain features 134 may be vectors or matrices representing time domain features determined for each feature of the feature data 122 and for each epoch of the video frame. The frequency domain features 132 may also be graph data, for example, as shown in Figures 7A-7B and 8A-8B.

[0034] In exemplary embodiments, the frequency domain features 132 may include the kurtosis of the power spectral density, the distortion of the power spectral density, the average power spectral density at 0.1–1 Hz, the average power spectral density at 1–3 Hz, the average power spectral density at 3–5 Hz, the average power spectral density at 5–8 Hz, the average power spectral density at 8–15 Hz, the total power spectral density, the maximum value of the power spectral density, the minimum value of the power spectral density, the average power spectral density, and the standard deviation of the power spectral density.

[0035] In an exemplary embodiment, the time-domain feature 134 may be the kurtosis, the mean of the feature signal, the median of the feature signal, the standard deviation of the feature signal, the maximum value of the feature signal, and the minimum value of the feature signal.

[0036] In step 210 of process 200 shown in Figure 2A, the sleep state classification component 140 may process frequency domain features 132 and time domain features 134 to determine a sleep prediction for each video frame of the video data 104. The sleep state classification component 140 may determine a marker representing a sleep state for each video frame of the video data 104. The sleep state classification component 140 may classify each video frame into one of three sleep states: wakefulness, non-REM, and REM. The wakefulness state may be non-sleep, or similar to non-sleep. The sleep state classification component 140 may use frequency domain features 132 and time domain features 134 to determine the sleep state marker for the video frame. In some embodiments, the sleep state classification component 140 may use window-based processing for the video frames described above. For example, to determine the sleep state markers of an instant video frame, the sleep state classification component 140 may process data (time and frequency domain features 132, 134) for a set of video frames that occur before the instant video frame, and data for a set of video frames that occur after the instant video frame. The sleep state classification component 140 may also output frame prediction data 142, which may be a vector or matrix of sleep state markers for each video frame of the video data 104. The sleep state classification component 140 may also determine confidence scores associated with the sleep state markers, which may represent the likelihood of a video frame corresponding to an indicated sleep state, or the reliability of the sleep state classification component 140 in determining the sleep state markers of a video frame. The confidence scores may be included in the frame prediction data 142.

[0037] The sleep state classification component 140 may determine frame prediction data 142 from frequency domain features 132 and time domain features 134 using one or more ML models. In some embodiments, the sleep state classification component 140 may use gradient boosting ML techniques (e.g., XGBoost techniques). In other embodiments, the sleep state classification component 140 may use random forest ML techniques. In yet another embodiment, the sleep state classification component 140 may use neural network ML techniques (e.g., multilayer perceptrons (MLPs)). In yet another embodiment, the sleep state classification component 140 may use logistic regression techniques. In yet another embodiment, the sleep state classification component 140 may use singular value decomposition (SVD) techniques. In some embodiments, the sleep state classification component 140 may use one or more combinations of the aforementioned ML techniques. The ML techniques may be trained to classify video frames of video data about a subject into sleep states, as described in relation to Figure 3 below.

[0038] In some embodiments, the sleep state classification component 140 may use additional or alternative data / features (e.g., video data 104, ellipse data 112, frame feature data 122, etc.) to determine the frame prediction data 142.

[0039] The sleep state classification component 140 may be configured to recognize transitions between one sleep state and another based on variations between frequency domain features 132 and time domain features 134. For example, frequency domain and time domain signals for a region of interest may differ in time and frequency for awake, non-REM, and REM states. As another example, frequency domain and time domain signals for the width-to-length ratio (ratio of major axis length to minor axis length) of an interest may differ in time and frequency for awake, non-REM, and REM states. In some embodiments, the sleep state classification component 140 may use one of several features (e.g., body region of the interest or width-to-length ratio) to determine the frame prediction data 142. In other embodiments, the sleep state classification component 140 may use a combination of features from several features (e.g., body region of the interest and width-to-length ratio) to determine the frame prediction data 142.

[0040] In step 212 of process 200 shown in Figure 2A, the post-classification component 150 may perform post-classification processing to determine sleep state data 152 representing the target sleep state and transitions between sleep states during the duration of the video. The post-classification component 150 may determine the sleep state data 152 by processing frame prediction data 142, which includes sleep state markers (and corresponding confidence scores) for each video frame. The post-classification component 150 may use a transition model to determine the transition from the first sleep state to the second sleep state.

[0041] The transitions between wakefulness, non-REM, and REM states are not random but generally follow expected patterns. For example, generally, a subject transitions from wakefulness to non-REM, and then from non-REM to REM. The post-classification component 150 may recognize these transition patterns and be configured to use a transition probability matrix and probability of occurrence for a given state. The post-classification component 150 may also serve as a validation component for frame prediction data 142 determined by the sleep state classification component 140. For example, in some cases, the sleep state classification component 140 may determine that the first video frame corresponds to a wakeful state and the subsequent second video frame corresponds to a REM state. In such a case, the post-classification component 150 may update the sleep state of the first or second video frame based on the fact that the probability of a transition from wakefulness to REM is low, particularly in the short period covered by the video frame. The post-classification component 150 may use a window-based processing of video frames to determine the sleep state of a video frame. In some embodiments, the post-classification component 150 may also take into account the duration of the sleep state before transitioning to another sleep state. For example, the post-classification component 150 may determine whether the sleep state of the video frame, as determined by the sleep state classification component 140, is accurate, based on how long the non-REM state lasts for the subject in the video data 104 before transitioning to the REM state. In some embodiments, the post-classification component 150 may employ various techniques, such as statistical models (e.g., Markov models, hidden Markov models, etc.) or probabilistic models. The statistical or probabilistic model may model the dependencies between sleep states (wakeful, non-REM, and REM states).

[0042] The post-classification component 150 may process the frame prediction data 142 to determine the duration of one or more sleep states (wakeful state, non-REM state, REM state) of the subject represented in the video data 104. The post-classification component 150 may process the frame prediction data 142 to determine the frequency (number of sleep states occurring in the video data 104) of one or more sleep states (wakeful state, non-REM state, REM state) of the subject represented in the video data 104. The post-classification component 150 may process the frame prediction data 142 to determine changes in one or more sleep states of the subject. The sleep state data 152 may include the duration of one or more sleep states of the subject, the frequency of one or more sleep states of the subject, and / or changes in one or more sleep states of the subject.

[0043] The classified component 150 can output sleep state data 152, which may be a vector or matrix containing sleep state markers for each video frame of the video data 104. For example, the sleep state data 152 may include a first marker "awake" corresponding to the first video frame, a second marker "awake" corresponding to the second video frame, a third marker "non-REM" corresponding to the third video frame, a fourth marker "REM" corresponding to the fourth video frame, and so on.

[0044] System 105 may transmit sleep state data 152 to device 102 for display. The sleep state data 152 may be presented as graph data, for example, as shown in Figures 11A-D.

[0045] As described herein, in some embodiments, the automated sleep state system 100 may determine multiple body regions of interest using multiple features (determined by the feature extraction component 120), where each body region corresponds to a video frame of video data 104, and the automated sleep state system 100 may determine sleep state data 152 based on changes in the multiple body regions in the video.

[0046] As described herein, in some embodiments, the automatic sleep state system 100 may determine a plurality of aspect ratios using a plurality of features (determined by the feature extraction component 120), each aspect ratio of the plurality of aspect ratios corresponding to a video frame of video data 104, and the automatic sleep state system 100 may determine sleep state data 152 based on the changes in the plurality of aspect ratios in the video.

[0047] In some embodiments, the automated sleep state system 100 may detect the transition from a non-REM state to a REM state based on changes in the target body region or body shape, and the changes in the body region or body shape may be the result of muscle relaxation. Such transition information may be included in the sleep state data 152.

[0048] The correlation between other features derived from the video data 104 and the target sleep state, which can be used with the automated sleep state system 100, is described below in the Examples section.

[0049] In some embodiments, the automated sleep state system 100 may be configured to determine the subject's respiration / respiratory rate by processing video data 104. The automated sleep state system 100 may determine the subject's respiratory rate by processing a plurality of features (determined by the feature extraction component 120). In some embodiments, the automated sleep state system 100 may use the respiratory rate to determine the subject's sleep state data 152. In some embodiments, the automated sleep state system 100 may determine the respiratory rate based on frequency domain and / or time domain features determined by the spectral analysis component 130.

[0050] The respiratory rate of the subject may vary between sleep states and may be detected using features derived from the video data 104. For example, the subject's body region and / or width-to-length ratio may change over a period of time such that the signal representation (time or frequency) of the body region and / or width-to-length ratio may be a consistent signal of 2.5–3 Hz. Such a signal representation may resemble a ventilation waveform. The automated sleep state system 100 may process the video data 104 to extract features representing changes in body shape and / or chest size that correlate with / correspond to the subject's respiration. Such changes may be visible in the video and may be extracted as time-domain and frequency-domain features.

[0051] During non-REM states, the subject may have a specific respiratory rate, for example, 2.5–3 Hz. The automated sleep state system 100 may be configured to recognize a specific correlation between respiratory rate and sleep state. For example, the amplitude-to-length ratio signal may be more pronounced / obvious in non-REM states than in REM states. As a further example, the amplitude-to-length ratio signal may change more significantly during REM states. The exemplary correlations described above may result from the subject's respiratory rate changing more during REM states than in non-REM states. Another exemplary correlation may be low-frequency noise captured in the amplitude-to-length ratio signal during non-REM states. Such correlations may be due to the subject's actions / movements to adjust its sleep posture during non-REM states, while the subject may not move during REM states due to muscle relaxation.

[0052] At least the width-to-length ratio signal (and other signals for other features) derived from the video data 104 illustrates that the video data 104 captures the visual movement of the subject's abdomen and / or chest, which can be used to determine the subject's respiratory rate.

[0053] Figure 2B is a flowchart of a process 250 for determining sleep state data 152 for multiple subjects represented in a video, according to an embodiment of the present disclosure. One or more of the steps of process 250 may be performed in a different order / sequence than that shown in Figure 2B. One or more steps of process 250 may be performed by components of system 105 shown in Figure 1.

[0054] In step 252 of process 250 shown in Figure 2B, system 105 may receive video data 104 representing images of multiple objects (for example, as shown in Figure 6C).

[0055] In step 254, the segmentation component 110 may perform instance segmentation processing using the video data 104 to identify individual objects represented in the video. The segmentation component 110 may process the video data 104 using instance segmentation techniques to generate segmentation masks that identify individual objects in the video data 104. The segmentation component 110 may generate a first segmentation mask for a first object, a second segmentation mask for a second object, and so on, and each segmentation mask may indicate which pixels in the video frame correspond to each object. The segmentation component 110 may also determine which pixels in the video frame correspond to the background / non-object. The segmentation component 110 may process the video data 104 using one or more machine learning models to determine first segmentation data showing a first set of pixels in the video frame corresponding to a first object, second segmentation data showing a second set of pixels in the video frame corresponding to a second object, and so on.

[0056] The segmentation component 110 may track each segmentation mask for individual objects using markers such as "object 1" and "object 2" (e.g., text markers, numerical markers, or other data). The segmentation component 110 may assign each marker to a segmentation mask determined from various video frames of the video data 104, and thus track a set of pixels corresponding to individual objects across multiple video frames. The segmentation component 110 may be configured to track individual objects across multiple video frames even when objects move, change position, or change location. The segmentation component 110 may also be configured to identify individual objects when objects are in close proximity to each other, as shown in Figure 6C, for example. In some cases, objects may prefer to sleep in close proximity to each other or nearby, and instance segmentation techniques can identify individual objects even when this occurs.

[0057] Instance segmentation techniques may involve the use of computer vision techniques, algorithms, and models. Instance segmentation may involve identifying each target instance within an image / video frame and assigning a label to each pixel of the video frame. Instance segmentation may use object detection techniques to identify all objects in the video frame, classify individual objects, and localize each target instance using a segmentation mask.

[0058] In some embodiments, the system 105 may identify and track individual objects from a group of objects based on several indicators of the object, such as body size, body shape, and body / hair color.

[0059] In step 256 of process 250, the segmentation component 110 may determine the ellipse data 112 for each individual object using a segmentation mask for each individual object. For example, the segmentation component 110 may determine the first ellipse data 112 using a first segmentation mask for a first object, the second ellipse data 112 using a second segmentation mask for a second object, and so on. The segmentation component 110 may determine the ellipse data 112 in a manner similar to that described above in relation to process 200 shown in Figure 2A.

[0060] In step 258 of process 250, the feature extraction component 120 may use each of the ellipse data 112 to determine multiple features for individual objects. These multiple features may be frame-based features, that is, they may be for each individual video frame of the video data 104, and may be provided as frame feature data 122. The feature extraction component 120 may use the first ellipse data 112 to determine a first frame feature data 122 corresponding to a first object, use the second ellipse data 112 to determine a second frame feature data 122 corresponding to a second object, and so on. The feature extraction component 120 may determine the frame feature data 122 in the same manner as described above in relation to process 200 shown in Figure 2A.

[0061] In step 260 of process 250, the spectral analysis component 130 may perform spectral analysis using multiple features (in the same manner as described above in relation to process 200 shown in Figure 2A) to determine frequency-domain features 132 and time-domain features 134 for individual objects. The spectral analysis component 130 may determine a first frequency-domain feature 132 for a first object, a second frequency-domain feature 132 for a second object, a first time-domain feature 134 for a first object, a second time-domain feature 134 for a second object, and so on.

[0062] In step 262 of process 250, the sleep state classification component 140 may process the respective frequency domain features 132 and time domain features 134 for each individual subject to determine the sleep prediction for each subject for each video frame of the video data 104 (in the same manner as described above in relation to process 200 shown in Figure 2A). For example, the sleep state classification component 140 may determine first frame prediction data 142 for a first subject, second frame prediction data 142 for a second subject, and so on.

[0063] In step 264 of process 250, the post-classification component 150 may perform post-classification processing (in the same manner as described above in relation to process 200 shown in Figure 2A) to determine sleep state data 152 representing the duration of the video and the sleep state of each subject during the transition between sleep states. For example, the post-classification component 150 may determine first sleep state data 152 for a first subject, second sleep state data 152 for a second subject, and so on.

[0064] In this way, using instance segmentation techniques, system 105 may identify multiple objects in the video and determine the sleep state data of each individual object using the corresponding feature data (and other data). By being able to identify each object even when they are close to one another, system 105 can determine the sleep state of multiple objects housed together (i.e., multiple objects contained in the same enclosure). One advantage of this is that objects can be observed in natural conditions and natural environments, which may involve cohabitation with other objects. In some cases, the behavior of other objects may also be identified / studied based on the cohabitation of objects (e.g., the effect of cohabitation on sleep state, whether cohabitation causes objects to follow the same / similar sleep patterns, etc.). Another advantage is that sleep state data for multiple objects can be determined by processing the same / single video, which can reduce the resources used (e.g., time, computational resources, etc.) compared to the resources used to process multiple separate videos, each representing one object.

[0065] Figure 3 conceptually illustrates the components and data that may be used to constitute the sleep state classification component 140 shown in Figure 1. As described herein, the sleep state classification component 140 may include one or more ML models for processing features derived from the video data 104. The ML models may be trained / configured using various types of training data and training techniques.

[0066] In some embodiments, spectral training data 302 may be processed by a model building component 310 to train / configure a trained classifier 315. In some embodiments, the model building component 310 may also process EEG / EMG training data to train / configure the trained classifier 315. The trained classifier 315 may be configured to determine sleep state labels for video frames based on one or more features corresponding to the video frames.

[0067] The spectral training data 302 may include frequency-domain and / or time-domain signals for one or more features of an object represented in the video data used for training. These features may correspond to features determined by the feature extraction component 120. For example, the spectral training data 302 may include frequency-domain and / or time-domain signals corresponding to body regions of an object in the video. The frequency-domain and / or time-domain signals may be annotated / labeled with corresponding sleep states. The spectral training data 302 may also include frequency-domain and / or time-domain signals for other features such as the width-to-length ratio of the object, the width of the object, the length of the object, the position of the object, Hu image moments, and other features.

[0068] The EEG / EMG training data 304 may also be electroencephalogram (EEG) data and / or electromyogram (EMG) data corresponding to the subject used to train / construct the sleep state classification components 140. The EEG data and / or EMG data may be annotated / labeled with the corresponding sleep state.

[0069] The spectral training data 302 and EEG / EMG training data 304 may correspond to the same sleep state. The model building component 310 may correlate the spectral training data 302 and EEG / EMG training data 304 to train / configure the trained classifier 315 and identify sleep states from the spectral data (frequency domain features and time domain features).

[0070] The training dataset may be unbalanced because subjects experience more non-REM states than REM states during sleep. To train / construct the trained classifier 315, a balanced training dataset may be generated that includes equal / similar numbers of REM states, non-REM states, and wakefulness states.

[0071] subject Some aspects of the present invention involve determining sleep state data of a subject. As used herein, the term “subject” may refer to humans, non-human primates, cattle, horses, pigs, sheep, goats, dogs, cats, birds, rodents, or other suitable vertebrates or invertebrates. In certain embodiments of the present invention, the subject is a mammal, and in certain embodiments of the present invention, the subject is a human. In some embodiments, the subject used in the methods of the present invention is a rodent, including but not limited to mice, rats, gerbils, hamsters, etc. In some embodiments of the present invention, the subject is a normal, healthy subject, and in some embodiments, the subject is known to have a disease or condition, is at risk of having a disease or condition, or is suspected of having a disease or condition. In certain embodiments of the present invention, the subject is an animal model of a disease or condition. For example, in some embodiments of the present invention, for lack of limit, the subject is a mouse that forms an animal model of sleep apnea.

[0072] In non-limiting examples, subjects evaluated by the methods and systems of the present invention may be subjects having, suspected of having, one or more of the following conditions: sleep apnea, insomnia, narcolepsy, brain injury, depression, mental illness, neurodegenerative disease, restless legs syndrome, Alzheimer's disease, Parkinson's disease, neurological conditions that can alter the status of sleep states, and metabolic disorders or conditions that can alter sleep states, and / or animal models relating to such conditions. A non-limiting example of a metabolic disorder or condition that can alter sleep states is a high-fat diet. Additional physical conditions may also be evaluated using the methods of the present invention, non-limiting examples of which include obesity, overweight, the effects of administered drugs, and / or the effects of alcohol consumption. Additional diseases and conditions, including, but not limited to, sleep disorders resulting from chronic diseases, drug abuse, injuries, etc., may also be evaluated using the methods of the present invention.

[0073] Furthermore, the methods and systems of the present invention may be used to evaluate subjects or test subjects who do not have one or more of the following conditions: sleep apnea, insomnia, narcolepsy, brain injury, depression, mental illness, neurodegenerative disease, restless legs syndrome, Alzheimer's disease, Parkinson's disease, neurological conditions that can alter sleep state status, and metabolic disorders or conditions that can alter sleep state. In some embodiments, the methods of the present invention are used to evaluate the sleep state of subjects who are obese, overweight, and do not consume alcohol. Such subjects can serve as control subjects, and the results of evaluations using the methods of the present invention can be used as control data.

[0074] In some embodiments of the methods of the present invention, the subject is a wild-type subject. As used herein, the term “wild-type” means the phenotype and / or genotype of a species in a typical form as it occurs in nature. In certain embodiments of the present invention, the subject is a non-wild-type subject, for example, a subject having one or more genetic modifications compared to the wild-type genotype and / or phenotype of the subject species. In some examples, the difference in the subject's genotype / phenotype compared to the wild-type is due to hereditary (germline) mutations or acquired (somatic) mutations. Factors that may result in a subject exhibiting one or more somatic mutations include, but are not limited to, environmental factors, toxins, ultraviolet radiation, spontaneous errors in cell division, radiation, maternal infection, and chemicals, as well as teratogenic events.

[0075] In certain embodiments of the methods of the present invention, the subject is a genetically modified organism, also referred to as a genetically engineered subject. The genetically engineered subject may include pre-selected and / or intentional genetic modifications, and thus exhibit one or more genotypes and / or phenotypic traits that differ from those in the unmanipulated subject. In some embodiments of the present invention, by using conventional genetic engineering techniques, it is possible to create a genetically engineered subject that exhibits genotype and / or phenotypic differences compared to an unmanipulated subject of the same species. As a non-limiting example, a genetically engineered mouse may have functional gene products present in the mouse at deficient or reduced levels, and the phenotype of the genetically engineered mouse may be evaluated using the methods or systems of the present invention, and the results may be compared with results obtained from a control (control result).

[0076] In some embodiments of the present invention, subjects may be monitored using the automated sleep state determination method or system of the present invention to detect the presence or absence of sleep disorders or conditions. In certain embodiments of the present invention, test subjects, which are animal models of sleep disorders, may be used to evaluate the test subjects' response to sleep disorders. In addition, test subjects, which are animal models of sleep and / or activity disorders, may be administered candidate therapeutic agents or methods monitored using the automated sleep state determination method and / or system of the present invention, and the results can be used to determine the effectiveness of candidate therapeutic agents for treating the conditions.

[0077] As described elsewhere in this specification, the methods and systems of the present invention may be configured to determine the sleep state of an object regardless of the object's physical characteristics. In some embodiments of the present invention, one or more of the object's physical characteristics may be pre-identified characteristics. For example, though not intended to be limiting, pre-identified physical characteristics may be one or more of body type, build, coat color, sex, age, and disease or pathological phenotype.

[0078] Testing and screening of control and candidate compounds Results obtained with respect to a subject using the method or system of the present invention can be compared with results with respect to a control. The method of the present invention can also be used to evaluate phenotypic differences between a subject and a control. Thus, some embodiments of the present invention provide a method for determining the presence or absence of one or more changes in sleep states in a subject compared to a control. Some embodiments of the present invention involve using the method of the present invention to identify phenotypic features of a disease or pathological condition, and in certain embodiments of the present invention, automated phenotypic analysis is used to evaluate the effect of a candidate therapeutic compound on a subject.

[0079] Results obtained using the methods or systems of the present invention can be advantageously compared with controls. In some embodiments of the present invention, one or more subjects can be evaluated using the methods of the present invention, and then the subjects can be re-tested after being administered a candidate therapeutic compound. In relation to subjects evaluated using the methods or systems of the present invention, the terms “subject” and “test subject” may be used herein and are interchangeable. In certain embodiments of the present invention, results obtained using the methods of the present invention for evaluating test subjects are compared with results obtained from methods performed on other test subjects. In some embodiments of the present invention, the results of test subjects are compared with the results of sleep state assessments performed on test subjects at different times. In some embodiments of the present invention, results obtained using the methods of the present invention for evaluating subjects are compared with control results.

[0080] Where used herein, the control result may be a predetermined value that can take various forms. It may be a single cutoff value, such as a median or mean. This may be established based on a comparison group, such as subjects evaluated using the system or method of the present invention under similar conditions to the test subjects, where the test subjects are administered the candidate therapeutic agent and the comparison group is not. Another example of a comparison group may include subjects known to have the disease or condition and a group that does not have the disease or condition. Another comparison group may include subjects with a family history of the disease or condition and subjects from the group that does not have such a family history. The predetermined value may be set, for example, so that the tested population is grouped evenly (or unevenly) based on the test results. Those skilled in the art will be able to select appropriate control groups and values ​​for use in the comparison method of the present invention.

[0081] Subjects evaluated using the method or system of the present invention can be monitored for changes in one or more sleep state characteristics occurring under the test conditions compared to control conditions. In a non-limiting example, changes in a subject may include, but are not limited to, one or more sleep state characteristics such as the duration of a sleep state, the time interval between two sleep states, the number of one or more sleep states during a sleep period, the ratio of RM sleep states to NRM sleep states, or the time before entering a sleep state. The method and system of the present invention can be used on a test subject to evaluate the effects of the disease or condition being tested, and can also be used to evaluate the effectiveness of a candidate therapeutic agent. As a non-limiting example of the method of the present invention for evaluating the presence or absence of changes in one or more sleep state characteristics of a subject as a means of identifying the effectiveness of a candidate therapeutic agent, a subject known to have a disease or condition affecting the subject's sleep state is evaluated using the method of the present invention. The subject is then administered the candidate therapeutic agent and evaluated again using the method. The presence or absence of changes in the subject's results indicates, respectively, the presence or absence of an effect of the candidate therapeutic agent on the disease or condition affecting the sleep state.

[0082] In some embodiments of the present invention, it will be understood that a subject can serve as a control of itself by, for example, being evaluated two or more times using the method of the present invention and comparing the results obtained from two or more of the different evaluations. The methods and systems of the present invention may be used to evaluate the progression or regression of a disease or condition of a subject by using embodiments of the methods or systems of the present invention to identify and compare changes in phenotypic features, such as sleep state characteristics, of the subject over time using two or more evaluations of the subject.

[0083] Exemplary devices and systems One or more components of the automated sleep state system 100 may implement an ML model that can take many forms, including an XgBoost model, a random forest model, a neural network, a support vector machine, or other models, or a combination of any of these models.

[0084] Various machine learning techniques may be used to train and operate a model to perform various steps described herein, such as determining segmentation masks, determining elliptic data, determining feature data, and determining sleep state data. The model may be trained and operated according to various machine learning techniques. Such techniques may include, for example, neural networks (deep neural networks and / or recurrent neural networks, etc.), inference engines, trained classifiers, etc. Examples of trained classifiers include support vector machines (SVMs), neural networks, decision trees, AdaBoost ("Adaptive Boost") in combination with decision trees, and random forests. Focusing on SVM as an example, an SVM is a supervised learning model that has an association learning algorithm that analyzes data and recognizes patterns in the data, and this supervised learning model is commonly used for classification and regression analysis. Given a set of training examples, each marked as belonging to one of two categories, the SVM training algorithm constructs a model that assigns new examples to one of the categories or the other, making it a non-stochastic binary linear classifier. A more complex SVM model may be constructed with a training set that identifies three or more categories, and the SVM determines which category is most similar to the input data. The SVM model may map examples of distinct categories so that they are separated by distinct gaps. New examples are then mapped into the same space and predicted to belong to a category based on which side of the gap they fall on. The classifier may issue a “score” indicating which category the data best matches. The score may provide an index of how closely the data matches a category.

[0085] A neural network may contain several layers, from an input layer to an output layer. Each layer is configured to take a specific type of data as input and to output a different type of data. The output from one layer is introduced as input to the next layer. Although the values ​​of the input / output data in a particular layer are unknown until the neural network actually operates at runtime, the data describing the neural network describes the structure, parameters, and operation of the neural network's layers.

[0086] One or more of the intermediate layers of a neural network may also be known as hidden layers. Each node in a hidden layer is connected to each node in the input layer and to each node in the output layer. If the neural network contains multiple intermediate layers, each node in a hidden layer will be connected to each node in the next higher layer and the next lower layer. Each node in the input layer represents a potential input to the neural network, and each node in the output layer represents a potential output from the neural network. Each connection from one node to another in the next layer may be associated with a weight or score. The neural network may output a single output or a weighted set of possible outputs. Different types of neural networks may be used, e.g., recurrent neural networks (RNNs), convolutional neural networks (CNNs), deep neural networks (DNNs), long-short-term memory (LSTMs), and / or others.

[0087] The processing performed by a neural network is determined by the learned weights of each node's input and the network's structure. Given a specific input, the neural network determines the output layer by layer until the output layer of the entire network is computed.

[0088] Connection weights can be initially learned by the neural network during training, associating a given input with a known output. The training data set provides various training examples into the network. In each example, the weights of the correct connections from input to output are typically set to 1, and all connections are weighted to 0. When the training data examples are processed by the neural network, the input may be sent to the network and compared with the associated output to determine how to compare the network's performance to the target performance. Training techniques such as backpropagation may be used to update the neural network's weights and reduce errors that occur when the neural network processes the training data.

[0089] To apply machine learning techniques, the machine learning process itself needs to be trained. In this case, to train machine learning components such as either the first or second model, it is necessary to establish a "ground truth" regarding the training examples. In machine learning, the term "ground truth" refers to the accuracy of classification on a training set for supervised learning techniques. Various techniques, including backpropagation, statistical learning, supervised learning, semi-supervised learning, stochastic learning, or other known techniques, may be used to train the model.

[0090] Figure 4 is a block diagram conceptually illustrating a device 400 that may be used with the system. Figure 5 is a block diagram conceptually illustrating exemplary components of a remote device such as system 105 that may assist in processing video data, identifying the actions of an object, etc. System 105 may include one or more servers. As used herein, “server” may refer to a conventional server as understood in a server / client computing structure, but may also refer to a number of different computing components that may assist in the operations described herein. For example, a server may include one or more physical computing components (such as a rack server) that are physically and / or via a network connected to other devices / components and capable of performing computing operations. A server may also include one or more virtual machines that emulate a computer system and run on one device or across multiple devices. A server may also include other combinations of hardware, software, firmware, or similar for performing the operations described herein. The server may be configured to operate using one or more of the following computing technologies: client-server model, computer bureau model, grid computing technology, fog computing technology, mainframe technology, utility computing technology, peer-to-peer model, sandbox technology, or other computing technologies.

[0091] Multiple systems 105 may be included in the overall system of this disclosure, such as one or more systems 105 for determining elliptic data, one or more systems 105 for determining frame features, one or more systems 105 for determining frequency domain features, one or more systems 105 for determining time domain features, one or more systems 105 for determining frame-based sleep marker prediction, and one or more systems 105 for determining sleep state data. When in operation, each of these systems may include computer-readable instructions and computer-executable instructions present on each device 105, as further described below.

[0092] Each of these devices (400 / 105) may include one or more controllers / processors (404 / 504) which may include a central processing unit (CPU) for processing data and computer-readable instructions, and memory (406 / 506) for storing data and instructions for each device. The memory (406 / 506) may individually include volatile random access memory (RAM), non-volatile read-only memory (ROM), non-volatile magnetoresistive memory (MRAM), and / or other types of memory. Each device (400 / 105) may also include data storage components (408 / 508) for storing data and controller / processor executable instructions. Each data storage component (408 / 508) may individually include one or more non-volatile storage types, such as magnetic storage, optical storage, and solid-state storage. Each device (400 / 105) may also be connected to removable or external non-volatile memory and / or storage (such as removable memory cards, memory key drives, or network storage) via its respective input / output device interface (402 / 502).

[0093] Computer instructions for operating each device (400 / 105) and its various components may be executed by the controller / processor (404 / 504) of each device, using memory (406 / 506) as temporary "working" storage at runtime. Computer instructions for devices may be stored non-temporarily in non-volatile memory (406 / 506), in storage (408 / 508), or in external devices. Alternatively, some or all of the executable instructions may be embedded in the hardware or firmware on each device, in addition to or instead of software.

[0094] Each device (400 / 105) includes an input / output device interface (402 / 502). Various components may be connected via the input / output device interface (402 / 502), as will be further described below. In addition, each device (400 / 105) may include an address / data bus (424 / 524) for transmitting data between the components of each device. Each component within a device (400 / 105) may also be connected directly to other components, in addition to (or instead of) being connected to other components via the bus (424 / 524).

[0095] Referring to Figure 4, device 400 may include an input / output device interface 402 for connecting to various components, such as an audio output component, including a speaker 412, a wired or wireless headset (not shown), or other components capable of outputting audio. Device 400 may also include a display 416 for displaying content. Device 400 may further include a camera 418.

[0096] The input / output device interface 402 may be connected to one or more networks 199 via antenna 414, via wireless local area network (WLAN) (such as WiFi) radio, Bluetooth, and / or wireless network radio, such as radios capable of communicating with wireless communication networks such as Long-Term Evolution (LTE) networks, WiMAX networks, 3G networks, 4G networks, and 5G networks. Wired connections such as Ethernet may also be supported. The system may be distributed across the network environment via network 199. The I / O device interfaces (402 / 502) may also include communication components that enable the exchange of data between devices such as different physical servers in a collection of servers or other components.

[0097] The components of device 400 or system 105 may include their own dedicated processor, memory, and / or storage. Alternatively, one or more components of device 400 or system 105 may utilize the I / O interface (402 / 502), processor (404 / 504), memory (406 / 506), and / or storage (408 / 508) of device 400 or system 105, respectively.

[0098] As described above, multiple devices may be employed within a single system. In such a multi-device system, each device may contain different components for performing different aspects of the system's processing. Multiple devices may contain overlapping components. The components of device 400 and system 105 described herein are illustrative and may be configured as standalone devices or may be included as a whole or in part as components of a larger device or system.

[0099] The concepts disclosed herein may be applied in a number of different devices and computer systems, including, for example, general-purpose computing systems, video / image processing systems, and distributed computing environments.

[0100] The above-described aspects of this disclosure are illustrative. They have been selected to illustrate the principles and uses of this disclosure and are not intended to be exhaustive or limit the disclosure. Many modifications and variations of the disclosed aspects may be apparent to those skilled in the art. Those skilled in the art in the computer and speech processing fields will recognize that the components and process steps described herein may be interchangeable with other components or steps, or with combinations of components or steps, and that the advantages and benefits of this disclosure may still be achieved. Furthermore, it will be apparent to those skilled in the art that this disclosure may be carried out without some or all of the specific details and steps disclosed herein.

[0101] The disclosed system configuration may be implemented as a computer method or as a manufactured article such as a memory device or a non-temporary computer-readable storage medium. The computer-readable storage medium may be computer-readable and may contain instructions for causing a computer or other device to perform the processes described herein. The computer-readable storage medium may be implemented by volatile computer memory, non-volatile computer memory, hard drives, solid-state memory, flash drives, removable disks, and / or other media. In addition, the system components may be implemented as firmware or hardware. [Examples]

[0102] Example 1. Development of a mouse sleep state classifier model method Animal husbandry, surgery, and experimental setup Sleep studies were conducted on 17 C57BL / 6J (The Jackson Laboratory, Bar Harbor, Maryland) male mice. C3H / HeJ (The Jackson Laboratory, Bar Harbor, Maryland, USA) mice were also imaged non-surgically for characteristic examination. All mice were acquired at 10–12 weeks of age. All animal studies were conducted in accordance with the National Institutes of Health guidelines for the management and use of laboratory animals and were approved by the University of Pennsylvania Animal Experimentation Board. The study methods were performed as previously described [Pack, AI et al. Physiol. Genomics. 28(2):232-238 (2007); McShane, BB et al., Sleep. 35(3):433-442 (2012)].

[0103] In short, mice were individually housed in standard, lidless mouse cages (6x6 inches). Each cage was extended to 12 inches in height to prevent mice from jumping out. This design allowed for simultaneous assessment of mouse behavior via video and sleep / wake phases via EG / EMG recording. Animals were given free access to water and food and followed a 12-hour light-dark cycle. During the light phase, the lux level at the bottom of the cage was 80 lux. For EEG recording, four silver ball electrodes were placed in the skull (two in the frontal region and two in the parietal-temporal region). For EMG recording, two silver wires were sutured to the dorsal nuchal muscle. All lead wires were subcutaneously placed in the center of the skull and connected to a plastic socket pedestal (Plastics One, Torrington, Connecticut) fixed to the skull with dental cement. Electrodes were implanted under general anesthesia. After surgery, animals were given a 10-day recovery period before recording.

[0104] EEG / EMG acquisition For EEG / EMG recording, the raw signals were read and amplified (20,000x) using Grass Gamma Software (Astro-Med, West Warwick, Rhode Island). The EEG signal filter settings were a low cutoff frequency of 0.1 Hz and a high cutoff frequency of 100 Hz. The EMG signal filter settings were a low cutoff frequency of 10 Hz and a high cutoff frequency of 100 Hz. The recordings were digitized at 256 Hz samples / second / channel.

[0105] Acquisition of video High-quality video data was recorded under both day and night conditions using a Raspberry Pi 3 Model B (Raspberry Pi Foundation, Cambridge, UK) night vision setup. A SainSmart (SainSmart, Las Vegas, Nevada) infrared night vision surveillance camera was used, accompanied by infrared LEDs to illuminate the scene when there was no visible light. The camera was mounted 18 inches above the floor of the home cage, providing a top-down view of the mice for observation. During the day, the video data was in color. At night, the video data was in black and white. Video was recorded at a resolution of 1920 x 1080 pixels and 30 frames per second using v4l2-ctl capture software. For information on the v412-CTL software, see, for example, www.kernel.org / doc / html / latest / userspace-api / media / v4l / v4l2.html or alternatively, the abridged version: www.kernel.org / .

[0106] Synchronization of video and EEG / EMG data The video and EEG / EMG data were synchronized using the computer clock time. The EEG / EMG data acquisition computer was used as the source clock. Visual cues were added to the video at known points in time on the EEG / EMG computer. The visual cues typically lasted for 2-3 frames in the video, suggesting that the possible error in synchronization could be up to 100 milliseconds. Since the EEG / EMG data was analyzed at 10-second (10s) intervals, any possible errors in temporal alignment would be negligible.

[0107] EEG / EMG annotations for training data 24-hour synchronized video and EEG / EMG data were collected from 17 C57BL / 6J male mice aged 10–12 weeks from the Jackson Laboratory. Both EEG / EMG data and video were divided into 10-second epochs, each scored by a trained score recorder, and labeled as REM, non-REM, or awake stage based on EEG and EMG signals. A total of 17,700 EEG / EMG epochs were scored by human experts. Of these, 48.3%+ / -6.9% were annotated as awake, 47.6%+ / -6.7% as non-REM, and 4.1%+ / -1.2% as REM. In addition, SPINDLE's method was applied for a second annotation [Miladinovic, D. et al., PLoS Comput Biol. 15, e1006968 (2019)]. Similar to human experts, 52% of epochs were annotated as awake, 44% as non-REM, and 4% as REM. Since SPINDLE annotated 4-second (4s) epochs, three consecutive epochs were combined and compared to a 10-second epoch, and epochs were compared only when the three 4-second epochs remained unchanged. Where specific epochs correlated, the agreement between human annotations and SPINDLE was 92% (89% awake, 95% non-REM, 80% REM).

[0108] Data preprocessing Starting with video data, the aforementioned segmentation neural network architecture was applied to generate a mouse mask. [Webb JM and Fu YH., Curr. Opin. Neurobiol. 69:19-24 (2021)]. 313 frames were annotated to train the segmentation network. A 4x4 diamond dilation followed by a 5x5 diamond deflation filter was applied to the raw predicted segmentation. These steady-state operations were used to improve segmentation quality. Using the predicted segmentation and the resulting ellipse fit, various per-frame image measurement signals were extracted from each frame, as shown in Table 1.

[0109] All of these measurements (Table 1) were computed by applying OpenCV contour functions to a neural network predictive segmentation mask. The OpenCV functions used included fitEllipse, contourArea, arcLength, moments, and getHuMoments. For more information on the OpenCV software, see, for example, / / opencv.org. Using all the measured signal values ​​in an epoch, a set of 20 frequency and time-domain features was derived (Table 2). These were computed using standard signal processing approaches and can be found in exemplary code [github.com / KumarLabJax / MouseSleep].

[0110] Training of a classifier To address the inherent imbalance in the dataset—namely, a greater number of non-REM epochs compared to REM sleep—a balanced dataset was generated by randomly selecting an equal number of REM, non-REM, and wakefulness epochs. A cross-validation approach was used to evaluate the classifier's performance. All epochs from 13 animals in the balanced training dataset were randomly selected, and the imbalanced data from the remaining 4 animals was used for testing. The process was repeated 10 times to generate various accuracy measurements. This approach allowed us to observe the performance on actual imbalanced data while effectively leveraging the training of the classifier on balanced data.

[0111] Prediction after processing To improve prediction quality by integrating larger amounts of temporal information, a Hidden Markov Model (HMM) approach was applied. The HMM model corrects erroneous predictions made by the classifier by integrating the probabilities of sleep state transitions, and thus yields more accurate prediction results. The hidden states of the HMM model are sleep stages, while the observables are obtained from the probability vector results of the XgBoost algorithm. The transition matrix was empirically calculated by applying the Viterbi algorithm [Viterbi AJ (April 1967) IEEE Transactions on Information Theory vol. 13(2): 260-269] to the training set sequence of sleep states, and then estimating the sequence of stages by considering the sequence of out-of-bag class votes in XgBoost. In this study, the transition matrix is ​​a 3x3 matrix T={S_ij}, where S_ij represents the probability of transitioning from state S_i to state S_j. T (Table 2).

[0112] Classifier performance analysis Performance was evaluated using several metrics of accuracy and classification performance: precision, recall, and F1 score. Precision was defined as the ratio of the epochs in which a given sleep stage was classified by both the classifier and the human score recorder to all epochs in which the classifier assigned that sleep stage. Recall was defined as the ratio of the epochs in which a given sleep stage was classified by both the classifier and the human score recorder to all epochs in which the human score recorder classified that given sleep stage. F1 was a combination of precision and recall, measuring the harmonic mean of recall and precision. The mean and standard deviation of the accuracy and performance matrix were calculated from 10-fold cross-validation.

[0113] result Experimental Design As shown in the schematic diagram in Figure 6A, the objective of the research described herein was to quantify the validity of using only video data to classify the sleep states of mice. The experimental paradigm was designed to utilize the current standard criterion for sleep state classification, EEG / EMG recordings, as markers for training and evaluating a visual classifier. In total, synchronized EEG / EMG and video data were recorded for 17 mice (24 hours per animal). The data was divided into 10-second epochs. Each epoch was manually scored by a human expert. Simultaneously, features from the video data that could be used in a machine learning classifier were designed. These features were built on frame-by-frame measurements describing the visual appearance of the animals in individual video frames (Table 1). Signal processing techniques were then applied to the frame-by-frame measurements to integrate the temporal information and generate a set of features for use in the machine learning classifier (Table 2). Finally, the human-labeled datasets were split by providing individual animals to training and validation datasets (80:20, respectively). Using the training dataset, a machine learning classifier was trained to classify 10-second epochs of video into three states: wakefulness, non-REM sleep, and REM sleep. The provided set of animals was used in the validation dataset to quantify the classifier's performance. When separating the validation set from the training set, the entire animal data was provided to ensure that the classifier generalized well across animals, rather than learning to predict well only for the animals shown. [Table 1] [Table 2]

[0114] Features per frame Computer vision techniques were applied to extract detailed visual measurements of the mouse in each frame. The primary computer vision technique used was pixel segmentation related to mouse pixels and background pixels (Figure 6B). The segmentation neural network was trained as an approach that works well under light and dark conditions, as well as in dynamic and challenging environments such as the moving mats found in the mouse arena [Webb, JM and Fu, YH., Curr. Opin. Neurobiol. 69:19-24 (2021)]. Segmentation also allowed for the removal of EEG / EMG cables emanating from devices on each mouse's head, thus not affecting visual measurements that used information about head movement. The segmentation network predicted pixels that were mouse-only, and therefore the measurements were based solely on mouse movement, and not on the movement of wires connected to the mouse's skull. Frames randomly sampled from all the footage were annotated using the aforementioned network to achieve this high-quality segmentation and elliptic fit [Geuther, BQ et al., Commun. Biol. 2:124 (2019)] (Figure 6B). The neural network required only 313 annotated frames to achieve good performance in segmenting mice. An example of the segmentation network's performance (not shown) is shown, where pixels predicted not to be mice are colored red and pixels predicted to be mice are colored blue on the original footage. Following segmentation, 16 measurements from the neural network-predicted segmentation describing the shape and position of the mice were calculated (Table 1). These included the large, small, and ratios of the mice from the elliptic fit describing the shape of the mice. The mouse position (x, y) and the change in x, y (dx, dy) were extracted for the center of the elliptic fit. We also calculated seven Hu image moments (HU0-6) of segmented mouse regions (m00), periphery, and rotation-invariant regions [Scammell, TE et al., Neuron. 93(4):747-765(2017)].Hu image moments are a numerical description of mouse segmentation obtained through the integration and linear combination of central image moments [Allada, R. and Siegel, JM Curr. Biol. 18(15):R670-R679 (2008)].

[0115] Time-frequency characteristics Next, time- and frequency-based analysis was performed for each 10-second epoch using the features per frame. This analysis allowed for the integration of temporal information by applying signal processing techniques. As shown in Table 3, six time-domain features (kurtosis, mean, median, standard deviation, maximum, and minimum for each signal) and 14 frequency-domain features (kurtosis of power spectral density, distortion of power spectral density, mean power spectral density at 0.1–1Hz, 1–3Hz, 3–5Hz, 5–8Hz, and 8–15Hz, total power spectral density, maximum, minimum, mean, and standard deviation of power spectral density) were extracted for the features per frame of each epoch, resulting in a total of 320 features (16 measurements × 20 time-frequency features) for each 10-second epoch. [Table 3]

[0116] These spectral window features were visually examined to determine whether they differed between awake, REM, and non-REM states. Figures 7A-B show representative epoch examples of the m00 (domain, Figure 7A) and wl_ratio (width-to-minor axis ratio of an ellipse, Figure 7B) features, which differ in time and frequency domains for awake, non-REM, and REM states. The raw signals of m00 and wl_ratio show distinct oscillations in non-REM and REM states (left panel, Figures 7A and 7B), which can be seen in the FFT (center panel, Figures 7A and 7B) and autocorrelation (right panel, Figures 7A and 7B). A single dominant frequency was present in non-REM epochs, while a broader peak was present in REM states. In addition, the FFT peak frequency varied slightly between non-REM (2.6 Hz) and REM (2.9 Hz), and generally, more regular and consistent oscillations were observed in non-REM epochs than in REM epochs. Therefore, the initial investigation of features revealed differences between sleep states and gave us confidence that useful indicators for use in a visual sleep classifier were encoded in the features.

[0117] breathing rate Previous studies in both humans and rodents have demonstrated that respiration and movement differ between sleep stages [Stradling, JR et al., Thorax. 40(5):364-370 (1985); Gould, GA et al., Am. Rev Respir. Dis. 138(4):874-877 (1988); Douglas, NJ et al., Thorax. 37(11):840-844 (1982); Kirjavainen, T. et al., J. Sleep. Res. 5(3):186-194 (1996); Friedman, L. et al., J. Appl. Physiol. 97(5):1787-1795 (2004)]. Examining the characteristics of m00 and wl_ratio revealed a consistent signal of 2.5–3 Hz appearing as a ventilation waveform (Figures 7A–7B). Video analysis revealed changes in body shape and chest size due to respiration, which may have been captured by time-frequency features. To visualize this signal, continuous wavelet transform (CWT) spectrograms were performed on the wl_ratio feature (Figure 8A, top panel). To summarize the data from these CWT spectrograms, dominant signals were identified (Figure 8A, respective bottom left panels), and histograms of the dominant frequencies of the signals were plotted (Figure 8A, respective bottom right panels). The mean and variance of the frequencies included in the dominant signals were calculated from the corresponding histograms.

[0118] Previous studies have demonstrated that C57BL / 6J mice have a respiratory rate of 2.5–3 Hz during non-REM states [Friedman, L. et al., J.Appl. Physiol. 97(5):1787-1795(2004); Fleury Curado, T. et al., Sleep. 41(8):zsy089 (2018)]. In a study of a long sleep period (10 minutes) including both REM and non-REM states, the wl_ratio signal was clearly present in both, although it was more pronounced in non-REM states than in REM states (Figure 8B). In addition, since REM states induce higher and more variable respiratory rates than non-REM states, the signal varied more within the range of 2.5–3.0 Hz during REM states. Low-frequency noise in this non-REM state signal was also observed due to greater mouse movement, such as adjusting its sleep posture. This suggests that the wl_ratio signal captures visual movement in the mouse's abdomen.

[0119] Verification of respiratory rate To confirm that the signals observed in the m00 and wl_ratio features during REM and non-REM epochs were abdominal movements and correlated with respiratory rate, genetic validation tests were performed. C3H / HeJ mice have an awakening respiratory frequency approximately 30% lower than that of C57BL / 6J mice, and it has been previously demonstrated that the correlations for C57BL / 6J and C3H / HeJ range from 4.5 Hz vs. 3.18 Hz [Berndt, A. et al., Physiol. Genomics. 43(1):1-11 (2011)], 3.01 vs. 2.27 Hz [Groeben, H. et al., Br. J. Anaesth. 91(4):541-545 (2003)], and 2.68 Hz vs. 1.88 Hz [Vium, Inc., Breathing Rate Changes Monitored Non-Invasively 24 / 7. (2019)]. Unequipped C3H / HeJ mice (5 males, 5 females) were video recorded, and sleep epochs were identified by applying the classic sleep / wake movement empirical rule (distance traveled) [Pack, AI et al. Physiol. Genomics. 28(2):232-238 (2007)]. Epochs were conservatively selected within the minimum 10% quantile relative to movement. Annotated C57BL / 6J EEG / EMG data were used to confirm that the movement-based cutoff could accurately identify sleep periods. Using annotated EEG / EMG data from C57BL / 6J mice, this cutoff was found to primarily distinguish between non-REM and REM epochs (Figure 9A). The selected epochs in the annotated data consisted of 90.2% non-REM, 8.1% REM, and 1.7% wakefulness epochs. Therefore, as expected, this movement-based cutoff method correctly distinguished between sleep / wake but not between REM / non-REM. From these low-activity sleep epochs, the mean value of the dominant frequency in the wl_ratio signal was calculated. This measurement was chosen for its sensitivity to chest region movement. The distribution of the mean dominant frequency for each mouse was plotted, and a consistent distribution was observed among the animals.For example, C57BL / 6J animals had an average frequency oscillation range of 2.2–2.8 Hz, while C3H / HeJ animals had a range of 1.5–2.0 Hz, and the respiratory rate of C3H / HeJ was approximately 30% lower than that of C57BL / 6J. This is a statistically significant difference between the two strains, namely C57BL / 6J and C3H / HeJ (p<0.001, Figure 9B), and is within a range similar to those previously reported [Berndt, A. et al., Physiol. Genomics. 43(1):1-11(2011); Groeben, H. et al., Br. J. Anaesth 91(4):541-545 (2003); Vium, Inc., Breathing Rate Changes Monitored Non-Invasively 24 / 7. (2019)]. Therefore, using this genetic validation method, it was concluded that the observed signal is strongly correlated with respiratory rate.

[0120] In addition to genetically determined overall changes in respiratory frequency, sleep respiration has been shown to be more organized and less variable during non-REM sleep than during REM sleep in both humans and rodents [Mang, GM et al., Sleep. 37(8):1383-1392 (2014); Terzano, MG et al., Sleep. 8(2):137-145 (1985)]. It was hypothesized that the detected respiratory signal would show greater variability in REM epochs than in non-REM epochs. EEG / EEG-annotated C57BL / 6J data were examined to determine if there were changes in the variability of the CWT peak signal across epochs spanning REM and non-REM states. Using only C57BL / 6J data, epochs were divided by non-REM and REM states, and the variability of the CWT peak signal was observed (Figure 9C). Non-REM states showed a smaller standard deviation of this signal, while REM states had a broader and higher peak. Non-REM states appeared to include multiple distributions, likely indicating a subdivision of non-REM sleep states [Katsageorgiou, VM. et al., PLoS Biol. 16(5):e2003663 (2018)]. To confirm that this unusual shape of the non-REM distribution was not an artifact of combining data from multiple animals, we plotted the data for each animal, and each mouse showed an increase in the standard deviation from non-REM to REM states (Figure 9D). Individual animals also exhibited this elongated non-REM distribution. Both of these experiments indicated that the observed signal was a respiratory rate signal. These results suggested good classifier performance.

[0121] classification Finally, a machine learning classifier was trained to predict sleep states using 320 visual features. For validation, all data from the animals were provided to avoid any bias that might be introduced by correlated data in the images. Ten-fold cross-validation was performed by shuffling which animals were provided for the calculation of training and test accuracy. A balanced dataset was created as described above in the Materials and Methods section of this specification, and several classification algorithms, including XgBoost, Random Forest, MLP, Logistic Regression, and SVD, were compared. Performance varied significantly among the classifiers (Table 4). Both XgBoost and Random Forest achieved good accuracy on the provided test data. However, the Random Forest algorithm achieved 100% training accuracy, indicating that it overfitted the training data. Overall, the best-performing algorithm was the XgBoost classifier. [Table 4]

[0122] The transitions between wakefulness, non-REM, and REM sleep are not random but generally follow predictable patterns. For example, wakefulness generally transitions to non-REM sleep, which then transitions to REM sleep. Hidden Markov models are ideal candidates for modeling the dependencies between sleep states. The transition probability matrix and occurrence probabilities in a given state are learned using training data. By adding the HMM model, we observed a 7% improvement in the overall classifier accuracy, from 0.839+ / -0.022 to 0.906+ / -0.021 (Figure 10A, +HMM).

[0123] To improve the classifier's performance, Hu moment measurements were incorporated from segmentation into the input features for classification [Hu, MK. IRE Trans Inf Theory. 8(2):179-187 (1962)]. These image moments provided a numerical description of mouse segmentation through integration and linear combination of central image moments. The addition of Hu moment features resulted in a slight increase in overall accuracy and improved classifier robustness, with the variability in cross-validation performance decreasing from 0.906+ / -0.021 to 0.913+ / -0.019 (Figure 10A, +Hu moment).

[0124] EEG / EMG scoring was performed by trained human experts, but there was often inconsistency among trained annotators [Pack, AI et al. Physiol. Genomics. 28(2):232-238 (2007)]. In fact, two experts generally agreed only 88–94% of the time for REM and non-REM [Pack, AI et al. Physiol. Genomics. 28(2):232-238 (2007)]. Using recently published machine learning methods, EEG / EMG data was scored to supplement data from human score recorders [Miladinovic, D. et al., PLoS Comput. Biol. 15(4):e1006968 (2019)]. SPINDLE annotations were compared with human annotations and found to agree 92% of the time over all epochs. Next, only epochs agreed by both human and machine-based methods were used as markers for visual classifier training. Training the classifier using only SPINDLE and human-matched epochs increased the accuracy by another 1% (Figure 10A, +filter annotation). Thus, the final classifier achieved a 3-state classification accuracy of 0.92+ / -0.05.

[0125] The classification features used were examined to determine which were most important. Mouse region and motion measurements were identified as the most important features (Figure 10B). While not intended to be restrictive, this result is likely observed because motion is the only feature used in the binary sleep-wake classification algorithm. In addition, three of the top five features were low-frequency (0.1–1.0 Hz) power spectral density (Figures 7A and 7B, FFT columns). Furthermore, it was observed that arousal epochs had the most power at low frequencies, REM sleep had low power at low frequencies, and non-REM sleep had the least power at low-frequency signals.

[0126] Using the highest-performing classifier, good performance was observed (Figure 10C). In the matrix shown in Figure 10C, the rows represent sleep states assigned by human score recorders, and the columns represent stages assigned by the classifier. Awakening had the highest accuracy rate among the classes, at 96.1%. By observing the area outside the diagonal of the matrix, it can be seen that the classifier performed better in distinguishing wakefulness from any sleep state than in distinguishing between sleep states, and that distinguishing non-REM from REM was a difficult task.

[0127] The final classifier achieved an overall accuracy of 0.92+ / -0.05 on average. The predicted accuracy for the arousal stage was 0.97+ / -0.01, with an average fitted recall of 0.98. The predicted accuracy for the non-REM stage was 0.92+ / 0.04, with an average fitted recall of 0.93. The predicted accuracy for the REM stage was 0.88+ / -0.05, with an average fitted recall of 0.535. The lower fitted recall for REM was due to the very small proportion of epochs labeled as REM (4%).

[0128] In addition to prediction accuracy, performance metrics including precision, recall, and F1 score were measured to evaluate the model from 10-fold cross-validation (Figure 10D). Considering the imbalanced data, precision-recall was a better indicator of classifier performance [Powers, DMW, J. Mach. Learn. Technol. 2(1):37-63 (2011) ArXiv:2010.16061; Saito, T. and Rehmsmeier, M. PLoS ONE. 10(3):e0118432 (2015)]. Precision measures the proportion of positive items correctly predicted, while recall measures the proportion of actual positives correctly identified. The F1 score was a weighted mean of precision and recall.

number

number

number

[0129] The final classifier performed well for both awake and non-REM states. However, it showed the worst performance for REM states, with a precision of 0.535 and an F1 of 0.664. The majority of misclassified states were between non-REM and REM. Because REM states were a minority class (only 4% of the dataset), even a relatively small false-positive rate would result in a large number of false positives that would overwhelm the rare true positives. For example, 9.7% of REM periods were misclassified as non-REM by the visual classifier, and 7.1% of predicted REM periods were actually non-REM (Figure 10C). While these misclassification errors may seem small, the imbalance between REM and non-REM could disproportionately affect the classifier's precision. Despite this, the classifier was still able to correctly identify 89.7% of the REM epochs present in the validation dataset.

[0130] Within the context of other existing alternatives to EEG / EMG recording, this model performed superiorly. Table 5 compares the performance of each previously reported model with that of the classifier model described herein. Note that each previously reported model used a different dataset with different characteristics. In particular, the piezo system was evaluated on a balanced dataset that could show higher precision due to its lower probability of false positives. The classifier approach developed herein outperformed all approaches for predicting arousal and non-REM states. REM prediction was a more challenging task for all approaches. Of the machine learning approaches, the model described herein achieved the highest accuracy. Figures 11A and 11B show a visual performance comparison of our classifier with manual scoring by human experts (hypnograms). The x-axis is time, consisting of continuous epochs, and the y-axis corresponds to three stages. For each subfigure, the upper panel represents the human scoring results, and the lower panel represents the classifier scoring results. Hypnograms show accurate transitions between stages, along with the frequency of isolated false positives (Figure 11A). Visual and human scoring for a single animal over 24 hours are also plotted (Figure 11B). Raster plots show better overall correlations between state classifications (Figure 11B). Next, all C57BL / 6J animals are compared between human EEG / EMG scoring and visual scoring (Figures 11C, D). High correlations are observed across all states, leading us to conclude that our visual classifier scoring results are consistent with human scores. [Table 5]

[0131] To improve classifier performance, various data augmentation approaches were also attempted. The proportion of different sleep states over 24 hours was remarkably unbalanced (48% wakefulness, 48% non-REM, and 4% REM). Typical augmentation techniques used for time series data include jittering, scaling, rotation, reordering, and cropping. These methods can be applied in combination with each other. It has been previously shown that classification accuracy can be increased by augmenting the training data by combining four data augmentation techniques [Rashid, KM and Louis, J. Adv Eng Inform. 42:100944 (2019)]. However, since the features extracted from the time series were dependent on spectral composition, it was decided to augment the size of the training dataset using a dynamic time stretching-based approach to improve the classifier [Fawaz, HI, et al., arXiv:1808:02455]. After data augmentation, the size of the dataset increased by approximately 25% (from 14K epochs to 17K epochs). Adding data through the augmentation algorithm was observed to decrease the prediction accuracy. The mean predicted accuracy for awake, non-REM, and REM states was 77%, 34%, and 31%, respectively. While we do not wish to be bound by any particular theory, the performance after data augmentation may be due to the introduction of more noise from the REM state data and a decrease in classifier performance. Performance was demonstrated with 10-fold cross-validation. The results of this data augmentation application are shown in Figure 12 and Table 6 (Table 6 shows the numerical results of data augmentation). This data augmentation approach did not improve classifier performance and was therefore not pursued further. Overall, the visual sleep state classifier was able to accurately identify sleep states using only visual data. Including HMM, Hu moments, and highly accurate labeling improved performance, while data augmentation using dynamic time stretching and motion amplification did not. [Table 6]

[0132] Consideration Sleep disorders are characteristic of many diseases, and high-throughput studies in model organisms are crucial for the discovery of new therapeutics [Webb, JM and Fu, YH., Curr. Opin. Neurobiol. 69:19-24 (2021); Scammell,TE et al., Neuron. 93(4):747-765 (2017); Allada, R. and Siegel, JM Curr. Biol. 18(15):R670-R679 (2008)]. Sleep studies in mice are difficult to conduct on a large scale due to the time investment required for surgery, recovery time, and scoring of recorded EEG / EMG signals. The system described herein provides a low-cost alternative to EEG / EMG scoring of mouse sleep behavior, enabling researchers to conduct larger-scale sleep experiments that would previously have been prohibitively expensive. While previous systems have been proposed to conduct such experiments, they have only been shown to adequately distinguish between wakeful and sleep states. The systems described herein are built upon these approaches and can also distinguish between REM and non-REM sleep states.

[0133] The system described herein achieves highly sensitive measurement of movement and posture in sleeping mice. This system has been shown to observe features correlated with mouse respiratory rate using only visual measurements. Previously published systems capable of achieving this level of sensitivity include plethysmography [Bastianini, S. et al., Sci. Rep. 7:41698 (2017)] or piezo systems [Mang, GM et al., Sleep. 37(8):1383-1392 (2014); Yaghouby, F., et al., J. Neurosci. Methods. 259:90-100 (2016)]. In addition, this specification shows that, based on the features used, this novel system may be able to identify subclusters of non-REM sleep epochs, which could further elucidate the structure of mouse sleep.

[0134] In conclusion, the high-throughput, non-invasive, computer vision-based methods described herein for determining sleep states in mice are useful to society.

[0135] Equivalents While several embodiments of the present invention are described and illustrated herein, those skilled in the art will readily conceive of various other means and / or structures for carrying out the function and / or obtaining the results and / or one or more advantages described herein, and each of such variations and / or modifications will be considered within the scope of the present invention. More generally, those skilled in the art will readily understand that all parameters, dimensions, materials and configurations described herein are illustrative, and that the actual parameters, dimensions, materials and / or configurations will depend on the specific one or more applications in which the teachings of the present invention are used. Those skilled in the art will understand many equivalents to the specific embodiments of the present invention described herein, or will be able to verify such equivalents using only routine experiments. Therefore, it will be understood that the embodiments described herein are presented only illustratively, and that the present invention may be carried out in ways other than those specifically described and described in the claims, within the scope of the appended claims and its equivalents. The present invention covers the individual features, systems, articles, materials and / or methods described herein. In addition, any combination of two or more such features, systems, articles, materials, and / or methods is included within the scope of the invention, provided that they are not mutually inconsistent. It will be understood that all definitions defined and used herein govern dictionary definitions, definitions in documents referenced by reference, and / or the common meanings of the defined terms.

[0136] As used herein in this specification and in the claims, the indefinite articles "a" and "an" will be understood to mean "at least one" unless explicitly stated to the contrary. As used herein in this specification and in the claims, the phrase "and / or" will be understood to mean "either or both" of the elements thus joined, that is, "either or both" of elements that exist as a combination in some cases and separately in other cases. Unless explicitly stated otherwise, other elements may exist, at their discretion, other than those specifically identified by the phrase "and / or," whether related to or unrelated to the specifically identified elements.

[0137] Conditional language used herein, in particular “can,” “could,” “might,” “may,” “eg,” and similar terms, is generally intended to convey that a particular embodiment includes certain features, elements, and / or steps, while other embodiments do not, unless otherwise specifically stated or understood in the context in which they are used. Therefore, such conditional language is not generally intended to mean that features, elements, and / or steps are required in any way with respect to one or more embodiments, or that one or more embodiments necessarily include logic for determining whether these features, elements, and / or steps are included in or performed within any particular embodiment, with or without other inputs or prompts. “comprising,” “including,” “having,” and similar terms are synonymous, used comprehensively and in an open-ended manner, and do not exclude additional elements, features, actions, operations, and similar items. Furthermore, the term "or" is used in its inclusive (and not exclusive) sense; for example, to connect a list of elements, the term "or" can mean one, some, or all of the elements in the list.

[0138] All references, patents, patent applications, and publications cited or referred to in this application are incorporated herein by reference in their entirety.

Claims

1. A method by which a computer performs an action. Receiving video data representing the target video, Using the aforementioned video data, determine a number of features corresponding to the subject, Determining a time-domain feature for each of the plurality of features in the video data, which is a time-domain feature based on the signal of the target feature in the video data, A frequency domain feature for each of the plurality of features in the video data, wherein the frequency domain feature is determined based on the signal of the target feature in the video data. A method performed by a computer, comprising using a machine learning classifier to process the time-domain features and the frequency-domain features to determine sleep state data for the subject.

2. The computer method according to claim 1, further comprising using a machine learning model to process the video data to determine segmentation data representing a first set of pixels corresponding to the subject and a second set of pixels corresponding to the background.

3. The method performed by a computer according to claim 2, further comprising processing the segmentation data to determine elliptic fit data corresponding to the object.

4. A method performed by a computer according to claim 2, wherein determining the plurality of features includes processing the segmentation data to determine the plurality of features.

5. The method performed by a computer according to claim 1, wherein the plurality of features include a plurality of visual features for each video frame of the video data.

6. The process further includes determining a time-domain feature for each of the plurality of visual features, wherein determining the time-domain feature includes determining one of kurtosis data, mean data, median data, standard deviation data, maximum data, and minimum data. The method performed by a computer according to claim 5, wherein the plurality of features include the time-domain features.

7. The process further includes determining frequency domain features for each of the plurality of visual features, wherein determining the frequency domain features includes determining one of the following: kurtosis of the power spectral density, skewness of the power spectral density, mean power spectral density, total power spectral density, maximum data, minimum data, mean data, and standard deviation of the power spectral density. The method performed by the computer according to claim 5, wherein the plurality of features include the frequency domain features.

8. A computer-based method according to claim 1, comprising using a machine learning classifier to process the plurality of features to determine a sleep state for a video frame of the video data, wherein the sleep state is one of a wakefulness, REM sleep, and non-REM sleep.

9. The computer method according to claim 1, wherein the sleep state data indicates the duration of a sleep state, the duration of one or more of the wakefulness, REM state, and non-REM state and / or the frequency intervals of one or more sleep states, and changes in one or more sleep states.

10. The process involves determining multiple body regions of the target using the aforementioned multiple features, wherein each of the multiple body regions corresponds to a video frame of the video data. A method performed by a computer according to claim 1, further comprising determining the sleep state data based on the changes in the plurality of body regions in the video.

11. The process involves determining multiple aspect ratios using the aforementioned features, wherein each of the aforementioned aspect ratios corresponds to a video frame of the video data. A method performed by a computer according to claim 1, further comprising determining the sleep state data based on the changes in the plurality of width ratios in the video.

12. Determining the aforementioned sleep state data A computer-based method according to claim 1, comprising detecting a transition from a non-REM state to a REM state based on a change in the body region or body shape of the subject, wherein the change in the body region or body shape is a result of muscle relaxation.

13. The process involves determining multiple aspect ratios for the aforementioned object, wherein the aspect ratios of the multiple aspect ratios correspond to the video frames of the video data. Using the aforementioned multiple width ratios, the time-domain features are determined, The process involves determining frequency-domain features using the aforementioned multiple width ratios, wherein the time-domain features and frequency-domain features represent the movement of the abdomen of the subject. A method performed by a computer according to claim 1, further comprising determining the sleep state data using the time-domain features and the frequency-domain features.

14. The method performed by a computer according to claim 1, wherein the video captures the object in its natural state, the natural state of the object includes the absence of invasive detection means within or on the object, and the invasive detection means includes one or both of an electrode attached to the object and an electrode inserted into the object.

15. The method performed by a computer according to claim 1, wherein the video is a high-resolution video.

16. Using a machine learning classifier, process the multiple features to determine multiple sleep state predictions for each video frame of the video data, A computer-based method according to claim 1, further comprising using a transition model to process the plurality of sleep state predictions and determine the transition between a first sleep state and a second sleep state, wherein the transition model is a hidden Markov model.

17. The video is a video of two or more objects, including at least a first object and a second object, and the method is Processing the aforementioned video data to determine first segmentation data that shows a first set of pixels corresponding to the first object, Processing the aforementioned video data to determine a second segmentation data that shows a second set of pixels corresponding to the second object, Using the first segmentation data, determine a set of first features corresponding to the first target, Using the aforementioned first set of features, first sleep state data for the first subject is determined, Using the second segmentation data, determine a second set of features corresponding to the second target, A method performed by a computer according to claim 1, further comprising determining second sleep data for the second subject using the second set of features.

18. The method executed by a computer according to claim 1, wherein the subject is a rodent, and optionally a mouse.

19. The computer-executed method according to claim 1, wherein the subject is a genetically modified subject.

20. At least one processor, At least one memory containing instructions and A system including, where the instruction, when executed by the at least one processor, the system Receiving video data representing the target video, Using the aforementioned video data, determine a number of features corresponding to the subject, Determining a time-domain feature for each of the plurality of features in the video data, which is a time-domain feature based on the signal of the target feature in the video data, A frequency domain feature for each of the plurality of features in the video data, wherein the frequency domain feature is determined based on the signal of the target feature in the video data. A system that uses a machine learning classifier to process the time-domain features and frequency-domain features to determine sleep state data for the subject.