Action detection using machine learning models

A multi-model, multi-orientation video processing approach enhances the detection and analysis of complex behaviors in subjects, addressing automation and adaptability challenges in behavioral neuroscience, enabling accurate disease assessment and therapeutic evaluation.

JP7767406B2Active Publication Date: 2025-11-11JACKSON LAB THE
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2023516504
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-09-16
Filing Date
2021-09-16
Publication Date
2025-11-11
Estimated Expiration
2041-09-16

AI Technical Summary

Technical Problem

Current methods for analyzing neuronal function and behavior in complex environments are time-consuming and lack automation, particularly in behavioral neuroscience, due to high costs of dataset curation and stringent implementation requirements, and existing systems are limited by measurements and environmental constraints, leading to ineffective action recognition and transferability issues.

Method used

A computer-implemented method using multiple machine learning models and video data processing techniques, including rotation and reflection of frames, to enhance the detection of behavioral actions such as grooming in subjects, particularly mice, by generating multiple probabilities and labels for improved accuracy and robustness across varying physical characteristics.

Benefits of technology

The method provides accurate and automated detection of complex behaviors, enabling assessment of diseases and therapeutic effects, with improved robustness and adaptability to variations in subjects, and facilitating the analysis of genetic variations and therapeutic agent efficacy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007767406000002
    Figure 0007767406000002
  • Figure 0007767406000003
    Figure 0007767406000003
  • Figure 0007767406000004
    Figure 0007767406000004
Patent Text Reader

Abstract

The systems and methods described herein provide techniques for detecting a subject's behavior by processing video data using one or more trained models configured to detect the subject's behavior. The described systems process sets of frames from the video data using different trained models. The systems further process different orientations of the set of frames. The various outputs from the different trained models and from processing the different orientations of the set of frames may be combined to make a final determination as to whether a subject is exhibiting a particular behavior during a particular frame.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Related Applications This application claims the benefit under 35 U.S.C. § 119(e) of U.S. Provisional Patent Application No. 63 / 078,952, filed September 16, 2020, the disclosure of which is incorporated herein by reference in its entirety.

[0002] The present invention relates in some aspects to the use of neural networks to detect actions of a subject, in part for use in assessing genetic variation.

[0003] Government subsidies This invention was made with government support under DA041668 and DA048634 awarded by the National Institutes of Health. The government has certain rights in this invention. [Background technology]

[0004] Behavior, the primary output of the nervous system, is complex, hierarchical, dynamic, and highly dimensional (Gomez-Marin, et al., 2014 Nature Neuroscience 17(11):1455-1462). Current approaches to analyzing neuronal function require analyzing behavior with high temporal and spatial resolution. Achieving this is a time-consuming task, and its automation remains a challenging problem in behavioral neuroscience. While some efforts have been made to integrate methods from the fields of computer vision and modern neural network approaches, few aspects of behavioral biology research utilize neural network approaches. This lack of application is often due to the high cost of curating and annotating datasets or the stringent implementation requirements. Nevertheless, action recognition within complex environments remains an unsolved challenge in the machine learning community, and the transferability of proposed solutions to behavioral neuroscience remains unresolved.

[0005] Behavioral neuroscientists have traditionally used several methods to classify mouse behavior. Simple behaviors, such as rearing, can be classified using physical measurement devices that detect when a mouse exceeds a certain height. For more complex psychiatric constructs, such as motivation, manipulative behavioral paradigms have been used. Open-source systems such as JAABA [Kabra et al., 2013 Nature Methods. 2013;10(1):64.] have been used by researchers to train their own machine learning classifiers for complex behaviors using movement and other measurements [Van den Boom, et al., 2017 J. Neuroscience Methods 289:48-56]. These systems are inherently limited by the measurements available. The available standard measurements include only center of mass tracking, which severely limits the types of behaviors that can be reliably classified. For mice, more modern systems integrate floor vibrometry and depth imaging techniques to enhance behavior detection [Quinn et al., 2003 J. Neurosci. Methods; 130(1):83-92 and Hong et al., 2015 PNAS; 112(38):E5351-E5360]. While some efforts have been made to automate grooming annotation using machine learning classifiers, prior techniques are not robust to animal coat color, lighting conditions, and equipment location [Van den Boom, et al., 2017 J. Neuroscience Methods 289:48-56].Recent advances in computer vision also provide general-purpose solutions for markerless tracking in laboratory animals [Mathis et al., 2018 Nature neuroscience. 2018;21(9):1281; Pereira et al., 2019 Nature methods. 2019;16(1):117-125], however, examples such as human action detection leaderboards suggest that pose estimation approaches, while powerful, routinely underperform end-to-end solutions that utilize raw video input for action classification [Feichtenhofer et al., 2019 Proceedings of the IEEE International Conference on Computer Vision; 2019. pp. 6202-6211; Choutas et al., 2018 Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018. pp. 7024-7033].

[0006] Previous attempts to directly manipulate visual data have been limited to unsupervised behavioral clustering approaches [Todd et al., 2017 Physical Biology 14(1):015002]. These efforts have included limitations on environmental contrast, which is crucial for the algorithms to function. Various prior systems rely heavily on aligning data from a top-down view [Berman et al., 2014 J. of The Royal Society Interface 11(99):20140672; Wiltschko et al., 2015 Neuron 88(6):1121-1135]. While these approaches can cluster similar video segments, the behavioral labeling of the resulting clusters is still determined by the user, thereby severely limiting their effectiveness. Summary of the Invention

[0007] According to one aspect of the present invention, a computer-implemented method is provided, the method including: receiving video data representing a video capturing movement of a subject; identifying a first set of frames from the video data; determining a rotated set of frames by rotating the first set of frames; processing the first set of frames using a first trained model configured to identify a likelihood of a subject exhibiting a predetermined behavioral action; determining a first probability of a subject exhibiting the predetermined behavioral action in a first frame of the first set of frames, the first frame corresponding to a time duration of the video data, based on the processing of the first set of frames by the first trained model; processing the rotated set of frames using the first trained model; determining a second probability of a subject exhibiting the predetermined behavioral action in a second frame of the rotated set of frames, the second frame corresponding to a time duration of the first frames, based on the processing of the rotated set of frames by the first trained model; and identifying a label of the first frame using the first probability and the second probability, wherein the first label indicates that the subject exhibits the predetermined behavioral action. In some embodiments, the method also includes processing the first set of frames using a second trained model configured to identify likelihood of an object exhibiting a predetermined behavioral action; determining a third probability of the object exhibiting the predetermined behavioral action in the first frames based on the processing of the first set of frames by the second trained model; processing the rotated set of frames using the second trained model; determining a fourth probability of the object exhibiting the predetermined behavioral action in the second frames based on the processing of the rotated set of frames by the second trained model; and identifying a first label using the first probability, the second probability, the third probability, and the fourth probability.In certain embodiments, the method also includes determining a set of reflected frames by reflecting the first set of frames; processing the set of reflected frames using the first trained model; determining a third probability of the subject exhibiting a predetermined behavioral action in a third frame of the set of reflected frames, the third frame corresponding to the first frame, based on the processing of the set of reflected frames by the first trained model; and identifying a first label using the first probability, the second probability, and the third probability. In some embodiments, the predetermined behavioral action includes a grooming behavior. In some embodiments, the predetermined behavioral action includes only one or more grooming behaviors. In certain embodiments, the subject is a mouse. In certain embodiments, the subject is a mammal, optionally a rodent, and the predetermined behavior includes a grooming behavior including at least one of paw licking, unilateral face washing, bilateral face washing, and flank licking. In some embodiments, the first set of frames represents a portion of video data for a time period, and the first frame is a last time frame of the time period.In some embodiments, the method also includes identifying a second set of frames from the video data, rotating the second set of frames to determine a second rotated set of frames, processing the second set of frames using the first trained model, and determining a third probability of an object exhibiting a predetermined behavioral action in a third frame of the second set of frames based on the processing of the second set of frames by the first trained model, processing the second rotated set of frames using the first trained model, and determining a fourth probability of an object exhibiting the predetermined behavioral action in a fourth frame of the rotated set of frames, the fourth frame corresponding to the third frame, based on the processing of the second rotated set of frames by the first trained model, and identifying a second label for the fourth frame using the third probability and the fourth probability, wherein the first label indicates that the object exhibits the predetermined behavioral action. In certain embodiments, the method also includes generating an ethogram representing the predetermined behavioral action of the object over a period of time using at least the first label and the second label. In certain embodiments, the first trained model is a machine learning classifier. In some embodiments, the method also includes, prior to receiving the video data, receiving training data including a first plurality of video frames and a second plurality of video frames, each of the first plurality of video frames associated with a positive label indicating that a subject is exhibiting a predetermined behavioral action and each of the second plurality of video frames associated with a negative label indicating that the subject is exhibiting a behavioral action that is not the predetermined behavioral action, and processing the training data using the first set of model parameters and the model data of the first classifier to determine a first trained model. In some embodiments, the first plurality of frames and the second plurality of frames represent movement of a plurality of subjects, one of the plurality of subjects including one or more pre-identified physical characteristics.In certain embodiments, the pre-identified physical characteristics are one or more of body shape, body size, coat color, sex, age, and disease or disorder phenotype. In some embodiments, the disease or disorder is a genetic disease, injury, or infectious disease. In some embodiments, the first plurality of frames and the second plurality of frames represent the movements of a plurality of mouse subjects, where a single mouse has coat color, sex, body shape, and size. In certain embodiments, the subject is a mammal. In some embodiments, the subject is a genetically engineered subject, optionally a genetically engineered rodent. In some embodiments, the subject is a genetically engineered mouse.

[0008] According to another aspect of the present invention, a computer-implemented method is provided, the method including receiving video data representing a video capturing movement of a subject; identifying a first set of frames from the video data; processing the first set of frames using a first trained model configured to identify likelihoods of subjects exhibiting a predetermined behavioral action; determining a first probability of a subject exhibiting the predetermined behavioral action in a first frame of the first set of frames based on the processing of the first set of frames by the first trained model; processing the first set of frames using a second trained model configured to identify likelihoods of subjects exhibiting the predetermined behavioral action; determining a second probability of a subject exhibiting the predetermined behavioral action in the first frames based on the processing of the first set of frames by the second trained model; and identifying a first label for the first frame using the first probability and the second probability, wherein the first label indicates that the subject exhibits the predetermined behavioral action. In particular embodiments, a method includes determining a set of rotated frames by rotating a first set of frames; processing the set of rotated frames using a first trained model; determining a third probability of an object exhibiting a predetermined behavioral action in a second frame of the set of rotated frames, the second frame corresponding to the first frame, based on the processing of the set of rotated frames by the first trained model; processing the set of rotated frames using a second trained model; determining a fourth probability of an object exhibiting the predetermined behavioral action in the second frame based on the processing of the set of rotated frames by the second trained model; and identifying a first label using the first probability, the second probability, the third probability, and the fourth probability.In certain embodiments, the method also includes determining a set of reflected frames by reflecting the first set of frames, processing the set of reflected frames using the first trained model, determining a third probability of an object exhibiting a predetermined behavioral action in a second frame of the set of reflected frames, the second frame corresponding to the first frame, based on the processing of the set of reflected frames by the first trained model, processing the set of reflected frames using the second trained model, determining a fourth probability of an object exhibiting the predetermined behavioral action in the second frame based on the processing of the set of reflected frames by the second trained model, and identifying a first label using the first probability, the second probability, the third probability, and the fourth probability. In some embodiments, the first model and the second model are neural network models, and the first model is initialized using a first set of parameters and the second trained model is initialized using a second set of parameters different from the first set of parameters. In some embodiments, the method also includes processing the first set of frames using a third trained model configured to identify a likelihood of the object exhibiting the predetermined behavioral action; determining a third probability of the object exhibiting the predetermined behavioral action in the first frames based on the processing of the first set of frames by the third trained model; processing the first set of frames using a fourth trained model configured to identify a likelihood of the object exhibiting the predetermined behavioral action; determining a fourth probability of the object exhibiting the predetermined behavioral action in the first frames based on the processing of the first set of frames by the fourth trained model; and identifying a first label using the first probability, the second probability, the third probability, and the fourth probability. In some embodiments, the object is a mammal. In certain embodiments, the predetermined behavioral action includes grooming behavior. In some embodiments, the predetermined behavioral action includes other than grooming behavior.In certain embodiments, the subject is a mammal and the predetermined behavioral action comprises a grooming behavior, the grooming behavior being at least one of: licking a paw, washing one side of the face, washing both sides of the face, and licking the flanks. In some embodiments, the subject is a rodent and the predetermined behavioral action comprises a grooming behavior, the grooming behavior being at least one of: licking a paw, washing one side of the face, washing both sides of the face, and licking the flanks. In some embodiments, the first set of frames represents a portion of video data for a period of time, and the first frame is the last time frame of the period. In certain embodiments, the method also includes identifying a second set of frames from the video data, processing the second set of frames using the first trained model, and determining a third probability of the subject exhibiting the predetermined behavioral action in a third frame of the second set of frames based on the processing of the second set of frames by the first trained model, processing the second set of frames using the second trained model, and determining a fourth probability of the subject exhibiting the predetermined behavioral action in the third frame based on the processing of the second set of frames by the second trained model, and using the third and fourth probabilities to identify a second label for the third frame, the second label indicating that the subject exhibits the predetermined behavioral action. In some embodiments, the method also includes generating an ethogram representing the predetermined behavioral action of the subject over a period of time using at least the first and second labels. In some embodiments, the first and second trained models are machine learning classifiers.In certain embodiments, the method also includes, prior to receiving the video data, receiving training data including a first plurality of video frames and a second plurality of video frames, each of the first plurality of video frames associated with a positive label indicating that the subject is exhibiting a predetermined behavioral action and each of the second plurality of video frames associated with a negative label indicating that the subject is exhibiting a behavioral action that is not the predetermined behavior; processing the training data using a first set of model parameters and first classifier model data to determine a first trained model; and processing the training data using a second set of model parameters and second classifier model data to determine a second trained model. In certain embodiments, the first plurality of frames and the second plurality of frames represent the movement of a plurality of objects, one of the plurality of objects including one or more pre-identified physical characteristics. In some embodiments, the pre-identified physical characteristics are one or more of body shape, body size, coat color, sex, age, and a disease or disorder phenotype. In some embodiments, the disease or disorder is a genetic disease, injury, or infectious disease. In certain embodiments, the first plurality of frames and the second plurality of frames represent the movements of a plurality of mouse subjects, each mouse having a different coat color, sex, body shape, and size. In some embodiments, the subject is a rodent, optionally a mouse. In certain embodiments, the subject is a genetically engineered subject.

[0009] According to another aspect of the present invention, there is provided a method for assessing a predetermined behavioral action in a subject, wherein the predetermined behavioral action comprises grooming behavior, including at least one of paw licking, unilateral face washing, bilateral face washing, and flank licking, and wherein the assessing means comprises any embodiment of the computer-implemented method described above. In some embodiments, the subject has a predetermined behavior-related disease or disorder, and optionally is an animal model of the predetermined behavior-related disease or disorder. In certain embodiments, the subject is a genetically engineered subject. In some embodiments, the subject is a rodent, optionally a mouse. In some embodiments, the mouse is a genetically engineered mouse. In certain embodiments, the method also comprises administering a candidate therapeutic agent to the subject, assessing the subject's predetermined behavior after administration of the candidate therapeutic agent, and comparing the assessment after administration with a control assessment of the predetermined behavior, wherein a change in the predetermined behavior after administration compared to the control predetermined behavior identifies an effect of the administered candidate therapeutic agent on the predetermined behavior. In some embodiments, the change comprises one or more of onset, increase, cessation, and decrease of the subject's predetermined behavior. In certain embodiments, the candidate therapeutic agent is administered to the subject before assessing the predetermined behavior. In some embodiments, the candidate therapeutic agent is administered to the subject concurrently with assessing the predetermined behavior. In some embodiments, the control assessment of the predetermined behavior is an assessment of the predetermined behavior in a control subject monitored with a computer-implemented method. In some embodiments, the control subject is an animal model of the predetermined behavior-related disease or disorder. In certain embodiments, the predetermined behavior-related disease or disorder is a genetic disease, injury, or infectious disease. In some embodiments, the predetermined behavior-related disease or disorder is bipolar disorder, dementia, depression, hyperactivity disorder, anxiety disorder, developmental disorder, sleep disorder, Alzheimer's disease, Parkinson's disease, or physical injury. In some embodiments, the control subject does not receive the candidate therapeutic agent. In certain embodiments, the control subject is administered a dose of the candidate therapeutic agent that differs from the dose of the candidate therapeutic agent administered to the subject. In some embodiments, the control results are results of previous monitoring of the subject using a computer-implemented method, optionally, the previous monitoring of the subject occurs before administration of the candidate therapeutic agent.In some embodiments, the subject monitoring identifies a predetermined behaviorally related disease or disorder in the subject. In some embodiments, the subject monitoring identifies the effectiveness of a candidate therapeutic agent for treating the predetermined behaviorally related disease or disorder.

[0010] According to another aspect of the present invention, there is provided a method for identifying the effectiveness of a candidate therapeutic agent for treating a predetermined behavior-related disease or disorder in a subject, the method comprising administering a candidate therapeutic agent to the subject and monitoring one or more predetermined behavioral actions of the subject, wherein the monitoring means comprises a computer-implemented method of any of the preceding methods, wherein the predetermined behavioral actions comprise grooming behaviors including at least one of paw licking, one-side face washing, both sides face washing, and flank licking, and wherein results of the monitoring indicative of a change in the subject's predetermined behavior identify the effectiveness of the candidate therapeutic agent for treating the predetermined behavior-related disease or disorder. In certain embodiments, the subject has a predetermined behavior-related disease or disorder and, optionally, is an animal model of the predetermined behavior-related disease or disorder. In some embodiments, the subject is an animal model of the predetermined behavior-related disease or disorder. In certain embodiments, the behavior-related disease or disorder is a genetic disease, injury, or infectious disease. In some embodiments, the behavior-related disease or disorder is bipolar disorder, dementia, depression, hyperactivity disorder, anxiety disorder, developmental disorder, sleep disorder, Alzheimer's disease, Parkinson's disease, or physical injury. In some embodiments, the subject is a genetically engineered subject. In some embodiments, the subject is a rodent, optionally a mouse. In certain embodiments, the mouse is a genetically engineered mouse. In some embodiments, the candidate therapeutic agent is administered to the subject prior to monitoring the predetermined behavior. In some embodiments, the candidate therapeutic agent is administered to the subject simultaneously with monitoring the predetermined behavior. In certain embodiments, the subject's monitored predetermined behavior is compared to control monitoring of the predetermined behavior, wherein the control monitoring comprises monitoring the predetermined behavior of the control subject using a computer-implemented method. In certain embodiments, the control subject is an animal model of a disease or disorder. In some embodiments, the control subject does not receive the candidate therapeutic agent. In some embodiments, the control subject is administered a dose of the candidate therapeutic agent that differs from the dose of the candidate therapeutic agent administered to the subject. In certain embodiments, the control monitoring comprises monitoring the subject's predetermined behavior using a computer-implemented method prior to administration of the candidate therapeutic agent. In some embodiments, monitoring the subject identifies the effectiveness of the candidate therapeutic agent for treating a predetermined behavior-related disease or disorder. [Brief explanation of the drawings]

[0011] For a more complete understanding of the present disclosure, reference is now made to the following descriptions taken in conjunction with the accompanying drawings, in which:

[0012] [Figure 1] FIG. 1 is a conceptual diagram of a system for determining the behavior of a subject, according to an embodiment of the present disclosure. [Figure 2] FIG. 2 is a conceptual diagram illustrating a process for analyzing video data of a subject to determine subject behavior, according to an embodiment of the present disclosure. [Figure 3] FIG. 3 is a conceptual diagram illustrating how different sets of frames can be processed using different trained models, according to an embodiment of the present disclosure. [Figure 4] FIG. 4 is a conceptual diagram showing how the different predictions can be processed to determine a final prediction regarding the subject exhibiting the behavior. [Figure 5] FIG. 5 is a conceptual diagram illustrating a system that employs multiple trained models to process different sets of frames. [Figure 6] FIG. 6 is a conceptual diagram illustrating components for training / constructing a machine learning (ML) model to determine whether a subject is exhibiting an action in a frame of video data. [Figure 7] FIG. 7 is a conceptual diagram of a process for analyzing a subject's video data 104 to detect subject behavior, according to an embodiment of the present disclosure. [Figure 8A]Figures 8A-D provide photographs and graphs related to the annotation of mouse grooming behavior. Figure 8A shows that mouse grooming involves a wide variety of postures. Paw licking, face washing, flank licking, and other syntactic features all contribute to this visually diverse behavior. Figure 8B shows images of grooming ethograms for six videos by five different trained annotators (observers 1-5, shown sequentially from top to bottom for each ethogram in each video). Overall, there was very high agreement between annotators. Figures 8C-D show results quantifying the overlap in agreement between individual annotators. The average agreement between all annotators was 89.13%. [Figure 8B] Same as above. [Figure 8C] Same as above. [Figure 8D] Same as above. [Figure 9A] Figures 9A-C show schematic diagrams providing additional details of annotator mismatch. Figure 9A shows how annotation mismatch types were categorized into three classes of errors: missed bouts, misaligned frames, and skipped breaks. Figures 9B-C show quantification of the error types in Figures 8A-D. Figure 9B shows the total frames included in each error category. 37.5% of the frames were missed bouts, 50% were misaligned, and 12.4% were skipped breaks. Error types were not evenly distributed across annotators, with annotator 2 accounting for the most missed bout frames, annotator 4 accounting for the most misaligned frames, and annotator 1 accounting for the most skipped break frames. FIG. 9C shows a similar distribution when counting the number of erroneous multi-frame occurrences, where 19.2% of calls were missed bouts, 74.6% were misaligned, and 6.3% were skipped interruptions. [Figure 9B] Same as above. [Figure 10A]Figures 10A-C show pie charts and schematic diagrams illustrating embodiments of the present invention. Figure 10A shows the pie chart distribution of a mouse grooming dataset and the performance of a machine learning algorithm applied to this dataset. A total of 2,637,363 frames were annotated across 1,253 video clips, annotated by two different annotators to create this dataset. The outer ring represents the distribution of the training dataset, and the inner ring represents the distribution of the validation dataset. Within each ring, gray-shaded areas with triangular markers indicate annotator agreement on grooming, gray-shaded areas without markers indicate annotator agreement on non-grooming, and light gray-shaded areas indicate annotator disagreement. Figure 10B provides a visual illustration of the implemented classification approach. A sliding window of frames was passed through the neural network to analyze the entire video. Dark gray shading indicates time spent grooming. Medium gray shading with stars indicates time spent non-grooming. Figure 10C shows how the network takes a video input and generates a single-frame grooming prediction. [Figure 10B] Same as above. [Figure 10C] Same as above. [Figure 11A]Figures 11A-E provide graphs, drawings, and photographs illustrating the dataset validation study. Figure 11A shows the agreement between annotators during dataset creation compared to the accuracy of the algorithm predicting on this dataset. The machine learning model was compared to only annotations where the annotators agreed. Figure 11B provides receiver operating characteristic (ROC) curves for three machine learning techniques trained on the training set and applied to the validation set. The final neural network model approach achieved the highest area under the curve (AUC) value, 0.9843402. The following graphs are used: No Marker, Final Neural Network Approach, Open Circle, Neural Network 20% Training Set, Open Square, JAABA 20% Training Set. Figure 11C provides a visual illustration of the proposed consensus solution. A 32x consensus approach was used, training four separate models, each with eight frame perspectives. To combine these predictions, all 32 predictions were averaged. While a single perspective from a single model may be incorrect, the average prediction using this consensus improved accuracy. Figure 11D-E provides images of example frames where the model correctly predicted grooming and non-grooming behaviors. [Figure 11B] Same as above. [Figure 11C] Same as above. [Figure 11D] Same as above. [Figure 11E] Same as above. [Figure 12]Figure 12 provides a graph showing the results of additional ROC curve subsets. This graph shows the performance of networks using different training data set sizes. The results show that as the number of training data samples increased, the ROC curve performance increased by up to one point. Graphed curves: open circle, four-model consensus neural network with temporal filtering; no marker, four-model consensus neural network; medium gray with asterisk, neural network 100% training set; triangle marker, neural network 65% training set; open square marker, neural network 20% training set; light gray without marker, neural network 10% training set. [Figure 13A]Figures 13A-E provide graphs of the results validating the algorithm's performance divided by video. The images show clusters of validation video ROC performance by machine learning approach. Figure 13A shows results where the majority of validation videos have good performance. Figure 13B shows results where some validation videos show slightly reduced performance. Figure 13C shows validation video results where the JAABA approach performs slightly worse than our neural network approach. Figure 13D provides results for two validation videos that performed poorly using both machine learning approaches. Inspecting these videos, frames annotated for grooming were visually difficult to classify. Figure 13E provides results for seven validation videos that performed well using the neural network but clearly performed poorly using JAABA. All videos that did not contain frames annotated for positive grooming did not have ROC curves and are therefore not shown. In most graphs, the traces of the four-model consensus neural network with temporal filtering and the JAABA 20% training set overlap. In all graphs except for graphs 54, 87, 90, 95, 108, 124, and 143, the top / leftmost traces of the non-overlapping traces or non-overlapping regions of traces represent the time-filtered 4-model consensus neural network, and the bottom / rightmost traces represent the JAABA 20% training set. In graphs 54, 87, 90, 95, 108, 124, and 143, the top / leftmost non-overlapping trace regions represent the JAABA 20% training set, and the bottom / rightmost non-overlapping trace regions represent the time-filtered 4-model consensus neural network. [Figure 13B] Same as above. [Figure 13C] Same as above. [Figure 13D] Same as above. [Figure 13E] Same as above. [Figure 14A]Figures 14A-B provide graphs showing the comparative results of different consensus styles and temporal smoothing. Figure 14A shows the ROC performance using different consensus styles. All consensus styles provided nearly identical results. Graphed curves: open circle, final neural network approach; no marker; single model; star, voting consensus; triangle, average pre-softmax function consensus; open square, maximum pre-softmax function consensus; diamond, average consensus; and "X", maximum consensus. Figure 14B shows the results of the temporal filter analysis. [Figure 14B] Same as above. [Figure 15A] Figures 15A-C provide diagrams, graphs, and tables related to embodiments of the present invention. Figure 15A provides an exemplary grooming ethogram for a single animal. Time is on the x-axis, with shaded bars indicating the time during which the animal engaged in grooming behavior. Summaries were calculated for ranges of 5, 20, and 55 minutes. Figure 15B provides a visual illustration of how grooming pattern phenotypes were defined. Figure 15C provides a table summarizing all behavioral indices analyzed. Phenotypes were categorized into four groups, including grooming amount, grooming pattern, open-field anxiety, and open-field activity. [Figure 15B] Same as above. [Figure 15C] Same as above. [Figure 16A]Figures 16A-H provide images of covariate analysis results for subjects with various characteristics, such as gender and season. Figure 16A shows results comparing male and female subjects. Figure 16B shows results comparing testing season for male and female subjects. Figure 16C shows results comparing time of testing day for male and female subjects. Figure 16D shows results comparing age at testing for male and female subjects. Figure 16E shows results comparing subjects housed in different rooms, the original room. For each room, the strains shown on the graph from left to right are C57BL / 6J F, C57BL / 6J M, C57BL / 6NJ F, and C57BL / 6NJ M. Figure 16F shows results comparing subjects tested under different lighting levels, lux, for male and female subjects. Figure 16G shows results comparing subjects tested by different testers for male and female subjects. FIG. 16H shows the results for male and female subjects comparing subjects tested with white noise during testing with subjects without white noise. [Figure 16B] Same as above. [Figure 16C] Same as above. [Figure 16D] Same as above. [Figure 16E] Same as above. [Figure 16F] Same as above. [Figure 16G] Same as above. [Figure 16H] Same as above. [Figure 17A]Figures 17A-F provide images or results of a lineage survey of grooming phenotypes with representative ethograms. Figure 17A shows the lineage survey results for total grooming time. The lines showed a smooth gradient of time spent grooming, with the wild-type line showing enrichment at the higher end. Figure 17B provides a representative ethogram showing lines with high and low total grooming time. Figure 17C provides the lineage survey results for number of grooming bouts. Figure 17D shows a comparative ethogram for two lines with different numbers of bouts but similar total time spent grooming. Figure 17E provides the lineage survey results for mean grooming bout duration. Figure 17F provides a comparative ethogram for two lines with different mean bout lengths but similar total time spent grooming. [Figure 17B] Same as above. [Figure 17C] Same as above. [Figure 17D] Same as above. [Figure 17E] Same as above. [Figure 17F] Same as above. [Figure 18A] Figures 18A-B show graphs depicting the total grooming and average bout length of different strains over time. Figure 18A shows results indicating that the total grooming time of the wild-type and classical strains is significantly different (*p<0.05, Mann-Whitney test). Figure 18B shows results indicating that the wild-type strain had significantly longer grooming bouts (**p<0.01, Mann-Whitney test). In both graphs, the BTBR strain is indicated by a triangle. [Figure 18B] Same as above. [Figure 19A]Figures 19A-C provide plots of results showing the association of grooming phenotypes. Figure 19A shows the results of a lineage study comparing total grooming time and number of bouts. The wild-type and BTBR lines showed enrichment for having high grooming but low bout numbers. Figure 19B provides the results of a lineage study comparing total grooming time and average bout duration. Most lines that groom more have longer average bout lengths. Figure 19C provides the results of a lineage study comparing number of bouts and average bout duration. [Figure 19B] Same as above. [Figure 19C] Same as above. [Figure 20-1] Figure 20 provides a graph panel showing clustering for systematic investigation of grooming patterns over time. Three distinct grooming patterns of mice in the open field were revealed by k-means clustering of grooming patterns over time. Grooming duration in 5-minute bins is shown over the course of the open field experiment (solid line) and data from individual mice (gray dots). [Figure 20-2] Same as above. [Figure 21] Figure 21 shows the results of k-means clustering. The first two major components from this clustering account for 81.7% of the variance. Three clusters were identified, allowing each lineage to be assigned to one of three grooming behaviors. [Figure 22A]Figures 22A-D provide graphs, plots, and heat maps showing genotype / phenotype correlations in grooming and open field data. Figure 22A provides calculated phenotype heritability estimates. Triangles, activity; filled squares, anxiety; diamonds, grooming pattern; and points, grooming amount. Figure 22B is a graph showing LD block size, i.e., average genotype correlations of single nucleotide polymorphisms (SNPs) at different genomic distances. Figure 22C provides a Manhattan plot of all phenotypes combined, with shading according to the peak SNP cluster and all SNPs in the same LD block shaded according to the peak SNP. Minimum p-value across all phenotypes for each SNP. Figure 22D is a heat map of all significant SNPs across all phenotypes. Each row (SNP) is shaded according to its assigned cluster in k-means clustering. Shading from the k-means clusters is used in Figure 22C. [Figure 22B] Same as above. [Figure 22C] Same as above. [Figure 22D] Same as above. [Figure 23] Figure 23 provides Manhattan plots for individual phenotypes. The plots were prepared from the linear mixed model (LMM) results for each SNP genotype using a Wald test for each phenotype. [Figure 24A]Figures 24A-G provide data related to mammalian phenotype ontology enrichment. Figure 24A shows the results for "Nervous System Phenotype" with p=7.5x10-4, MP:0003631, and 178 genes. Figure 24B shows the results for "Preweaning Lethality" with p=3.5x10-3, MP:0010770, and 189 genes. Figure 24C shows the results for "Abnormal Embryonic Development" with p=5.5x10-3, MP:0001672, and 62 genes. Figure 24D shows the results for "No Abnormal Phenotype Detected" with p=6.1x10-3, MP:0002169, and 102 genes. Figure 24E shows the results for "Normal Phenotype" with p=6.5x10-3, MP:0002873[ / g5], and 102 genes. Figure 24F shows the results of "embryonic growth retardation" with p=1.1x10-2, MP:0003984, and 41 genes. Figure 24G shows the results of "prenatal growth retardation" with p=1.4x10-2, MP:0010865, and 54 genes. [Figure 24B] Same as above. [Figure 24C] Same as above. [Figure 24D] Same as above. [Figure 24E] Same as above. [Figure 24F] Same as above. [Figure 24G] Same as above. [Figure 25-1] Figure 25 provides a diagram showing human-mouse trait relationships through a weighted bipartite network of PheWAS results. The width of the edge between gene nodes (filled circles) and psychotropic trait nodes is proportional to the association strength (-log10(p-value)). The size of the node is proportional to the number of associated genes or traits, and the color of the trait node corresponds to the subchapter level in the psychotropic domain. Eight modules were identified and visualized using Gephi 0.9.2 software. [Figure 25-2] Same as above. [Figure 25-3] Same as above. [Figure 25-4] Same as above. [Figure 25-5] Same as above. [Figure 25-6] Same as above. [Figure 25-7] Same as above. [Figure 25-8] Same as above. [Figure 25-9] Same as above. [Figure 26] FIG. 26 is a block diagram conceptually illustrating exemplary components of an apparatus according to an embodiment of the present disclosure. [Figure 27] FIG. 27 is a block diagram conceptually illustrating exemplary components of a server, in accordance with an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0013] The present invention relates, in part, to a method and system for classifying behavioral activity, such as mouse grooming behavior, using machine learning models (e.g., neural networks). Grooming represents a highly biologically significant form of stereotyped or patterned behavior, consisting of a range of actions from very small to very large. Grooming is an innate behavior conserved across animal species, including mammals. In rodents, a significant amount of waking behavior, between 20% and 50%, consists of grooming [Van deWeerd H. et al., 2001 Behavioral Processes 53(1-2):11-20; Spruijt BM. et al., 1992 Physiology Reviews 72(3):825-852; and Bolles RC. 1960 Journal of Comparative and Physiological Psychology 53(3):306]. Grooming serves many adaptive functions, including hair and body care, stress reduction, arousal reduction, social function, thermoregulation, nociception, and other functions [Spruijt BM et al., 1992 Physiology Reviews 72(3):825-852; Kalueff AV et al., Neurobiology of Grooming Behavior. Cambridge University Press; 2010; and Fentress JC 1988 Annals of the New York Academy of Sciences 525(1):18-26]. Although the neural circuits regulating grooming behavior have been studied, their function remains largely unknown. Patterned activity, including but not limited to grooming behavior, is an endophenotype for many psychiatric disorders. For example, high levels of stereotypic behavior are observed in autism spectrum disorder (ASD), whereas Parkinson's disease demonstrates an inability to generate patterned behavior [Kalueff AV et al., Neurobiology of Grooming Behavior. Cambridge University Press; 2010].Certain embodiments of the present invention can be used for the accurate, automated analysis of behaviors to assess diseases and conditions. The term "condition" is used interchangeably herein for the identification and use of therapeutics to treat diseases and conditions associated with one or more predetermined behaviors, also referred to herein as behavioral pattern activity, (the term "disorder") associated with behavioral pattern activity.

[0014] This disclosure describes techniques for detecting subject behavior by analyzing video data capturing the subject's behavior. The system may process the video data using multiple machine learning (ML) models to generate multiple predictions regarding whether one or more frames of the video data indicate a subject exhibiting a defined behavior. These ML models may be configured to detect specific behaviors using training data including videos capturing the subject's movements, where the training data includes a label for each video frame that identifies whether the subject is exhibiting the specific behavior. These ML models may be configured using a large training dataset. Based on the configuration of the ML models, the system may be configured to detect different behaviors.

[0015] Each ML model processing the video data may be configured with different initialization parameters or settings, allowing the ML models to have variations in terms of certain model parameters (learning rate, weights, batch size, etc.), and therefore result in different predictions (regarding the subject's behavior) when processing the same video frames.

[0016] The system may also process different representations of video data capturing the subject's movements. The system may determine different representations of the video data by changing the orientation of the video. For example, one orientation may be determined by rotating the video 90 degrees to the left, another orientation may be determined by rotating the video 90 degrees to the right, and yet another orientation may be determined by reflecting the video along a horizontal or vertical axis. The system may process video frames in the originally captured orientation and other different orientations. Based on the processing of the different orientations, the system may determine different predictions (regarding the subject's behavior) for the same time period.

[0017] The different predictions determined as described above can be used to make a final decision on whether the subject is exhibiting an action in the video frame. Aspects of the present disclosure provide improved subject action detection. For example, the use of different predictions from multiple ML models and from processing different orientations of the video data results in robust predictions. The system configuration allows for the detection of subject actions in different subjects, even when they vary in terms of physical characteristics (e.g., size, body shape, color, etc.). Furthermore, the techniques described herein can be used to detect different actions.

[0018] FIG. 1 conceptually illustrates a system 100 that may be used to detect subject behavior using video data. System 100 may include an image capture device 101, a device 102, and one or more systems 150 connected across one or more networks 199. Image capture device 101 may be part of, included in, or connected to another device (e.g., device 2600), which may be a camera, a high-speed video camera, or other type of device capable of capturing images and video. Device 101 may include motion detection sensors, infrared sensors, temperature sensors, atmospheric condition detection sensors, and other sensors configured to detect various characteristics / environmental conditions in addition to or instead of an image capture device. Device 102 may be a laptop, desktop, tablet, smartphone, or other type of computing device capable of outputting data and may include one or more components described in connection with device 2600 below.

[0019] Image capture device 101 may capture a video (or one or more images) of an object and may transmit video data 104 representing the video to system 150 for processing as described herein. The video may capture the movement of an object (including non-movement of the object) over a period of time. System 150 may include one or more components shown in FIG. 1 and may be configured to process video data 104 to determine the object's behavior over time. System 150 may determine output labels 130 associated with one or more frames of video data 104, where the output labels may indicate whether the object is exhibiting the behavior in the respective frame. Data representing output labels 130 may be transmitted to device 102 for output to a user to view the results of processing video data 102.

[0020] The components of system 150 are described in detail below. The various components may be located on the same or different physical devices. Communication between the various components may occur directly or via network 199. Communication between device 101, system 150, and device 102 may occur directly or via network 199. One or more of the components shown as part of system 150 may be located on device 102 or on a computing device (e.g., device 2600) connected to image capture device 102.

[0021] System 150 may use video data 104 to determine multiple sets of frames, where different sets may represent different orientations of the video data. Set of frames 112 may be the original orientation of video data 104 captured by image capture device 101. Rotated set of frames 114 may be a rotated orientation of video data 104; for example, set of frames 112 may be rotated 90 degrees to the left to generate rotated set of frames 114. Reflected set of frames 116 may be a reflected orientation of video data 104; for example, set of frames 112 may be reflected across a horizontal axis (or rotated 180 degrees) to generate reflected set of frames 116. Rotated set of frames 118 may be another rotated orientation of video data 104; for example, set of frames 112 may be rotated 90 degrees to the right to generate rotated set of frames 118. In other embodiments, sets of frames 114, 116, and 118 may be generated by manipulating (original) set of frames 112 in other ways (e.g., reflected across a vertical axis, rotated by another number of degrees, etc.) In other embodiments, more or fewer orientations of the video data may be processed.

[0022] The sets of frames 112, 114, 116, and 118 may correspond to the same time period of the video data 104. For example, the first set of frames 112, the first rotated set of frames 114, the first reflected set of frames 116, and the first set of frames 118 may correspond to a first time period, the second set of frames 112, the second rotated set of frames 114, the second reflected set of frames 116, and the second set of frames 118 may correspond to a second time period, etc.

[0023] System 150 may include a video processing component 120 that may be configured to process sets of frames 112, 114, 116, and 118 to determine final predicted data 125. Details of how video processing component 120 may process video data 104 are described below in conjunction with FIGS. 2-6. Final predicted data 125 may indicate during which frames of video data 104 the subject is exhibiting a behavior and during which frames of video data 104 the subject is not exhibiting a behavior. In some embodiments, video processing component 120 may be configured to detect when the subject is exhibiting grooming behaviors that may include paw licking, one-sided face washing, both-sided face washing, and flank licking.

[0024] In some embodiments, the final prediction 125 may be used to determine the output labels 130 so that data associating the output labels 130 with the video frames may be transmitted to the device 102 (or multiple devices 102). In other embodiments, the final prediction 125 may be used to determine an ethogram representing the subject's behavior during the period of the video data 104.

[0025] FIG. 2 is a conceptual diagram illustrating how different sets of frames are processed using trained models. Video processing component 120 may employ one or more trained models, such as trained model 210. In some embodiments, trained model 210 may process set of frames 112 to generate predicted data 220. Predicted data 220 may be a probability or likelihood of exhibiting a target behavior in the video represented by set of frames 112. Trained model 210 may process set of rotated frames 114 to generate predicted data 222, which may be a probability or likelihood of exhibiting a target behavior in the video represented by set of rotated frames 114. Trained model 210 may process set of rotated frames 116 to generate predicted data 224, which may be a probability or likelihood of exhibiting a target behavior in the video represented by set of rotated frames 116. Trained model 210 may process set of rotated frames 118 to generate predicted data 226, which may be a probability or likelihood of exhibiting a target behavior in the video represented by set of rotated frames 118. In this way, the same trained model 210 may process different orientations of the video data and generate different predictions for the motion of the same captured object.

[0026] FIG. 3 is a conceptual diagram illustrating how different sets of frames are processed using different trained models. The video processing component 210 may employ different trained models 212. In some embodiments, the trained models 212 may process the set of frames 112 to generate predicted data 240. The predicted data 240 may be a probability or likelihood of a subject exhibiting a behavior during the video represented by the set of frames 112. The trained models 212 may process the set of rotated frames 114 to generate predicted data 242, which may be a probability or likelihood of a subject exhibiting a behavior during the video represented by the set of rotated frames 114. The trained models 212 may process the set of rotated frames 116 to generate predicted data 244, which may be a probability or likelihood of a subject exhibiting a behavior during the video represented by the set of rotated frames 116. The trained models 212 may process the set of rotated frames 118 to generate predicted data 248, which may be a probability or likelihood of a subject exhibiting a behavior during the video represented by the set of rotated frames 118. In this way, different trained models 212 may process different orientations of the video data to generate different additional predictions for the same captured object's motion. The probabilities may be values ​​ranging from 0.0 to 1.0, or values ​​ranging from 0 to 100, or another numeric range.

[0027] Each of the predicted data 220, 222, 224, 226, 240, 242, 244, and 246 may be a data vector including multiple probabilities (or scores), each probability corresponding to a respective one of the frames 112, 114, 116, 118 of the set and indicating the likelihood of exhibiting the target behavior during the corresponding frame. For example, the predicted data may include a first probability corresponding to a first frame of the video data 104, a second probability corresponding to a second frame of the video data 104, etc. Each of the predicted data 220, 222, 224, 226, 240, 242, 244, and 246 may be a different probability of exhibiting the target behavior during the video represented in the set of frames 112.

[0028] In some embodiments, the set of frames 112 may include a number of video frames (e.g., 16 frames), each frame being a duration of the video for a period of time (e.g., 30 milliseconds, 30 seconds, etc.). Each of the trained models 210 and 212 may be configured to process the set of frames 112 to determine the probability of exhibiting the target behavior in the last frame of the set of frames. For example, if there are 16 frames in the set of frames, the output of the trained models may indicate whether the target exhibits the behavior in the 16th frame of the set of frames. The trained models 210 and 212 may be configured to make a prediction of the last frame using contextual information from other frames in the set of frames. In other embodiments, the output of the trained models 210 and 212 may determine the likelihood of exhibiting the target behavior in another frame of the set of frames (e.g., an intermediate frame, an eighth frame, a first frame, etc.).

[0029] 4 is a conceptual diagram illustrating how different predictions may be processed to determine a final prediction regarding a subject exhibiting a behavior. The video processing component 120 may include an aggregation component 230 to process different predictions determined by different trained models (e.g., 210 and 212) using different sets of frames (e.g., 112, 114, 116, and 118) to determine the final prediction data 125. The aggregation component 230 may be configured to aggregate, aggregate, or otherwise combine the different prediction data 220, 222, 224, 226, 240, 242, 244, and 246 to determine the final prediction data 125.

[0030] In some embodiments, the aggregation component 230 may average the probabilities of each frame represented in the predicted data 220, 222, 224, 226, 240, 242, 244, and 246, and the final predicted data 125 may be a data vector of average probabilities for each frame in the video data 104. In some embodiments, the video processing component 120 may determine the output label 130 for a frame based on the corresponding average probability of the frame satisfying a condition (e.g., if the probability exceeds a threshold probability / value).

[0031] In other embodiments, the summation component 230 may sum the probabilities of each frame represented in the prediction data 220, 222, 224, 226, 240, 242, 244, and 246, and the final prediction data 125 may be a data vector of summed probabilities for each frame in the video data 104. In some embodiments, the video processing component 120 may determine the output label 130 for a frame based on the frame's corresponding summed probability satisfying a condition (e.g., if the probability exceeds a threshold probability / value).

[0032] In some embodiments, the aggregation component 230 may be configured to select the maximum value (e.g., highest probability) from the predicted data 220, 222, 224, 226, 240, 242, 244, and 246 for each frame as the final predicted data 125 for the frame. In other embodiments, the aggregation component 230 may be configured to determine the median value from the predicted data 220, 222, 224, 226, 240, 242, 244, and 246 for each frame as the final predicted data 125 for the frame.

[0033] In some embodiments, system 150 may determine binned value representations of the probabilities in final predicted data 125. For example, probabilities falling within a first range of values ​​(e.g., 0 to 0.33) may be assigned a "low" value / label, probabilities falling within a second range of values ​​(e.g., 0.34 to 0.66) may be assigned a "medium" value / label, and probabilities falling within a third range of values ​​(e.g., 0.67 to 1.0) may be assigned a "high" value / label. In other embodiments, the binned value representations may be "low," "high low," "medium," "low high," and "high."

[0034] In some embodiments, the system 150 may apply a temporal smoothing filter over a number of frames (e.g., 46 frames). The temporal smoothing filter may be applied after the system 150 processes the video data 104 to identify frames in which the subject exhibits behavior. Temporal smoothing may be used to correct potentially inaccurate isolated predictions, such as, but not limited to, a single frame predicted as "not grooming" within a grooming bout (where the subject exhibits grooming behavior for a period of time, and frames from that period may be labeled as grooming, but one or two frames may be labeled as not grooming). In some embodiments, the temporal smoothing is a rolling average of a functional 46 frames that suppresses outlier predictions. Alternative temporal smoothing strategies may also be selected and used in combination with the methods of the present disclosure. FIG. 5 is a conceptual diagram illustrating a system employing multiple trained models to process different sets of frames. 2 and 3 illustrate the use of two trained models 210 and 212, as shown in Figure 5, video processing component 120 may, in some embodiments, employ four trained models 210, 212, 214, and 216. In other embodiments, video processing component 120 may employ fewer or more than four trained models.

[0035] Each of the trained models 210, 212, 214, and 216 may process a respective set of frames 112, 114, 116, and 118 and generate 32 different predictions corresponding to the frame and indicating whether the subject exhibited the behavior in that frame. As described above, the aggregation component 220 may combine the 32 different predictions to determine the final prediction 125.

[0036] 6 conceptually illustrates components for training / configuring a machine learning (ML) model to determine whether a subject exhibits an action in a frame of video data. As described above, system 150 may employ multiple trained models (e.g., 210, 212, 214, and 216). Each of the trained models may be trained separately and may use different initialization parameters, resulting in different trained models. Depending on how trained models 210, 212, 214, and 216 are trained, each of them may output different predictions / probabilities for a frame.

[0037] The model construction component 610 may train one or more ML models to determine whether a user input results in an error and when to rephrase the user input. The model construction component 610 may train one or more ML models during offline operation to generate one or more trained models. The model construction component 610 may train one or more ML models using a training dataset.

[0038] In some embodiments, the ML model is a neural network (e.g., a convolutional neural network, a recurrent neural network, a deep learning network, etc.). The model construction component 610 may be provided with different initialization parameters to determine the different trained models 210, 212, 214, and 216. The initialization parameters may relate to and define one or more of initial weights corresponding to one or more layers of the neural network, biases for one or more layers of the neural network, a learning rate of the ML model, etc. The model construction component 610 may be provided with different algorithms or data to determine the different trained models 210, 212, 214, and 216. Such algorithms or data may relate to and define one or more of an optimization algorithm, a loss function, a batch size, training data, the order in which the training data is processed, the number of epochs, etc.

[0039] In some embodiments, the ML model is a classifier (which may be a neural network-based classifier). The classifier may be configured to perform binary classification and determine whether a subject exhibits a behavior in a video frame. The classifier may be configured to perform multi-class or polynomial classification and determine whether a subject exhibits a behavior from two or more behavioral classes / categories in a video frame. The classifier may be configured to perform multi-label classification and determine whether a subject exhibits one or more behaviors in a video frame.

[0040] In some embodiments, system 150 may include one or more ML models, including, but not limited to, one or more classifiers, one or more neural networks, one or more probability graphs, one or more decision trees, etc. In other embodiments, system 150 may include a rule-based engine, one or more statistics-based algorithms, one or more mapping functions, or other types of functions / algorithms for detecting behavior of interest.

[0041] The training dataset 602 may be video data representing the movements of multiple different objects. In some embodiments, the objects may be mice, and the training dataset 602 may be videos representing the movements of different mice. The training dataset 602 may include videos of mice of different body sizes, body shapes, coat colors, etc. In this manner, the trained models 210, 212, 214, and 216 may be configured to detect the behavior of mouse objects regardless of the physical characteristics of the mice. The training dataset 602 may include several hours of video data so that the ML models are sufficiently trained. The training dataset 602 may include labeled data (which may also be referred to herein as behavioral actions, behavioral activities, predetermined behavioral activities, or predetermined behavioral actions) that identify which frames in the video exhibit behaviors. In some embodiments, the training dataset 602 includes labeled data that identifies when mouse objects exhibit grooming behaviors. In other embodiments, the training dataset 602 may include labeled data that identifies when objects exhibit another predetermined behavioral action. The methods of the present disclosure can be used to identify and evaluate grooming behaviors, and can also be used to detect other "active" behaviors, which may be behaviors that involve movement. Non-limiting examples of other behaviors that may be detected and evaluated using embodiments of the present disclosure are rearing behaviors, running behaviors, jumping behaviors, cognitive behaviors, conscious behaviors, consummatory behaviors, elimination behaviors, emotional behaviors, impulsive behaviors, motor behaviors, motivational behaviors, play behaviors, reproductive behaviors, social behaviors, stress-related behaviors, rhythmic behaviors, and behavioral regulation.

[0042] Using the training data set 602 and a first set of initialization parameters, data, and algorithms, the model construction component 610 may construct the first trained model 210. Using the training data set 602 and a second set of initialization parameters, data, and algorithms, the model construction component 610 may construct the second trained model 212. Using the training data set 602 and a third set of initialization parameters, data, and algorithms, the model construction component 610 may construct the third trained model 214. Using the training data set 602 and a fourth set of initialization parameters, data, and algorithms, the model construction component 610 may construct the fourth trained model 216. Once constructed, the trained models 210, 212, 214, and 216 may be stored for use during runtime operations when the video data 104 is processed.

[0043] 7 is a flowchart illustrating a process 700 for analyzing a subject's video data 104 to detect subject behavior, according to an embodiment of the present disclosure. The process steps illustrated in FIG. 7 may be performed by system 150. In other embodiments, one or more of the process steps may be performed by a computing device associated with device 102 or image capture device 101.

[0044] System 150 receives video data capturing the movement of a subject (702). System 150 identifies a set of frames (e.g., 16 frames) from the video data for processing (704). System 150 uses the set of frames to determine a set of rotated frames (706). For example, system 150 may rotate the original set of frames 90 degrees to determine a corresponding set of rotated frames. System 150 uses the set of frames to determine a set of reflected frames (708). For example, system 150 may reflect the original set of frames across a horizontal axis to determine a corresponding set of reflected frames. System 150 processes the set of frames, the set of rotated frames, and the set of reflected frames (710) using one or more trained models configured to detect subjects exhibiting a predetermined behavioral action. System 150 uses the output of the one or more trained models to determine that the subject exhibits a behavioral action (712).

[0045] subject Some aspects of the present invention involve the use of automated phenotyping methods with a subject. As used herein, the term "subject" may refer to a human, non-human primate, cow, horse, pig, sheep, goat, dog, cat, pig, bird, rodent, or other suitable vertebrate or invertebrate organism. In certain embodiments of the present invention, the subject is a mammal, and in certain embodiments of the present invention, the subject is a human. In some embodiments, the methods of the present invention may be used with rodents, including, but not limited to, mice, rats, gerbils, hamsters, and the like. In some embodiments of the present invention, the subject is a normal, healthy subject, and in some embodiments, the subject is known to have, is at risk for, or is suspected of having a disease or condition. The terms "subject" and "test subject" may be used interchangeably herein.

[0046] By way of non-limiting example, a subject evaluated using the systems and / or methods of the present invention may be a subject that is an animal model of a disease or condition, such as one or more models of bipolar disorder, dementia, depression, hyperactivity disorder, anxiety disorder, developmental disorder, sleep disorder, Alzheimer's disease, Parkinson's disease, physical injury, etc. Additional models of diseases and disorders that can be evaluated using the methods and / or systems of the present invention are known in the art, see, e.g., Barrot M. Neuroscience 2012;211:39-50; Graham, DM, Lab Anim (NY) 2016;45:99-101; Sewell, RDE, Ann Transl Med 2018;6:S42. 2019 / 01 / 08; and Jourdan, D., et al., Pharmacol Res 2001;43:103-110, the contents of which are incorporated herein by reference in their entireties.

[0047] In some embodiments, a subject may be monitored using the activity determination methods or systems of the present invention to detect the presence or absence of an activity disorder or condition. In certain embodiments of the present invention, a subject that is an animal model of an activity and / or mobility condition may be used to assess the subject's response to a condition. Furthermore, a subject that is an animal model of a mobility and / or activity condition may be administered a candidate therapeutic agent or method that is monitored using the activity monitoring methods and / or systems of the present invention, and the results may be used to determine the effectiveness of the candidate therapeutic agent in treating the condition. The terms "activity" and "action" may be used interchangeably herein.

[0048] In some embodiments of the methods of the present invention, the subject is a wild-type subject. As used herein, the term "wild-type" refers to a phenotype and / or genotype that is typical of a species occurring in nature. In certain embodiments of the present invention, the subject is a non-wild-type subject, for example, a subject that has one or more genetic modifications compared to the wild-type genotype and / or phenotype of the subject's species. In some cases, the difference in the subject's genotype / phenotype compared to the wild-type is due to inherited (germline) mutations or acquired (somatic) mutations. Factors that can cause a subject to exhibit one or more somatic mutations include, but are not limited to, environmental factors, toxins, ultraviolet radiation, spontaneous errors occurring in cell division, radiation, maternal infection, chemicals, and other teratogenic events.

[0049] In certain embodiments of the methods of the present invention, the subject is a genetically modified organism, also referred to as a genetically engineered subject and / or engineered subject. An engineered subject may contain preselected and / or intentional genetic modifications and, as such, exhibit one or more genotypic and / or phenotypic traits that differ from those of a non-engineered subject. In some embodiments of the present invention, conventional genetic engineering techniques can be used to generate engineered subjects that exhibit genotypic and / or phenotypic differences compared to non-engineered subjects of the same species. As a non-limiting example, genetically engineered mice in which functional gene products are absent or present at reduced levels can be used to evaluate the phenotype of the genetically engineered mice, and the methods or systems of the present invention can be used to compare the results with those obtained from a control (control results).

[0050] As described elsewhere, the trained models of the present invention may be configured to detect a subject's behavior regardless of the subject's physical characteristics. In some embodiments of the present invention, one or more of the subject's physical characteristics may be pre-identified characteristics. For example, and without intending to be limiting, the pre-identified physical characteristics may be one or more of body shape, body size, coat color, sex, age, and disease or disorder phenotype.

[0051] Diseases and Disorders The methods and systems of the present invention can be used to assess the activity and / or behavior of subjects known to have, suspected of having, or at risk of having a disease or condition. In some embodiments, the disease and / or condition is associated with abnormal levels of activity or behavior. In a non-limiting example, a subject who may have anxiety or be an animal model of anxiety may have one or more activities or behaviors associated with anxiety that can be detected using embodiments of the methods of the present invention. The results of assessing a subject can be compared to control results of the assessment, such as a control subject who does not have anxiety, a control subject who is not an animal model of anxiety, or a control obtained from multiple subjects without the condition. Differences in the results between the subject and the control can be compared. Some embodiments of the methods of the present invention can be used to identify subjects with a disease or condition associated with abnormal activity and / or behavior. The terms "behavior" and "predetermined behavior" may be used interchangeably herein.

[0052] The methods and systems of the present invention can be used to assess, determine, and / or monitor one or more predetermined behaviors in a subject. Results from such assessments can be used to assess and / or monitor a predetermined behavior-related disease or disorder. The term "predetermined behavior-related disease or disorder" may be used interchangeably herein with the term "predetermined behavior-related disease or condition." The methods and systems of the present invention can be used to determine, assess, and / or monitor behaviors that reflect the physical characteristics of a predetermined behavior-related disease or disorder. As used herein, the term "predetermined behavior-related disease or disorder" refers to a disease, condition, or disorder that can be characterized by one or more predetermined behaviors that can be assessed using the methods or systems of the present invention. In non-limiting examples, embodiments of the methods of the present invention can be used to assess a subject known to have or suspected of having Parkinson's disease, or to assess a subject that is an animal model of Parkinson's disease. One or more predetermined behaviors associated with Parkinson's disease are assessed in the subject, and the results identify the subject's Parkinson's disease status. It will be understood that the results obtained by assessing a subject can be compared to a control assessment, thereby identifying the subject's Parkinson's disease status. Method embodiments of the present invention can be used to determine the presence or absence of a given behavior-related disease or disorder in a subject, and can be used to determine and / or monitor the onset, progression, and / or regression of a given behavior-related disease or disorder in a subject.

[0053] The onset, progression, and / or regression of a disease or condition associated with abnormal activity and / or behavior can also be assessed and tracked using embodiments of the methods of the present invention. For example, in certain embodiments of the methods of the present invention, two, three, four, five, six, seven, or more assessments of a subject's activity and / or behavior are performed at different time points. Comparison of two or more assessment results from different time points can indicate differences in the subject's activity and / or behavior. An increase in the determined level or type of activity can indicate the onset and / or progression of the subject's disease or condition associated with the assessed activity. A decrease in the determined level or type of activity can indicate the regression of the subject's disease or condition associated with the assessed activity. A determination that the subject's activity has stopped can indicate the cessation of the subject's disease or condition associated with the assessed activity.

[0054] Certain embodiments of the methods of the present invention can be used to evaluate the effectiveness of a therapy for treating a disease or condition associated with abnormal activity and / or behavior. For example, a subject may be provided with a candidate therapy and a method of the present invention, which is used to determine whether or not the subject has a change in activity associated with a disease or condition. A decrease in abnormal activity after providing the candidate therapy can indicate the effectiveness of the candidate therapy for the disease or condition.

[0055] Non-limiting examples of activity or behavior-related diseases and conditions that can be assessed using the methods of the present invention include bipolar disorder, depression, anxiety, eating disorders, hyperactivity disorders, drug addition, obsessive-compulsive disorder, schizophrenia, Alzheimer's disease, Parkinson's disease, sleep disorders, and the like.

[0056] Assays and Screening of Control and Candidate Compounds Using the activity monitoring methods or systems of the present invention, results obtained for a subject can be compared with control results. The methods of the present invention can also be used to assess differences in phenotype between a subject and a control. Thus, some aspects of the present invention provide methods for determining the presence or absence of changes in activity in a subject compared to a control. Some embodiments of the present invention include using embodiments of the methods of the present invention to identify phenotypic characteristics of a disease or condition, and in certain embodiments of the present invention, automated phenotyping is used to evaluate the effect of a candidate therapeutic compound on a subject.

[0057] Results obtained using the methods and / or systems of the present invention can be advantageously compared to a control. In some embodiments of the present invention, one or more subjects can be evaluated using the methods of the present invention, and then the subjects can be re-examined after administering a candidate therapeutic compound to the subjects. The terms "subject" and "test subject" may be used herein in connection with a subject being evaluated using the methods or systems of the present invention, and the terms "subject" and "test subject" are used interchangeably herein. In certain embodiments of the present invention, results obtained using a method for assessing one or more activities in a subject are compared to results obtained from the method performed on another subject. In some embodiments of the present invention, results from a subject are compared to assay results performed on the subject at a different time. In some embodiments of the present invention, results obtained using the methods of the present invention to evaluate a subject are compared to control results.

[0058] As used herein, a control result may be a predetermined value that can take various forms. It may be a single cutoff value, such as a median or mean. It may be established based on a comparison group, such as subjects evaluated using the methods and / or systems of the present invention under conditions similar to those of the subjects, where the subjects are administered a candidate therapeutic agent and the comparison group is not administered the candidate therapeutic agent. Another example of a comparison group may include subjects known to have a disease or condition and a subject or group of subjects without the disease or condition. Another comparison group may be subjects with a family history of the disease or condition and subjects from a group without such a family history. The predetermined value can be configured, for example, by dividing the tested population into equal (or unequal) groups based on the results of the test. Those skilled in the art will be able to select appropriate control groups and values ​​for use in the comparison methods of the present invention.

[0059] Subjects evaluated using the methods or systems of the present invention can be monitored for changes occurring under test conditions compared to control conditions. By way of non-limiting example, changes occurring in a subject can include, but are not limited to, one of the following: frequency of movement, licking behavior, response to external stimuli, etc. The methods and systems of the present invention can be used with subjects to assess symptoms of a disease or disorder in the subject, and can also be used to evaluate the effectiveness of candidate therapeutic agents.

[0060] In some embodiments, the methods and / or systems of the present invention are used to assess a predetermined behavioral action in a subject and to evaluate the efficacy and / or effectiveness of a candidate therapeutic agent in the subject. In certain embodiments, the method includes administering a candidate therapeutic agent to the subject, assessing the predetermined behavior in the subject after administration of the candidate therapeutic agent, and comparing the post-administration assessment with a control assessment of the predetermined behavior, wherein a change in the predetermined behavior after administration compared to the control predetermined behavior identifies the effect of the administered candidate therapeutic agent on the predetermined behavior. In some embodiments, the subject has a predetermined behavior-related disease or disorder. In some embodiments, the subject is an animal model of a predetermined behavior-related disease or disorder.

[0061] As a non-limiting example of a method of the present invention for assessing the presence or absence of a change in a subject as a means of identifying the effectiveness of a candidate therapeutic agent, a subject known to have a behavior-related condition is assessed using a method of the present invention. The subject is then administered a candidate therapeutic agent and assessed again using the method. The presence or absence of a change in the subject's outcome indicates the presence or absence of an effect of the candidate therapeutic agent on the condition, respectively.

[0062] It will be appreciated that in some embodiments of the invention, a subject may serve as its own control, for example, by being assessed more than once using a method of the invention and comparing the results obtained at two or more of the different assessments. The methods and systems of the invention may be used to assess the progression or regression of a subject's disease or condition using two or more assessments of the subject, using embodiments of the method or system of the invention, by identifying and comparing changes in the subject's phenotypic characteristics over time. [Example]

[0063] Materials and Methods for Examples 1-8 Dataset annotations Data were selected for grooming by training a preliminary JAABA classifier and then annotating by clipping video chunks based on predictions across a wide variety of videos. The initial JAABA classifier was trained on 13 short clips manually enriched for grooming activities. This classifier was intentionally weak, designed to simply prioritize video clips that were informative for annotation selection. 150-frame video time segments surrounding grooming activity predictions were clipped to mitigate the possibility of a highly imbalanced dataset. 1,253 video clips were generated with a total of 2,637,363 frames. Each video had variable duration, depending on the grooming prediction length. The shortest video clip contained 500 frames, while the longest contained 23,922 frames. The median video clip length was 1,348 frames.

[0064] From this, seven annotators were trained. From this pool of seven trained annotators, two annotators were assigned to fully annotate each video clip. If there was confusion about a particular frame or sequence of frames, the annotators could request additional opinions. For frames that were difficult to annotate, the annotators were asked to provide a "grooming" or "not grooming" annotation for each frame, with the intention of obtaining different annotations from each annotator. Training and validation were performed using only frames on which the annotators agreed, reducing the total number of frames to 2,487,883.

[0065] Neural Network Model The neural network followed a typical feature encoder structure, except that it used 3D convolution and pooling instead of 2D convolution. The model started with a 16x112x112x1 input video segment, where "16" indicates the input's temporal dimension and "1" indicates the color depth (monochrome). Each applied convolution layer was zero-padded to maintain the same height and width dimensions. Furthermore, each convolution layer was followed by batch normalization and rectified linear unit (reLU) activation. The first applied was two successive 3D convolution layers with a kernel size of 3x3x3 and a filter count of 4. The second applied was a max-pooling layer of shape 2x2x2, resulting in a new tensor shape of 8x64x64x4. This two-fold iteration of 3D convolution and max-pooling, doubling the filter depth each time, was repeated three more times, resulting in a 1x8x8x32 tensor shape. Two final 3D convolutions with a kernel size of 1x3x3 and a filter depth of 64 were applied, resulting in a 1x8x8x64 tensor shape. The network was then flattened to produce a 64x64 tensor. After flattening, two fully connected layers were applied, each with a filter depth of 128, batch normalization, and reLU activation. Finally, another fully connected layer with a filter depth of only 2 and softmax activation was added. This final layer was used as the output probability for grooming and non-grooming prediction.

[0066] Neural Network Training Four individual neural networks were trained using the same training set and four independent initializations. During training, we randomly sampled video chunks from the dataset whose final frame contained annotations agreed upon by the annotators. Because a 16-frame duration was sampled, this refers to the annotation of the 16th frame. If the selected frame did not previously have 15 frames of video, the tensor was padded with zero initialized frames. Random rotations and data reflections were applied, increasing the effective dataset size by 8x. The loss function used in the networks was categorical cross-entropy loss, which compares the softmax prediction from the network with a one-hot vector with the correct classification. We used the Adam optimizer with an initial learning rate of 10.5. A learning rate decay schedule was applied to halve the learning rate if five epochs followed without an increase in validation accuracy. A stopping criterion was also adopted if validation accuracy did not improve by more than 1% after 10 epochs. During training, we assembled example video clips with a batch size of 128. Typical training runs were performed after 13-15 epochs, followed by 23-25 ​​epochs without any additional improvement.

[0067] JAABA training The Janelia Automatic Animal Behavior Annotator (JAABA) classifier was trained using two different approaches. The first approach was to use guidelines provided by the software developers. This involved training the classifier interactively and iteratively. The data selection approach involved annotating some data and prioritizing new annotations if the algorithm was uncertain or was about to make an incorrect prediction. This interactive training continued until the algorithm no longer improved with k-fold cross-validation.

[0068] The second approach was to subset the large annotated dataset to fit JAABA and train on a consensus annotation. Initially, we attempted to utilize the entire training dataset, but the machine did not have enough RAM to process the entire training dataset. The workstation used had 96 GB of available RAM. A custom program script was written to convert the annotation format and populate the JAABA classifier file with the annotations. To ensure the data was entered correctly, we inspected the annotations within the JAABA interface. Once this file was created, we were able to train the JAABA classifier using the JAABA interface. After training, the model was applied to a validation dataset and compared to the neural network model. This was repeated with training datasets of various sizes.

[0069] Defining grooming behavior indicators This section describes the various grooming behavior metrics used in the following analyses. Following the approach described in [Kalueff AV, et al., Neurobiology of Grooming Behavior. Cambridge University Press; 2010], a grooming bout was defined as a continuous time duration of uninterrupted grooming lasting more than 3 seconds. Short interruptions (less than 10 seconds) were allowed, but locomotor activity was not allowed for this merge of time segments spent grooming. Specifically, pauses occurred when the mouse's movement did not exceed twice its average body length. To reduce data complexity, grooming duration, bout number, and average bout duration were summarized into 1-minute segments. To capture the total number of bouts per time duration, grooming bouts were assigned to the time segment in which the bout began. In the rare cases where multiple fractional bouts occurred, this allowed a 1-minute time segment to contain grooming durations worth more than 1 minute.

[0070] From this, the total duration of grooming calls across all grooming bouts was calculated as the total duration of grooming calls. Note that uncoupled grooming segments shorter than 3 seconds were excluded because they were not considered bouts. Furthermore, the total number of bouts was counted. Once the number of bouts and total duration were determined, the average bout duration was calculated by dividing the two. Finally, the data were binned into one-minute time segments, and a line was fitted to the data. The positive slope for total grooming duration suggested that the longer an individual mouse remained in the open field test, the more time it spent grooming. The negative slope for total grooming duration suggested that mice spent more time grooming at the beginning of the open field test than at the end. This is usually because mice choose to spend more time on other activities, such as sleeping, rather than grooming. The positive slope for the number of bouts suggested that the longer a mouse remained in the open field test, the more grooming bouts it initiated.

[0071] Genome-wide association analysis Phenotypes obtained by machine learning algorithms for several strains were used to study the association between genome and strain behavior. A subset of 10 individuals from each strain and sex combination was randomly selected from the tested mice to ensure equal group sample sizes. Genotypes for different strains were obtained from the Mouse Phenome Database (http: / / phenome.jax.org / genotypes). Mouse diversity array (MDA) genotypes were used to infer diallelic genomes from the parental genomes. SNPs with at least 10% MAF and up to 5% missing data were used, resulting in 222,967 SNPs out of 470,818 SNPs genotyped by the MDA array. The LMM method in the GEMMA software package (Zhou and Stephens, 2012 Nature Genetics 44(7):821-824) was used for each phenotype in the GWAS to calculate p-values ​​using Wald tests. Using the Leave One Chromosome Out (LOCO) method, each chromosome was tested using a kinship matrix calculated using other chromosomes to avoid proximal contamination. Initial results showed a broad peak on chromosome 7 around the Tyr gene, a well-known coat color locus, across most phenotypes. To control for this phenomenon, genotype at SNP rs32105080 was used as a covariate when running GEMMA. Sex was also used as a covariate. To assess SNP heritability, GEMMA was used without the LOCO method. The kinship matrix was assessed using the GEMMA LMM output of all SNPs in the genome and the proportion of explained phenotypic variance. PVE and PVESE were used as chip heritability and their standard errors.

[0072] To determine linkage intervals, LD decay was calculated. 100 snp pairs were selected within a 2.5 MB region and correlation coefficients were calculated. These correlations were binned to 100,000 bp and set at a threshold r of 0.2. 2To determine the peak region for each phenotype GWAS, SNPs were sorted according to their p-values. Then, for each SNP, a peak region centered around this SNP was determined by adding other SNPs with high correlation (r2 > 0.2) to the peak SNP. Peaks were limited to 10 million base pairs or less from the first selected peak SNP. These regions were used to find neighboring SNPs and non-correlated SNPs in the genome. Peak SNPs from all phenotypes were aggregated, and the peaks were clustered into clusters using the k-means algorithm implemented in R using the p-values ​​from all phenotype GWAS results. After observing the results, seven clusters were selected.

[0073] To combine the 24 phenotypes tested, phenotypes were taken from the same group and all phenotypes were taken, and for each SNP, the minimum p-value from the phenotypes within the group was taken.

[0074] Significant peaks from each phenotype and aggregated peak areas from all phenotypes assigned to the same cluster were tested for Gene Ontology (GO) enrichment using INRICH [Lee et al., 2012 Bioinformatics. 2012 04;28(13):1797-1799]. The interval used for gene enrichment was the peak area described above for each peak SNP. GO annotations associated with each gene were retrieved from EnsEBML via the Biomart and biomaRt R interfaces [Durinck et al., 2009 Nature Protocols. 2009;4:1184-1191].

[0075] The GWAS run is wrapped in an R package called mousegwas, available at github: / / github.com / TheJacksonLaboratory / mousegwas, which includes a singularity container definition file and a nextflow pipeline for reproducing the results.

[0076] Example 1 Grooming Mouse grooming Behaviors vary widely in both temporal and spatial scales, from subtle spatial movements such as whisking, blinking, or trembling to large spatial movements such as turning or walking, and vary in duration from milliseconds to minutes. The goal of these studies was to develop a classifier that generalizes to the complex behaviors observed in mice. We decided to classify grooming because it is conserved across species and is a highly interesting and neurobiologically important behavior [Kalueff et al., Neurobiology of Grooming Behavior. Cambridge University Press; 2010]. Grooming behaviors consist of a variety of constructs, from small or microscopic movements (paw licking) to medium-sized movements (unilateral and bilateral face washing) and large movements (flank licking) (see Figure 8A). Rare constructs, such as genital and tail grooming, are also present. Grooming durations can vary from subseconds to minutes. We reasoned that a successful approach to classifying grooming behaviors would be important to the neurobiological community and could serve as a prototype for other behaviors.

[0077] Grooming annotation

[0078] Our approach to grooming annotation was to classify all frames in a video as being in one of two states: either the mouse was grooming or not grooming. We specified that a frame should be annotated as grooming if the mouse was performing any of the grooming constructs, regardless of whether the mouse was performing a grooming syntactic sequence. This explicitly included individual paw licks as grooming, even though they do not constitute grooming bouts. Scratching was not a grooming construct. This included a wide variety of postures and action durations, contributing to diverse visual appearances. Variability in human annotations was investigated by having five trained annotators label the same six 5-minute videos (totaling 30 minutes, Figure 8C). To assist the human raters, raters were provided with three (3) videos of the mouse from the top and side (Figure 8B). Each annotator was given the same instructions for labeling behaviors (see methods described above in this specification). Strong agreement between annotators (average 89.1%) was observed. Closer examination of these disagreements between annotators revealed that misclassifications fell into three classes: missed bouts, skipped interruptions, and misalignments (Figures 8A–D and 9A–C). Missed bouts were called when disagreements occurred within the consensus of non-grooming calls. Similarly, skipped interruptions were called when disagreements occurred within the consensus of grooming calls. Finally, misalignments were called when both annotators agreed that grooming had either started or ended, but did not agree on the exact frame at which this occurred.

[0079] The most frequent type of error was misalignment, accounting for 50% of the total annotated duration of mismatched frames and 149.75% of mismatched calls (Figures 8A-D and 9A-C). The 89% agreement observed was consistent with previous work annotating mouse grooming behavior [Kyzar et al., 2011 Behavioral Brain Research. 225(2):426-431]. From this, we constructed a large annotation dataset to train machine learning algorithms. While most machine learning competitions attempting to solve tasks similar to those described herein have varied widely in dataset size, the work described here leverages network performance in these competitions for dataset design. Networks in these competitions performed well when each class contained at least 10,000 annotated frames [Girdhar et al., 2019 Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; p. 244-253]. When the number of annotations within a class exceeded 100,000, network performance on this task achieved MAP scores above 0.7 [Girdhar et al., 2019 Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; pp. 244-253; Zhang et al., 2019 arXiv preprint arXiv:190412993]. In deep learning approaches, model performance can benefit from additional annotations [Sun et al., 2017 Proceedings of the IEEE International Conference on Computer Vision; pp. 843-852].

[0080] To ensure success, the study described herein specified annotating over 2 million frames for either grooming or non-grooming. The goal was to balance this dataset for grooming behavior by selecting video segments based on a tracking heuristic, prioritizing segments at slower speeds, since mice cannot groom while walking. Additionally, video frames were cropped to center on the mouse and a tracker was used to reduce visual clutter [Guther et al., 2019 Communications Biology. 2019;2(1):124]. This mouse-centered cropping followed a video tube approach, as described by Feichtenhofer et al., 2019 Proceedings of the IEEE International Conference on Computer Vision, pp. 6202–6211. From a pool of seven validated annotators, two annotations with 94.3% agreement between annotations were obtained for 1,253 video segments totaling 2,637,363 frames (Figure 10A).

[0081] Example 2 Proposed Neural Network Solution A neural network classifier was trained using a large annotated dataset. Of the 1,253 video segments, 153 video clips were submitted for validation. Using this partition, we were able to achieve a similar distribution of frame-level classifications between the training and validation sets (Figure 10A). The machine learning approach described herein took video input data and generated an ethogram output for grooming behavior (Figure 10B). Functionally, the neural network model took 16 112 × 112 frame inputs and applied multiple layers of 3D convolution, 3D pooling, and fully connected layers to generate a prediction for the last frame only (Figure 10C). To predict the ethogram for the completed video, the process involved sliding a 16-frame window across the video. The neural network approach disclosed herein was compared to a previously established machine learning approach, JAABA, for annotating laboratory animal behavior [Kabra et al., 2013 Nature Methods. 2013;10(1):64].

[0082] Our neural network approach achieved the best validation performance, both in terms of accuracy and true positive rate (TPR), at a false positive rate (FPR) of 5%. The neural network achieved 93.7% accuracy and 91.9% TPR at a 5% FPR (Figure 11A). In comparison, the JAABA-trained classifier achieved poorer performance, with 84.9% accuracy and 64.2% TPR at a 5% FPR (Figure 11A). Due to memory limitations in JAABA, it was possible to train using only 20% of the training set. Additionally, the neural network was trained using the same training set used by JAABA, and the neural network still outperformed JAABA (Figure 11B). Using training datasets of different sizes, we observed improvements in validation performance as dataset size increased (Figure 12). Finally, the authors of JAABA also recommend an interactive training protocol. Even poorer performance was observed using the interactive training protocol. This is likely due to the significant difference in size of the annotated datasets used in training (475,000 187 frames vs. 17,000 frames). Our neural network approach performed comparably to human annotators, given our previous observations of 89% agreement in Figures 8B-C.

[0083] Receiver operating characteristic (ROC) curve performance was examined for each video, revealing that performance was not uniform across all videos (Figure 13A-E). The majority of validation videos were properly annotated by both the neural network and JAABA. However, two videos showed poor performance for both algorithms, and seven videos showed significant improvement using our neural network on the JAABA-trained classifier. Visual inspection of the two videos for which both algorithms performed poorly suggested that these particular video clips did not provide sufficient visual information to annotate grooming. During development of the final neural network solution, two forms of consensus modality were applied to improve single-model performance (Figure 11C).

[0084] Each trained model made slightly different predictions due to random initialization. By training multiple models and merging the predictions, a slight improvement in validation performance was achieved. Additionally, the input image was also altered for different predictions. The rotation and reflection of the input image were visually distinct for the neural network. By training four models and applying eight rotation and reflection transformations to the input, 32 distinct predictions were achieved for each frame. These individual predictions were merged by averaging the probability predictions. This consensus modality improved the ROC area under the curve (AUC) from 0.975 to 0.978.

[0085] Other approaches to merging the 32 predictions were attempted, including selecting the maximum value or applying voting (median prediction). Averaging the predicted probabilities achieved the best performance (Figure 14A). Finally, a time-smoothing filter was applied to the predictions of 46 frames. 46 frames were identified as the optimal window for the rolling average (Figure 14B), resulting in a final accuracy of 93.7% (ROC AUC 0.984). Because the network could only predict half a second's worth of information, we investigated the prediction of extreme grooming bouts in a large, non-human-annotated phylogenetic dataset. While most predictions of long bouts (>2 min) were real, there were several false positives in which mice were resting in a grooming-like posture. To mitigate these rare false positives, we implemented a heuristic to adjust the predictions. Our results showed that grooming behavior typically results in an ellipse-fit shape change (W / L) of 2.5 x 10. -4 When the mouse was resting, the standard deviation of the shape change (W / L) was 2 x 10 -5 Knowing that resting mouse postures can be visually similar to grooming postures, the prediction is that the standard deviation of shape change (W / L) over a 31-frame window should not exceed 5 × 10 -5Time segments less than 1 / 2 second were assigned to the "non-grooming" prediction. Of all frames with this difficult-to-annotate pose, 12% were classified as grooming. This suggests that this is not a case of network failure, but rather a limitation of the network when given only half a second's worth of information to make a prediction.

[0086] This approach was determined to be able to handle varying mouse postures and physical appearances, such as hair color and weight. Good performance was observed across a wide range of postures and hair colors (Figure 11D-E). Even nude mice, which have a completely different appearance from other mice, achieved good performance. Visually, a small number of frame orientations and instances where the model made inaccurate predictions were observed. Nevertheless, the consensus classifier made correct predictions.

[0087] Example 3 Defining grooming behavior indicators A study was conducted and various grooming behavior indices were designed to describe both the amount of grooming and the grooming pattern. A single grooming bout was defined as continuous time spent grooming without an interruption of more than 3 seconds (see Kalueff et al., Neurobiology of Grooming Behavior, Cambridge University Press, 2010). Short interruptions (less than 10 seconds) were allowed, but locomotor activity was not tolerated for this merge of time segments spent grooming. Specifically, pauses occurred when the mouse's movement did not exceed twice its average body length. From this, grooming ethograms were obtained for each mouse (Figure 15A). Using the ethograms, total grooming duration was calculated by summing the total duration of grooming calls across all grooming bouts. When the number of bouts and total duration were known, the average bout duration was calculated by dividing the two. For measurement purposes, 5-, 20-, and 55-minute summaries of these measurements were calculated. 5 and 20 minutes were included because these are typical open-field assay durations. Various grooming pattern indices were calculated using the 1-minute binned data (Figure 15B). A linear slope was fitted to the results to discover the temporal patterning of grooming over the 55-minute assay (GrTimeSlope55min). The positive slope for total grooming duration inferred that the longer an individual mouse remained in the open field assay, the more time it spent grooming. The negative slope for total grooming duration inferred that mice spent more time grooming at the beginning of the open field assay than at the end. This is typically because mice choose to spend more time on other activities, such as sleeping, rather than grooming. The positive slope for the number of bouts inferred that the longer a mouse remained in the open field assay, the more grooming bouts it initiated.Using the 5-minute binned data, we designed additional metrics to describe grooming patterns by selecting the time period during which mice spent the most time grooming (GrPeakMidBin) and the duration of grooming within that time period (GrPeakVal). We also calculated the ratio between these values ​​(GrPeakSlope). Finally, by looking at the lineage-level average of grooming, we were able to identify how long a lineage remained at its peak grooming (GrPeakLength). We compared various open field measures, including both grooming behavior and classic open field measures (Figure 15C). These phenotypes were categorized into four groups. Grooming volume described how much an animal groomed. Grooming pattern indices described how an animal changed its grooming behavior over time. The open field anxiety measure was a traditional open field phenotype validated to measure anxiety. Open field activity was a traditional open field phenotype describing the animal's overall activity.

[0088] Example 4 Gender and environmental covariate analysis of grooming behavior Using this trained classifier, we conducted a study to determine whether gender and environmental factors influenced the expression of grooming behavior in the open field. This analysis was performed using data collected over a 29-month period for two mouse strains, C57BL / 6J and C57BL / 6NJ. These two strains are substrains of the same strain in 1951 and are two of the most widely used strains in mouse research [Bryant et al., 2018 In: Molecular-genetic and statistical technology for behavioral and neural research Elsevier; pp. 165-190]. C57BL / 6J is the mouse reference strain, and C57BL / 6NJ is used by the International Mouse Phenotyping Consortium (IMPC) to generate a large amount of phenotypic data [Brown S, & Moore M, 2012 Towards a Encyclopaedia of Mammalian Gene Function: the International Mouse Phenotyping Consortium]. We analyzed 775 C57BL / 6J (317F, 267 458M) and 563 C57BL / 6NJ (240F, 323M) mice under a wide range of experimental conditions over two and a half years. Across all of these novel exposures to the open field, their grooming behavior was quantified for the first 30 minutes (Figure 16A-H, 669 hours of data total). Data were analyzed for the effects of sex, season, time of day, age, mouse origin, light level, tester, and white noise. To achieve this, stepwise linear model selection was applied to model these covariates. Both forward and backward model selection results were consistent. After identifying significant covariates, a second round of model selection, including a sex interaction term, was applied. Model selection identified sex, strain, origin, time of day, and season as significant, whereas age, weight, presence of white noise, and tester were not significant under the described test conditions. Furthermore, interactions between sex and both origin and season were identified as significant covariates. The study results are shown in Table 1. [Table 1]

[0089] The results showed a strain effect (Figure 16A, p=0.0268 C57BL / 6J vs. C57BL / 6NJ). Although the effect size was small, C57BL / 6NJ groomed more than C57BL / 6J. Furthermore, a sex difference was observed (Figure 16A, p<2.2x10 -16 Males vs. females). Males groomed more than females in both strains. Because sex had a strong effect, the interaction term was included with other covariates in a second pass of model selection. The model identified season as a significant covariate (Figure 16B, p = 0.004). Surprisingly, the model also identified an interaction between sex and season (p = 0.024). Female mice of both strains showed increased grooming during the summer and decreased grooming during the winter. Males did not show this trend, visually confirming the sex-season interaction. Tests were performed between 8:00 AM and 4:00 PM. To determine whether the time of day the tests were performed affected grooming behavior, the data were divided into two groups: morning (8:00 AM - 12:00 PM) and afternoon (12:00 PM - 4:00 PM). A clear effect of time of day was observed (Figure 16C, p = 0.00015). Mice tested in the morning groomed more overall. Mice of different ages were tested, ranging from 6 to 26 weeks of age. At the start of each test, mice were weighed and found to have weights ranging from 16 g to 42 g. No significant effects of age (Figure 16D, r = -0.065, p = 0.119) or weight (r = 0.206, p = 0.289) on grooming duration were observed.

[0090] Grooming levels of mice from internally shipped production were compared to an assay room (B2B) with mice born and raised in a room adjacent to the assay room. Six production rooms exclusively supplied C57BL / 6J (AX4, AX29, AX1, MP23, MP14, MP15), three rooms exclusively supplied C57BL / 6NJ (MP13, MP16, AX5), and one room supplied both strains (AX8). All shipped mice were housed in B2B at least one week prior to assay. A significant effect was observed based on room of origin (Figure 16E, p = 5.357 × 10 -13). For example, C57BL / 6J males from AX4 and AX29 groomed less than males in other rooms, including B2B. Shipped C57BL / 6NJ males in all rooms appeared to groom at lower levels compared to B2B. It was concluded that the room of origin and shipping location significantly influenced grooming behavior. Two light levels were also tested: 350–450 lux and 500–600 lux of white light (5600K). Results showed a significant effect of light level on grooming behavior (Figure 16F, p = 0.04873). Females of both strains groomed more in the darker areas, whereas males appeared unaffected. Despite this, the model did not include a light-sex interaction, suggesting that other covariates better accounted for visual interactions with sex here. Open field assays were performed by two male testers, with the majority of tests performed by tester 2. Both testers carefully followed the testing protocol, which was designed to minimize tester variability. No significant effects were observed between testers (Figure 16G, p = 0.65718). Finally, white noise was frequently added to previous open-field assays to generate uniform background noise levels and mask experimenter-generated noise [Gould, T. Mood and Anxiety Related Phenotypes in Mice, 2009 Neuromethods 42. DOI 10.1007 / 978-1-60761-303-9_1, Humana Press]. Although the effects of white noise have not been extensively studied in mice, existing data suggest that higher levels of white noise increase locomotion [Weyers P, et al., 1994 Behavioral Processes 31(2-3):257-267]. The effect of white noise (70 db) on grooming behavior in C57BL / 6J and C57BL / 6NJ mice was tested using the methods of the present invention, and a test was performed that identified significant differences in the duration of time spent grooming (Figure 16H).Although stratification appeared to exist in both C57BL / 6J 316 and C57BL / 6NJ females, other cofactors may better explain this result. Together, these results indicate that environmental factors such as season, time of day, and the mice's home room can influence grooming behavior and act as environmental confounds in any grooming study. Age, weight, light level, tester, and white noise were also investigated, and it was determined that these cofactors did not influence grooming behavior under these experimental conditions.

[0091] Example 5 Differences in grooming behavior Next, we used the grooming classifier to investigate grooming behavior in inbred mouse strains. Tested animals included 43 standard strains, 8 wild-type strains, and 11 diallel F1 hybrid mice from Jackson Laboratory mouse production. These were tested over a 31-month period and mostly consisted of single-mouse shipments from Jackson Laboratory production. In addition to C57BL / 6J and C57BL6 / NJ, an average of 8 males and 8 females from each strain were tested, with the animals averaging 11 weeks of age. Each mouse was tested for 55 minutes in an open field as previously described [Geuther BQ, et al., 2019 Communications Biology. 2(1):124]. The dataset consisted of 2,457 animals and 2,252 hours of video. Video data was classified into grooming behavior and measures of open-field activity and anxiety. Behavioral measures were extracted as described in Figure 15. To visualize phenotypic variance, each animal was plotted across all strains with the corresponding strain mean and one standard deviation range and ethograms of selected strains (Figure 17A-F). The study included distinguishing between classical laboratory strains and wild-type inbred strains.

[0092] Amount and pattern of grooming in genetically diverse mice A large, continuous variance in total grooming time, mean grooming bout length, and number of grooming bouts was observed in the 55-minute open-field assay (Figure 17A-F). Total grooming time varied from 2-3 min in strains such as 129X1 / SvJ and BALB / cByJ to 12 min in strains such as SJL / J and PWD / PhJ, representing an approximately six-fold difference in grooming time. Strains such as 129X1 / SvJ and C57BR / cdJ had fewer than 10 bouts, whereas MA / MyJ had nearly 40 bouts. Furthermore, bout duration varied from 5 seconds to approximately 50 seconds in BALB / cByJ and PWD / PhJ, respectively. To visualize the relationship between the phenotypic lineage means and 1 SD range correlation plots (Figure 19A-C) were created. There was a positive correlation between total grooming time and bout number, as well as between total grooming time and mean bout time. Overall, strains with higher total grooming time had increased bout numbers and longer bout durations. However, there appeared to be no relationship between bout number and average bout duration, suggesting that bout length was constant regardless of how many bouts occurred (Figure 19A-C). In general, for classically inbred strains, C57BL / 6J and C57BL6 / NJ fell roughly in the middle. The study included examining patterns of grooming over time by constructing rates of change in 5-minute bins for each strain (Figure 20). Using k-means clustering, three clusters of grooming patterns were defined based on the rate of increase in grooming over time, total grooming level, time of peak grooming, and duration of peak grooming (Figure 21).

[0093] Type 1 consisted of 13 strains with an inverted-U grooming pattern. These strains exhibited a rapid increase in grooming in the open field, peaked, and then began to decrease in grooming, typically resulting in an overall negative grooming gradient. Animals from these strains were often found asleep by the end of the 55-minute open-field assay. These strains included highly groomed strains such as CZECHII / EiJ, MORF / EiJ, and less groomed strains such as 129X1 / SvJ and I / LnJ.

[0094] Type 2 consisted of 12 heavily groomed lines that did not decrease grooming by the end of the assay. Type 2 reached peak grooming early and remained at this level for the duration of the assay (e.g., PWD / PhJ, SJL / J, and BTBR). Others in this group reached peak grooming later and remained stable (e.g., DBA / 2J, CBA / J). The defining feature of this group was that high levels of grooming were maintained throughout the assay.

[0095] Type 3 comprised most strains (30) and showed a steady increase in grooming until the end of the assay. Overall, these were the moderate-to-low grooming strains in this group, with a consistent low positive or flat slope. We conclude that under these experimental conditions, there are at least three broad, continuous, but observable classes of grooming patterns in mice.

[0096] Example 6 Grooming patterns of wild-type versus classical strains We compared grooming patterns between classical and wild-type laboratory strains. Classical laboratory strains are derived from a limited genetic stock originating from Japanese and European mouse breeders [Keeler CE. Laboratory Mouse: Its Origin, Heredity, and Culture; Cambridge, Harvard University Press, 1931; Morse HC. Origins of Inbred Mice: Proceedings of a Workshop, Bethesda, Maryland, February 14-16, Acad. Press, 1978; and Silver, LM. Mouse Genetics: Concepts and Applications. Oxford University Press, 1995]. Classical inbred laboratory mouse strains represent 95% of the genome of Mus musculus domesticus (Mm domesticus) and only 5% of the genome of Mus musculus [Yang et al., 2011 Nature Genetics. 2011;43(7):648]. To overcome the limited genetic diversity of classical inbred strains, we established new wild-type inbred strains [Guenet JL, & Bonhomme F. 2003 Trends in Genetics 19(1):24-31 and Koide T, et al., Experimental Animals. 2011;60(4):347-354]. Surprisingly, these results showed that most wild-type strains groomed at significantly higher levels and had longer average bout lengths than classical inbred strains. Five of the top 16 grooming strains were wild-type (PWD / PhJ, WSB / EiJ, CZECHII / EiJ, MSM / MsJ, and MORF / EiJ) (Figure 17A). Furthermore, wild-type strains had significantly longer grooming bouts, and six of the 16 strains from this group had the longest average grooming bouts. Both total grooming time and average bout length were significantly different between the classical and wild-type strains (Figures 18A-B).These high-grooming strains represent the Mm domesticus and Mm musculus subspecies and were precursors of the classical laboratory strains [Yang et al., 2011 Nature Genetics. 2011;43(7):648]. These wild-type strains also represented a much larger proportion of the natural genetic diversity in mouse populations than the number of classical strains tested. This led us to conclude that the high levels of grooming observed in the wild-type strains represent normal levels of grooming behavior in mice. This suggested that classical laboratory strains, at least as observed in these experimental conditions, may select against low grooming behavior.

[0097] BTBR Grooming Pattern Experiments were also conducted to closely examine the grooming patterns of the BTBR strain, which has been proposed as a model for certain characteristics of autism spectrum disorder (ASD), a complex neurodevelopmental disorder that leads to impaired communication, repetitive behaviors, and social interaction [Association AP, et al. Diagnostic and Statistical Manual of Mental Disorders (DSM-5®). American Psychiatric Pub; 2013]. Compared to C57BL / 6J mice, BTBR mice have been shown to have higher levels of repetitive behaviors, hyposociability, abnormal vocalizations, and behavioral maladjustment [McFarlane HG, et al., 2008 Genes, Brain and Behavior 7(2):152-163; Silverman JL, et al., 2010 Neuropsychopharmacology 35(4):976-989; Moy SS, et al., 2007 Behavioral Brain Research 176(1):4-20; Scattoni ML, et al., 2008 PloS One 3(8)]. Repetitive behaviors are often assessed by self-grooming behaviors, and medications effective in reducing symptoms of repetitive behaviors in ASD have been determined to also reduce grooming in the BTBR without affecting overall activity levels, providing some level of construct validity [Silverman JL, et al., 2010 Science Translational Medicine 4(131):131ra51 and Amodeo, DA, et al., 2017 Gene, Brain and Behavior 16(3):342-351].

[0098] The results of the study described herein identified that total grooming time in BTBR was higher than that of C57BL / 6J, but not exceptionally high compared to all strains (Figures 17A-F). C57BL / 6J groomed for approximately 5 minutes over a 55-minute open-field session, whereas BTBR groomed for approximately 12 minutes (Figure 17A). Several classical inbred strains, such as SJL / J, DBA / 1J, and CBA / CaJ, exhibited similarly high grooming activity. The grooming pattern of BTBR belonged to type 2, including five other strains (Figure 20). One distinguishing factor for BTBR was that they had longer average grooming bouts early in the open field (Figures 18A-B). However, again, they were not exceptionally high in terms of average bout length (Figures 18A-B). Strains such as SJL / J, PWD / PhJ, MORF / EiJ, and NZB / BINJ exhibited similar, long-term bouts early on. BTBR showed high levels of grooming with long grooming bouts, but this behavior was similar to several wild-type and classical inbred laboratory strains and was not considered exceptional. Because social interactions and other features of ASD were not measured, the results did not argue against BTBR as a model for ASD.

[0099] Example 7 Grooming for mouse GWAS This study investigated the genetic architecture underlying complex mouse grooming and open-field behaviors and related them to human traits. A genome-wide association study (GWAS) was conducted using data from 51 classical inbred strains and 11 diallel F1 hybrid strains. Eight wild-type strains were not included because they are highly divergent and may distort mouse GWAS analyses. Twenty-four phenotypes were classified into four categories: (1) open-field activity, (2) anxiety, (3) grooming pattern, and (4) quantity (Figure 15A-C). A linear mixed model (LMM) implemented in Genome-wide EZcient Mixed Model Association (GEMMA) was used for this analysis [Zhou X. & Stephens M. 2012 Nature Genetics 44(7):821-824]. First, the heritability of each phenotype was calculated by determining the proportion of phenotypic variance explained by the typed genotype (PVE) (Figure 22A). Heritability ranged from 6% to 68%, with 22 / 24 traits representing heritability estimates greater than 20%, a reasonable estimate for behavioral traits in mice and humans [Valdar W, et al., 2006 Genetics 174(2):959-984 and Bouchard Jr TJ. 2004 Current Directions in Psychological Science 13(4):148-151], making them suitable for GWAS analysis (Figure 22A).

[0100] Each phenotype was analyzed using GEMMA, taking into account the resulting Wald test p-value. To correct for the multiple (222,966) SNPs tested and to account for correlations between SNP genotypes, an empirical threshold p-value was obtained by shuffling the values ​​of one normally distributed phenotype (OFDistTraveled20m) and taking the minimum p-value for each permutation. This process yielded a p-value of 1.4 x 10, which reflects a corrected p-value of 0.05. -5The p-value threshold was set at r [Belmonte M. & Yurglun-Todd D. 2001 IEEE Transactions on Medical Imaging 20(3):243-248]. To avoid calling multiple correlated neighboring SNPs, correlated SNPs were clustered under the same peak. 2 A correlation coefficient of >=0.2 was chosen, which resulted in a large peak area but seemed a reasonable compromise between capturing LD blocks and avoiding overly expanded blocks (Figure 22B).

[0101] The GWAS analysis yielded 22 peaks that exceeded the permutation threshold p-value (Figure 23). Overall, open field activity had 15 significant peaks, while anxiety had 10, grooming pattern had 76, and grooming amount had 51 peaks, leading to a combined total of 130 peaks across all tested phenotypes (Figure 9C). Pleiotropy was observed at the same locus significantly associated with multiple phenotypes. Pleiotropy was expected because many phenotypes are correlated and individual traits may be controlled by similar genetic structures. For example, grooming times of 55 and 20 minutes (GrTime55 and GrTime20) are correlated traits, so pleiotropy was expected. We also predicted that some loci that regulate open field activity phenotypes may also regulate grooming.

[0102] To better understand the pleiotropic structure of the GWAS results, we created a heatmap of significant SNPs across all phenotypes. These were then clustered to identify sets of SNPs that regulated groups of phenotypes (Figure 22D). Phenotypes were clustered into five subgroups consisting of grooming pattern (I), open field activity (II), open field anxiety (III), grooming length (IV), and grooming number and amount (V) (top x-axis in Figure 22D). Seven clusters of SNPs that regulated combinations of these phenotypes were identified (y-axis in Figure 22D). For example, clusters A and G were composed of pleiotropic SNPs that regulated grooming length (IV) and grooming time, but SNPs in cluster G also regulated bout number and amount (V). SNP cluster D regulated open field activity and anxiety phenotypes. Cluster E contained SNPs that regulated grooming, open-field activity, and anxiety phenotypes; however, most SNPs only had significant p-values ​​for either the open-field or grooming phenotype, but not both, indicating that independent genetic architecture contributes significantly to these phenotypes. The associated regions within the GWAS (Figure 22C) were shaded to mark one of seven SNP clusters (Figure 22D). These clusters ranged from 13 to 35 SNPs; the smallest, Cluster F, was nearly pleiotropic for grooming activity, and the largest, Cluster G, was pleiotropic for most grooming-related phenotypes. To prioritize genes, the associated genes were ordered based on their degree of pleiotropy.

[0103] These highly pleiotropic genes included several known to regulate grooming, striatal function, neurodevelopment, and even language. Mammalian Phenotype Ontology enrichment showed a significant correlation for 178 genes (p = 7.5 × 10 -4 ) and subsequently showed "nervous system development" as the most important module with preweaning lethality (p=3.5×10 -3, 189 genes) and abnormal embryonic development (p=5.5×10 -3, 62 genes) were shown (see Figure 24A-G). Pathway analysis was performed using the KEGG and Reactome databases using pathwAX [Ogris, C. et al., 2016 Nucleic Acids Res Jul 8;44(W1):W105-9]. This analysis revealed 14 enriched disease pathways, including Parkinson's disease (9.68E-09), Huntington's disease (1.07E-06), nonalcoholic fatty liver disease (9.31E-06), and Alzheimer's disease (1.15E-05) as the most significantly enriched. Enriched pathways included oxidative phosphorylation (6.42E-08), ribosome (0.00000102), RNA transport (0.00000315), and ribosome biogenesis (0.00000465). Reactome-enriched pathways included mitochondrial translation termination and elongation (2.50E-19 and 5.89E-19, respectively) and ubiquitin-specific processing proteases (1.86E-08). The most highly pleiotropic gene was Sox5, which was associated with 11 grooming and open field phenotypes. Sox5 is widely associated with neuronal differentiation, patterning, and stem cell maintenance [Lefebvre V. 2010 The International Journal of Biochemistry & Cell Biology 42(3):429-432]. Its dysregulation in humans has been implicated in Ram-Schaffer syndrome and ASD, both neurodevelopmental disorders [Kwan KY. In: International Review of Neurobiology, vol. 113 Elsevier; 2013. p. 167-205 and Zawerton A, et al., 2020 Genetics in Medicine 22(3):524-537]. 102 genes were associated with 10 phenotypes and 105 genes were associated with 9 phenotypes. The analysis was restricted to genes with at least six significantly associated phenotypes, resulting in 860 genes.Other genes included FoxP1, which is associated with striatal function and language regulation [Bowers, JM & Konopka, G., 2012 Disease Markers 33(5):251-60]. Grin2b, a regulator of glutamate signaling, and Ctnnb1, a regulator of Wnt signaling, were identified. Combined, this analysis implicated genes known to regulate nervous system function and development, as well as genes known to regulate neurodegenerative diseases, as regulators of grooming and open-field behavior. GWAS analysis also defined the genetic architecture of grooming and open-field behavior in mice.

[0104] Example 8 PheWAS Other studies have been conducted to link 860 genes associated with mouse open field and grooming phenotypes (see Example 7) with human phenotypes. It was hypothesized that mice and humans share a common underlying genetic and neurological structure, but that different phenotypes arise in each organism. Thus, for example, dysregulation of a specific pathway in mice may lead to an excessive grooming phenotype, while in humans, perturbation of the same pathway may manifest as neuroticism or obsessive-compulsive disorder. The relationship between phenotypes in these organisms can be revealed through the identification of a common underlying genetic structure.

[0105] To link mouse grooming gene circuits to human phenotypes, we conducted PheWAS using the Psychiatric Genetics GWAS catalog. The first step involved identifying human orthologs of 860 mouse grooming and open field genes with at least six degrees of pleiotropy. For each human ortholog, PheWAS summary statistics were downloaded from gwasATLAS (http: / / atlas.ctglab.nl / ) [Watanabe K, et al., 2019 Nature Genetics 51(9):1339-1348]. Currently, gwasATLAS contains 4756 GWASs from 473 unique studies across 3302 unique traits, categorized into 28 domains. Studies focused on associations with gene-level p-values ​​≤ 0.001 in the psychiatric domain. To further visualize and cluster these associations, the relationships between genes and psychiatric traits were represented by a weighted reciprocal network, in which the width of the edge between the gene node and the psychiatric trait node was proportional to the strength of the association [-log10(p-value)]. The size of the node was proportional to the number of associated genes or traits, and the shading of the trait node corresponded to the subchapter level of the psychiatric domain. To identify modules within this network, we applied an improved community detection algorithm that maximizes weighted modularity in the weighted reciprocal network [Dormann C. F. & Strauss R. 2014 Methods in Ecology and Evolution 5(1):90-98]. This network was used as an input to detect communities, creating a ranked list of communities based on their modularity scores, with the top communities representing promising candidates for further study.

[0106] [Newman ME. & Girvan M. 2004 Physical Review E 69(2):026113]. This analysis yielded eight gene-phenotype modules (Figure 25). These modules contained 15 to 32 individual phenotypes and 41 to 103 genes. At the subchapter level, modules were enriched for temperament and personality phenotypes, mental and behavioral disorders (schizophrenia, bipolar disorder, dementia), addictions (alcohol, tobacco, cannabinoids), obsessive-compulsive disorder, anxiety, and sleep.

[0107] Surprisingly, the results identified the same genes that showed high levels of pleiotropy in the mouse GWAS and the resulting human PheWAS. FOXP1 was the most pleiotropic gene, with 35 associations, followed by SOX5 with 33 associations. This network was used as input to detect communities to generate a ranked list of communities based on their modularity scores, with the top communities representing promising candidates for further study. The modularity scores of the eight modules ranged from 0.028 to 0.083, with module 1 ranking at the top with a modularity score of 0.103. Furthermore, the Simes test was used to combine the p-values ​​of the genes to obtain an overall p-value for each psychiatric trait association. The median association value [-log10(Simes p-value)] was then calculated for each detected community for prioritization. Similarly, module 1 ranked at the top of the eight modules (median LOD = 5.29). Module 1 primarily consisted of temperament and personality phenotypes, including neuroticism, mood swings, and irritability traits. Genes in this module exhibited high levels of pleiotropy in both human PheWAS and mouse GWAS. These genes included SOX5, which was associated with 33 human phenotypes, RANGAP1, GRIN2B, and others. This module contained eight of the 10 most pleiotropic genes from the PheWAS analysis. Genes in this module included SOX5, the second most pleiotropic gene in the PheWAS study, with 33 significant associations; RANGAP1 with 31 significant associations; and EP300 with 23 significant associations. In conclusion, the PheWAS analysis linked genes that upregulate grooming and other open-field behaviors to human phenotypes. These human phenotypes include personality traits, addiction, and schizophrenia. Furthermore, similar genes were found to be highly pleiotropic in both mouse GWAS and human PheWAS analyses.

[0108] Summary and Discussion of Examples 1-8 Grooming is an ethologically conserved, neurobiologically significant behavior of interest to the behavioral genetics community. Often used as an endophenotype for several psychiatric disorders and a prime example of stereotyped pattern behavior, the ability to automatically quantify grooming behavior is a necessary tool [Spruijt BM et al., 1992 Physiology Reviews 72(3):825-852; Kalueff AV et al., Neurobiology of Grooming Behavior. Cambridge University Press; 2010; and Kalueff AV et al., 2016 Nature Reviews Neuroscience 17(1):45]. Furthermore, the ability to detect grooming behavior, with its highly variable posture and duration, serves as a prototype for other behaviors. The work described herein presents a neural network approach toward automatic model biobehavioral classification and ethogram generation.

[0109] This approach was implemented on mouse grooming behavior, a complex behavior that poses challenges to existing automated systems. The system and method described herein achieved human-level performance. This grooming behavior classifier was used to analyze large datasets. The resulting data demonstrated the stability of grooming as a behavioral indicator by performing covariate analysis. Mouse GWAS and human PheWAS studies have also been conducted to understand the genetic architecture underlying grooming and open-field behavior in laboratory mice and link them to human traits.

[0110] While the machine learning community has implemented a wide variety of solutions for human action detection, few applications have been applied to animal behavior. This may be due to a variety of reasons, including the widespread availability of human action datasets and the stringent performance requirements for biobehavioral research. It has been observed that the cost of achieving this stringent performance is prohibitive, requiring extensive annotation. In many cases, previous experimental paradigms were short or small enough that the cost could be reduced to simply annotating the data without automation.

[0111] Other machine learning approaches have been applied to this automated annotation of behavioral data. A 3D convolutional neural network was observed to outperform the JAABA classifier when trained on the same training dataset. This improvement was not uniform across all samples, but instead localized to specific types of grooming bouts. This suggests that while the JAABA classifier is powerful and has utility for smaller, more uniform datasets, experiments and behaviors with diverse representations require more powerful machine learning approaches. The grooming classifier was used to determine the genetic and environmental factors regulating this behavior. A study was conducted on a large dataset with two reference strains, C57BL / 6J and C67BL / 6N, collected over an 18-month period to evaluate the influence of several factors that varied over time in the dataset, including sex, strain, age, time of day, season, tester, room of origin, white noise, and weight. All mice were housed under identical conditions for at least one week prior to testing. Strong effects of sex, testing time, and even season were observed.

[0112] Tester effects have been widely observed in both mice and rats in previous open-field studies [Walsh RN. & Cummins RA. 1976 Psychological Bulletin 83(3):482; McCall RB. et al., 1969 Developmental Psychology 1(6p1):771; and Bohlen M. et al., 2014 Behavioral Brain Research 272:46-54]. Recent studies have shown that male experimenters, or even male clothing, elicit stress responses and increased thigmotaxis in mice [Sorge RE. et al., 2014 Nature Methods 11(6):629]. However, experimenter effects are not always observed across open-field studies [Lewejohann L. et al., 2006 Genes, Brain, and Behavior 5(1):64-72]. The results of the study described herein demonstrate that the room of origin had a strong influence on grooming. We compared the grooming of mice shipped from the Jackson Laboratory's production room with that of mice bred in a room adjacent to the phenotyping room. Shipped mice were observed to exhibit variable total grooming duration compared to mice that did not experience shipping. There was no clear direction of effect, and in some cases the effect size was high (Z > 1). Mice of the same strain shipped from different rooms exhibited either higher or lower total grooming. We hypothesized that the change in grooming was due to stress, which has previously been demonstrated to alter this behavior [Kalueff AV. et al., 2016 Nature Reviews Neuroscience 17(1):45]. Presumably, all external mice had similar experiences of shipping from the production room to the testing area where they were equally housed for at least one week prior to testing. Therefore, potential differential stress experiences may be due to the room of origin where the mice were born and held until shipping. This raises a caveat to using this behavior as an endophenotype.

[0113] A large-scale lineage study was conducted to characterize grooming behavior in laboratory mice. Three grooming patterns were identified under the test conditions. Type 1 consisted of mice that increased and decreased grooming over the course of a 55-minute open-field test. Strains in this group were often asleep by the end of the assay, demonstrating a low level of alertness toward the end of the assay. We hypothesized that these strains used grooming as a successful form of non-arousal behavior, previously observed in rats, birds, and monkeys. [Spruijt BM et al., 1992 Physiology Reviews 72(3):825-852 and Delius JD, 1970 Psychologische Forschung 33(2):165-188. doi:10.1007 / BF00424983]. Similar to types 1 and 2, this group rapidly increased grooming and reached peak grooming, but did not appear to decrease grooming during the assay. These strains were hypothesized to require longer periods of non-arousal or to have some kind of deficit under the assay conditions. The Type 3 strain, which grew for the duration of the assay, demonstrated that it had not yet reached peak grooming under the assay conditions. BTBR is a member of the Type 2 group with early and prolonged high levels of grooming, possibly indicating hyperarousal or an inability to non-arousal. BTBR has previously been shown to have high arousal and altered dopamine function, which may lead to persistently high levels of grooming. It is speculated that other strains in the Type 2 grooming class may also exhibit endophenotypic characteristics of ASD.

[0114] The study results showed that wild-type strains had distinct grooming patterns compared with classical strains. Wild-type strains groomed significantly more and had longer grooming bouts than classical strains. Grooming clustering analysis revealed that most wild-type strains belonged to types 1 or 2, whereas most classical strains belonged to type 3. In addition to Mm domesticus, the wild-type inbred strains tested represented Mm musculus, Mm castaneous, and Mm molossinus subspecies. Despite the existence of dozens of classical inbred strains, approximately 5 million SNPs exist between classical inbred laboratory strains such as C57BL / 6J and DBA2J [Keane TM. et al., 2011 Nature. 477(7364):289-294]. In fact, over 97% of the genomes of classical strains can be explained by fewer than 10 haplotypes, representing a small number of classes in which all strains are identical by descent to a common ancestor [Yang H. et al., 2011 Nature Genetics 43(7):648]. In contrast, wild-type inbred strains such as CAST / EiJ and 599 PWK / PhJ have over 17 million SNPs compared to B6J, and WSB / EiJ has 8 million SNPs. Thus, the seven wild-type strains tested represent much more genetic diversity in natural mouse populations than many classical inbred laboratory strains. The behavior observed in the wild-type strains is likely representative of the behavior of natural mouse populations.

[0115] Classic laboratory strains originated in mouse breeders in China, Japan, and Europe before being co-selected for biomedical research [Morse HC. 1978 Proceedings of a Workshop, Bethesda, Maryland, February 14-16, 1978. Acad. Press; 1978 and Silver LM. Mouse Genetics: Concepts and Applications. Oxford University Press, 1995]. As a result, despite the existence of hundreds of classical strains, genetic variance within these strains is limited [Yang H. et al., 2011 Nature Genetics 43(7):648]. To overcome these limitations of classical strains, wild-type strains have been specifically developed [Poltorak A. et al., 2018 Mammalian Genome 29(7-8):577-584]. Mouse breeders breed mice for visual and behavioral characteristics, often for display in animal shows. Mouse breeders judge mice on their "condition and temperament" and suggest that "it is no use to make them appear rude in their coats, or in any other condition than the perfect condition of the mice" [Davies C. Fancy Mice: Their Varieties and Management as Pets or for Show, Including the Latest Scientific Information as to Breeding for Colour, L.U. Gill; 1912]. Much like dogs and horses, "the best individuals, regardless of relationship, should be bred to each other, so long as they remain large, robust, and in a disease-free condition" [Davies C. Fancy Mice: Their Varieties and Management as Pets or for Show, Including the Latest Scientific Information as to Breeding for Colour, L.U. Gill; 1912].It is reasonable to assume that the normal level of grooming behavior observed in wild mice is indicative of poor hygiene or parasites such as lice, tics, fleas, and mites. High grooming may be interpreted as poor condition, leading mouse breeders to select mouse strains with low grooming behavior. This selection may explain the low grooming observed in classical strains.

[0116] Using pedigree survey data, we conducted a mouse GWAS to identify xx gene loci that regulate heritable variation in open-field and grooming behavior. Results indicated that the majority of grooming traits are moderately to highly heritable. In the study exemplified herein, we performed a detailed analysis of 862 pleiotropic genes, revealing genes and pathways known to regulate neuronal development and function. These associated regions were identified as belonging to one of seven clusters that regulate the combined open-field and grooming phenotype. A previous study using the BXD recombinant inbred strain panel identified a significant locus on chromosome 4 that regulates grooming and open-field activity [Delprato A. et al., 2017 Genes, Brain and Behavior 16(8):790-799]. Furthermore, as described herein, we performed a PheWAS using these genes to identify psychiatric traits associated with these genes. This approach enabled us to link mouse and human phenotypes through the underlying genetic structure. This approach linked human temperament and personality traits, schizophrenia, and bipolar disorder traits to mouse open field and grooming phenotypes. Grooming can be used as a model for human grooming disorders, such as trichotillomania. However, grooming is regulated by the basal ganglia and other brain regions, including ASD, schizophrenia, and Parkinson's disease [Kalueff AV. et al., 2016 Nature Reviews Neuroscience 17(1):45]. The findings herein link grooming to temperament and personality traits, schizophrenia, and bipolar disorder, among others. GWAS results have advanced our understanding of the genetic architecture of grooming behavior.

[0117] In conclusion, the experiments and studies described herein demonstrate a neural network-based machine learning approach for behavior detection in mice and its application to grooming behavior. This tool has been and will continue to be used to characterize grooming behavior and its underlying genetic architecture in laboratory mice and other mammals. This approach to grooming can be implemented using a standard open-field apparatus and should be useful to the behavioral neuroscience community.

[0118] Device and System Examples One or more of the trained models of system 150 can take many forms, including a neural network. A neural network can include several layers, from an input layer to an output layer. Each layer is configured to take a particular type of data as input and output another type of data. The output from one layer is received as input to the next layer. While the values ​​of the input / output data for a particular layer are not known until the neural network actually operates at runtime, data describing the neural network describes the structure, parameters, and operation of the layers of the neural network.

[0119] One or more of the intermediate layers of a neural network may also be known as a hidden layer. Each node in the hidden layer is connected to each node in the input layer and each node in the output layer. If a neural network includes multiple intermediate networks, each node in the hidden layer connects to each node in the next higher and next lower layer. Each node in the input layer represents a potential input to the neural network, and each node in the output layer represents a potential output of the neural network. Each connection from one node to another node in the next layer may be associated with a weight or score. A neural network may output a single output or a weighted set of predicted outputs.

[0120] In one embodiment, the neural network may be a fully connected convolutional neural network (CNN), which has a normalized version of a multilayer perceptron. In a fully connected network, each neuron in one layer is connected to every neuron in the next layer. A typical method of normalization involves adding some form of magnitude measure of weight to the loss function. CNNs take a different approach to normalization, exploiting hierarchical patterns in the data and using smaller, simpler patterns to assemble more complex patterns.

[0121] In one aspect, a neural network may be constructed with recurrent connections such that the outputs of a hidden layer of the network feed back into the hidden layer for the next set of inputs. Each node in the input layer connects to each node in the hidden layer. Each hidden node connects to each node in the output layer. The outputs of the hidden layer are fed back into the hidden layer for processing the next set of inputs. A neural network incorporating recurrent connections may be referred to as a recurrent neural network (RNN).

[0122] In some embodiments, the neural network may be a long short-term memory (LSTM) network. In some embodiments, the LSTM may be a bidirectional LSTM. A bidirectional LSTM operates on inputs from two time directions: one from a past state to a future state and one from a future state to a past state, where the past state may correspond to features of the video data for a first time frame and the future state may correspond to features of the video data for a second, subsequent time frame.

[0123] Processing by a neural network is determined by the learned weights on each node input and the structure of the network: given a particular input, the neural network determines its output one layer at a time until the output layer of the entire network has been calculated.

[0124] Connection weights can be initially learned by a neural network during training, where a given input is associated with a known output. In a set of training data, various training examples are fed to the network. Each example typically sets the weight of the correct connection from input to output to 1 and gives all connections a weight of 0. Once examples of the training data have been processed by the neural network, inputs can be sent to the network and compared with their associated outputs to determine how the network performance compares to target performance. Training techniques such as backpropagation may be used to update the neural network's weights to reduce errors made by the neural network when processing the training data.

[0125] Various machine learning techniques may be used to train and operate models for performing various steps described herein, such as user-recognition feature extraction, encoding, user-recognition scoring, and user-recognition confidence determination. Models may be trained and operated according to various machine learning techniques. These techniques may include, for example, neural networks (such as deep neural networks and / or recurrent neural networks), inference engines, trained classifiers, and the like. Examples of trained classifiers include support vector machines (SVMs), neural networks, decision trees, AdaBoost (short for adaptive boosting) combined with decision trees, and random forests. Focusing on SVMs as an example, SVMs are supervised learning models with associated learning algorithms that analyze data, recognize patterns within the data, and are commonly used in classification and regression analysis. Given a set of training examples, each marked as belonging to one of two categories, the SVM training algorithm builds a model that assigns new examples to one category or the other, making it a non-probabilistic binary linear classifier. More complex SVM models may be built with a training set that identifies three or more categories, and the SVM determines the category that is most similar to the input data. The SVM model may map examples of distinct categories so that they are separated by a clear gap. New examples are then mapped into that same space and predicted to belong to a category based on which side of the gap they fall on. The classifier may issue a "score" indicating the category to which the data most closely matches. The score may indicate how closely the data matches the category.

[0126] To apply machine learning techniques, the machine learning process itself must be trained. In this case, to train a machine learning component, such as one of the first or second models, a "ground truth" of training examples must be established. In machine learning, the term "ground truth" refers to the classification accuracy of a training set for supervised learning techniques. Various techniques can be used to train a model, including backpropagation, statistical learning, supervised learning, semi-supervised learning, probabilistic learning, or other known techniques.

[0127] FIG. 26 is a block diagram conceptually illustrating a device 2600 that may be used with the system. FIG. 27 is a block diagram conceptually illustrating exemplary components of a remote device, such as system 150, that may assist in processing video data, detecting subject activity, and the like. System 150 may include one or more servers. As used herein, server may refer to a traditional server, as understood in a server / client computing architecture, but may also refer to several different computing components that may assist in the operations discussed herein. For example, a server may include one or more physical computing components (such as a rack server) that are connected physically and / or via a network to other devices / components and can perform computing operations. A server may also include one or more virtual machines that emulate a computer system and run on one or across multiple devices. A server may also include other combinations of hardware, software, firmware, and the like for performing the operations discussed herein. A server may be configured to operate using one or more of a client-server model, a computer bureau model, grid computing technology, fog computing technology, mainframe technology, utility computing technology, a peer-to-peer model, sandbox technology, or other computing technologies.

[0128] Multiple systems 150 may be included in the overall system of the present disclosure, such as one or more systems 150 for determining different orientations for frames, one or more systems 150 for running a first trained model to process different sets of frames, one or more systems 150 for running a second trained model to process different sets of frames, one or more systems 150 for aggregating results of the different trained models, and one or more systems 150 for training / constructing the different trained models. In operation, each of these systems may include computer-readable and computer-executable instructions residing on the respective device 150, as discussed further below.

[0129] Each of these devices (2600 / 150) may include one or more controller / processors (2604 / 2704), each of which may include a central processing unit (CPU) for processing data and computer-readable instructions, and memory (2606 / 2706) for storing the respective device's data and instructions. The memory (2606 / 2706) may individually include volatile random access memory (RAM), nonvolatile read-only memory (ROM), nonvolatile magnetoresistive memory (MRAM), and / or other types of memory. Each device (2600 / 150) may also include a data storage component (2608 / 2708) for storing data and controller / processor-executable instructions. Each data storage component (2608 / 2708) may individually include one or more nonvolatile storage types, such as magnetic storage, optical storage, solid-state storage, etc. Each device (2600 / 150) may also be connected to removable or external non-volatile memory and / or storage (removable memory cards, memory key drives, network storage, etc.) via a respective input / output device interface (2602 / 2702).

[0130] The computer instructions for operating each device (2600 / 150) and its various components may be executed by the respective device's controller / processor (2604 / 2704), using the memory (2606 / 2706) as temporary "working" storage during execution. The device's computer instructions may be stored in a non-transitory manner in non-volatile memory (2606 / 2706), storage device (2608 / 2708), or an external device. Alternatively, some or all of the executable instructions may be embedded in the respective device's hardware or firmware in addition to or instead of software.

[0131] Each device (2600 / 150) includes an input / output device interface (2602 / 2702). Various components may be connected through the input / output device interfaces (2602 / 2702), as discussed further below. Additionally, each device (2600 / 150) may include an address / data bus (2624 / 2724) for transmitting data between the components of the respective device. Each component within a device (2600 / 150) may also be directly connected to other components in addition to (or instead of) being connected to other components across a bus (2624 / 2724).

[0132] 26, device 2600 may include an input / output device interface 2602 that connects to various components, such as an audio output component, such as a speaker 2612, a wired or wireless headset (not shown), or other component capable of outputting sound. Device 2600 may further include a display 2616 for displaying content. Device 2600 may further include a camera 2618.

[0133] Via antenna 2614, input / output device interface 2602 may connect to one or more networks 199 via wireless network radio, such as a wireless local area network (WLAN) (e.g., WiFi) radio, Bluetooth, and / or a radio capable of communicating with wireless communication networks, such as a Long Term Evolution (LTE) network, a WiMAX network, a 3G network, a 4G network, a 5G network, etc. Wired connections, such as Ethernet, may also be supported. Via network 199, the system may be distributed across a network environment. I / O device interface (2602 / 2702) may also include communication components that enable data to be exchanged between devices, such as different physical servers, in a collection of servers or other components.

[0134] The components of device 2600 or system 150 may include their own dedicated processors, memory, and / or storage devices, or one or more of the components of device 2600 or system 150 may utilize the I / O interfaces (2602 / 2702), processors (2604 / 2704), memory (2606 / 2706), and / or storage devices (2608 / 2708) of device 2600 or system 150, respectively.

[0135] As noted above, multiple devices may be employed in a single system. In such a multi-device system, each of the devices may include different components for performing different aspects of the system's processing. Multiple devices may include overlapping components. The components of device 2600 and system 150 described herein are exemplary and may reside as standalone devices or may be included in whole or in part as components of a larger device or system.

[0136] The concepts disclosed herein may be applied within a number of different devices and computer systems, including, for example, general-purpose computing systems, video / image processing systems, and distributed computing environments.

[0137] The above-described aspects of the present disclosure are meant to be illustrative. They have been selected to illustrate the principles and applications of the present disclosure and are not intended to be exhaustive or to limit the present disclosure. Many modifications and variations of the aspects of the present disclosure will be apparent to those skilled in the art. Those skilled in the art of computers and speech processing should recognize that the components and process steps described herein may be interchangeable with other components or steps, or combinations of components or steps, and still achieve the benefits and advantages of the present disclosure.

[0138] Moreover, it should be apparent to one skilled in the art that the present disclosure may be practiced without some or all of the specific details and steps disclosed herein.

[0139] Aspects of the disclosed system may be implemented as a computer method or as an article of manufacture, such as a memory device or non-transitory computer-readable storage medium. The computer-readable storage medium may be readable by a computer and may contain instructions for causing a computer or other device to perform the processes described in this disclosure. The computer-readable storage medium may be implemented by volatile computer memory, non-volatile computer memory, a hard drive, a solid-state memory, a flash drive, a removable disk, and / or other media. Furthermore, components of the system may be implemented as firmware or hardware.

[0140] equivalent While several embodiments of the present invention have been described and illustrated herein, those skilled in the art will readily envision various other means and / or structures for performing the functions and / or obtaining the results and / or one or more advantages described herein, and each of such variations and / or modifications is deemed to be within the scope of the present invention. More broadly, those skilled in the art will readily appreciate that all parameters, dimensions, materials, and configurations described herein are exemplary, and that the actual parameters, dimensions, materials, and / or configurations will depend on the specific application or applications for which the teachings of the present invention are used. Those skilled in the art will recognize, or be able to ascertain using no more than routine experimentation, many equivalents to the specific embodiments of the invention described herein. Accordingly, it is understood that the foregoing embodiments are offered by way of example only, and that within the scope of the appended claims and their equivalents, the invention may be practiced otherwise than as specifically described and claimed. The present invention is directed to each individual feature, system, article, material, and / or method described herein. Furthermore, any combination of two or more such features, systems, articles, materials, and / or methods is included within the scope of the present invention, provided that such features, systems, articles, materials, and / or methods are not mutually inconsistent. All definitions defined and used herein should be understood to govern dictionary definitions, definitions in documents incorporated by reference, and / or ordinary meanings of the defined terms.

[0141] As used herein in the specification and claims, the indefinite articles "a" and "an" should be understood to mean "at least one" unless clearly indicated to the contrary. As used herein in the specification and claims, the word "and / or" should be understood to mean "either or both" of the elements so conjoined, i.e., elements conjunctively present in some cases and contingently present in other cases. Except where expressly indicated otherwise, other elements, whether related or unrelated to the elements specifically identified, may optionally be present other than the elements specifically identified by the "and / or" clause.

[0142] Conditional language used herein, particularly "can," "could," "might," "may," and the like, is generally intended to convey that certain embodiments include certain features, elements, and / or steps, while other embodiments do not, unless specifically stated otherwise or understood otherwise within the context of use. Thus, such conditional language is not generally intended to imply that features, elements, and / or steps are in any way required by one or more embodiments, or that one or more embodiments necessarily include logic for determining whether those features, elements, and / or steps are included in or performed in any particular embodiment, with or without other input or prompt. Terms such as "comprising," "including," and "having" are synonymous and used in an inclusive, open-ended manner and do not exclude additional elements, features, acts, operations, etc. Also, the term "or" is used in its inclusive sense (and not its exclusive sense), so that, for example, to connect a list of elements, the term "or" means one, some, or all of the elements in the list.

[0143] All references, patents, and patent applications and publications cited or referred to in this application are hereby incorporated by reference in their entirety.

Claims

1. receiving video data representing a video capturing movement of an object among a plurality of objects, the video being video of a period including a plurality of time durations, each frame of the video data corresponding to a respective one of the plurality of time durations; identifying a first set of frames from the video data; determining a set of rotated frames by rotating the first set of frames; processing the first set of frames using a first trained model configured to identify a likelihood of the one subject exhibiting a predetermined behavioral action; determining a first probability of the one subject exhibiting the predetermined behavioral action in a first frame of the first set of frames, the first frame corresponding to a time duration of the plurality of time durations, based on the processing of the first set of frames by the first trained model; processing the set of rotated frames using the first trained model; and determining a second probability of the one subject exhibiting the predetermined behavioral action in a second frame of the set of rotated frames, the second frame corresponding to the time duration of the first frame, based on the processing of the set of rotated frames by the first trained model; and identifying a first label for the first frame using the first probability and the second probability, the first label indicating that the one object exhibits the predetermined behavioral action.

2. processing the first set of frames using a second trained model configured to identify a likelihood of the one subject exhibiting the predetermined behavioral action; determining a third probability of the one subject exhibiting the predetermined behavioral action in the first frames based on the processing of the first set of frames by the second trained model; processing the set of rotated frames using the second trained model; and determining a fourth probability of the one subject exhibiting the predetermined behavioral action in the second frame based on the processing of the rotated set of frames by the second trained model; and The computer-implemented method of claim 1 , further comprising: identifying the first label using the first probability, the second probability, the third probability, and the fourth probability.

3. determining a set of reflected frames by reflecting the first set of frames; processing the set of reflected frames using the first trained model; and determining a third probability of the one subject exhibiting the predetermined behavioral action in a third frame of the set of reflected frames, the third frame corresponding to the first frame, based on the processing of the set of reflected frames by the first trained model; The computer-implemented method of claim 1 , further comprising identifying the first label using the first probability, the second probability, and the third probability.

4. 4. The method of claim 1, wherein the predetermined behavioral action comprises a grooming behavior, the grooming behavior comprising at least one of paw licking, one-sided face washing, both-sided face washing, and flank licking.

5. The computer-implemented method of claim 1 , wherein the first set of frames represents a portion of the video data for a period of time, and the first frame is a last time frame of the period.

6. identifying a second set of frames from the video data; determining a second rotated set of frames by rotating the second set of frames; processing the second set of frames using the first trained model; and determining a third probability of the one subject exhibiting the predetermined behavioral action in a third frame of the second set of frames based on the processing of the second set of frames by the first trained model; and processing the second set of rotated frames using the first trained model; and determining a fourth probability of the one subject exhibiting the predetermined behavioral action in a fourth frame of the set of rotated frames, the fourth frame corresponding to the third frame, based on the processing of the second set of rotated frames by the first trained model; 2. The computer-implemented method of claim 1, further comprising: using the third probability and the fourth probability to identify a second label for the fourth frame, the second label indicating that the one object exhibits the predetermined behavioral action.

7. 7. The computer-implemented method of claim 6, further comprising generating an ethogram representing the predetermined behavioral action of the one subject over a period of time using at least the first label and the second label.

8. before receiving the video data, receiving training data including a first plurality of video frames and a second plurality of video frames, each of the first plurality of video frames associated with a positive label indicating that the one subject is exhibiting the predetermined behavioral action, and each of the second plurality of video frames associated with a negative label indicating that the one subject is exhibiting a behavioral action that is not the predetermined behavioral action; 10. The computer-implemented method of claim 1, further comprising: processing the training data using a first set of model parameters and first classifier model data to determine the first trained model.

9. 9. The computer-implemented method of claim 8, wherein the first plurality of video frames and the second plurality of video frames represent the movement of a plurality of objects, and wherein the one object in the plurality of objects includes one or more pre-identified physical characteristics.

10. 10. The computer-implemented method of claim 9, wherein the pre-identified physical characteristics are one or more of body shape, body size, coat color, sex, age, and disease or disorder phenotype.

11. The computer-implemented method of claim 10 , wherein the disease or disorder is a genetic disease, injury, or infectious disease.

12. 9. The computer-implemented method of claim 8, wherein the first plurality of video frames and the second plurality of video frames represent movements of the plurality of objects that are mice, and wherein one mouse in the plurality of objects that are mice has a coat color, sex, body shape, and size.

Citation Information

Patent Citations

  • Long-term and continuous animal behavioral monitoring

    CN111225558A

  • Method and system for automated behavior classification of test subjects

    US20180225516A1