System and method for analyzing the motion of an object to separate it into sub-second modules
By using 3D depth cameras and AR-HMM models, the subjectivity and short-term recognition problems of animal behavior analysis in the prior art are solved, and the automation of animal behavior and multi-time scale structural characterization are realized, and the accuracy and reliability of the analysis are improved.
Patent Information
- Application Number
- CN202211082452.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2016-03-18
- Filing Date
- 2017-03-16
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2037-03-16
AI Technical Summary
Existing animal behavior analysis methods rely on subjective assessment of human observers, making it difficult to objectively identify and classify animal behavior, especially behavior modules within short time scales, and existing systems cannot effectively characterize the multi-time scale organizational structure of animal behavior.
Animal video images were obtained by using 3D depth cameras, and animal behavior modules were identified through Bayesian inference and AR-HMM model. Combined with principal component analysis and neural network dimensionality reduction, we automatically identify and classify animal behaviors, and identify subsecond-level behavior modules and transfer structures.
The objective and automated identification and classification of animal behavior are realized, the sub-second organizational structure of animal behavior is revealed, and the accuracy and repeatability of behavior analysis are improved.
Smart Images

Figure CN115500818B_ABST
Abstract
Description
[0001] This application is a divisional application of the patent application with application number 201780031046.1, filed on March 16, 2017, and with the invention name “Automatic Classification of Animal Behavior”. Technical Field
[0002] The present invention relates to systems and methods for identifying and classifying animal behavior, human behavior, or other behavioral metrics. Background Art
[0003] The following description includes information that may be helpful in understanding the present invention. However, it is not an admission that any of the information provided herein is prior art or relevant to the presently claimed invention, or that any publication specifically or implicitly cited is prior art.
[0004] Quantifying animal behavior is an essential first step in research ranging from drug discovery to understanding the biology of neurodegenerative diseases. It is typically done manually; trained observers observe the animal's behavior in person or on videotape and record all moments of interest.
[0005] Behavioral data from a single experiment may involve hundreds of mice, spanning hundreds of hours of video, requiring a team of observers, which inevitably reduces the reliability and reproducibility of the results. Moreover, the question of what constitutes a "behavior of interest" is largely left to human observers: while it is trivial for human observers to assign anthropomorphic names to specific behaviors or series of behaviors (i.e., "rearing," "sniffing," "sniffing," "walking," "freezing," "eating," etc.), there are almost certainly behavioral states generated by mice that are relevant to mice that defy simple human categorization.
[0006] In more advanced applications, videos can be analyzed semi-automatically by computer programs. However, the brain produces behavior that unfolds smoothly over time but is composed of distinct motor patterns. Individual sensory neurons that trigger actions can complete behavior-related computations in as little as a millisecond, and the neural populations that mediate behavior exhibit dynamics that evolve on timescales of tens to hundreds of milliseconds. [1-8] This rapid neural activity interacts with slower neuromodulatory systems to produce behaviors that are organized simultaneously at multiple timescales. [9]Ultimately understanding how neural circuits generate complex behaviors—especially the spontaneous or innate behaviors expressed by freely behaving animals—requires a clear and well-defined framework for characterizing how behavior is organized at timescales relevant to the nervous system. Summary of the Invention
[0007] While evolution has shaped behaviors that enable animals to achieve specific goals (such as finding food or a mate), it is unclear how these behaviors are organized over time, especially on rapid timescales. However, a powerful approach to characterizing behavioral structure comes from ethnology, which proposes that the brain builds coherent behaviors by expressing stereotyped modules of simpler actions in specific sequences.
[10] For example, both supervised and unsupervised classification methods have identified underlying behavioral modules expressed during exploration in C. elegans, D. melanogaster larvae, and D. melanogaster adults. [11-16] These experiments have revealed an underlying structure for behavior in these organisms, which in turn reveals strategies used by invertebrate brains to adapt behavior to environmental changes. In the case of C. elegans, navigation to olfactory cues is mediated at least in part by neural circuits that serve to adjust transition probabilities connecting behavioral modules into time-dependent sequences; thus, the worm nervous system can generate seemingly novel sensory-driven behaviors (such as positive chemotaxis) by reordering a core set of behavioral modules. [17-19] Similar observations have been made for sensory-driven behaviors in fly larvae.
[11] .
[0008] These insights into the temporal structure underlying behavior come from the ability to quantify morphological changes in worms and flies and use this data to identify behavioral modules. [11-16] However, gaining a similar understanding of the overall behavioral organization of mammals has been difficult. While researchers have divided the innate exploration, grooming, social patterns, aggression, and reproductive behaviors of mice into potential modules, this approach of breaking down mammalian behavior into components relies on human-defined definitions of what constitutes a meaningful behavioral module (e.g., running, mating, fighting). [20-25] ,Thus, this approach is largely limited by human perception and intuition.,In particular, human perception has difficulty identifying modules that only span,short timescales.
[0009] To systematically describe the structure of animal behavior and understand how the brain changes this structure to achieve adaptation, three key problems need to be solved. First, when trying to modularize mouse behavior, it is not clear which behavioral features are important to measure. Although most current methods track two-dimensional parameters such as the position, speed, or shape of the mouse's up-down or left-right profile, [20,22-24,26-28] , but mice exhibit complex three-dimensional postural dynamics that are difficult to capture but may provide important insights into the organization of behavior. Second, if behavior evolves in parallel at several timescales, it is unclear how to objectively identify the relevant spatiotemporal scales at which behavior is modularized. Finally, effective representations of behavior need to accommodate the fact that behavior is both stereotyped (a prerequisite for modularity) and variable (an unavoidable feature of noisy neural and motor systems).
[29] .
[0010] This variability presents a significant challenge for algorithms tasked with identifying the number and content of behavioral modules expressed in a given experiment, or assigning any given instance of observed action to a specific behavioral module. Furthermore, identifying the spatiotemporal scales at which natural behavior is organized has been a clear challenge in ethology, and thus, most efforts to date to explore the underlying structure of behavior have relied on ad hoc definitions of what constitutes a behavioral module and have focused on specific behaviors rather than systematically considering behavior as a whole. It is currently unclear whether the spontaneous behavior exhibited by animals has a definable underlying structure that can be used to characterize action as it evolves over time.
[0011] Furthermore, existing computerized systems for animal behavior classification match the parameters used to describe the observed behavior to a manually annotated and curated database of parameters. Consequently, in both manual and existing semi-automated scenarios, subjective assessments of the animal's behavioral state are built into the system—human observers must determine in advance what constitutes a particular behavior. This leads to biased assessments of the behavior and limits the assessment to specific behaviors that researchers can distinguish with human perception, making it limited, especially for behaviors that occur over short timescales. Furthermore, the video acquisition systems deployed in these semi-supervised forms of behavioral analysis (which almost always acquire data in two dimensions) are optimized only for specific behaviors, thereby limiting throughput and increasing wasted experimental effort due to alignment errors.
[0012] summary
[0013] Despite these challenges, the inventors have discovered systems and methods for automatically identifying and classifying animal behavior modules by processing video recordings of animals. In accordance with the principles of the present invention, monitoring methods and systems use hardware and custom software capable of classifying animal behavior. The classification of animal behavioral states is determined by quantitatively measuring the animal's posture in three dimensions using a depth camera. In one embodiment, a 3D depth camera is used to acquire an animal video image stream having area information and depth information. The background image (empty test area) is then removed from each of the multiple images to generate a processed image having bright and dark areas. The outlines of the bright areas in the multiple processed images are found, and parameters are extracted from the area image information and depth image information within these outlines to form multiple multidimensional data points, each data point representing the posture of the animal at a specific time. The posture data points can then be clustered so that the point clusters represent the animal behavior.
[0014] This data can then be fed into a model-free algorithm or into a computational model to characterize the structure of natural behavior. In some embodiments, these systems fit behavioral models using methods in Bayesian inference that allow for unsupervised identification of the optimal number and identity of behavioral modules from a given data set. Behavioral modules are defined based on the structure of the 3D behavioral data itself (rather than using a priori definitions of what should constitute measurable action units), thereby identifying previously unexplored sub-second laws that define the time scale used when organizing behavior, generating key information about the composition and structure of behavior, providing an understanding of the nature of behavioral changes, and enabling people to objectively discover subtle changes in patterned actions.
[0015] Example application of a video of mice exploring the wilderness
[0016] In one example, the inventors measured how the body shape of a mouse changed as it freely explored a circular open field. The inventors used a depth sensor to capture the three-dimensional (3D) postural dynamics of the mouse, and then quantified how the mouse's posture changed over time by centering and aligning the mouse's images along the inferred axis of its spine.
[0017] By plotting these 3D data over time, they revealed that mouse behavior is characterized by periods of slowly evolving postural dynamics, punctuated by rapid transitions that separate these periods; this pattern appears to break the behavioral imaging data into chunks consisting of a small number of frames, typically lasting 200 to 900 milliseconds. This suggests that mouse behavior can be organized on two distinct timescales: the first, defined by the rate at which the mouse's posture can change within a given chunk, and the second, defined by the rate of transitions between chunks.
[0018] To characterize mouse behavior within these blocks and determine how behavior might differ between blocks, it is first necessary to estimate the timescale over which these blocks are organized. In some embodiments, to identify approximate boundaries between blocks, behavioral imaging data are submitted to a changepoint algorithm designed to detect sudden changes in data structure over time. In one example, this method automatically identified potential boundaries between blocks and revealed that the average block duration was approximately 350ms.
[0019] In addition, the inventors performed autocorrelation and spectral analysis, which provides complementary information about the time scale of behavior. The temporal autocorrelation in mouse posture largely dissipated within 400ms (τ=340±58ms), and almost all of the frequency components that distinguished the behavioral mice from the dead mice were concentrated between 1Hz and 6Hz (measured by spectral ratio or Wiener filter, average value 3.75±0.56Hz); These results indicate that most of the dynamics of mouse behavior occur within the time scale of 200ms to 900ms.
[0020] Furthermore, visual inspection of the block-by-block patterns of behavior exhibited by mice revealed that each block appeared to encode brief behavioral motifs that were separated from subsequent behavioral motifs by rapid transitions (e.g., right or left turns, darts, pauses, rearing). Taken together, these findings reveal a previously unappreciated subsecond organization of mouse behavior—during normal exploration, mice express brief motor motifs that appear to switch rapidly from one to another in succession.
[0021] The discovery that behavior is naturally broken down into short motor themes suggests that each of these themes is a behavioral module—a stereotyped and reusable unit of behavior that the brain arranges into sequences to build more complex action patterns. Next, we disclose systems and methods for identifying multiple instances of the same stereotyped, sub-second behavioral theme.
[0022] Processing algorithms and methods for identifying modules in video data
[0023] To identify similar modules, the mouse behavioral data can first be subjected to dimensionality reduction using, for example, (1) principal component analysis and (2) neural networks (e.g., multi-layer perceptrons). For example, using principal component analysis (PCA), the first two principal components can be plotted. Each block in the postural dynamics data corresponds to a continuous trajectory through the PCA space; for example, the individual blocks associated with the mouse's spine in an elevated state correspond to a particular sweep in the PCA space. Using template matching methods to scan the behavioral data for matching subjects, several additional examples of this sweep were identified in different animals, indicating that each of these PCA trajectories can represent an individual instance in which a stereotyped behavioral module is reused.
[0024] Given this evidence of sub-second modularity, the inventors designed a series of computational models—each describing a different underlying structure of mouse behavior—trained these models on 3D behavioral imaging data, and determined which models predicted or identified the underlying structure of mouse behavior. In particular, the inventors employed computational inference methods optimized to automatically identify structure in large datasets, including Bayesian non-parametric approaches and Gibbs sampling.
[0025] Each model differs in whether it considers behavior to be continuous or modular, the possible content of the modules, and the transition structure that governs how modules are sequenced over time. To compare model performance, the models were tested to predict the content and structure of real mouse behavioral data, which the models had not exposed. Among the alternatives, the best quantitative predictions were achieved by a model that assumed that mouse behavior consists of individual modules (each capturing a brief episode of 3D body movement) that switch from one to another on the subsecond timescale identified by our model-free analysis of postural dynamics data.
[0026] AR-HMM model
[0027] One model represents each behavioral module as an AR (vector autoregressive) process, which is used to capture the stereotyped trajectory through the PCA space. Furthermore, within this model, an HMM (Hidden Markov Model) is used to represent the switching dynamics between different modules. Together, this model is referred to here as the "AR-HMM."
[0028] In some embodiments, AR-HMM makes predictions about mouse behavior based on its ability to discover (within the training data) a set of behavioral modules and transition patterns that provide the most concise explanation of the overall structure of mouse behavior as it evolves over time. Thus, a trained AR-HMM can be used to reveal the identity of behavioral modules and their transition structure from a behavioral dataset, thereby revealing the underlying organization of mouse behavior. After training, the AR-HMM can assign each frame of trained behavioral data to one of the modules it has discovered, thereby revealing when any given module was expressed by the mouse during a given experiment.
[0029] Consistent with AR-HMMs recognizing inherent block structure in 3D behavioral data, the module boundaries identified by the AR-HMM adhered to the inherent block structure embedded within the postural dynamics data. Furthermore, the distribution of module durations identified by the model was similar to that of block durations identified by change point analysis; however, the module boundaries identified by the AR-HMM improved upon the approximate boundaries proposed by change point analysis (78% of module boundaries were within 5 frames of the change point). Importantly, the ability of the AR-HMM to identify behavioral modules depended on the inherent sub-second organization of mouse postural data; shuffling frames constituting behavioral data in small blocks (i.e., less than 300 milliseconds) significantly degraded the model's performance, whereas shuffling behavioral data in larger blocks had little effect. These results indicate that the AR-HMM recognizes the inherent sub-second block structure of the behavioral data.
[0030] In addition, the specific behavioral modules identified by the AR-HMM encode a set of different, reused motion themes. For example, the PCA trajectories assigned to a behavioral module by the model track similar paths through the PCA space. Consistent with each of these trajectories encoding similar action themes, collating and inspecting 3D movies associated with multiple data instances of this specific module confirmed that it encodes a stereotypical theme of behavior that would be called upright by a human observer. In contrast, data instances drawn from different behavioral modules track different paths through the PCA space. Furthermore, visual inspection of the 3D movies assigned to each of these modules showed that each module encodes a repetitive and coherent three-dimensional motion pattern that can be distinguished and labeled with descriptors (e.g., "walk," "hesitate," and "low rear" modules).
[0031] To quantitatively and comprehensively assess the distinctiveness of each behavioral module identified by the AR-HMM, we performed a cross-likelihood analysis, which revealed that data instances associated with a given module were assigned to that module rather than to any other behavioral module in the parsed dataset. In contrast, the AR-HMM failed to identify any well-separated modules in a synthetic mouse behavioral dataset lacking modularity, demonstrating that the modularity found in the real behavioral data is a characteristic of the dataset itself, rather than an artifact of the model. Furthermore, restarting the model training process from a random starting point returned identical or highly similar groups of behavioral modules, consistent with the AR-HMM tracking and identifying the inherent modular structure of the behavioral data. Together, these data suggest that mouse behavior is fundamentally organized into distinct subsecond modules when viewed through the lens of the AR-HMM.
[0032] Furthermore, if AR-HMMs identify the behavioral modules and transitions that underlie mouse behavior, then synthetic behavioral data generated by trained AR-HMMs can provide a reasonable replica of real postural dynamics data. AR-HMMs appear to capture the richness of mouse behavior, as synthetic behavioral data (in the form of spinal dynamics, or 3D movies of behaving mice) are qualitatively indistinguishable from behavioral data generated by real animals. Thus, mouse postural dynamics data have an inherent structure that is organized at a subsecond timescale and well parsed by AR-HMMs into defined modules; furthermore, optimal identification of these modules and effective prediction of behavioral structure require explicit modeling of modularity and switching dynamics.
[0033] SLDS SVAE model
[0034] In order to reduce redundant dimensions and make the modeling computationally tractable, various techniques can be used to reduce the dimensionality of each image. One method of dimensionality reduction is principal component analysis, which reduces the dimensionality to a linear space. However, the inventors have found that simply reducing the dimensionality to a linear space will not be able to accommodate the various variations in mice that are unrelated to behavior. This includes variations in the size of the mice, the breed of the mice, etc.
[0035] Therefore, the inventors have discovered that by using certain types of neural networks, such as multilayer perceptrons, one can effectively reduce the dimensionality of images. Furthermore, these reduced-dimensional images provide an effective way to develop models that are independent of the size of the mouse or other animal and can account for other variations unrelated to behavior. For example, one can employ neural networks that reduce the dimensionality to a three-dimensional image manifold.
[0036] Therefore, based on these algorithms, the inventors developed SVAE generative models and corresponding variational family algorithms. As an example, the inventors focused on a specific generative model on time series based on switched linear dynamic systems (SLDS) (Murphy, 2012; Fox et al., 2011), which illustrates how SVAE can combine discrete latent variables with rich probabilistic dependencies and continuous latent variables.
[0037] The systems and methods of the present invention can be applied to a wide variety of animal species, such as animals in animal models, humans in clinical trials, and humans in need of diagnosis and / or treatment of a particular disease or disorder. These animals include, but are not limited to, mice, dogs, cats, cows, pigs, goats, sheep, rats, horses, guinea pigs, rabbits, reptiles, zebrafish, birds, fruit flies, worms, amphibians (e.g., frogs), chickens, non-human primates, and humans.
[0038] The systems and methods of the present invention can be used in a variety of applications, including but not limited to: drug screening; drug classification; genetic classification; disease research including early detection of disease onset; toxicology research; side effect research; learning and memory process research; anxiety research; and consumer behavior analysis.
[0039] The systems and methods of the present invention are particularly useful for disorders that affect a subject's behavior, including neurodegenerative diseases such as Parkinson's disease, Huntington's disease, Alzheimer's disease, and amyotrophic lateral sclerosis; and neurodevelopmental disorders such as attention deficit hyperactivity disorder, autism, Down syndrome, Mendelsohn syndrome, and schizophrenia.
[0040] In certain embodiments, the system and method of the present invention can be used to study how a known drug or test compound can change the behavioral state of an object. This can be achieved by comparing the behavioral representations obtained before and after a known drug or test compound is administered to an object. The term "behavioral representation" as used herein refers to a set of sub-second behavioral modules and their transfer statistics determined using the system or method of the present invention. In a non-limiting manner, behavioral representations can be in the form of a matrix, table, or heat map.
[0041] In some embodiments, the system and method of the present invention can be used for drug classification. The system and method of the present invention can create multiple reference behavior representations based on existing drugs and the diseases or disorders they treat, wherein each reference behavior representation represents a class of drugs (e.g., antipsychotics, antidepressants, stimulants or inhibitors). The test behavior representation can be compared with multiple reference behavior representations, and if the test behavior representation is similar to one of the multiple reference behavior representations, the test compound is determined to belong to the same class of drugs represented by the specific reference behavior representation. In a non-limiting manner, the test compound can be a small molecule, an antibody or its antigen-binding fragment, a nucleic acid, a polypeptide, a peptide, a peptidomimetic, a polysaccharide, a monosaccharide, a lipid, a glycosaminoglycan or a combination thereof.
[0042] In some embodiments, this can include a system for automatically classifying an animal's behavior as belonging to a class of drugs relative to a list of alternatives. For example, to develop the system, one can provide a training set of mice under many different drug conditions and construct a linear or nonlinear classifier to discover which combinations and ranges of features constitute membership in a particular drug class. Once training is complete, the classifier is immediately fixed, allowing one to apply it to previously unseen mice. Potential classifier algorithms can include logistic regression, support vector machines with linear basis kernels, support vector machines with radial basis function kernels, multilayer perceptrons, random forest classifiers, or k-nearest neighbor classifiers.
[0043] Similar to drug classification, in some embodiments, the systems and methods of the present invention can be used for gene function classification.
[0044] In certain embodiments of drug screening, a known existing drug for treating a specific disease or disorder can be administered to the first test subject. Then, the system and method of the present invention can be used for the first test subject to obtain a reference behavior representation, and the reference behavior representation includes a group of behavioral modules that can characterize the therapeutic effect of the drug on the first test subject. Subsequently, the test compound can be administered to a second test subject of the same animal type as the first test subject. Then, the system and method of the present invention can be used for the second test subject to obtain the test behavior representation. If the test behavior representation is similar to the reference behavior representation, the test compound is determined to be effective in treating a specific disease or disorder. If the test behavior representation is dissimilar to the reference behavior representation, the test compound is determined to be invalid in treating a specific disease or disorder. It should be noted that the first and second test subjects can each be a group of test subjects, and the behavior representation obtained can be an average behavior representation.
[0045] Similar to drug screening, in some embodiments, the systems and methods of the present invention can be used for gene therapy screening.Gene therapy can include nucleic acid delivery and gene knockout.
[0046] In certain embodiments, system and method of the present invention can be used for the research of disease or disorder.For example, system and method of the present invention can be used for finding the new behavior module of the object with specific disease or disorder.For example, system and method of the present invention can come early diagnosis disease or disorder by the reference behavior performance of the people suffering from disease or disorder or the object in disease or disorder process by identifying.If also observe reference behavior representation or its significant part in the object of suspected suffering from this disease or disorder, then this object is diagnosed as suffering from this disease or disorder.Therefore, object can be carried out early clinical intervention.
[0047] Furthermore, in some embodiments, the systems and methods of the present invention can be used to study consumer behavior, such as how consumers respond to scents (e.g., perfumes). The systems and methods of the present invention can be used to identify reference behavioral representations that represent a positive response to a scent. In the presence of a scent, a person who exhibits a reference behavioral representation, or a significant portion thereof, is determined to have a positive response to the scent. Reference behavioral representations that represent a negative response to a scent can also be identified and used to measure a person's response. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate embodiments of the present invention and, together with the description, serve to explain and illustrate the principles of the invention. These drawings are intended to illustrate the main features of the exemplary embodiments in a schematic manner. These drawings are not intended to show every feature of actual embodiments or to show relative sizes of elements, and are not drawn to scale.
[0049] Figure 1 Depicted is a diagram of a system designed for capturing video data of animals, according to various embodiments of the present invention.
[0050] Figure 2A Depicted are flow diagrams illustrating processing steps performed on video data according to various embodiments of the present invention.
[0051] Figure 2B Depicted are flow diagrams illustrating processing steps performed on video data according to various embodiments of the present invention.
[0052] Figure 3 Depicted are flow diagrams illustrating analysis performed on video data output from processing steps, according to various embodiments of the present invention.
[0053] Figure 4 Depicted is a flow chart illustrating the implementation of an AR-HMM algorithm according to various embodiments of the present invention.
[0054] 5A depicts a graph showing the proportion of frames explained by each module (Y-axis), sorted by usage (X-axis), plotted against module group, according to various embodiments of the present invention.
[0055] 5B depicts a graph showing modules (X-axis) sorted by usage (Y-axis) with Bayesian credible intervals represented therein, according to various embodiments of the present invention.
[0056] Figure 6A -6E depicts the impact of the physical environment on module usage and spatial expression patterns according to various embodiments of the present invention. Figure 6A : Modules identified by AR-HMM and sorted by usage (n=25 mice, 500 min total, data from circular open field). Figure 6B : Hinton plot of observed binary model probabilities, which depicts the probability that any pair of modules is observed as an ordered pair. Figure 6C: Module usage sorted by context. The dark line represents the average usage of the animals, and the light line represents the bootstrap estimate (n=100). The labeled modules discussed in the text and shown in Figure 6D: square = circular thigmotaxis, circle = rosette, diamond = square thigmotaxis, cross = square gallop. Figure 6D: The occupancy chart of mice in the circular open field (left, n=25, a total of 500 minutes) represents the average spatial position across all trials. The occupancy chart depicts the deployment of the circular thigmotaxis module (center, the average orientation of the area represented by the arrow in the experiment) and the circular enriched rosette module (right, the orientation of each animal represented by the arrow). Figure 6E: The occupancy chart of mice in the box (left, n=15, a total of 300 minutes) represents the cumulative spatial position across all experiments. Occupancy diagrams depict the square-enriched haptotaxis module (center, average orientation across experiments, indicated by arrowhead regions) and the square-flying module (right, orientation of individual animals indicated by arrowheads).
[0057] Figure 7 Depicted are histograms showing the average velocities of modules that are differentially upregulated and interconnected after TMT exposure "freezing" compared to all other modules in the dataset, according to various embodiments of the present invention.
[0058] Figures 8A-8E depict how odor avoidance changes transition probability according to various embodiments of the present invention. Figure 8A: Occupancy plots under control conditions (n=24, 480 minutes total) and exposed to the unimolecular fox-derived odorant trimethylthiazoline (TMT, 5% dilution in the vehicle DPG, n=15, 300 minutes total) in the lower left quadrant (arrow). Figure 8B: Module usage plot sorted by "TMT-ness". The dark line depicts the average usage, and the light line depicts the bootstrap estimate. Labeled modules discussed in this specification and Figure 8E: square = sniffing in the TMT quadrant, circle = freezing away from TMT. Figure 8C Left and middle: Behavioral state diagrams of mice transitioning to explore the box under control conditions (blank) and TMT exposure, where modules are depicted as nodes (with usage proportional to each node's diameter) and binary model transition probabilities are depicted as directed edges. The two-dimensional layout is designed to minimize the overall distance between all connected nodes and is seeded using spectral clustering to emphasize neighborhood structure. Figure 8C: State diagram depicting the difference between blank and TMT. Differences in usage are represented by resized colored circles (blue for upregulation, red for downregulation, and black for blank usage). Changed binary model probabilities are represented by the same color code. Figure 8D: Mountain plot of the joint probability of module expression and spatial location plotted with respect to the TMT corners (X-axis); note the "bump" two-thirds of the way through the graph appears because the two corners are equidistant from the odor source. Figure 8E: Occupancy map showing the spatial location of mice after TMT exposure when emitting a scout sniff module (left) or a hesitation module.
[0059] Figures 9A-9C Plots are shown of how the AR-HMM according to various embodiments of the present invention distinguish between wild-type, heterozygous, and homozygous mice. Figure 9A: Utilization plot of the modules exhibited by mice sorted by "mutant-property" (n=6+ / +, n=4+ / -, n=5- / -, open field assay, 20 minute test). The black line depicts the average utilization across animals, and the blurred line depicts the bootstrap estimate. Figure 9B: State diagram depicting the baseline OFA behavior of + / + animals as in Figure 9C; state diagrams of the differences between the + / + and + / - genotypes (middle) and between the + / + and - / - genotypes (right) as in Figure 9C. Figure 9C: Illustration of the "stagger" module, in which the animal's hind limbs are raised above the shoulder girdle and the animal moves forward with a staggering gait.
[0060] Figures 10A-10B Figure 10A shows how photostimulation perturbations of the motor cortex, according to various embodiments of the present invention, generate novel morphological and physiological modules. Figure 10A: Mountain curves depict the probability of expression of each behavioral module (each behavioral module is assigned a unique color on the Y-axis) as a function of time (X-axis), where two seconds of light stimulation begins at time zero (each curve is an average of 50 trials). Note that due to the experimental structure (in which mice are sequentially exposed to increasing light levels), moderate changes in baseline behavioral patterns are captured before light onset across conditions. Stars indicate two modules expressed during baseline conditions that were also upregulated at medium power (11 mW) but not high power (32 mW); crosses indicate the hesitation module that was upregulated at light termination. Figure 10B: Average position of example mice for the two modules induced under the highest stimulation condition (arrows indicate orientation over time). Note that these curves are taken from a single animal and represent the complete data set (n=4); due to variability in viral expression, the threshold power required to elicit behavioral changes varied between animals, but all expressed the spinning behavior identified in Figure 10A.
[0061] Figures 11A-11C illustrate how depth imaging, according to various embodiments of the present invention, reveals block structure in mouse postural dynamics data. Figure 11A plots the three-dimensional posture of a mouse captured in a circular field using a standard RGB camera (left) and a 3D depth camera (right, mouse height is plotted in color, mm = mm above the ground). Figure 11B depicts an arrow representing the inferred axis of the animal's spine; all mouse images are centered and aligned along this axis to enable quantitative measurement of postural dynamics over time during free behavior. Visualization of the postural data reveals inherent block structure in the 3D postural dynamics. Compression of the preprocessed and spine-aligned data using a random projection technique reveals sporadic, sharp shifts in the postural data over time. Similar data structure is observed in the raw data and in the height of the animal's spine (top figure, spine height at any given location is colored, mm = mm above the ground). When the animal is standing upright (at the beginning of the data stream), its cross-sectional profile relative to the camera is smaller; when the animal is on all fours, its profile is larger. Figure 11C A change point analysis identifying potential boundaries between these blocks is shown (normalized probabilities of change points indicated in the trace at the bottom of the behavioral data). Block duration distributions are shown by plotting the duration of each block determined by the change point analysis (n=25, 500 minutes of imaging, mean=358 ms, SD 495 ms). The average block duration is plotted in black, whereas the duration distribution associated with each individual mouse is plotted in gray. Figure 11C , middle and right. Autocorrelation analysis shows that the decorrelation rate of mouse posture slows down after about 400ms (left, average value plotted in dark blue, decorrelation of individual mice plotted in light blue, τ = 340 ± 58ms). By plotting the spectral power ratio between behaving mice and dead mice (right, average value plotted in black, individual mice plotted in gray), it is shown that most behavioral frequency components are expressed between 1 and 6 Hz (average = 3.75 ± 0.56 Hz).
[0062] Figure 12A-1 2D rendering of how the postural dynamics data of a mouse may contain reused behavioral modules according to various embodiments of the present invention. Figure 12ADepicts how projections of mouse posture data into principal component (PC) space (bottom) reveal that individual patches identified in the posture data encode re-used trajectories. Following principal component analysis of the mouse posture data, the values of the top two PCs at each time point were plotted as a two-dimensional graph (with point density plotted in color). Trajectories in PC space (white) were identified by tracing paths associated with patches highlighted by change point analysis (top). By searching the posture data using a template matching procedure, other examples of similar trajectories in PC space (with changes from blue to red indicating changes over time) encoding that patch were identified, indicating that the template patch represents a re-used motion theme. Figure 12B Depicts the identification of individual behavioral modules by modeling mouse posture data using AR-HMM. AR-HMM parses behavioral data into a finite set of identifiable modules (top – labeled “Tags,” each module is uniquely color-coded). Multiple data instances associated with a single behavioral module each have a stereotyped trajectory through PCA space (bottom left, green trajectory); multiple trajectories define a behavioral sequence (bottom center). Depicting a side view of the mouse (inferred from depth data, bottom right) reveals that each trajectory in the behavioral sequence encodes a different elemental action (time within the module is represented by increasingly darker lines from the beginning to the end of the module). Figure 12C depicts an isometric view of 3D imaging data associated with the walking, hesitating, and low upright modules. Figure 12D depicts a cross-likelihood analysis, which depicts the probability that a data instance assigned to a particular module will be effectively modeled by another module. Cross-likelihoods were calculated for the open-field dataset, and the likelihood that any given data instance assigned to a particular module would be accurately modeled by different modules was heat-plotted (in nats, where enats is the likelihood ratio); note that the high likelihood diagonal and all off-diagonal lines are relatively low likelihoods. Plotting the same metric on a model trained on synthetic data with an autocorrelation structure that matches the real mouse data but lacks modularity reveals that the AR-HMM fails to recognize modules when underlying modularity is absent in the training data.
[0063] Figures 13A-13B plot patch and autocorrelation structure in mouse depth imaging data according to various embodiments of the present invention. Figure 13A plots the presence of patch structure in random projection data, spine data, and raw pixel data derived from aligned mouse pose dynamics. Figure 13BIt is shown that living mice exhibit a clear block structure in the imaging data (left panel), while dead mice do not (right panel). Compression does not significantly affect the autocorrelation structure of the mouse posture dynamics data. The raw pixels, PCA data and random projections representing the same depth data set (left panel) all decorrelate at roughly the same rate, indicating that data compression does not affect the fine timescale correlation structure in the imaging data. If the mouse posture evolves into a Levy flight (middle panel) or a random walk (right panel), no such correlation structure is observed, indicating that living mice express a specific sub-second autocorrelation structure associated with switching dynamics.
[0064] Figure 14 The variance explained after removing dimensions using principal component analysis according to various embodiments of the present invention is plotted. The plot comparing the variance explained (Y-axis) to the number of PCA dimensions included (X-axis) shows that 88% of the variance is captured by the first 10 principal components; this number of dimensions was used by the AR-HMM for data analysis.
[0065] Figure 15 plots comparative modeling of mouse behavior according to various embodiments of the present invention. A series of computational behavioral models were constructed, each exemplifying a different hypothesis about the behavioral infrastructure, and all trained on mouse behavioral data (in the form of the first 10 principal components extracted from the aligned depth data). These models include the Gaussian model (which proposes that mouse behavior is a single Gaussian model in posture space), GMM (Gaussian mixture model, which proposes that mouse behavior is a Gaussian mixture model in posture space), Gaussian HMM (Gaussian hidden Markov model, which proposes that the behaviors created by the model, each behavior is a Gaussian model in posture space, temporally related to each other with definable transition statistics), GMM HMM (Gaussian mixture model hidden Markov model, which proposes that the behaviors created by the model, each behavior is a mixture of Gaussian models in posture space, temporally related to definable transition statistics), AR model (which proposes that mouse behavior is a single, continuous autoregressive trajectory in posture space), AR MM (which proposes that mouse behavior is built from modules, where each module encodes an autoregressive trajectory in posture space and transfers from one to another randomly), and AR sHMM (which proposes that mouse behavior is built from modules, where each module encodes an autoregressive trajectory in posture space and transfers from one to another with definable transition statistics). The performance of these models in predicting the structure of behavioral data for mice that have not been exposed to these models is shown on the Y-axis (measured in likelihood units and normalized to the performance of the Gaussian model), and the ability of each model to predict behavior on a frame-by-frame basis is shown on the X-axis (top). This graph was split into three slices at different time points, demonstrating that the best AR HMM outperforms alternative models on time scales where the switching dynamics inherent in the data come into play (e.g., over 10 frames, error bars are SEM).
[0066] Figure 16 The duration distributions of qualitatively similar blocks and modules according to various embodiments of the present invention are plotted. The percentage of blocks / modules of a given duration (Y-axis) plotted against the block duration (X-axis) shows that the duration distributions of blocks identified by the change point algorithm and behavioral modules identified by the model are roughly similar. Although these distributions are not identical, they are expected to be similar because the change point algorithm identifies local changes in the data structure, while the model identifies modules based on their content and their transition statistics; note that the model does not have direct access to the "local break" metric used by the change point algorithm.
[0067] Figure 17 The data is plotted showing how shuffling behavior at fast time scales degrades AR-HMM performance according to various embodiments of the invention.
[0068] Figure 18 Plotted are visualizations of mouse behavior generated by models according to various embodiments of the present invention, each model being trained on behavioral data (left) and then allowing the generation of an "ideal" version of mouse behavior (right); here, the output is a visualization of the shape of the animal's spine over time. The individual modules identified by each model are represented by a color code below each model (labeled "markers").
[0069] Figure 19 Plotted how sparse module interconnections are according to various embodiments of the present invention. Without thresholding, the average module is interconnected with 16.85 ± 0.95 other modules; this moderate interconnectivity is significantly reduced even with moderate thresholding (X-axis, thresholding applied to binary probabilities), consistent with the sparse temporal interconnections between individual behavioral modules.
[0070] Figure 20 Determined filtering parameters according to various embodiments of the present invention are plotted. To filter the data from Kinect, we used an iterative median filtering approach, in which we iteratively apply a median filter across space and time; this approach has been shown to effectively preserve data structure while removing noise. To identify the optimal filter settings, we imaged mice dead in different rigor mortis poses; an ideal filter setting would distinguish mice in different poses, but not data from the same mouse. The filter settings are represented as ((pixels), (frames)), where each number in brackets refers to the iterative setting for each round of filtering. To evaluate filter performance, we calculated the within / between pose correlation ratio (Y-axis), where the average spatial correlation across all frames in the same pose is divided by the average spatial correlation across all frames in different poses. This reveals that optical filtering (using settings ((3), (3,5))) optimizes the distinguishability in the data.
[0071] Figure 21 The parameters of the algorithm for identifying change points according to various embodiments of the present invention are plotted. Clear optimal values for sigma (σ) and H (two panels on the left) were identified by grid scanning by optimizing for the change point ratio (the number of change points identified in live mice relative to dead mice, Y axis). This change point ratio is highly insensitive to K; therefore, a setting of 48 (at the observed maximum) was selected.
[0072] Figure 22A graphical model of an AR-HMM according to various embodiments of the present invention is depicted. For time indices t = 1, 2, …, T, the shaded nodes labeled y_t represent preprocessed 3D data sequences. Each such data node y_t has a corresponding state node x_t that assigns the data frame to a behavioral mode. Other nodes represent parameters for managing transitions between modes (i.e., the transition matrix π) and the autoregressive dynamic parameters for each mode (i.e., the parameter set θ).
[0073] Figure 23 Plotted are image descriptions of dimensionality reduction using a neural network to form an image manifold according to various embodiments of the present invention.
[0074] Figure 24 Plotted are image representations of structured variational autoencoders according to various embodiments of the invention.
[0075] Figure 25 Plotted are applications of structured variation autoencoders according to various embodiments of the present invention. Figure 4 This is an example of generation using mouse video data. Figure 5 shows an example of filtering and generating images from 1D bouncing data. Figure 6 compares the point problem using natural gradient updates (bottom trend line) and standard gradient updates (top). Figure 7 is a 2D grid in mouse image manifold coordinates.
[0076] In the drawings, for ease of understanding and convenience, the same reference numerals and any acronyms identify elements or operations having the same or similar structure or function. To facilitate discussion of any particular element or operation, the most significant digit or digits in the reference number refer to the figure number that first introduces the element. DETAILED DESCRIPTION
[0077] In some embodiments, characteristics such as size, shape, relative positions, etc. used to describe and claim certain embodiments of the present invention are understood to be modified by the term "about."
[0078] Various examples of the present invention will now be described. The following description provides specific details for a thorough understanding of these examples and to enable illustration of these examples. However, those skilled in the relevant art will appreciate that the present invention can be practiced without many of these details. Similarly, those skilled in the relevant art will also appreciate that the present invention may include many other distinct features not described in detail herein. In addition, some well-known structures or functions will not be shown or described in detail below to avoid unnecessarily obscuring the relevant description.
[0079] The terms used below should be interpreted in their broadest, most reasonable manner, even when used in conjunction with the detailed description of certain specific examples of the present invention. In fact, certain terms may even be emphasized below; however, any term that is intended to be interpreted in any limiting manner will be openly and explicitly defined in this detailed description section.
[0080] Although this specification contains many specific implementation details, these details should not be interpreted as limitations on the scope of any invention or what may be claimed, but rather should be understood as descriptions of the characteristics of particular implementations of particular inventions. Certain features described in this specification in the context of separate implementations may also be implemented in combination in a single implementation. Conversely, various features described in the context of a single implementation may also be implemented individually in multiple implementations or in any suitable subcombination. Furthermore, although features may be described as functioning in certain combinations and even as originally claimed, in some cases, one or more features from a claimed combination may be deleted from that combination, and a claimed combination may involve a subcombination or a variation of a subcombination.
[0081] Similarly, although operations may be described in a particular order in the accompanying drawings, this should not be understood as requiring that the operations be performed in the particular order shown, or in sequential order, or that all illustrated operations be performed, in order to achieve the desired results. In certain circumstances, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system components in the above-described implementations should not be understood as requiring such separation in all implementations, and it should be understood that the described program components and systems may generally be integrated together in a software product or packaged into multiple software products.
[0082] Overview
[0083] The present inventors have discovered systems and methods for automatically and objectively identifying and classifying animal behavior modules by processing video data of animals. These systems can classify animal behavioral states by quantitatively measuring, processing, and analyzing three-dimensional animal postures or posture trajectories using a depth camera. These systems and methods avoid the need for a priori definition of what constitutes a measurable action unit, thereby making the classification of behavioral states objective and unsupervised.
[0084] In one aspect, the present invention relates to a method for analyzing the motion of an object to divide it into sub-second modules, the method comprising: (i) using a computational model to process three-dimensional video data representing the motion of the object to divide the video data into at least one group of sub-second modules and transition periods between the at least one group of sub-second modules; and (ii) assigning the at least one group of sub-second modules to a category representing an animal behavior type.
[0085] Figure 1 An embodiment of a process that can be utilized by the system for automatically classifying video frames or sets of frames into behavioral modules is shown. For example, the system can include a camera 100 and a tracking system 110. In some embodiments, the camera 100 can be a 3D depth camera, and the tracking system 110 can project structured infrared light into the test field 10. An infrared receiver on the tracking system can determine the location of the target based on parallax. In some embodiments, the camera 100 can be connected to the tracking system 110, or in some embodiments, they can be separate components.
[0086] Camera 100 may output data related to video images and / or tracking data from tracking system 110 to computing device 113. In some embodiments, computing device 113 performs pre-processing of the data locally before sending the data over network 120 for analysis by server 130 and storage in database 160. In other embodiments, the data may be processed and fitted locally on computing device 113.
[0087] In one embodiment, a 3D depth camera 100 is used to obtain a video image stream of an animal 50 having area information and depth information. A background image (a blank experimental area) is then removed from each of the multiple images to generate a processed image having bright and dark areas. The outlines of the bright areas in the multiple processed images can be obtained, and parameters can then be extracted from the area image information and depth image information within the outlines to form a plurality of multidimensional data points, each of which represents the posture of the animal at a specific time. The posture data points can then be clustered so that the clusters represent the animal's behavior.
[0088] The pre-processed depth camera video data can then be input into various models to classify the video data into sub-second "modules" and transition periods that describe repeating units of behavior that are grouped together to form coherent behaviors observable to the human eye. The output of the model that classifies the video data into modules can output several key parameters, including: (1) the number of behavioral modules observed in a given experimental data set (i.e., the number of states), (2) parameters describing the pattern of motion expressed by the mouse associated with any given module (i.e., state-specific autoregressive dynamic parameters), (3) parameters describing how frequently any particular module transitions to any other module (i.e., the state transition matrix), and (4) the assignment of each video frame to a behavioral module (i.e., the sequence of states associated with each data sequence). In some embodiments, these latent variables are defined by a generative probabilistic process and estimated simultaneously using a Bayesian inference algorithm.
[0089] Camera setup and initialization
[0090] Various methods can be used to record and track video images of an animal 50 (e.g., a mouse). In some embodiments, the recorded video can be recorded in three dimensions. Various devices can be used for this function, for example, the experiments disclosed herein utilized Microsoft's Kinect for Windows. In other embodiments, the following additional devices can be used: (1) a stereo vision camera (which can include two or more two-dimensional camera groups calibrated to produce depth images), (2) a time-of-flight depth camera (e.g., CamCube, PrimeSense, Microsoft Kinect 2), a structured lighting depth camera (e.g., Microsoft Kinect 1), and X-ray video.
[0091] The camera 100 and tracking system 110 can project structured infrared light onto the imaging field 10 and calculate the three-dimensional position of objects in the imaging field 10 based on parallax. Microsoft's Windows Kinect has a minimum working distance of 0.5 meters (in close-up mode); by quantifying the number of missing depth pixels within the imaging field, the optimal sensor position can be determined. For example, the inventors have found that, depending on the ambient light conditions and the test material, the optimal sensor position for the Kinect is between 0.6 and 0.75 meters from the test field.
[0092] Data collection
[0093] The data output from the camera 100 and the tracking system 110 can be received and processed by the computing device 113, which processes the depth frames and saves them in a suitable format (e.g., binary or other format). In some embodiments, the data from the camera 100 and the tracking system 110 can be output directly to the server 130 via the network 120, or can be temporarily buffered and / or sent to the associated computing device 113 via a USB or other connection, which temporarily stores the data before sending it to the centralized server 130 via the network 120 for further processing. In other embodiments, the data can be processed by the associated computer 113 without being sent over the network 120.
[0094] For example, in some embodiments, the data output from the Kinect can be sent to a computer via a USB port that uses custom Matlab or other software to connect to the Kinect via the official Microsoft.NET API, which retrieves depth frames at a rate of 30 frames per second and saves them in raw binary format (16-bit signed integers) to an external hard drive or other storage device. Because USB 3.0 has enough bandwidth to allow real-time data streaming to an external hard drive or computing device with memory. However, in some embodiments, the network may not have enough bandwidth to allow real-time remote data streaming.
[0095] Data preprocessing
[0096] In some embodiments, after the raw images of the video data are saved and / or stored in a database or other memory, various pre-processing steps may be performed to separate the animals in the video data and orient the images of the animals along a common axis for further processing. In some embodiments, the orientation of the head may be used to orient the images in a common direction. In other embodiments, the inferred orientation of the spine may be included.
[0097] For example, tracking the evolution of the posture of an imaged mouse over time requires: identifying the mouse within a given video sequence; segmenting the mouse from the background (in this case, the equipment the mouse is exploring); orienting separate images of the mouse along the axis of its spine; correcting the image for perspective distortion; and then compressing the image for processing by the model.
[0098] Separating animal video data
[0099] Figure 2AThe process that the system can perform to separate regions of interest and subtract background images to separate video data of an animal 50 is shown. First, to isolate the experimental field where the mouse is performing the behavior action, the system can first identify a region-of-interest (ROI) 210 for further analysis. In other embodiments, the region of interest 210 can include the entire field of view 10 of the recorded video data. To separate the regions, the outer edge of any imaging field can be manually traced; pixels outside the ROI 210 can be set to zero to prevent false target detection. In other embodiments, the system can automatically define the ROI 210 using various methods. In some embodiments, the system can filter the raw imaging data using an iterative median filter, which is well suited for removing correlated noise from the sensor, such as in the Kinect.
[0100] After selecting the region of interest 210, the original image can be cropped to the region of interest 215. Then, the missing pixel values can be input 225. After that, the X, Y, and Z positions can be calculated 230 for each pixel, and the pixel positions can be resampled. Thus, the image can be resampled to real-world coordinates. The system then calculates a median real-world background image 240 and subtracts the median real-world background image 245 from the real-world image.
[0101] To subtract the background image of the venue from the video data, various techniques may be performed, including, for example, subtracting the median value of a portion of a set time period (e.g., 30 seconds) of the video data. For example, in some embodiments, the first 30 seconds of data in any imaging stream may be subtracted from all video frames, and any spurious values less than zero may be reset to zero.
[0102] To further ensure that the analysis is focused on the animal, the system can binarize the image (or perform a similar process using thresholding) and eliminate any objects that do not survive a certain number of iterations of the morphological opening operation. Once this is done, the system can then perform Figure 2B Additional processing is shown. Thus, the background subtracted image (mouse video data) 250 can be filtered and artifacts 255 can be removed. In some embodiments, this may involve iterative median filtering.
[0103] The animal in the image data can then be identified by defining it as the largest object in the arena that survives the subtraction and masking process or by blob detection 260. An image of the mouse can then be extracted 265.
[0104] Identifying animal orientation
[0105] The centroid of the animal (e.g., mouse) can then be identified as the centroid of the pre-processed image or reduced by other suitable methods 270; an ellipse can then be fitted to its outline 285 to detect its overall orientation. In order to correctly orient the mouse 280, various machine learning algorithms (e.g., random forest classifiers) can be trained on a set of manually oriented extracted mouse images. In the case of an image, the orientation algorithm then returns an output indicating whether the mouse's head is correctly oriented.
[0106] Once the position is identified, additional information 275 can be extracted from the video data, including the location of the animal's centroid, head and tail positions, orientation, length, width, height, and their first-order derivatives with respect to time. Characterization of the animal's postural dynamics requires correction for perspective distortion in the x and y axes. This distortion can be corrected by first generating an array of (x, y, z) coordinates for each pixel in real-world coordinates, and then resampling these coordinates to make them fall on a uniform grid in the (x, y) plane using Delaunay triangulation.
[0107] Output to model-based or model-free algorithms
[0108] like Figure 3 As shown, in some embodiments, the output of the orientation-corrected image will be a principal component analysis (PCA) time series 310 or other statistical method for reducing data points. In some embodiments, the data is run through a model fitting algorithm 315 (e.g., the AR-HMM algorithm or SLDS SVAE algorithm disclosed herein), or it can be run through a model-free algorithm 320 as disclosed to identify behavioral modules 300 contained within the video data. In addition, in some embodiments, a PCA time series is not performed. In some embodiments, a multilayer perceptron is used to reduce dimensionality.
[0109] In an embodiment with a model-free algorithm 320, various combinations of algorithms may be utilized with the goal of separating sub-second behavioral modules with similar directional profiles and trajectories. Some examples of these algorithms are disclosed herein, however, additional algorithms for segmenting data into behavioral modules may be envisioned.
[0110] Image dimensionality reduction
[0111] In some embodiments, both include a model-free algorithm 320 or a model-fitting 315 algorithm, and the information captured in each pixel is often highly correlated (adjacent pixels) or uninformative (pixels at the border of the image that never represent the body of the mouse). In order to both reduce redundant dimensions and make modeling computationally tractable, various techniques can be used to reduce each image in dimension. For example, a 5-level wavelet decomposition can be performed to convert the image into a representation in which each dimension captures and aggregates information at a single spatial scale; in this conversion, some dimensions may explicitly encode fine edges at a scale of a few millimeters, while other dimensions encode broad variations at a spatial scale of centimeters.
[0112] However, this wavelet decomposition will expand the dimensionality of the image. To reduce this dimensionality, various techniques can be applied.
[0113] In some embodiments, random projections can be used to reduce the dimensionality of the data. Random projections are a method of generating new dimensions from an original signal of dimension D_orig by randomly weighting each original dimension and then summing each dimension according to that weight, resulting in a single number per data point. This process can be repeated multiple times with new random weights to produce a set of "randomly projected" dimensions. The Johnson–Lindenstrauss lemma states that the distances between points in the original data set of dimension D_orig are preserved in the randomly projected dimension D_proj, where D_proj <D_orig。
[0114] In other embodiments, principal component analysis can be applied to these vectors to project the wavelet coefficients into ten dimensions, and the inventors still found that >95% of the total variance was captured. For example, principal components can be constructed using a canonical dataset of 25 6-week-old C57 BL / 6 mice, each recorded for 20 minutes, and all datasets are projected into this common posture space. Therefore, the output of PCA can then be input into a modeling algorithm for model identification.
[0115] However, PCA reduces the dimensionality to a linear space. The inventors have found that reducing the dimensionality to a linear space will not accommodate various variations in mice that are not related to behavior. This includes variations in mouse size, mouse breed, etc.
[0116] Therefore, the inventors have discovered that using certain types of neural networks, such as multilayer perceptrons, one can effectively reduce the dimensionality of images while maintaining more robust representations. For example, as disclosed herein, the inventors propose a structured variational autoencoder for dimensionality reduction. Furthermore, these reduced-dimensional images provide an effective method for developing models that are independent of the size of the mouse or other animal and can account for other variations unrelated to behavior. For example, some neural networks can be used that reduce the dimensionality to a ten-dimensional image manifold.
[0117] Model-free algorithms: identifying behavioral module lengths
[0118] In some embodiments with a model-free algorithm 320, in order to assess the time scale of animal behavior self-similarity (which reflects the rate at which an animal transitions from one movement mode to another), an autocorrelation analysis can be performed. Because some data smoothing is required to eliminate sensor-specific noise, when the autocorrelogram is calculated as the statistical correlation between the time-lagged versions of the signal, a decreasing autocorrelogram will result, even for animals with postmortem stiffness (e.g., mice). Therefore, the correlation distance between all 10 dimensions of the mouse posture data can be used as a comparator between the time-lagged versions of the time series signal in question, resulting in a flat autocorrelation function with a value of ~1.0 for dead animals and a decreasing autocorrelation function for behaving animals (e.g., mice). The rate at which this autocorrelogram of the behaving mouse decreases is a measure of the fundamental time scale of the behavior, which can be characterized as the time constant τ of the exponential decay curve. τ can be fitted using the Levenberg-Marquardt algorithm (nonlinear least squares method) using the SciPy optimization package.
[0119] In certain embodiments, power spectral density (PSD) analysis can be performed on mouse behavior data to further analyze its temporal structure. For example, a Wiener filter can be used to identify the temporal frequencies that must be increased in the signal extracted from the dead mouse to best match the behavior mouse. This can be achieved simply by using the ratio of the PSD of the behavior mouse to the PSD of the dead mouse. In certain embodiments, the Welch periodogram method can be used to calculate PSD, which uses the average PSD of the sliding window in the entire signal.
[0120] Model-free algorithm: Locating the change points during the transfer period
[0121] In some embodiments where a model is not used to identify module 320, various methods can be used to determine the change point of the transfer period. By drawing the random projection of the mouse depth image over time, obvious stripes are produced, and each stripe is a potential change point over time. In order to automatically identify these change points (these change points represent the potential boundaries between the obvious block structures in the random projection data), a simple change point identification technique known as a filtered derivative algorithm can be used. For example, an algorithm for calculating the derivative of a normalized random projection with a lag of k=4 frames can be used. For each time point, for each dimension, the algorithm can determine whether the signal has crossed a certain threshold value h=0.15mm. Then, the binary change point indicator signal can be summed in each random projection dimension of D=300, and then the resulting 1D signal can be smoothed using a Gaussian filter with a core standard deviation of sigma=0.43 frames. Then, the change point can be identified as the local maximum of the smoothed 1D time series. The process depends in part on the specific values of the parameters k, h, and sigma (σ); for example, values that maximize the number of change points in behaving mice while producing no change points in dead mice can be used.
[0122] Model-free algorithms: identifying similar or repeated modules
[0123] In some embodiments, where the data is analyzed without using model 320, certain algorithms may be used to identify similar and repetitive modules. Thus, a set of repetitive modules may be identified as words or syllables of the animal's behavior. Thus, to determine whether any reasonably long segments of behavior (greater than only a few frames) have ever "repeated" (independent of the underlying model of the behavior), the system may use a template matching process to identify similar trajectories in the PCA or MLP manifold space. For example, to identify similar trajectories, the system and method may calculate the Euclidean distance between a target segment, a "template," and each possible segment of equal length (generally defined by approximate block boundaries identified by change point analysis). Other similar methods may also be used to identify modules, including other statistically based methods.
[0124] In some embodiments, a collection of similar modules will be selected as the most similar segments, while ignoring segments found to be offset by less than 1 second from each other (to ensure that we selected behavioral segments that were temporally separated from each other and also occurred in separate mice).
[0125] Data Modeling
[0126] In other embodiments, systems and methods for identifying behavioral modules in video data using data models 315 can be employed. For example, the data models can implement the well-established paradigm of generative probabilistic modeling, which is commonly used to model complex dynamic processes. Such models are generative in the sense that they describe a process that can synthetically generate observational data through the model itself, and they are probabilistic because the process is mathematically defined in terms of sampling probability distributions. Furthermore, by fitting an interpretable model to the data, the data is "parsed" in a manner that reveals that the latent variable structure assumed by the model gave rise to the data (including parameters describing the number and identity of states and parameters describing transitions between states).
[0127] In some embodiments, model 315 can be expressed using a Bayesian framework. The Bayesian framework provides a natural way to express hierarchical models of behavioral organization, priors, or regularizers that reflect known or observed constraints on motion patterns within 3D data and a consistent representation of uncertainty. The framework also enables significant and sophisticated computational machinery for reasoning about the key parameters of any model. Within the Bayesian framework, data corrects the posterior distribution of latent variables for a specific model structure (e.g., the spatiotemporal properties of states and their possible transitions) and the prior distribution of latent variables.
[0128] In the following, a model-based approach for characterizing behavior is defined in two steps: first, a mathematical definition of the generative model and priors used, and second, a description of the inference algorithm.
[0129] Example model for identifying behavioral modules—AR-HMM
[0130] In some embodiments, the system can use a discrete-time hidden Markov model 315 (HMM) to identify behavioral modules. The HMM contains a series of random processes for modeling sequential and time series data. The HMM model assumes that at each time point (for example, for each frame of imaging data), the mouse is in a discrete state (Markov state) that can be labeled. Each Markov state represents a brief three-dimensional movement theme that the animal takes while in that state. Because the observed three-dimensional behavior of the mouse depends on the specific movement pattern expressed by the animal in the recent period, ideally, each Markov state will predict its future behavior based on the mouse's recent posture dynamics. Therefore, each Markov state consists of a hidden discrete component (for identifying the animal's behavioral pattern) and several delays of the observation sequence (for predicting the animal's short-term behavior based on the behavioral pattern). This model structure is commonly referred to as an SVAR (switching vector-autoregressive) model and an autoregressive HMM (AR-HMM).
[0131] Figure 4 An example is provided of how the AR-HMM algorithm converts input data (spinal aligned depth imaging data 305 that has been dimensionally reduced 405 using PCA 310) into a fitted model, where the fitted model describes the number of behavioral modules and the trajectories they encode in PCA space, a module-specific duration distribution that governs the persistence of any trajectory within a given module, and a transition matrix that describes how these individual modules connect to each other over time.
[0132] In addition, the AR-HMM can be configured to assign a label to each frame of the training data, where the label associates the frame with a given behavioral module. After preprocessing and dimensionality reduction 405, the imaging data is decomposed into a training set 415 and a test set 410. The training set 415 is then submitted to the AR-HMM 315. After randomly initializing the parameters of the model 315 (here, the autoregressive parameters describing the trajectory of each module through the PCA space, the transition matrix describing the probability of temporal interconnections between modules, the duration distribution parameters describing the likely duration of any instance of a given module, and the label assigned to each frame of the imaging data and associating the frame with a specific module), the AR-HMM attempts to fit the model 315 by varying one parameter while keeping the other parameters constant. The AR-HMM alternates between two main updates: the algorithm 315 first attempts to partition the imaging data into modules that give a fixed set of transition statistics and a fixed description of the AR parameters describing any given module. Then, the algorithm switches to a fixed partitioning and updates the transition matrix and AR parameters 455. AR-HMM 315 uses a similar approach to assign any given frame of imaging data to a given module. It first calculates the probability that a given module is the "correct" module, which is proportional to a measure of how well the corresponding autoregressive parameters 455 of the state describe the data at that time index and how well the resulting state transitions are consistent with the transition matrix 450.
[0133] In the second step, the AR-HMM 315 changes the autoregressive parameters 455 and the transition parameters 450 to better fit the assigned data, thereby updating each behavioral module and the transition model between modules. The product of this process is the described parameters 455, which are then evaluated for their quality in describing the behavior using a likelihood measure of the data derived from training 475.
[0134] By identifying discrete latent states 445 associated with 3D pose sequence data, the HMM model 315 can identify segments of data that exhibit similar short-timescale motion dynamics and explain these segments in terms of reused autoregressive parameters. For each observation sequence, there exists a sequence of unobserved states: if the discrete state at time index t is x_t=i, then the probability that the discrete state x_(t+1) takes value j is a deterministic function of i and j and is independent of all previous states. In symbolic terms,
[0135] p(x t+1 |x t ,x t-1 ,x t-2 ,...,x1)=p(x t+1 |x t )
[0136] p(x t+1 =j|x t =i) =π ij
[0137] where π is the transition matrix 450, where the (i, j) element is the probability of transitioning from state i to state j. In some embodiments, the dynamics of the discrete states can be fully parameterized by the transition matrix, where the transition matrix is assumed to be invariant over time. One of the tasks of the inference algorithm (described below) is to reason about sequences of discrete states and the possible values of the transition matrix governing the deployment of these sequences, thereby reasoning about a series of reusable behavioral modules and the transition patterns governing how these modules are connected over time.
[0138] In the case of a sequence of discrete states, the corresponding sequence of 3D pose data can be modeled as a conditional vector autoregressive (VAR) process. Each state-specific vector autoregressive can capture the short-timescale motion dynamics specific to the corresponding discrete state; in other words, each behavioral module can be modeled as its own autoregressive process. More precisely, given a discrete state x_t of the system at any time index t, the value of the observed data vector at time point y_t is distributed according to a state-specific noise regression of the K previous values of the observation sequence y_(t-1),…,y_(tK). The inference algorithm can also be responsible for inferring the most likely values of the autoregressive dynamic parameters for each state and the number of lags used in the dynamics.
[0139] In some embodiments, these switching autoregressive dynamics define the core of the AR-HMM. However, because different animal populations or experimental conditions are expected to result in behavioral differences, when considering two or more such experimental conditions, the model can be constructed hierarchically: different experimental conditions can be allowed to share the same state-specific VAR dynamic library, but learn their own transition patterns and any unique VAR dynamic patterns. This simple extension enables the model to reflect parameter changes caused by experimental changes. In addition, the combined Bayesian inference algorithm used directly extends this hierarchical model.
[0140] To employ Bayesian inference methods, a unified representation can be used to treat the unknown quantities (including the transition matrix 450 and the autoregressive parameters 455 describing each state 445) as latent random variables. Specifically, a weak prior distribution 465 can be placed on these quantities, and their posterior distributions 465 adjusted on the observed 3D imaging data can be studied. For the autoregressive parameters, a prior containing a lasso-like penalty can be used to encourage uninformative lag exponents so that their corresponding regression matrix coefficients tend to zero.
[0141] A hierarchical Dirichlet process 435 prior can be used for the transition matrix 450 to regularize the number of discrete hidden states 445. In addition, the transition matrix 450 also previously includes a sticky bias, which is a single non-negative number that controls the tendency of discrete states to transition to themselves. Since this parameter controls the time scale of the inference transition dynamics, it can be set so that the output of the model inference algorithm matches (as closely as possible) the model-free duration distribution determined by the change point analysis disclosed herein (or other methods of identifying module lengths) and the autocorrelogram generated from the preprocessed and unmodeled 3D pose data. In some embodiments, for example, this parameter can be tuned to define a prior on the time scale of the behavior.
[0142] In some embodiments, by removing certain parts of the model structure, a simpler model than the AR-HMM model can be used. For example, by removing the discrete transition dynamics captured in the transfer matrix and replacing them with a mixture model, an alternative model can be generated in which the distribution over each discrete state does not depend on its previous state. This is the case where the animal has a set of behavioral modules to choose from, and the probability of expressing any given one of them does not depend on the order in which they appear. This simplification leads to an autoregressive mixture model (AR-MM).
[0143] Alternatively, replacing the conditional autoregressive dynamics with simple state-specific Gaussian emissions yields a Gaussian output HMM (G-HMM); this model explores the hypothesis that each behavioral module is best described by a simple posture rather than a dynamic trajectory. Applying these two simplifications yields a Gaussian mixture model (G-MM), where behavior is simply a sequence of postures over time, where the probability of expressing any given posture does not depend on previous postures. Removing the switching dynamics results in a pure autoregressive (AR) or linear dynamical system (LDS) model, where behavior is described as a trajectory in posture space without any reuse of the discrete behavioral model.
[0144] Analysis of behavioral modules
[0145] In some embodiments, the system may provide an indication of the relationships between behavior modules, describe the most commonly used behavior modules, or perform other useful analysis of the behavior modules.
[0146] For example, to represent grammatical relationships between syllables in a behavior, the probability (e.g., a bigram) that two syllables appear one after another (a "bigram" of a module) can be calculated as a fraction of all observed bigrams. In some embodiments, to calculate this value for each module pair (i, j), for example, a square n×n matrix A can be used, where n is the number of total modules in the token sequence. The systems and methods can then scan the token sequence saved at the final iteration of Gibbs sampling, incrementing the entry A[i, j] each time the system identifies a syllable i that directly precedes syllable j. At the end of the token sequence, the system can divide by the number of total bigrams observed.
[0147] To visually organize modules that are specifically upregulated or selectively expressed due to manipulation, the system can assign a selectivity index to each module. For example, where p(condition) represents the condition's percentage usage of the module, the system can sort modules by comparing the circular field to the square box using (p(circle) - p(square) / (p(circle) + p(square)). In the comparison between an odorless odor and a fox odor (TMT), the system can sort modules using (p(tmt) - p(odorless)) / (p(tmt) + p(odorless)).
[0148] Statechart visualization
[0149] The system can also output the syllable bigram probability and syllable usage rate of n syllables on the graph G = (V, E), where each node i∈V = {1,2,...,n} corresponds to syllable i and each directed edge (i,j)∈E = {1,2,...,n} 2 \{{i,i}:i∈V} corresponds to a bigram. The graph can be output as a set of circular nodes and directed arcs, such that the size of each node is proportional to the usage of the corresponding letter, and the width and opacity of each arc are proportional to the probability of the corresponding bigram within the minimum and maximum ranges shown in the legend. In order to arrange each graph in a reproducible non-(pseudo)random manner (up to global rotation of the drawing), the system can use a spectral layout algorithm to initialize the positions of the nodes and a Fructherman-Reingold iterative force-directed layout algorithm to fine-tune the node positions; both algorithms we use are available in the NetworkX software package.
[0150] Overview of Main Inference Algorithms
[0151] In some embodiments, an inference algorithm can be applied to the model 315 to estimate parameters. For example, Gibbs sampling can be used to perform approximate Bayesian inference, i.e., a Markov Chain Monte Carlo (MCMC) inference algorithm. In the MCMC paradigm, the inference algorithm constructs approximate samples based on the posterior distribution of interest, and these samples are used to calculate the average or serve as a proxy for the posterior mode. The sample sequence generated by the algorithm is distributed in areas of high posterior probability while escaping from low posterior probability or poor local optimal areas. In the main AR-HMM model, the latent variables of interest include: vector autoregressive parameters, hidden discrete state sequences, and transition matrices (e.g., autoregressive parameters that define the posture dynamics within any given behavioral module, module sequences, and transition probabilities between any given module and any other module). The application of the MCMC inference algorithm to the 3D imaging data generates a set of samples of these latent variables for the AR-HMM.
[0152] The Gibbs sampling algorithm has a natural alternating structure that is directly analogous to the alternating structures of expectation maximization (EM) and variational mean field algorithms. When applied to an AR-HMM after initialization with random samples from a prior, the algorithm can alternate between two main updates: first, it can resample the sequence of hidden discrete states that gives the transition matrix and autoregressive parameters, and second, it can resample the parameters that give the hidden states.
[0153] In other words, algorithm 315 first attempts to segment the imaging data into modules 300 using a fixed set of transition statistics and a fixed specification of the AR parameters used to describe any given module. The algorithm then switches to the fixed segmentation and updates the transition matrix 450 and AR parameters 455. To assign each 3D pose video frame to one of the behavior modes 300 in the first step of this process, a state label 445 for a particular time index is randomly sampled from a set of possible discrete states, where the probability of sampling a given state is proportional to how well the state's corresponding autoregressive parameters describe the data at that index and how consistent the resulting state transition is with the transition matrix 450. In the second step, once the data subsequences have been assigned to states, the autoregressive and transition parameters are resampled to fit the assigned data, thereby updating the dynamic model for each behavior mode and the model of transitions between models. This process, implemented by the Gibbs sampling algorithm, is noisy, allowing the algorithm to avoid local maxima that could prevent it from effectively exploring parameter space.
[0154] Example
[0155] The following discloses examples of specific implementations of the models described herein for performing the disclosed examples. Variations of these models can be implemented to identify behavioral models.
[0156] Prior on the transfer matrix
[0157] The sticky HDP prior is placed on the transfer matrix π with concentration parameters α,γ>0 and stickiness parameter κ>0, where
[0158]
[0159]
[0160] Among them, when i=j, δ ij is 1, otherwise it is 0, and π i Denotes the i-th row of π. Gamma priors are placed on α and γ, setting α to Gamma(1, 1 / 100) and γ to Gamma(1, 1 / 100).
[0161] Generation of discrete state sequences
[0162] In the case of the transfer matrix, the prior x of the discrete state sequence is
[0163]
[0164] Here, x1 is generated by a stable distribution under π.
[0165] Priors on autoregressive parameters
[0166] Sample the autoregressive parameters for each state i=1,2,... from the Matrix Normal Inverse-Wishart prior yes:
[0167] (A,b),∑~MNIW(v0,S0,M0,K0)
[0168] or equivalently
[0169] ∑~InvWishart(v0,S0)
[0170]
[0171] in represents the Kronecker product, and (A,b) represents the matrix formed by appending b as a column to A. In addition, a block ARD prior with respect to K0 is used to encourage the uninformative lag to shrink to zero:
[0172] K0=diag(k1,...,k KD )
[0173] Generation of principal components of 3D pose sequences
[0174] In the case of autoregressive parameters and a discrete state sequence, the data sequence y is generated according to an affine autoregression:
[0175]
[0176] in, A vector representing K lags:
[0177]
[0178] The surrogate model is a special case of AR-HMM and is constructed by adding constraints. In particular, the Gaussian output HMM (G-HMM) corresponds to the constraint A for each state index (i) = 0. Similarly, autoregressive mixtures (AR-MM) and Gaussian mixtures (GMM) correspond to restricting the transfer matrix to be constant across rows in AR-HMM and G-HMM, respectively, for i and i', π ij =π i'j =π j .
[0179] Specific implementation of the inference algorithm in the example
[0180] As described above, the Gibbs sampling inference algorithm alternates between two main phases: updating the data segments into modules while keeping the transition matrix and autoregressive parameters fixed, and updating the transition matrix and autoregressive parameters while keeping the segments fixed. Mathematically, based on [ ], updates the segments of the labeled sequence x sampled conditioned on the data y, the autoregressive parameters θ, and the values of the transition matrix π; that is, samples the conditional random variables x|θ,π,y. Similarly, in the case of segmented sampling of π|x and θ|x,y, the transition matrix and autoregressive parameters are updated, respectively.
[0181] For inference in AR-HMM, a weak limit approximation of the Dirichlet process is used, where the infinite model is approximated as a finite model. That is, some finite approximation parameters L, β, and π are chosen, and a finite Dirichlet distribution of size L is used for modeling.
[0182] β=Dir(γ / L,...,γ / L)
[0183] π k ~Dir(αβ1,...,αβ j +kδ kj ,...,αβ L )
[0184] Among them, π k Denotes the i-th row of the transition matrix. This finite representation of the transition matrix allows the state sequence x to be resampled as blocks and, for large L, provides an arbitrarily good approximation to the infinite Dirichlet process.
[0185] Using the weak limit approximation, the Gibbs sampler of the AR-HMM iteratively resamples the conditional random variable:
[0186] x|π,θ,y θ|x,y and β,π|x.
[0187] For simplicity, throughout this section, we suppress the notation for tuning hyperparameters and superscript notation for multiple observation sequences.
[0188] Sample x|π,θ,y
[0189] Sampling the state labels x given the dynamic parameters π and θ and the data y corresponds to segmenting the 3D video sequence and assigning each segment to a behavior model describing its statistics.
[0190] In the case of observation parameters θ and transfer parameters π, the hidden state sequence x is a Markov chain graph. The standard HMM reverse message passing recursion is
[0191]
[0192] For t = 1, 2, ..., T-1 and k = 1, 2, ..., K, where B T (k) = 1, and where y t+1:T =(y t+1 ,y t+2 ,...,y T ). Using these messages, all future states x 2:T The conditional distribution of the first state x1 of the marginalization is
[0193] p(x1=k|π,θ,y)∝p(x1=k|π)p(y1|x1=k,θ)B1(k)
[0194] It can be sampled effectively. In the case of , the conditional distribution of the second state x2 is
[0195]
[0196] Therefore, after passing the HMM message backward, the state sequence can be recursively sampled forward.
[0197] Sample θ|x,y
[0198] In the case of a state sequence x and a data sequence y, sampling the autoregressive parameters θ corresponds to updating the dynamic parameters of each mode to describe the 3D video data segment assigned to it.
[0199] To resample the observation parameter θ conditioned on a fixed sample of state sequences x and observations y, we can exploit the conjugation between the autoregressive likelihood and the MNIW prior. That is, the condition also follows the MNIW distribution:
[0200] p(A (k) ,∑ (k) |x,y,S0,v0,M0,K0)=p(A (k) ,∑ (k) |S n ,v n ,M n ,K _n )
[0201] Among them, (S n ,v n ,M n ,K n ) are the posterior hyperparameters, which are functions of the element y assigned to state k and the previous lagged observations:
[0202]
[0203]
[0204]
[0205] v n =v0+n
[0206] in
[0207]
[0208] n=#{t:x t =k}
[0209] Therefore, resampling θ|x,y involves three steps: collecting statistics from the data assigned to each state, forming a priori hyperparameters for each state, and updating the observed parameters for each state by simulating a graph drawn from the appropriate MNIW. n ,v n ,M n ,K n ) is performed as follows:
[0210] ∑~InvWishart(S n ,v n )
[0211] in,
[0212] Sample β,π|x
[0213] Given a state sequence x, sampling the transition parameters π and β corresponds to updating the transition probabilities between behavioral modules to reflect the transition patterns observed in the state sequence. Updating β promotes the removal of redundant behavioral patterns from the model, while updating each π ij Fit the observed transition from state i to state j.
[0214] The transfer parameters β and π extracted from the weak limiting approximation of the (viscous) HDP are resampled using an auxiliary variable sampling scheme. That is, β,\pi|x is generated by first sampling the auxiliary variable m|β,x. Then, β,\pi|x,m is generated by first sampling from the boundary β|m and then from the condition π|β,x.
[0215] The transition count matrix in the sampled state sequence x is
[0216] n kj =#{t:x t =k,x t+1 =j,t=1,2,...,T-1}.
[0217] For simplicity, the restriction condition is marked, and the auxiliary variable m={m kj :k,j=1,2,...,K} for sampling
[0218] in
[0219] where Bernoulli(p) denotes a Bernoulli random variable that takes the value 1 with probability p and 0 otherwise. Note that updates to the HDP-HMM without a sticky bias correspond to setting k=0 in these updates.
[0220] In the case of auxiliary variables, the update to β is the conjugate of the Dirichlet polynomial, where
[0221]
[0222] Where, for j = 1, 2, ..., K, The update for π|β,x is similar, with π k |β,x~Dir(αβ1+n k1 ,...,αβ j +n kj+kδ kj ,...,\alphaβ K +n kK ).
[0223] Application of the model in examples
[0224] Datasets from open field experiments, odor experiments, and genetic manipulation experiments were modeled jointly to improve statistical power. Because the neural implants associated with optogenetic experiments moderately changed the appearance of the animals, these data were modeled separately. In all experiments, the top 10 principal components were collected for each frame of each imaged mouse. The data were then divided again and assigned "training" or "test" labels at a training:test ratio of 3:1. Mice labeled "test" were retained from the training process and used to test their generalization performance by measuring the retention likelihood. This approach allows us to directly compare algorithms with compositions that reflect different behavioral infrastructures.
[0225] We trained models on the data using the procedure described in this paper; the models were robust to initialization settings and parameter and hyperparameter settings (except for κ, see below). Specifically, we found that the number of lags used in our AR observation distribution and the number of states used in our transition matrix with the HDP prior were robust to specific hyperparameter settings for both priors. We varied the hyperparameters of the sparse ARD prior by several orders of magnitude and saw no significant changes in the retention likelihood, the number of lags used, or the number of states used. We also varied the hyperparameters of our HDP prior by several orders of magnitude and again observed no changes in the number of states used or the retention likelihood. All jointly trained data shared the observation distribution, but each processing level had its own transition matrix. Each model was updated via 1000 iterations of Gibbs sampling; the model output was saved at the final iteration of Gibbs sampling; all further analyses were performed on this final update.
[0226] The “stickiness” of the duration distribution of our behavior modules (as defined by the model’s κ setting) affects the average duration of the behavior modules found by the AR-HMM; this allows us to control the time scale over which the behavior is modeled. As discussed in the main text, the autocorrelation, power spectral density, and change-point algorithms identify switching dynamics at a specific sub-second time scale (as encapsulated by the change-point duration distribution and reflected by the spectrogram and autocorrelation plot). Therefore, we empirically set the κ stickiness parameter of the time series model to optimally match the duration distribution found by change-point detection. To find the κ setting that best matches these distributions, we perform a dense grid search to minimize the Kolmogorov-Smirnov distance between the distribution of inter-change-point intervals and the posterior behavior module duration distribution.
[0227] Mouse strains, housing, and habits
[0228] Unless otherwise stated, all experiments were performed on 6-8 week old C57 / BL6 males (Jackson Laboratory). Mice from the rorβ and rbp4 strains were habituated and tested in an equivalent manner to the reference C57 / BL6 mice. Mice were brought into our colony at 4 weeks of age, where they were housed for two weeks in a reverse 12-hour light / 12-hour dark cycle. On the day of testing, mice were placed in a light-tight container in the laboratory, where they were habituated for 30 minutes in the dark before testing.
[0229] Example 1: Behavioral Assay: Innate Exploration
[0230] To address these possibilities, we first used AR-HMM to define the baseline architecture of mouse exploratory behavior in the open field and then explored how this behavioral template could be modified by different manipulations of the external world.
[0231] For the open field assay (OFA), mice were habituated as described above and then placed in the middle of an 18" diameter circular enclosure with 15" high walls (Plastics, USA) immediately prior to the start of 3D video recording. Animals were allowed to freely explore the enclosure during the 30-minute experimental session. Mice whose behavior was assessed in the square box were handled and measured in an identical manner to the OFA, except in the odor box described below.
[0232] AR-HMM identified ~60 reliably used behavioral modules (51 modules explained 95% of the imaging frames, and 65 modules explained 99% of the imaging frames, Figures 5A, 5B) from the circular open field dataset, which represents the exploratory behavior of normal mice in the laboratory ( Figure 6A , n = 25 animals, 20 min trial). Figure 5A shows the proportion of frames explained by each module (Y-axis), plotted relative to the module group and sorted by usage (X-axis). 95% of the frames were explained by 51 behavioral modules; 99% of the frames were explained by 62 behavioral modules in the open field dataset.
[0233] Figure 5B shows the modules (X-axis) sorted by usage (Y-axis) and indicates the Bayesian credible intervals. Note that all credible intervals are smaller than the SE calculated based on the bootstrap estimate (Figure 5B). As mentioned above, many of these modules encode behavioral components that can be described by humans (e.g., upright, walking, hesitating, turning).
[0234] The AR-HMM also measures the probability that any given module precedes or follows any other module; in other words, after model training, each module is assigned pairwise transition probabilities with every other module in the group; these probabilities summarize the sequence of modules expressed by the mouse during behavior. Plotting these transition probabilities as a matrix reveals that they are highly heterogeneous, with each module preferentially temporally connected to some modules and not to others ( Figure 6B The average node degree without thresholding was 16.82 ± 0.95, while after thresholding, the probability of a binary model was below 5%, 4.08 ± 0.10. This specific connectivity between module pairs constrains the module sequences observed in the dataset (8,900 out of ~125,000 possible trigrams), suggesting that certain module sequences are favored; this observation suggests that the mice's behavior is predictable, as knowing what the mouse is doing at any given moment provides an observer with an idea of what the mouse might do next. Information-theoretic analysis of the transition matrix confirmed the remarkable predictability of the mice's behavior, as the average per-frame entropy rate was lower relative to a uniform transition matrix (3.78 ± 0.03 bits without self-transitions, 0.72 ± 0.01 bits with self-transitions, and 6.022 bits with self-transitions), and the average mutual information between interconnected modules was significantly greater than zero (1.92 ± 0.02 bits without self-transitions, 4.84 ± 0.03 bits with self-transitions). This deterministic quality of mass may help ensure that mice emit coherent movement patterns; consistent with this possibility, frequently observed sequences of modules were found to encode different aspects of exploratory behavior upon examination.
[0235] The behaviors expressed by mice in the circular open field reflect scene-specific locomotor exploration patterns. We hypothesized that mice would adapt to changes in apparatus shape by locally altering behavioral structure to generate new postural dynamics to interact with specific physical features of the context; to test this hypothesis, we imaged mice in a smaller square box and then co-trained a model with the circular open field data and the square data, enabling a direct comparison of modules and transfers across the two conditions (n=25 mice in each case). While mice tended to explore the corners of the square box and the walls of the circular open field, the overall usage of most modules was similar between these apparatuses, consistent with exploratory behaviors in the locomotion arena sharing a single common feature ( Figure 6C AR-HMM also identified a small number of behavioral modules that were widely deployed in one context but negligible or completely absent in another, consistent with the idea that different physical contexts drive the expression of new behavioral modules ( Figure 6C Based on bootstrap estimates, all usage rate differences discussed below have p < 10 -3 ).
[0236] Interestingly, these “new” modules were deployed not only during physical interactions with specific features of the device (which would be predicted to induce novel postural dynamics), but also during periods of unconstrained exploration. For example, a circular arena-specific module encoded tactile behavior, in which mice moved near the walls of the arena with their body posture matching the curvature of the walls. This module was also expressed when the mice were near the center of the circular arena and without physical contact with the walls, demonstrating that the expression of this module was not simply a direct result of physical interaction with the walls but rather reflected the mice’s behavioral state in the curved arena. Although tactile behavior also occurred in the square box, the associated behavioral module encoded movement using a straight body and was used during straight trajectories in both the square and circular devices (Figures 6D to 6E, middle panels). Similarly, within the square box, mice expressed a context-specific module that encoded rapid movements from the center of the square to one of the adjacent corners; this movement pattern is likely a consequence of the small central open space of the square field, rather than a specific product of the physical constraints imposed on the mice.
[0237] It was found that some additional modules were preferentially expressed in one context or the other, and these upregulated modules seemed to encode behaviors deployed in a polycentric pattern specified by the shape of the playground. In the circular playground, for example, mice preferentially reared, where they paused near the center of the field with their bodies pointing outward; whereas in the smaller square box, mice preferentially reared up in the corners of the box (Figure 6E, data not shown). These results suggest that the behavior of mice (i.e., egocentric behavior) is regulated by the mouse's position in space (i.e., its center position). Taken together, these data suggest that mice adapt to new physical contexts, at least in part, by adding a limited set of context-specific behavioral modules (which encode postural dynamics appropriate to the context) to the baseline action pattern; these new modules (as well as other modules that are enriched in one context or the other) are differentially deployed in space to reflect changes in context.
[0238] Example 2. Behavioral Assay: Stimulus-Driven Innate Behavior - Response to Odor
[0239] Because mice exhibited the same underlying behavioral state (locomotor exploration) in both the circle and the square, we predicted that the changes in behavioral modules observed in this setting would be local and limited in scope. We therefore explored how the underlying structure of behavior changes when mice are exposed to sensory memories that drive global changes in behavioral states (within an otherwise constant physical environment), including the performance of novel and motivated actions.
[0240] To assess innate behavioral responses to volatile odors, we developed an odor delivery system that spatially isolates odors in specific quadrants of a square box. Each 12" x 12" box is constructed from 1 / 4" black matte acrylic (Altech Plastics) with 3 / 4" holes formed in a cross pattern at the bottom of the box and a 1 / 16" thick glass lid (Tru Vue). These holes are penetrated by PTFE tubing and connected to a vacuum manifold (Sigma Aldrich), which provides negative pressure to isolate the odor within the quadrant. Odors are injected into the box through a 1 / 2" NPT to 3 / 8" fitting (Cole-Parmer). Filtered air (1.0 L / min) is blown onto odor-soaked blotting paper (VWR) placed at the bottom of a Vacutainer syringe vial (Covidien). The odor-laden airflow then flows through corrugated PTFE tubing (Zeus) into one of four fittings at the corner of the odor box.
[0241] We validated the ability of the odor box to separate odors within designated quadrants by illuminating the box with a low-power handheld helium-neon laser to visualize vaporized odors or smoke. This approach enabled us to adjust the vacuum flow and odor flow to achieve odor separation, which was validated using a photoionization device (Aurora Scientific). To eliminate the possibility of cross-contamination between experiments, the odor box was soaked in a 1% Alconox solution (Aurora Scientific) overnight and then thoroughly cleaned with a 70% ethanol solution. Mice were habituated to the experimental chamber for 30 minutes before the start of the experiment. Under control conditions, dipropylene glycol in air (1.0 L / min) was delivered to each corner of the apparatus. A mouse was then placed in the center of the box and allowed to freely explore while 20 minutes of 3D video recording was acquired. The same group of animals were tested for odor response, and the experiment was subsequently repeated with odorized air delivered to one of the four quadrants. All 3D video recordings were performed in complete darkness. TMT was obtained from Phertech and used at a concentration of 5%.
[0242] Thus, mice exploring a box were exposed to the aversive fox odor, trimethylthiazoline (TMT), delivered to one quadrant of the box via an odorometer. This odor triggered complex and profound behavioral state changes, including odor detection, escape, and freezing behavior accompanied by increased levels of corticosteroids and endogenous opioids. Consistent with these known effects, mice sniffed the quadrant containing the odor and then avoided the quadrant containing the predator cues, and displayed prolonged periods of immobility (traditionally described as freezing behavior) ( Figure 7 ). Figure 7Shown are histograms depicting the average speed of "freezing" of modules that are differentially upregulated and interconnected after TMT exposure compared to all other modules in the dataset.
[0243] Surprisingly, this new set of behaviors was encoded by the same set of behavioral modules exhibited during normal exploration; several modules were upregulated or downregulated after TMT exposure, but no new modules were introduced or eliminated relative to controls (n = 25 animals in control conditions, n = 15 animals under TMT, with models trained on both datasets simultaneously). Instead, TMT altered the usage rate and transition probability between specific modules, resulting in re-preferred behavioral sequences encoding TMT-modulated behaviors (all usage rate and transition differences discussed below were P < 10 based on bootstrap estimates). -3 ).
[0244] Mapping the altered module shifts following TMT exposure defined two neighborhoods in the behavioral state map; the first neighborhood included modules moderately downregulated by TMT and an expanded set of interconnected modules, while the second neighborhood included modules upregulated by TMT and a clustered set of shifts. During normal behavior, these recently interconnected modules temporarily dispersed and appeared individually to encode distinct morphological forms of hesitation or balling. In contrast, under the influence of TMT, these modules were linked into new sequences that were found to encode freezing behavior when examined and quantified (mean sequence velocity was -0.14 ± 0.54 mm / s, compared to 34.7 ± 53 mm / s for other modules). For example, the most frequently expressed freezing ternary model was expressed 716 times after TMT exposure (over 300 minutes of imaging), compared to only 17 times under control conditions (over 480 minutes of imaging). The TMT-induced neighborhood structure imposed on these hesitant modules to induce freezing demonstrated that behavior could be modified through aggregated changes in transition probabilities. This local rewriting of transition probabilities was accompanied by an increase in the overall determinism of the mice's behavior—their overall behavioral patterns became more predictable due to TMT exposure (the entropy rate per frame decreased from 3.92 ± 0.02 bits to 3.66 ± 0.08 bits in the absence of self-transitions, and from 0.82 ± 0.01 bits to 0.64 ± 0.02 bits in the presence of self-transitions)—consistent with the mice executing a deterministic avoidance strategy.
[0245] Proximity to the odor source also governed the expression patterns of specific behavioral modules (Figures 8D to 8E). For example, a set of modules associated with freezing tended to be expressed in the quadrant farthest from the odor source, whereas the expression specificity of the detection upright module (whose overall usage was not altered by TMT) was enriched within the odor quadrant (Figures 8D to 8E). Together, these findings suggest two additional mechanisms by which the mouse nervous system can generate new adaptive behaviors. First, the transition structure between individual modules that are typically associated with different behavioral states (e.g., motor exploration) can be altered to generate new behaviors (e.g., freezing). Second, the spatial pattern of deployment of pre-existing modules and sequences can be adjusted to support motivated behaviors such as odor detection and avoidance. Thus, behavioral modules are not reused over time, but rather serve as flexibly interconnected components of behavioral sequences, the performance of which is dynamically regulated in time and space.
[0246] Example 3. Effects of genes and neural circuits on modules
[0247] As described above, the fine-scale timescale structure of behavior is selectively susceptible to changes in the physical or sensory environment that affect actions on timescales of minutes. Furthermore, the AR-HMM comprehensively encapsulated the behavioral patterns exhibited by mice (within the limitations of our imaging). These observations suggest that the AR-HMM, which provides a systematic window into mouse behavior at subsecond timescales, can both quantify overt behavioral phenotypes and reveal novel or subtle phenotypes induced after experimental manipulations that affect behavior across a range of spatiotemporal scales.
[0248] To explore how changes in individual genes, acting on a timescale spanning the mouse lifespan, affect rapid behavioral modules and shifts, we characterized the phenotype of a mouse mutant for the retinol-related orphan receptor 1β (Ror1β) gene, which is expressed in neurons of the brain and spinal cord; we chose this mouse for analysis because homozygous mutant animals exhibit the abnormal gait that we hoped to detect using AR-HMM. [37-40]. After imaging and modeling, it was found that littermate control mice were almost indistinguishable from fully inbred C57 / BL6 mice, while mutant mice expressed a unique behavioral module encoding a staggering gait (Figures 9A, 9C). This behavioral change was accompanied by the opposite behavior: the expression of five behavioral modules that encode normal forward movement at different speeds in wild-type and C57 mice was downregulated in Ror1 mutants (Figure 9A, average speed between modules = 114.6 ± 76.3 mm / sec). In addition, the expression of a group of four modules that encode brief hesitations and nodding was also upregulated (Figure 9A, average speed during the module = 8.8 ± 25.3 mm / s); this hesitation phenotype has not been previously reported in the literature. Interestingly, heterozygous mice (which do not have the reported phenotype) 37-40 ) appeared normal and exhibited wild-type wheel-running behavior40), and heterozygous mice were found to express a fully penetrant mutant phenotype: they overexpressed the same set of hesitant modules that were upregulated in the complete Ror1β mutant and failed to express the more dramatic staggering phenotype (Fig. 9A).
[0249] Thus, the AR-HMM describes the pathological behavior of Ror 1β mice as a combination of increased expression of a single, novel variant staggering module and a small set of physiological modules encoding hesitant behaviors; heterozygous mice express a defined subset of these behavioral abnormalities with penetrance that is not intermediate but equal to that observed in mutants. These results demonstrate that the sensitivity of the AR-HMM enables separation of severe and subtle behavioral abnormalities within the same litter, enables discovery of novel phenotypes, and facilitates comparisons between genotypes. These experiments also demonstrate that genotype-dependent variation in behavior (i.e., the result of indelible and lifelong changes in specific genes in the genome) can influence module expression and transfer statistics operating on millisecond timescales.
[0250] Example 4. Behavioral Assay: Optogenetics-Neuromotor Effects on Modules
[0251] Finally, we wanted to investigate whether the behavioral structure captured by AR-HMMs could provide insights into transient or unreliable behavioral changes. Therefore, we briefly triggered neural activity in motor circuits and investigated how stimulation at different intensity levels affected the immediate organization of behavior. 41-42 The optically-gated ion channel Channelrhodopsin-2 was unilaterally expressed in mice and behavioral responses were assessed before, during, and after 2 s of light-mediated activation of the motor cortex (n=4 mice to train the model separately from previous experiments).
[0252] Four adult male Rbp4-Cre (Jackson Laboratory) mice were anesthetized with 1.5% isoflurane and placed in a stereotaxic frame (Leica). A microinjection pipette (OD 10-15 μm) was inserted into the left motor cortex (Bregma coordinates: 0.5AP, -1ML, 0.60DV). Each mouse was injected with 0.5 μL of AAV5.EF1a.DIO.hChR2 (H134R) -eYFP.WPRE.hGH (~10 12 Infectious units / mL, Penn Vector Core), in the additional 10 minutes thereafter, the virus particles are allowed to spread from the injection site. After injection, a bare optical fiber with a zirconium oxide ceramic pin (OD200μm, a numerical aperture of 0.37) is inserted 100μm above the injection site and fixed to the skull using acrylic cement (LANG). Within 28 days after viral injection, the mice were placed in a circular sports arena, and the optical implant was connected to a laser pump (488nm, CrystaLaser) via a patch tape and a rotary joint (Doric Lenses). The laser was directly controlled by a PC. After 20 minutes of familiarity with the sports arena, light stimulation was started. Laser power, pulse width, pulse interval, and inter-string interval were controlled by custom software (NI Labview software). Each string of laser pulses consisted of 30 pulses (pulse width: 50ms) at 15Hz. The interval between consecutive strings was set to 18 seconds. For each laser intensity, 50 strings of laser were emitted. During the experiment, the animals were gradually exposed to higher laser intensities.
[0253] At the lowest power level, no light-induced behavioral changes were observed, whereas at the highest power level, the AR-HMM identified two behavioral modules whose expression was reliably induced by light (Figure 10A). Neither of these modules is expressed during normal mouse locomotion; inspection revealed that they encode two forms of rotational behavior (which differed in length and angle of rotation) in which mice traced out a semicircle or donut in space (Figure 10B). Although the induction of novel behavioral variants following intense unilateral motor cortex stimulation is not surprising, it is noteworthy that the AR-HMM not only identified these behaviors as novel but also encapsulated them into two distinct behavioral modules. However, we noted that approximately 40% of the time, the overall behavioral pattern did not return to baseline within a few seconds after light onset. This deviation from baseline was not due to persistent expression of the module triggered at light onset; instead, mice frequently exhibited hesitation modules (average velocity during the module = 0.8 ± 7 mm / sec) upon light onset, as if "resetting" after involuntary movements.
[0254] The behavioral changes induced by high-intensity optogenetic stimulation are reliable because the animals emit one of the two spinning modules in essentially every experiment. We then explored whether the sensitivity of the AR-HMM could quantitatively analyze more subtle changes in behavior, such as those occurring in the intermediate mechanisms of motor cortex stimulation that cause unreliable emission of specific behavioral modules. Therefore, we reduced the level of light stimulation until one of the two new morphological behavioral modules was no longer detected, while the other was only expressed in 25% of the trials. Surprisingly, we could detect an upregulation of the second set of behavioral modules, each of which was expressed 25% of the time (Figure 10A). These modules are not new, but are typically expressed during physiological exploration and encode turning and nodding behaviors (data not shown). Although each of these individual light-regulated modules is unreliably emitted, overall, behavioral changes in all modules indicate that lower levels of neural activation reliably affect behavior, but mainly by inducing physiological rather than new behaviors (Figure 10A). In summary, the examination of stimulus-locked induction of behavioral modules and the persistence of stimulus module usage rates demonstrate that neurally induced behavioral changes can influence the subsecond structure of behavior. Furthermore, the identification of a set of physiologically expressed light-modulated behavioral modules, whose induction is not evident under strong stimulation conditions, demonstrates that AR-HMMs can reveal subtle relationships between neural circuits and the temporal structure of behavior.
[0255] Example 5: Dimensionality Reduction - Probabilistic Graphical Models and Variational Autoencoders
[0256] like Figure 3 As shown, after the direction of the correction image, method can be utilized and the dimension of data is reduced.For example, each image can be a 900-dimensional vector, so reducing dimension is very important for model analysis.In certain embodiments, the two information comprising model-free algorithm 320 or model fitting 315 algorithms obtained in each pixel are normally highly correlated (adjacent pixels) or information-free (the pixel on the image border never represents the body of mice).In order to reduce redundant dimension and make modeling computationally easy to handle, various techniques can be adopted to reduce the dimension of each image.
[0257] In some examples, in some embodiments, the output of the orientation-corrected image will be a principal component analysis time series 310 or other statistical method for reducing data points. However, PCA reduces the dimensionality to a linear space. The inventors have found that reducing the dimensionality to a linear space cannot accommodate various variations in mice that are unrelated to behavior. This includes variations in mouse size, mouse breed, etc.
[0258] Therefore, the inventors have discovered that using certain neural networks, such as multilayer perceptrons, can effectively reduce the dimensionality of images. Furthermore, these reduced-dimensional images provide an effective method for developing models that are independent of the size of mice or other animals and can explain other variations unrelated to behavior. For example, neural networks can be used that reduce the dimensionality to a ten-dimensional image manifold.
[0259] The inventors developed a new unsupervised learning framework that combines probabilistic graphical models with deep learning methods, combining their respective strengths for dimensionality reduction. Their approach uses image models to represent structured probability distributions and recent advances in deep learning to learn flexible feature models and bottom-up recognition networks. All components of these models are learned simultaneously using a single objective, resulting in the following scalable fitting algorithm that leverages natural gradient stochastic variational inference, graphical model message passing, and backpropagation with reparameterization techniques.
[0260] Unsupervised probabilistic modeling typically has two goals: first, to learn models flexible enough to represent complex, high-dimensional data, such as images or speech recordings; and second, to learn interpretable model structures that admit meaningful priors and generalize to new tasks. That is, simply learning the probability density of the data is often insufficient: one also wants to learn meaningful representations. Probabilistic graphical models (Koller & Friedman, 2009; Murphy, 2012) provide many tools for constructing such structured representations, but their capacity can be limited and may require extensive feature engineering before application to data. Alternatively, advances in deep learning have provided not only flexible and scalable generative models for complex data such as images, but also new techniques for automatic feature learning and bottom-up inference (Kingma & Welling, 2014; Rezende et al., 2014).
[0261] Considering the problem of learning unsupervised generative models for tracking depth videos of freely behaving mice, e.g. Figure 23 As shown. Learning interpretable representations of this data and studying how these representations change as the animal's genetics are edited or as its brain chemistry changes can create powerful behavioral characterization tools for neuroscience and high-throughput drug discovery (Wiltschko et al., 2015). Each frame of the video is a depth image of the mouse in a specific pose, so even though each image is encoded as 30 x 30 = 900 pixels, the data lies on a low-dimensional nonlinear manifold. A good generative model must not only learn this manifold, but also represent many other salient aspects of the data.
[0262] For example, from one frame to the next, corresponding manifold points should be close to each other, and indeed, trajectories along the manifold can follow very structured dynamics. To inform the structure of these dynamics, a natural assumption used in behavior and neurobiology (Wiltschko et al., 2015) is that mouse behavior consists of short, repetitive movements, such as galloping, rearing, and grooming bouts. Therefore, a natural representation would consist of discrete states, where each state captures the simple dynamics of a specific primitive action, a representation that would be difficult to encode in an unsupervised recurrent neural network model.
[0263] The two tasks of learning image manifolds and learning structured dynamical models are complementary: we want to learn the image manifold not just as a set, but also in terms of manifold coordinates, and the structured dynamical model fits the data well. Similar modeling challenges arise in speech (Hinton et al., 2012), where high-dimensional data lie near a low-dimensional manifold because they are generated by physical systems with relatively few degrees of freedom (Deng, 1999), but also include discrete latent dynamical structures of phonemes, words, and grammar (Deng, 2004).
[0264] To address these challenges, the inventors developed a graphical model for representing structured probability distributions and used ideas from variational autoencoders (Kingma & Welling, 2014) to learn not only nonlinear feature manifolds but also bottom-up recognition networks to improve inference. Thus, this approach enables the combination of flexible deep learning feature models with structured Bayesian priors (including non-parametric models).
[0265] This approach yields a single variational inference objective in which all components of the model are learned simultaneously. Furthermore, we develop a scalable fitting algorithm that combines several advances in efficient inference, including stochastic variational inference (Hoffman et al., 2013), graphical model message passing (Roller & Friedman, 2009), and backpropagation with reparameterization (Kingma & Welling, 2014). Consequently, our algorithm efficiently computes the natural gradients of a few variational parameters, exploiting their conjugate exponential family structure, enabling efficient second-order optimization (Martens, 2015), while using backpropagation to compute the gradients of all other parameters. The general approach is referred to as a structured variational autoencoder (SVAE). This SVAE is illustrated here using a graphical model based on switched linear dynamic systems (SLDS) (Murphy, 2012; Fox et al., 2011).
[0266] Natural Gradient Stochastic Variational Inference
[0267] Stochastic variational inference (SVI) (Hoffman et al., 2013) applies stochastic gradient ascent to the mean field variational inference objective by exploiting the exponential family conjugation to efficiently compute the natural gradient (Amari, 1998; Martens, 2015). Consider a model consisting of global latent variables and local latent variables, where θ and the local latent variables satisfy
[0268] and observed data
[0269]
[0270] Where ρ(θ) is the exponential family p(x n ,y n |θ) before the natural exponential family conjugation,
[0271] ln p(θ)=<η θ , t θ (θ)>-ln Z θ (η θ ) (2)
[0272] ln p(x n ,y n |θ)=<η xy (θ), t xy (x n ,y n )>-ln Z xy (η xy (θ)
[0273] = <t θ (θ), (t xy (x n ,y n ), 1)>. (3)
[0274] Consider the mean field family q(θ)q(x)=q(θ)Πn q(xn). Due to the conjugated structure, the optimal global mean field factor q(θ) is in the same family as the previous ρ(θ),
[0275]
[0276] Then, the mean-field objective of the global variational parameter for optimizing the local variational factor q(x) can be written as
[0277]
[0278] Furthermore, the natural gradient of the objective (5) decomposes into a sum of local expected sufficient statistics (I Hoffman et al., 2013):
[0279]
[0280] Among them, q*(x n ) is the local optimal local mean field factor. Therefore, we can calculate the local mean field factor by n Sampling, local mean field factor q(x n ) and compute proportional estimates with sufficient statistics to compute stochastic natural gradient updates for our global mean field objective.
[0281] 2.2. Variational Autoencoder
[0282] Variational autoencoders (VAE) (Kingma & Welling, 2014; Rezende et al., 2014) are recently proposed models and inference methods that combine neural network autoencoders (Vincent et al., 2008) with mean field variational Bayesian. Given a high-dimensional dataset (e.g., a collection of images), VAE trains the model based on a low-dimensional latent variable y. n and a nonlinear observation model with the following parameters for each observation y n To model:
[0283]
[0284]
[0285] in
[0286] h l (x n )=f(W l h l-1 (x n )+b l ), l=1, 2,...,L, (9)
[0287]
[0288]
[0289]
[0290] Because we will reuse this particular MLP structure, we introduce the following notation
[0291]
[0292] To approximate the posterior, the variational autoencoder uses the mean field family:
[0293]
[0294]
[0295] Figure 2. Graphical model of a variational autoencoder
[0296] The key insight of the variational autoencoder is to use the conditional variational density q(x n |y n ), where the parameters of the variational distribution depend on the corresponding data points. In particular, we can write q(x n \y n ) are set to n(y n \(p) and E(yrl;), where
[0297] (μ(y n ;φ),∑(y n ;φ))=MLP(y n ;φ) (15)
[0298] represents a set of MLP parameters. Therefore, the variational distribution q(x n |ij n ) acts like a random encoder, from the observed value to the distribution of the latent variable, and the forward model p(y n |x n ) acts like a stochastic decoder from latent variable values to observation distributions.
[0299] The resulting mean field objective represents a variational Bayesian version of the autoencoder. The variational parameters are the encoder parameters and the decoder parameters, and the objective is
[0300]
[0301] To optimize this objective efficiently, Kingma & Welling (2014) applied a reparameterization trick. To simplify notation and computation, first, we rewrite the objective as
[0302]
[0303] The term KL(q(x\y)\\p(x)) is the KL divergence between two Gaussians, and its gradient with respect to f can be computed in closed form. The stochastic gradient of the expectation term is computed, since the random variable can be parameterized as
[0304]
[0305] The expectation term can be rewritten using the Monte Carlo over approximation of the gradient,
[0306]
[0307] These gradient terms can also be computed using standard backpropagation. For scalability, the sum of the data points can also be approximated via Monte Carlo.
[0308] Generative Models and Variational Families
[0309] Therefore, based on these algorithms, the inventors developed a generative model for SVAE and a corresponding variational family. Specifically, we focus on a specific generative model for time series based on switched linear dynamic systems (SLDS) (Murphy, 2012; Fox et al., 2011), which illustrates how SVAE can combine discrete latent variables with continuous latent variables with rich probabilistic dependencies.
[0310] The approach presented here is applicable to a wide range of probabilistic graphical models and is not limited to time series. First, Section 3.1 presents a generative model that illustrates the combination of a graphical model representing latent structure with a flexible neural network to generate observations. Next, Section 3.2 presents a structured variational family that exploits the mean field approximation of structured graphs and a flexible recognition network.
[0311] 3.1. Switched Linear Dynamic Systems with Nonlinear Observation
[0312] Switching Linear Dynamical Systems (SLDS) represent data with continuous hidden states evolving according to a set of discrete linear dynamics. At each time instant, there is a discrete-valued hidden state that indicates the dynamical pattern, and a continuous-valued hidden state that evolves according to the linear Gaussian dynamics of that pattern:
[0313]
[0314] Figure 24 The graphical model for the SLDS generative model and the corresponding structured CRF variational family are shown.
[0315] The discrete hidden state evolves according to Markov dynamics,
[0316]
[0317] Generate the initial state respectively:
[0318] z1|π init ~π init , (18)
[0319]
[0320] Therefore, we reasoned about the latent variables and parameters of SLDS. In addition to the Markov switching between different linear dynamics, we also identified a set of recurring dynamic patterns, each of which is described as a linear dynamic system on the hidden state. The dynamic parameters can be expressed as:
[0321]
[0322] At each time, successive hidden states produce conditional Gaussian observations.
[0323]
[0324] In a typical SLDS (Fox et al., 2011), it can be written as
[0325]
[0326] However, to achieve flexible modeling of images and other complex features, the algorithm can allow the dependencies to be more general nonlinear models. In particular, we consider the following equation:
[0327]
[0328] Note that by construction, the density is in the exponential family. We can choose the prior p(0) as a natural exponential family conjugate prior, writing
[0329] ln p(θ)=<η θ , t θ (θ)>-ln Z θ (η θ ) (twenty three)
[0330] ln p(z, x|θ)=<η zx (θ), t zx (z,x)>-ln Z zx (η zx (θ)
[0331] = <t θ (θ), (t zx (z, x), 1)>. (24)
[0332] We can also use Bayesian nonparametric priors and generate discrete state sequences based on the Hierarchical Dirichlet Process (HDP) HMM (Fox et al., 2011). Although the Bayesian nonparametric case is not discussed further, the algorithm developed here is immediately extended to the HDP-HMM using the method in Johnson & Willsky (2014).
[0333] As a special case, this structure contains the generative model of the above VAE. Specifically, the VAE uses the same kind of MLP observation model, but each latent value x t are modeled as independent and identically distributed Gaussians, while the SVAE model proposed here allows for rich joint probabilistic structure. SLDS generative models also include Gaussian mixture models (GMMs), Gaussian emission discrete-state HMMs (G-HMMs), and Gaussian linear dynamic systems (LDSs) as special cases, so the algorithms developed here for SLDS are directly specialized for these models.
[0334] While using conditional linear dynamics within each state may seem limited, the flexible nonlinear observation distribution greatly expands the capacity of these models. Indeed, recent work on neural word embeddings (Mikolov et al., 2013) and neural image models (Radford et al., 2015) has demonstrated the ability to learn latent spaces where linear structure corresponds to meaningful semantics.
[0335] For example, addition and subtraction of word vectors can correspond to semantic relations between words, and translations in the latent space of an image model can correspond to rotations of objects. Thus, linear models in learned latent spaces can yield remarkable expressive power while enabling fast probabilistic inference, interpretable priors and parameters, and a host of other tools. In particular, linear dynamics allows one to learn or encode information about timescales and frequencies: the eigenvalue spectrum of each transition matrix A(k) directly represents its characteristic timescale, allowing us to control and interpret the structure of linear dynamics in ways that nonlinear dynamical models do not allow.
[0336] 3.2. Variational family and CRF recognition network
[0337] Here we describe a family of structured mean fields that enable variational inference on the posterior distribution of the generative model from Section 3.1. This family of mean fields demonstrates that SVAEs can not only exploit graphical models and exponential family structures, but also learn bottom-up inference networks. As shown below, these structures allow us to combine several effective inference algorithms including SVI, information passing, backpropagation, and reparameterization techniques.
[0338] In mean-field variational inference, one constructs tractable variational families by breaking dependencies in the posterior (Wainwright & Jordan, 2008). To construct a structured mean-field family for the generative model developed in Section 3.1, one can break the posterior dependencies between the dynamical parameters θ, the observation parameters, the discrete state sequence, and the continuous state sequence, and write the corresponding decomposition density as
[0339]
[0340] Note that this family of structured mean fields does not destroy dependencies between discrete states or between continuous states as in the naive mean field model, since these random variables are highly correlated in the posterior. By preserving joint dependencies across time, these structured factors provide a more accurate representation of the posterior while still allowing tractable inference via graphical model message passing (Wainwright & Jordan, 2008).
[0341] To exploit the bottom-up inference network, the factors can be parameterized as conditional random fields (CRFs) (Murphy, 2012). That is, using the fact that the optimal factor in the chain graph is Markov, we write it as the pair potential and node potential terms
[0342]
[0343] where the node potential is a function of the observations. Specifically, using the notation from Section 2.2, we choose each node potential to be a Gaussian factor, where the precision matrix and potential vector depend on the corresponding observations made through the MLP,
[0344]
[0345] These local recognition networks allow one to fit a regression of probabilistic guesses from each observation to the corresponding hidden state. Using graphical model inference, these local guesses can be synthesized with the dynamics model into a coherent joint factor over the entire state sequence.
[0346] This family of structured mean fields can be directly compared to the family of fully factorized models used in the variational autoencoders described above. That is, there is no graph structure between the latent variables of the VAE. SVAE generalizes VAEs by allowing the output of the recognition network to be arbitrary potentials in the graphical model (e.g., the node potentials considered here). Moreover, in SVAEs, some of the graphical model potentials are caused by the probabilistic model rather than the output of the recognition network; for example, the best pairwise potentials are caused by the variational factors of the dynamical parameters and the latent discrete states, as well as the forward generative model (see Section 4.2.1). Thus, SVAEs provide a way to combine bottom-up information from flexible inference networks with top-down information from other latent variables in the structured probabilistic model.
[0347] When p(θ) is chosen as a conjugate prior, as in Equation (23), the optimal factor q(θ) is in the same exponential family:
[0348]
[0349] To simplify notation, as in Section 2.2, we take the variational factors of the observation parameters to be singularly distributed, and then the mean-field objective in terms of the global variational parameters is
[0350]
[0351] where, as in Equation (5), the free parameters of the local variational factors are maximized. In Section 4, we show how to optimize this variational objective.
[0352] 4. Learning and Reasoning
[0353] This section discloses an efficient algorithm for computing the stochastic gradients of the SVAE objective in Equation (29). These stochastic gradients can be used in general optimization procedures such as stochastic gradient ascent or Adam (Kingma & Ba, 2015).
[0354] As disclosed, the SVAE algorithm is essentially a combination of SVI (Hoffman et al., 2013) and AEVB (Kingma & Welling, 2014), described in Sections 2.1 and 2.2, respectively. Leveraging SVI, the SVAE algorithm can exploit the exponential family conjugate structure, when available, to efficiently compute the natural gradient of some variational parameter. Because the natural gradient adapts to the geometry of the variational family and is invariant to model reparameterization (Amari & Nagaoka, 2007), natural gradient ascent provides an efficient second-order optimization method (Martens & Grosse, 2015; Martens, 2015). Leveraging AEVB, these algorithms are applicable to general nonlinear observation models and flexible bottom-up recognition networks.
[0355] The algorithm is divided into two parts. First, in Section 4.1, we present a general algorithm for computing the target gradient based on the results of the model-specific inference routine. Then, in Section 4.2, we present the inference routine for applying the model to SLDS.
[0356] SVAE Algorithm
[0357] Here, the results of the model inference subroutine are used to compute the stochastic gradient of the SVAE mean-field objective (29). The algorithm is summarized in Algorithm 1.
[0358] For scalability, the stochastic gradient used here is computed on a small dataset. To simplify notation, assume the dataset is a collection of N sequences, each of length T. It is possible to sample a sequence uniformly at random and compute the stochastic gradient from it. It is also possible to sample subsequences and compute stochastic gradients with controllable biases (Foti et al., 2014).
[0359] The SVAE algorithm computes natural and standard gradients. To compute these gradients, as described in Section 2.2, we split the object into
[0360]
[0361]
[0362] Note that only the second term depends on the variable dynamic parameters. Moreover, it is the KL divergence between two members of the same exponential family (Eqs. (23) and (28)), so, as in Hoffman et al. (2013) and Section 2.1, we can write the natural gradient of (30) as:
[0363]
[0364] where q(z) and q(x) are considered to be the locally optimal local mean field factors as shown in Equation (6). Therefore, by sampling the sequence index n uniformly at random, the unbiased estimate of the natural gradient is given by:
[0365]
[0366] We abbreviate it as:
[0367]
[0368] These expected sufficient statistics are computed efficiently using the model inference routine described in Section 4.2.
[0369]
[0370] \
[0371] Therefore, we must distinguish between the processes used by the model inference routine to compute these quantities. Performing this distinction efficiently for SLDS corresponds to backpropagation via message passing.
[0372] 4.2. Model Inference Subroutine
[0373] Because VAEs correspond to specialized SVAEs with finite latent probabilistic structure, the inference routine can be viewed as a generalization of the two steps in the AEVB algorithm (Kingma & Welling, 2014). However, the inference routine of an SVAE can often perform other computations: First, because SVAEs can include other latent random variables and graph structures, the inference routine can optimize local mean-field factors or perform message passing. Second, because SVAEs can perform stochastic natural gradient updates on global factors, the inference routine can also compute expected sufficient statistics.
[0374] To simplify the notation, the sequence index n can be deleted and y(n) can be used instead of y. The algorithm is summarized in Algorithm 2.
[0375] 4.2.1. Optimizing the local mean field factor
[0376] As with the SVI algorithm in Section 2.1, for a given data sequence y, we optimize the local mean-field factor. That is, for a fixed global parameter factor with natural parameters output by the recognition network and fixed node potentials, we optimize the variational objective over the local variational factors of the discrete hidden states and the local variational factors of the continuous hidden states. This optimization is performed efficiently by leveraging the SLDS exponential family table and the structured variational family.
[0377] 4.2.2. Sample, Expected Statistics, and KL
[0378] After optimizing the local variational factors, the model inference routine draws samples using the optimized factors, calculates the expected sufficient statistics, and calculates the KL divergence. The results of these inference calculations are then used to calculate the gradient of the SVAE objective.
[0379] 6. Experiment
[0380] 6.1.Bounce Points in ID
[0381] As a representative toy problem, consider a one-dimensional sequence of images where points bounce from one edge of the image to the other at a fixed speed. Figure 25 Figure 1 shows the inference results of an LDS SVAE applied to this problem. The top panel shows noisy image observations over time. The second panel shows the model inference on past and future images: conditioned on observations to the left of the vertical red line, the model is filtering, while conditioned on observations to the right of the vertical red line, the model is predicting. This figure demonstrates that the model, having learned appropriate low-dimensional representations and dynamics, is able to consistently predict the future.
[0382] One can also use point problems to illustrate the significant optimization advantage that Natural Gradient provides over variable dynamics parameters. In Figure 6, Natural Gradient updates are compared to standard gradient updates for three different learning rates. The Natural Gradient algorithm not only learns faster, but is also more stable: when Natural Gradient updates use a step size of 0.1, the standard gradient dynamics become unstable and terminate early at step sizes of 0.1 and 0.05. Although a step size of 0.01 produces stable standard gradient updates, training is several orders of magnitude slower than with the Natural Gradient algorithm.
[0383] 6.2. Mouse behavioral phenotype
[0384] The goal of behavioral phenotyping is to identify behavioral patterns and study how these patterns change when an animal's environment, genetics, or brain function changes. Here, the inventors use a 3D depth camera dataset from Wiltschko et al. (2015) to show how the SLDS SVAE can learn flexible but structured generative models for this video data.
[0385] The nonlinear observation model of VAE is key to learning the manifold of depth images of mice. Figure 7 (exist Figure 25 Figure 3 shows images corresponding to points on a random 2D grid in the latent space, illustrating how a nonlinear observation model can generate accurate images. SVAE learns this feature manifold while fitting the structured latent probabilities.
[0386] Figure 4 (exist Figure 25 The dynamic structure of some learning is illustrated in the reference (referencing the generative video completion task). The figure contains model-generated data and the corresponding real data in alternating rows. In the model-generated data, the data between the two red lines is generated without any corresponding observations, while the data outside the two red lines is generated conditionally.
[0387] in conclusion
[0388] This paper discloses a new class of models and corresponding inference algorithms that leverage probabilistic graphical models and flexible feature representations from deep learning. In the context of time series, this approach offers several new nonlinear models that can be used for inference, estimation, and even control. For example, by preserving the underlying linear structure in SVAEs, some dynamic programming control problems may remain tractable.
[0389] While this paper focuses on time series models, specifically SLDS and related models, the architecture proposed here is more general: the basic strategy of learning a flexible bottom-up inference network based on the CRF potential and then combining this bottom-up information with coherent probabilistic reasoning in structured models is likely to be relevant wherever graphical models prove useful. SVAEs also enable many other tools in probabilistic modeling to be combined with newer deep learning methods, including hierarchical modeling, structured regularization, automatic relevance determination, and easy handling of missing data.
[0390] References
[0391] 1 Fettiplace,R.&Fuchs,PAMechanisms of hair cell tuning.Annual review of physiology 61,809-834,(1999).
[0392] 2 Fettiplace,R.&Kim,K.X.The Physiology of MechanoelectricalTransduction Channels in Hearing.Physiological reviews 94,951-986,(2014).
[0393] 3 Gollisch,T.&Herz,A.M.V.Disentangling Sub-Millisecond Processeswithin an Auditory Transduction Chain.PLoS Biology 3,e8,(2005).
[0394] 4 Kawasaki,M.,Rose,G.&Heiligenberg,W.Temporal hyperacuity in singleneurons of electric fish.Nature 336,173-176,(1988).
[0395] 5 Nemenman,I.,Lewen,G.D.,Bialek,W.&de Ruyter van Steveninck,R.R.Neural Coding of Natural Stimuli:Information at Sub-MillisecondResolution.PLoS computational biology 4,e1000025,(2008).
[0396] 6 Peters,A.J.,Chen,S.X.&Komiyama,T.Emergence of reproduciblespatiotemporal activity during motor learning.Nature 510,263-267,(2014).
[0397] 7 Ritzau-Jost,A.,Delvendahl,I.,Rings,A.,Byczkowicz,N.,Harada,H.,Shigemoto,R.,Hirrlinger,J.,Eilers,J.&Hallermann,S.Ultrafast Action PotentialsMediate Kilohertz Signaling at a Central Synapse.Neuron 84,152-163,(2014).
[0398] 8 Shenoy,K.V.,Sahani,M.&Churchland,M.M.Cortical Control of ArmMovements:A Dynamical Systems Perspective.Annual review of neuroscience 36,337-359,(2013).
[0399] 9 Bargmann,C.I.Beyond the connectome:How neuromodulators shape neuralcircuits.BioEssays 34,458-465,(2012).
[0400] 10 Tinbergen,N.The study of instinct.(Clarendon Press,1951).
[0401] 11 Garrity,P.A.,Goodman,M.B.,Samuel,A.D.&Sengupta,P.Running hot andcold:behavioral strategies,neural circuits,and the molecular machinery forthermotaxis in C.elegans and Drosophila.Genes&;Development 24,2365-2382,(2010).
[0402] 12 Stephens,G.J.,Johnson-Kerner,B.,Bialek,W.&Ryu,W.S.Dimensionalityand Dynamics in the Behavior of C.elegans.PLoS computational biology 4,e1000028,(2008).
[0403] 13 Stephens,G.J.,Johnson-Kerner,B.,Bialek,W.&Ryu,W.S.From Modes toMovement in the Behavior of Caenorhabditis elegans.PLoS ONE 5,e13914,(2010).
[0404] 14 Vogelstein,J.T.,Vogelstein,J.T.,Park,Y.,Park,Y.,Ohyama,T.,Kerr,R.A.,Kerr,R.A.,Truman,J.W.,Truman,J.W.,Priebe,C.E.,Priebe,C.E.,Zlatic,M.&Zlatic,M.Discovery of brainwide neural-behavioral maps via multiscaleunsupervised structure learning.Science(New York,NY)344,386-392,(2014).
[0405] 15 Berman,G.J.,Choi,D.M.,Bialek,W.&Shaevitz,J.W.Mapping the structureof drosophilid behavior.(2013).
[0406] 16 Croll,N.A.Components and patterns in the behaviour of the nematodeCaenorhabditis elegans.Journal of zoology 176,159-176,(1975).
[0407] 17 Pierce-Shimomura,J.T.,Morse,T.M.&Lockery,S.R.The fundamental roleof pirouettes in Caenorhabditis elegans chemotaxis.Journal of Neuroscience19,9557-9569,(1999).
[0408] 18 Gray,J.M.,Hill,J.J.&Bargmann,C.I.A circuit for navigation inCaenorhabditis elegans.Proceedings of the National Academy of Sciences of theUnited States of America 102,3184-3191,(2005).
[0409] 19 Miller,A.C.,Thiele,T.R.,Faumont,S.,Moravec,M.L.&Lockery,S.R.Step-response analysis of chemotaxis in Caenorhabditis elegans.Journal ofNeuroscience 25,3369-3378,(2005).
[0410] 20 Jhuang,H.,Garrote,E.,Yu,X.,Khilnani,V.,Poggio,T.,Steele,A.D.&Serre,T.Automated home-cage behavioural phenotyping of mice.NatureCommunications 1,68,(2010).
[0411] 21 Stewart,A.,Liang,Y.,Kobla,V.&Kalueff,A.V.Towards high-throughputphenotyping of complex patterned behaviors in rodents:Focus on mouse self-grooming and its sequencing.Behavioural brain…,(2011).
[0412] 22 Ohayon,S.,Avni,O.,Taylor,A.L.,Perona,P.&Egnor,S.E.R.Automatedmulti-day tracking of marked mice for the analysis of social behavior.Journalof neuroscience methods,1-25,(2013).
[0413] 23 de Chaumont,F.,Coura,R.D.-S.,Serreau,P.,Cressant,A.,Chabout,J.,Granon,S.&Olivo-Marin,J.-C.Computerized video analysis of social interactionsin mice.Nature Methods 9,410-417,(2012).
[0414] 24 Kabra,M.,Robie,A.A.,Rivera-Alba,M.,Branson,S.&Branson,K.JAABA:interactive machine learning for automatic annotation of animalbehavior.Nature Methods 10,64-67,(2013).
[0415] 25 Weissbrod,A.,Shapiro,A.,Vasserman,G.,Edry,L.,Dayan,M.,Yitzhaky,A.,Hertzberg,L.,Feinerman,O.&Kimchi,T.Automated long-term tracking and socialbehavioural phenotyping of animal colonies within a semi-naturalenvironment.Nature Communications 4,2018,(2013).
[0416] 26 Spink,A.J.,Tegelenbosch,R.A.,Buma,M.O.&Noldus,L.P.The EthoVisionvideo tracking system--a tool for behavioral phenotyping of transgenicmice.Physiology&;behavior 73,731-744,(2001).
[0417] 27 Tort,A.B.L.,Neto,W.P.,Amaral,O.B.,Kazlauckas,V.,Souza,D.O.&Lara,D.R.A simple webcam-based approach for the measurement of rodent locomotionand other behavioural parameters.Journal of neuroscience methods 157,91-97,(2006).
[0418] 28 Gomez-Marin,A.,Partoune,N.,Stephens,G.J.,Louis,M.&Brembs,B.Automated tracking of animal posture and movement during exploration andsensory orientation behaviors.PLoS ONE 7,e41642,(2012).
[0419] 29 Colgan,P.W.Quantitative ethology.(John Wiley&;Sons,1978).
[0420] 30 Fox,E.B.,Sudderth,E.B.,Jordan,M.I.&Willsky,A.S.inProc.International Conference on Machine Learning(2008).
[0421] 31 Fox,E.B.,Sudderth,E.B.,Jordan,M.I.&Willsky,A.S.BayesianNonparametric Inference of Switching Dynamic Linear Models.IEEE Transactionson Signal Processing 59,(2011).
[0422] 32 Johnson,M.J.&Willsky,A.S.The Hierarchical Dirichlet Process HiddenSemi-Markov Model.Arxiv abs / 1203.3485,(2012).
[0423] 33 Teh,Y.W.,Jordan,M.I.&Beal,M.J.Hierarchical dirichletprocesses.Journal of the american…,(2006).
[0424] 34 Geman,S.&Geman,D.Stochastic Relaxation,Gibbs Distributions,and theBayesian Restoration of Images.IEEE Trans.Pattern Anal.Mach.Intell.6,721-741,(1984).
[0425] 35 Wallace,K.J.&Rosen,J.B.Predator odor as an unconditioned fearstimulus in rats:elicitation of freezing by trimethylthiazoline,a componentof fox feces.Behav Neurosci 114,912-922,(2000).
[0426] 36 Fendt,M.,Endres,T.,Lowry,C.A.,Apfelbach,R.&McGregor,I.S.TMT-induced autonomic and behavioral changes and the neural basis of itsprocessing.Neurosci Biobehav Rev 29,1145-1156,(2005).
[0427] 37 André,E.,Conquet,F.,Steinmayr,M.,Stratton,S.C.,Porciatti,V.&Becker-André,M.Disruption of retinoid-related orphan receptor beta changescircadian behavior,causes retinal degeneration and leads to vacillansphenotype in mice.The EMBO journal 17,3867-3877,(1998).
[0428] 38 Liu,H.,Kim,S.-Y.,Fu,Y.,Wu,X.,Ng,L.,Swaroop,A.&Forrest,D.An isoformof retinoid-related orphan receptorβdirects differentiation of retinalamacrine and horizontal interneurons.Nature Communications 4,1813,(2013).
[0429] 39 Eppig,J.T.,Blake,J.A.,Bult,C.J.,Kadin,J.A.,Richardson,J.E.&Group,M.G.D.The Mouse Genome Database(MGD):facilitating mouse as a model for humanbiology and disease.Nucleic Acids Research 43,D726-736,(2015).
[0430] 40 Masana,M.I.,Sumaya,I.C.,Becker-Andre,M.&Dubocovich,M.L.Behavioralcharacterization and modulation of circadian rhythms by light and melatoninin C3H / HeN mice homozygous for the RORbeta knockout.American journal ofphysiology.Regulatory,integrative and comparative physiology 292,R2357-2367,(2007).
[0431] 41 Glickfeld,L.L.,Andermann,M.L.,Bonin,V.&Reid,R.C.Cortico-corticalprojections in mouse visual cortex are functionally target specific.NatureNeuroscience 16,219-226,(2013).
[0432] 42 Mei,Y.&Zhang,F.Molecular tools and approaches foroptogenetics.Biological psychiatry 71,1033-1038,(2012).
[0433] 43 Lashley,K.S.(ed Lloyd A Jeffress)(Psycholinguistics:A book ofreadings,1967).
[0434] 44 Sherrington,C.The Integrative Action of the Nervous System.TheJournal of Nervous and Mental Disease,(1907).
[0435] 45 Bizzi,E.,Tresch,M.C.,Saltiel,P.&d'Avella,A.New perspectiveson spinal motor systems.Nature Reviews Neuroscience 1,101-108,(2000).
[0436] 46 Drai,D.,Benjamini,Y.&Golani,I.Statistical discrimination ofnatural modes of motion in rat exploratory behavior.Journal of neurosciencemethods 96,119-131,(2000).
[0437] 47 Brown,T.G.in Proceedings of the Royal Society of London Series B(1911).
[0438] 48 Crawley,J.N.Behavioral phenotyping of rodents.Comparative medicine53,140-146,(2003).
[0439] 49 Anderson,D.J.&Perona,P.Toward a science of computationalethology.Neuron 84,18-31,(2014).
[0440] 50 Berg,H.C.&Brown,D.A.Chemotaxis in Escherichia coli analysed bythree-dimensional tracking.Nature 239,500-504,(1972).
[0441] 51 Berg,H.C.Chemotaxis in bacteria.Annual review of biophysics andbioengineering 4,119-136,(1975).
[0442] 52 Berg,H.C.Bacterial behaviour.Nature 254,389-392,(1975).
[0443] 53 Hong,W.,Kim,D.-W.&Anderson,D.J.Antagonistic Control of Socialversus Repetitive Self-Grooming Behaviors by Separable Amygdala NeuronalSubsets.Cell 158,1348-1361,(2014).
[0444] 54 Lin,D.,Boyle,M.P.,Dollar,P.,Lee,H.,Lein,E.S.,Perona,P.&Anderson,D.J.Functional identification of an aggression locus in the mousehypothalamus.Nature 470,221-226,(2011).
[0445] 55 Swanson,L.W.Cerebral hemisphere regulation of motivatedbehavior.Brain research 886,113-164,(2000).
[0446] 56 Aldridge,J.W.&Berridge,K.C.Coding of serial order by neostriatalneurons:a";natural action";approach to movement sequence.The Journalof neuroscience:the official journal of the Society for Neuroscience 18,2777-2787,(1998).
[0447] 57 Aldridge,J.W.,Berridge,K.C.&Rosen,A.R.Basal ganglia neuralmechanisms of natural movement sequences.Canadian Journal of Physiology andPharmacology 82,732-739,(2004).
[0448] 58 Jin,X.,Tecuapetla,F.&Costa,R.M.Basal ganglia subcircuitsdistinctively encode the parsing and concatenation of action sequences.NaturePublishing Group 17,423-430,(2014).
[0449] 59 Tresch,M.C.&Jarc,A.The case for and against musclesynergies.Current opinion in neurobiology 19,601-607,(2009).
[0450] 60 Flash,T.&Hochner,B.Motor primitives in vertebrates andinvertebrates.Current opinion in neurobiology 15,660-666,(2005).
[0451] 61 Bizzi,E.,Cheung,V.C.K.,d'Avella,A.,Saltiel,P.&Tresch,M.Combining modules for movement.Brain Research Reviews 57,125-133,(2008).
[0452] 62 Tresch,M.C.,Saltiel,P.&Bizzi,E.The construction of movement by thespinal cord.Nature Neuroscience 2,162-167,(1999).
[0453] 63 Berwick, RC, Okanoya, K., Beckers, GJL & Bolhuis, JJ Songs to syntax: the linguistics of birdsong. Trends in cognitive sciences 15, 113-121, (2011).
[0454] 64 Wohlgemuth, MJ, Sober, SJ & Brainard, MS Linked control of syllablesequence and phonology in birdsong. Journal of Neuroscience 30, 12936-12949, (2010).
[0455] 65 Markowitz, JE, Ivie, E., Kligler, L. & Gardner, TJ Long-range Order in Canary Song. PLoS computational biology 9, e1003052, (2013).
[0456] 66 Fentress, JC & Stilwell, FPLetter: Grammar of a movement sequence in inbred mice. Nature 244, 52-53, (1973).
[0457] Selected Examples
[0458] While the above description and appended claims disclose various embodiments of the present invention, other alternative aspects of the present invention are disclosed in the following further embodiments.
[0459] 1. A method for analyzing the motion of an object to separate it into modules, the method comprising:
[0460] processing three-dimensional video data representing motion of the object using a computational model to partition the video data into at least one set of modules and at least one set of transition statistics between the modules; and
[0461] The at least one set of modules is assigned to a category representing a type of animal behavior.
[0462] 2. The method of embodiment 1, wherein the processing comprises the following steps: separating the object from the background in the video data.
[0463] 3. According to the method of embodiment 2, the processing further comprises the following steps: identifying the orientation of the features of the object across a set of frames of the video data with respect to a common coordinate system for each frame.
[0464] 4. According to the method of embodiment 3, the processing further includes the following steps: modifying the orientation of the object in at least a subset of the set of frames so that the features are oriented in the same direction relative to the coordinate system to output a set of aligned frames.
[0465] 5. According to the method of embodiment 4, the processing further includes the following steps: using principal component analysis (PCA) to process the aligned frames to output posture dynamics data, wherein the posture dynamics data represents the posture of the object through the principal component space in each aligned frame.
[0466] 6. According to the method of embodiment 4, the processing further includes the following steps: using a multi-layer perceptron (MLP) to process the aligned frames to output posture dynamics data, wherein the posture dynamics data represents the posture of the object in each aligned frame through the manifold space.
[0467] 7. According to the method of embodiment 5, the processing further includes the following steps: using a computational model to process the aligned frames to temporarily divide the posture dynamics data into different groups of modules, wherein all sub-second modules in a group of modules exhibit similar posture dynamics.
[0468] 8. The method of embodiment 7, wherein the model is a switched linear dynamic system (SLDS) model.
[0469] 9. The method of embodiment 7, wherein the multilayer perceptron is a structured variational autoencoder.
[0470] 10. The method of embodiment 6, wherein the model is trained using gradient descent and back propagation.
[0471] 11. The method of embodiment 7, wherein processing the aligned frames using the MLP occurs simultaneously with processing the frames using the computational model.
[0472] 12. The method of embodiment 5, further comprising the step of displaying a representation of each group of modules that appear in the three-dimensional video data with a frequency above a threshold.
[0473] 13. The method of embodiment 1, wherein the computational model comprises modeling the sub-second module as a vector autoregressive process, wherein the vector autoregressive process represents a stereotyped trajectory through a PCA space.
[0474] 14. The method of embodiment 1, wherein the computational model comprises using a hidden Markov model to model sub-second transition periods between modules.
[0475] 15. The method of embodiment 1, wherein the three-dimensional video data is first processed to output a series of points in a multi-dimensional vector space, wherein each of the points represents the three-dimensional posture dynamics of the object.
[0476] 16. The method of any one of embodiments 1-10, wherein the subject is an animal in an animal study.
[0477] 17. The method of any one of embodiments 1-10, wherein the subject is a human.
[0478] 18. A method for analyzing the motion of an object to separate it into modules, the method comprising:
[0479] pre-processing the three-dimensional video data representing the motion of the object to separate the object from the background;
[0480] identifying, over a set of frames of the video data, an orientation of features of the object relative to a common coordinate system for all frames;
[0481] modifying the orientation of the object in at least a subset of the set of frames so that the features are oriented in the same direction relative to the coordinate system to output a set of aligned frames;
[0482] processing the aligned frames using a multi-layer perceptron (MLP) to output pose dynamics data, wherein the pose dynamics data represents a pose of the object in each aligned frame in a three-dimensional graphics space;
[0483] processing the aligned frames to temporally segment the gesture dynamics data into different groups of sub-second modules, wherein all sub-second modules in a group of modules exhibit similar gesture dynamics; and
[0484] Representations of respective groups of modules that occur in the three-dimensional video data at a frequency above a threshold are displayed.
[0485] 19. The method of embodiment 18, wherein the step of processing the aligned frames is performed using a model-free algorithm.
[0486] 20. The method of embodiment 19, wherein the model-free algorithm includes computing an auto-correlogram.
[0487] 21. The method of embodiment 18, wherein the step of processing the aligned frames is performed using a model-based algorithm.
[0488] 22. The method of embodiment 21, wherein the model-based algorithm is an AR-HMM algorithm.
[0489] 23. The method of embodiment 21, wherein the model-based algorithm is a SLDS algorithm.
[0490] 24. The method of embodiment 18, wherein the multilayer perceptron is a SVAE.
[0491] 25. The method of embodiment 24, wherein the SVAE and MLP are trained using a variational inference objective and performing gradient ascent.
[0492] 26. The method of embodiment 25, wherein the SVAE and MLP are trained simultaneously.
[0493] 27. The method of any one of embodiments 18-22, wherein the subject is an animal in an animal study.
[0494] 28. The method of any one of embodiments 18-22, wherein the subject is a human.
[0495] 29. The method of any one of embodiments 18-22, wherein the object is analyzed over a period of time long enough for the size of the object to change.
[0496] 30. The method of embodiment 25, wherein the SVAE and MLP are trained using data based on different mouse or rat strains.
[0497] 31. A method for classifying a test compound, the method comprising:
[0498] after administering the test compound to a test subject, identifying a test behavioral representation comprising a set of modules in the test subject;
[0499] comparing the test behavioral representation to a plurality of reference behavioral representations, wherein each reference behavioral representation represents each class of drugs; and
[0500] If the test behavior representation is identified by the classifier as matching the reference behavior representation representing a class of drugs, then it is determined that the test compound belongs to the class of drugs.
[0501] 32. The method of embodiment 31, wherein the test behavior representation is identified by the following steps:
[0502] receiving three-dimensional video data representing motion of the test object;
[0503] processing the three-dimensional data using a computational model to divide the data into at least one set of modules and at least one set of transition periods between the modules; and
[0504] The at least one set of modules is assigned to a category representing a type of animal behavior.
[0505] 33. A method according to embodiment 32, wherein the computational model includes modeling the sub-second module as a vector autoregressive process, wherein the vector autoregressive process represents a trajectories through a principal component analysis (PCA) space.
[0506] 34. The method of embodiment 32, wherein the computational model comprises modeling the sub-second module as a simultaneously fitted SLDS, while the MLP learns a feature manifold for the sub-second module.
[0507] 35. The method of embodiment 34, wherein the MLP is a SVAE.
[0508] 36. The method of embodiment 33, wherein the computational model comprises modeling the transfer period using a hidden Markov model.
[0509] 37. A method according to any one of embodiments 31-36, wherein the three-dimensional video data is first processed to output a series of points in a multi-dimensional vector space, wherein each point represents the 3D posture dynamics of the test object.
[0510] 38. The method of any one of embodiments 31-37, wherein the test compound is selected from the group consisting of a small molecule, an antibody or antigen-binding fragment thereof, a nucleic acid, a polypeptide, a peptide, a peptidomimetic, a polysaccharide, a monosaccharide, a lipid, a glycosaminoglycan, and combinations thereof.
[0511] 39. The method of any one of embodiments 31-38, wherein the test subject is an animal in an animal study.
[0512] 40. A method for analyzing the motion of an object to separate it into modules, the method comprising:
[0513] receiving three-dimensional video data representing motion of the subject before and after administration of an agent to the subject;
[0514] pre-processing the three-dimensional video data to separate the object from the background;
[0515] identifying, over a set of frames of the video data, an orientation of features of the object relative to a common coordinate system for all frames;
[0516] modifying the orientation of the object in at least a subset of the set of frames so that the features are oriented in the same direction relative to the coordinate system to output a set of aligned frames;
[0517] processing the aligned frames using a multi-layer perceptron (MLP) to output pose dynamics data, wherein the pose dynamics data represents a pose of the object in each aligned frame in a three-dimensional feature manifold;
[0518] processing the aligned frames using a computational model to temporally segment the posture dynamics data into different groups of modules, wherein all sub-second modules in a group of sub-second modules exhibit similar posture dynamics;
[0519] determining the number of modules in each set of sub-second modules prior to administering the agent to the subject;
[0520] determining the number of modules in each set of sub-second modules after administering the agent to the subject;
[0521] comparing the number of modules in each set of sub-second modules before and after administering the agent to the subject; and
[0522] An indication of a change in frequency of expression of the number of modules in each group of modules before and after administration of the agent to the subject is output.
[0523] 41. The method of embodiment 40, wherein each group of sub-second modules is classified into predetermined behavioral modules based on comparison with reference data representing the behavioral modules.
[0524] 42. The method of embodiment 40 or 41, wherein said change in expression frequency of the number of modules in each group of modules before and after administration of said agent to said subject is compared to said reference data representing change in expression frequency of modules after exposure to an agent of a known class.
[0525] 43. The method of embodiment 42, further comprising the step of classifying the agent as one of a plurality of agents of known classes based on comparison with reference data representing changes in the frequency following exposure to agents of known classes.
[0526] 44. The method of any one of embodiments 40-42, wherein the agent is a pharmaceutically active compound.
[0527] 45. The method of any one of embodiments 40-42, wherein the agent is a visual stimulus or an auditory stimulus.
[0528] 46. The method of any one of embodiments 40-42, wherein the agent is an odorant.
[0529] 47. The method of any one of embodiments 40-46, wherein the subject is an animal in an animal study.
[0530] 48. The method of any one of embodiments 40-46, wherein the subject is a human.
[0531] Computer and hardware installation of the present invention
[0532] First, it should be understood that the disclosure herein can be implemented using any type of hardware and / or software, and can be a pre-programmed general-purpose computing device. For example, the system can be implemented using a server, a personal computer, a portable computer, a thin client, or any suitable device or devices. The present invention and / or its components can be a single device at a single location, or multiple devices at a single or multiple locations, where they are connected together using any appropriate communication protocol via any communication medium (e.g., cable, fiber optic cable, or wireless).
[0533] It should also be noted that for the purposes of illustrating and discussing the present disclosure herein as having multiple modules that perform specific functions. It should be understood that for the sake of clarity, these modules are merely schematically illustrated based on their functions and do not necessarily represent specific hardware or software. In this regard, these modules can be hardware and / or software that generally perform the specific functions discussed. Furthermore, these modules can be combined together in the present disclosure or divided into additional modules based on the specific functions desired. Therefore, the present disclosure should not be interpreted as limiting the present invention, but rather is merely understood as illustrating an example embodiment thereof.
[0534] A computing system may include a client and a server. The client and server are typically remote from each other and typically interact via a communication network. The relationship between the client and the server is generated by computer programs running on their respective computers and having a client-server relationship. In some implementations, the server sends data (e.g., an HTML page) to a client device (e.g., for displaying data to a user interacting with the client device and receiving user input from the user). Data generated at the client device (e.g., the results of user interactions) may be received from the client device at the server.
[0535] Implementations of the subject matter described in this specification can be implemented in a computing system that includes a back-end component (e.g., as a data server), or includes an intermediate component (e.g., an application server), or includes a front-end component (e.g., a client computer with a graphical user interface or a web browser through which a user can interact with implementations of the subject matter described in this specification), or includes any combination of one or more such back-end components, intermediate components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include local area networks ("LANs") and wide area networks ("WANs"), internetworks (e.g., the Internet), and point-to-point networks (e.g., ad hoc point-to-point networks).
[0536] The subject matter and implementation of the operations described in this specification can be implemented in digital electronic circuitry or computer software, firmware, or hardware, including the structures disclosed in this specification and their structural equivalents, or a combination of one or more structures thereof. The implementation of the subject matter described in this specification can be implemented as one or more computer programs (i.e., one or more modules of computer program instructions) encoded on a computer storage medium, or for execution by a data processing device or for controlling the operation of a data processing device. Alternatively or in addition, the program instructions can be encoded on an artificially generated propagated signal, such as a machine-generated electrical, optical, or electromagnetic signal, which is used to encode information for transmission to a suitable receiving device for execution by the data processing device. The computer storage medium can be a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or a combination of one or more of them, or be included therein. In addition, although a computer storage medium is not a propagated signal, a computer storage medium can be a source or destination of computer program instructions encoded in an artificially generated propagated signal. The computer storage medium can also be one or more independent physical components or media (e.g., multiple CDs, disks, or other storage devices) or be included therein.
[0537] The operations described in this specification may be implemented as operations performed by a "data processing apparatus" on data stored on one or more computer-readable storage devices or received from other sources.
[0538] The term "data processing device" encompasses various devices, apparatuses, and machines for processing data, including, for example, a programmable processor, a computer, a system on a chip, or a plurality of the aforementioned devices, or a combination of the aforementioned devices. These devices may include dedicated logic circuits, such as an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit). In addition to hardware, the device may also include code that creates an execution scenario for the computer program, such as processor firmware, a protocol stack, a database management system, an operating system, a cross-platform runtime scenario, a virtual machine, or a combination of one or more thereof. The device and execution scenario can implement a variety of different computing model infrastructures, such as web services, distributed computing, and grid computing infrastructure.
[0539] A computer program (also referred to as a program, software, software application, script, or code) may be written in any form of programming language, including compiled or interpreted languages, declarative languages, or procedural languages, and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, object, or other unit suitable for use in a computing context. A computer program may (but does not necessarily) correspond to a file in a file system. A program may be stored as part of a file that stores other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program, or in multiple coordinated files (e.g., files storing one or more modules, subroutines, or portions of code). A computer program may be deployed for execution on one computer or on multiple computers located at one site or distributed across multiple sites and interconnected by a communications network.
[0540] The processes and logic flows described in this specification can be performed by one or more programmable processors executing one or more computer programs to perform operations by operating on input data and generating output. These processes and logic flows can also be performed by special purpose logic circuitry, and devices can also be implemented as, for example, an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit).
[0541] The processes and logic flows described in this specification can be performed by one or more programmable processors executing one or more computer programs to perform operations by operating on input data and generating output. These processes and logic flows can also be performed by, and devices can be implemented as, special purpose logic circuitry, such as an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit).
[0542] For example, processors suitable for executing computer programs include general-purpose and special-purpose microprocessors, as well as any one or more processors of any type of digital computer. Typically, a processor will receive instructions and data from a read-only memory or a random access memory, or both. The essential elements of a computer are a processor for performing actions according to instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices (e.g., magnetic, magneto-optical, or optical disks) for storing data, or be operably connected to a storage device to receive data from the storage device, send data to the storage device, or both. However, a computer does not require such devices. In addition, a computer can be embedded in another device, such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a Global Positioning System (GPS) receiver, or a portable storage device (e.g., a Universal Serial Bus (USB) flash drive), etc. Devices suitable for storing computer program instructions and data include all forms of nonvolatile memory, media, and storage devices, including, for example, semiconductor memory devices (e.g., EPROM, EEPROM, and flash memory devices); magnetic disks (e.g., internal hard disks or removable disks); magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and memory can be supplemented by, or incorporated in, special purpose logic circuitry.
[0543] in conclusion
[0544] The various methods and techniques described above provide a variety of ways to implement the present invention. Of course, it should be understood that not all of the objects or advantages described can be achieved in accordance with any specific embodiment described herein. Thus, for example, those skilled in the art will recognize that these methods can be implemented in a manner that achieves or optimizes one or more of the advantages taught herein without necessarily achieving other objects or advantages described or suggested herein. Various alternatives are mentioned herein. It should be understood that some embodiments specifically include one, another, or several features, while other embodiments specifically exclude one, another, or several features, while still other embodiments circumvent specific features by including one, another, or several advantageous features.
[0545] In addition, those skilled in the art will recognize the applicability of various features from different embodiments. Similarly, the various elements, features and steps discussed above and other known equivalents of each element, feature or step can be used by those skilled in the art in different combinations to perform the method according to the principles described herein. In different embodiments, some elements, features and steps among the various elements, features and steps will be specifically included, while other steps will be specifically excluded.
[0546] While the present application has been disclosed in the context of certain embodiments and examples, those skilled in the art will appreciate that the embodiments of the present application extend to other alternatives and / or uses and modifications beyond the specifically disclosed embodiments and their equivalents.
[0547] In some embodiments, the terms "one", "an" and "the" and similar references used in describing specific embodiments of the present application (particularly in the context of certain claims below) may be interpreted as including both single and multiple. Reference to a numerical range here is merely a convenient method for individually referring to each individual value falling within the range. Unless otherwise indicated herein, each individual value is included in the specification as if it were individually listed here. All methods described herein may be performed in any suitable order, unless otherwise indicated herein or clearly contradictory to the context. Any and all examples or exemplary language (e.g., "such") are used. The content provided here about certain embodiments is only for better illustration of the application, rather than for limiting the scope of application claimed in other ways. Any language in this specification should not be interpreted as representing any non-claimed element that is essential to the application practice.
[0548] Certain embodiments of the present application are described herein. Variations of these embodiments will become apparent to those of ordinary skill in the art upon reading the foregoing description. It is expected that skilled artisans may employ such variations as appropriate and may implement the present application in a manner different from that specifically described herein. Therefore, many embodiments of the present application include all modifications and equivalents of the subject matter recited in the appended claims as permitted by applicable law. Furthermore, this application encompasses any combination of the above-described elements in all possible variations thereof, unless otherwise indicated herein or clearly contradicted by the context.
[0549] Certain embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. In some cases, the operations recited in the claims can be performed in a different order and still achieve desirable results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the particular order shown, or sequential order, to achieve desirable results.
[0550] All patents, patent applications, patent application publications, and other materials (e.g., articles, books, specifications, publications, documents, articles, etc.) cited herein are hereby incorporated by reference in their entirety for all purposes, except any prosecution history related thereto, any patent that is inconsistent or conflicting with this document, or any patent that would have the effect of limiting the broadest present or future scope of the claims of this document. For example, if there is any inconsistency or conflict between the description, definition, and / or usage of a term in any incorporated material and that in this document, the description, definition, and / or usage in this document shall control.
[0551] Finally, it should be understood that the embodiments of the application disclosed herein illustrate the principles of the embodiments of the present application. Other modifications that may be adopted within the scope of the present application. Therefore, by way of example and not limitation, alternative configurations of the embodiments of the present application may be adopted according to the teachings of this document. Therefore, the embodiments of the present application are not limited to the embodiments shown or described.
Claims
1. A system for analyzing the motion of an object to separate it into sub-second modules, the system comprising: a three-dimensional camera configured to output video image data representing motion of the object; a memory in communication with the 3D camera and comprising a non-volatile machine-readable storage medium having machine-executable instructions stored thereon; as well as a control system comprising one or more processors connected to the memory, the one or more processors configured to execute the machine-executable instructions, the machine-executable instructions being configured to cause the control system to: processing three-dimensional video data representing motion of the object received from the three-dimensional camera using a computational model comprising a multi-layer perceptron to output pose dynamics data, wherein the multi-layer perceptron comprises a structured variational autoencoder; dividing the gesture dynamics data into at least one set of sub-second modules and at least one set of transition statistics between the sub-second modules; and The at least one set of sub-second modules is assigned to a category representing a type of animal behavior.
2. The system of claim 1, wherein: Processing the three-dimensional video data includes separating the object from a background in the three-dimensional video data.
3. The system of claim 1, wherein: Processing the three-dimensional video data includes identifying, for a common coordinate system of each frame in a set of frames of the three-dimensional video data, an orientation of a feature of the object across the set of frames of the three-dimensional video data.
4. The system of claim 3, wherein: Processing the three-dimensional video data includes generating a set of aligned frames by modifying the orientation of the object in at least a subset of the set of frames so that the features are oriented in the same direction with respect to the coordinate system.
5. The system of claim 4, wherein: Processing the three-dimensional video data includes generating pose dynamics data by processing the set of aligned frames using the multilayer perceptron, wherein the pose dynamics data represents a pose of the object through manifold space for each aligned frame.
6. The system of claim 5, wherein: Processing the aligned frames includes temporally segmenting the gesture dynamics data into different groups of sub-second modules, wherein each sub-second module in a group of modules exhibits similar gesture dynamics.
7. The system of claim 1, wherein: Processing the three-dimensional video data includes generating a series of points in a multi-dimensional vector space using the three-dimensional video data, wherein each point in the series of points represents a three-dimensional pose dynamics of the object.
8. The system of claim 1, wherein: The subject is a human or an animal.
9. A method for analyzing the motion of an object to separate it into sub-second modules, the method comprising: processing the three-dimensional video data representing motion of the object using a computational model comprising a multilayer perceptron to output pose dynamics data, wherein the multilayer perceptron comprises a structured variational autoencoder; dividing the gesture dynamics data into at least one set of sub-second modules and at least one set of transition statistics between the sub-second modules; and The at least one set of sub-second modules is assigned to a category representing a type of animal behavior.
10. The method of claim 9, wherein: Processing the three-dimensional video data includes separating the object from a background in the three-dimensional video data.
11. The method of claim 9, wherein: Processing the three-dimensional video data includes identifying, for a common coordinate system of each frame in a set of frames of the three-dimensional video data, an orientation of a feature of the object across the set of frames of the three-dimensional video data.
12. The method of claim 11, wherein: Processing the three-dimensional video data includes generating a set of aligned frames by modifying the orientation of the object in at least a subset of the set of frames so that the features are oriented in the same direction with respect to the coordinate system.
13. The method of claim 12, wherein: Processing the three-dimensional video data includes generating pose dynamics data by processing the set of aligned frames using the multilayer perceptron, wherein the pose dynamics data represents a pose of the object through manifold space for each aligned frame.
14. The method of claim 13, wherein: Processing the aligned frames includes temporally segmenting the gesture dynamics data into different groups of sub-second modules, wherein each sub-second module in a group of modules exhibits similar gesture dynamics.
15. The method of claim 9, wherein: Processing the three-dimensional video data includes generating a series of points in a multi-dimensional vector space using the three-dimensional video data, wherein each point in the series of points represents a three-dimensional pose dynamics of the object.
16. The method of claim 9, wherein: The subject is a human or an animal.
17. A method for analyzing the motion of an object to separate it into modules, the method comprising: processing three-dimensional video data representing the motion of the object using a computational model to divide the video data into at least one set of modules and at least one set of transition statistics between the modules, wherein processing the three-dimensional video data further comprises: separating the object from a background in the video data; identifying, for a common coordinate system for each frame, an orientation of a feature of the object across a set of frames of the video data; modifying the orientation of the object in at least a subset of the set of frames so that the features are oriented in the same direction with respect to the coordinate system, thereby outputting a set of aligned frames; processing the aligned frames using a multilayer perceptron including a structured variational autoencoder to output pose dynamics data, wherein the pose dynamics data represents a pose of the object through manifold space for each aligned frame; and The at least one set of modules is assigned to a category representing a type of animal behavior.
18. The method of claim 17, wherein: The three-dimensional video data is first processed to output a series of points in a multi-dimensional vector space, wherein each point represents the three-dimensional pose dynamics of the object.
19. The method of claim 17, wherein: The subjects are animals in animal research.
20. The method of claim 17, wherein: The subject is a human.
21. A method for analyzing the motion of an object to separate it into modules, the method comprising: pre-processing the three-dimensional video data representing the motion of the object to separate the object from a background of the three-dimensional video data; identifying, over a set of frames of the three-dimensional video data, an orientation of a feature of the object relative to a common coordinate system of each frame in the set of frames; modifying the orientation of the object in at least a subset of the set of frames so that the features are oriented in the same direction relative to the coordinate system to output a set of aligned frames; processing aligned frames in the set of aligned frames using a deep learning model including a structured variational autoencoder to output pose dynamics data, wherein the pose dynamics data represents a pose of the object in each aligned frame in a three-dimensional graphics space; processing aligned frames in the set of aligned frames to temporally segment the gesture dynamics data into different groups of sub-second modules, wherein each sub-second module in the set of sub-second modules exhibits similar gesture dynamics; and A representation of each module in the set of sub-second modules that occurs at a frequency above a threshold in the three-dimensional video data is displayed.
22. The method of claim 21, wherein: The sub-second module of processing the aligned frames in the set of aligned frames to temporally segment the gesture dynamics data into different groups includes using a model-free algorithm.
23. The method of claim 22, wherein: Using the model-free algorithm includes computing an autocorrelogram.
24. The method of claim 21, wherein: The sub-second module that processes the aligned frames in the set of aligned frames to temporally segment the gesture dynamics data into different groups includes using a model-based algorithm.
25. The method of claim 24, wherein: The model-based algorithm is the AR-HMM algorithm.
Citation Information
Patent Citations
Automatically classifying animal behavior
CN108471943A