Long-term continuous animal behavior monitoring
Neural network-based systems address the challenge of tracking animals in complex environments by providing robust and scalable animal tracking, enabling accurate and continuous monitoring of animal behavior with reduced user intervention.
Patent Information
- Application Number
- JP2023102972
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-04-23
- Filing Date
- 2023-06-23
- Publication Date
- 2025-06-23
- Estimated Expiration
- 2038-08-07
AI Technical Summary
Existing technologies face challenges in accurately tracking animals in complex and dynamic environments, leading to data confusion and non-reproducible results in long-term animal monitoring experiments.
The development of systems and methods using neural networks that can robustly and scalably track animals, such as mice, in open fields by processing video data with high spatio-temporal resolution, and automatically adjusting to various environmental conditions without user intervention.
These systems enable continuous and accurate monitoring of animal behavior over long periods, reducing user involvement and improving data reliability, even in heterogeneous environments with different animal strains and coat colors.
Smart Images

Figure 0007696955000011 
Figure 0007696955000012 
Figure 0007696955000013
Abstract
Description
Technical Field
[0001]
[0001] This application claims the benefit of U.S. Provisional Patent Application No. 62 / 542,180, entitled "Long-Term and Continuous Animal Behavioral Monitoring," filed on Aug. 7, 2017, and U.S. Provisional Patent Application No. 62 / 661,610, entitled "Robust Mouse Tracking In Complex Environments Using Neural Networks," filed on Apr. 23, 2018. The entire contents of each of these applications are incorporated herein by reference.
Background Art
[0002]
[0002] Animal behavior can be understood as the output of the nervous system in response to internal or external stimuli. The ability to accurately track an animal can be beneficial as part of the process of classifying animal behavior. For example, changes in behavior are prominent features of aging, mental illness, or metabolic disease, and can reveal important information regarding the effects on an animal's physiological, neurocognitive, and emotional states.
Summary of the Invention
[0003]
[0003] Conventionally, experiments for evaluating animal behavior have been conducted non-invasively, and researchers interact directly with the animals. As an example, a researcher may remove an animal such as a mouse from its living environment (e.g., a cage) and transfer the animal to a different environment (e.g., a device such as a maze). Then, the researcher may position themselves near the new environment and observe the animal's working ability by tracking the animal. However, it is known that animals can exhibit different behaviors in a new environment or different behaviors towards an experimenter performing a test. This often leads to data confusion and results that are non-reproducible and misleading.
[0004]
[0004] To minimize human interference during behavioral monitoring experiments, low-invasive monitoring technologies have been developed. As an example, video monitoring used for monitoring animal behavior has been studied. However, there are still challenges with video monitoring. On one hand, the main hurdle remains being able to continuously capture video data with high spatio-temporal resolution over an extended period under a set of broad environmental conditions. Animal observational studies conducted over long periods such as several days, weeks, and / or months can generate large amounts of data that are costly to acquire and store. On another hand, even assuming that sufficient quality video data can be acquired and stored, it is economically infeasible for researchers to manually scrutinize the large amount of video footage generated during long-term observations and track animals over such long periods. This problem becomes more prominent when the number of animals to be observed increases, as may be required when conducting new drug screening or genomics experiments.
[0005]
[0005] To address this problem, computer-based technologies for analyzing captured videos of animal behavior have been developed. However, existing computer-based systems cannot accurately track different animals in complex and dynamic environments. As an example, existing computer-based technologies for tracking animals cannot accurately identify an animal from its background (e.g., the walls and / or floor of a cage, objects within the cage such as a water container) or distinguish between multiple animals. At best, if a given animal is not accurately tracked during the observation period, valuable observational data may be lost. In the worst case, if a given animal or part of it is incorrectly tracked or misidentified as another during the observation period Errors may be introduced into the classified actions from the acquired video data. Although techniques such as changing the coat color of animals are adopted to facilitate tracking, the behavior of animals may also change due to the change in the coat color of animals. As a result, existing video tracking methods implemented in complex and dynamic environments or genetically heterogeneous animals require a high level of user involvement, thus losing the above-mentioned advantages regarding video observation. For this reason, large-scale and / or long-term animal monitoring experiments are still impossible to achieve.
[0006]
[0006] As neuroscience and behavioral science enter the era of large amounts of behavioral data and computational ethology, better techniques for tracking animals are needed to facilitate the classification of animal behavior in semi-natural and dynamic environments over a long period of time.
[0007]
[0007] Therefore, systems and methods using neural networks that can provide robust and scalable tracking of animals (e.g., mice) in open fields have been developed. As an example, systems and methods are provided that facilitate the acquisition of video data of animal movements with high spatio-temporal resolution. This video data can be continuously captured over a long period of time under a set of extensive environmental conditions.
[0008]
[0008] The acquired video data can be adopted as the input of a convolutional neural network architecture for tracking. The neural network can be trained to perform tracking under multiple environmental conditions with high robustness and without user-involvement adjustment when a new environment or animal is presented. Examples of such experimental conditions can include different mouse strains, regardless of different cage environments, as well as different coat colors, body shapes, and behaviors. For this reason, embodiments of the present disclosure can facilitate the monitoring of the behaviors of a large number of animals over a long period of time under heterogeneous conditions by facilitating minimally invasive animal tracking.
[0009]
[0009] In certain embodiments, the disclosed video observation and animal tracking techniques may be employed in combination. However, it can be understood that each of these techniques can be employed alone or in any combination with each other or with other techniques.
[0010]
[0010] In one embodiment, a method of animal tracking is provided. The method may include receiving, by a processor, video data representing an observation of an animal, and executing, by the processor, a neural network architecture. The neural network architecture may be configured to receive an input video frame extracted from the video data, generate, based on the input video frame, an ellipse description of at least one animal, where the ellipse description is defined by predetermined ellipse parameters, and provide data including values characterizing the predetermined ellipse parameters for at least one animal.
[0011]
[0011] In another embodiment of this method, the ellipse parameters may be coordinates representing the position of the animal in a plane, the lengths of the major and minor axes of the animal, and the angle at which the animal's head is facing, the angle being defined with respect to the direction of the major axis.
[0012]
[0012] In another embodiment of this method, the neural network architecture may be an encoder-decoder segmentation network. The encoder-decoder segmentation network may be configured to predict a foreground-background segmentation image from the input video frame, predict whether an animal is present in the input video frame from the perspective of pixels based on the segmentation image, output a segmentation mask based on the prediction from the perspective of pixels, and fit an ellipse to the portion of the segmentation mask where an animal is predicted to be present to determine values characterizing the predetermined ellipse parameters.
[0013]
[0013] In another embodiment of this method, the encoder-decoder segmentation network may comprise a feature encoder, a feature decoder, and an angle predictor. The feature encoder may be configured to abstract an input video frame into a set of features with a small spatial resolution. The feature decoder may be configured to convert a set of features into the same shape as the input video frame and output a foreground-background segmentation image. The angle predictor may be configured to predict the angle at which the animal's head is facing.
[0014]
[0014] In another embodiment of this method, the neural network architecture may comprise a binning classification network configured to predict a heatmap of the most accurate value of each ellipse parameter of the ellipse description.
[0015]
[0015] In another embodiment of this method, the binning classification network may comprise a feature encoder configured to abstract an input video frame into a small spatial resolution, and the abstraction may be employed to generate a heatmap.
[0016]
[0016] In another embodiment of this method, the neural network architecture may comprise a regression network configured to extract features from an input video frame and directly predict a value characterizing each ellipse parameter.
[0017]
[0017] In another embodiment of this method, the animal may be a rodent.
[0018] In one embodiment, a system for animal tracking is provided. The system may include a data storage device that maintains video data representing observations of animals. The system may also include a processor configured to receive video data from the data storage device and implement a neural network architecture. The neural network architecture may be configured to receive input video frames extracted from the video data, generate an ellipse description of at least one animal based on the video frames, where the ellipse description is defined by predetermined ellipse parameters, and provide data including values characterizing the predetermined ellipse parameters for at least one animal.
[0018]
[0019] In another embodiment of the system, the ellipse parameters may be coordinates representing the position of the animal in a plane, the lengths of the major and minor axes of the animal, and the angle at which the animal's head is facing, the angle being defined with respect to the direction of the major axis.
[0019]
[0020] In another embodiment of the system, the neural network architecture may be an encoder-decoder segmentation network. The encoder-decoder segmentation network may be configured to predict a foreground-background segmented image from the input video frames, predict whether an animal is present in the input video frames based on the segmented image in terms of pixels, output a segmentation mask based on the prediction in terms of pixels, fit an ellipse to the portion of the segmentation mask where an animal is predicted to be present, and determine values characterizing the predetermined ellipse parameters.
[0020]
[0021] In another embodiment of this system, the encoder-decoder segmentation network may include a feature encoder, a feature decoder, and an angle predictor. The feature encoder may be configured to abstract an input video frame into a set of features with a small spatial resolution. The feature decoder may be configured to convert a set of features into the same shape as the input video frame and output a foreground-background segmentation image. The angle predictor may be configured to predict the angle at which the animal's head is facing.
[0021]
[0022] In another embodiment of this system, the neural network architecture may include a binning classification network. The binning classification network may be configured to predict a heatmap of the most accurate values of each ellipse parameter of the ellipse description.
[0022]
[0023] In another embodiment of this system, the binning classification network may include a feature encoder configured to abstract an input video frame into a small spatial resolution, and the abstraction may be employed to generate a heatmap.
[0023]
[0024] In another embodiment of this system, the neural network architecture may include a regression network configured to extract features from an input video frame and directly predict values characterizing each of the ellipse parameters.
[0024]
[0025] In another embodiment of this system, the animal may be a rodent.
[0026] In one embodiment, a non-transitory computer program product storing instructions is provided. The instructions, when executed by at least one data processor of at least one computing system, can execute a method including receiving video data representing observations of animals and executing a neural network architecture. The neural network architecture can be configured to receive input video frames extracted from the video data, generate an ellipse description of at least one animal based on the input video frames, where the ellipse description is defined by predetermined ellipse parameters, and provide data including values characterizing the predetermined ellipse parameters for at least one animal.
[0025]
[0027] In another embodiment, the ellipse parameters can be coordinates representing the position of the animal in a plane, the lengths of the major and minor axes of the animal, and the angle at which the animal's head is facing, the angle being defined with respect to the direction of the major axis.
[0026]
[0028] In another embodiment, the neural network architecture can be an encoder-decoder segmentation network. The encoder-decoder segmentation network can be configured to predict a foreground-background segmentation image from the input video frames, predict whether an animal is present in the input video frames from the perspective of pixels based on the segmentation image, output a segmentation mask based on the prediction from the perspective of pixels, and fit an ellipse to the portion of the segmentation mask where an animal is predicted to be present to determine values characterizing the predetermined ellipse parameters.
[0027]
[0029] In another embodiment, the encoder-decoder segmentation network may comprise a feature encoder, a feature decoder, and an angle predictor. The feature encoder may be configured to abstract an input video frame into a set of features with a small spatial resolution. The feature decoder may be configured to convert a set of features into the same shape as the input video frame and output a foreground-background segmentation image. The angle predictor may be configured to predict the angle at which the animal's head is facing.
[0028]
[0030] In another embodiment, the neural network architecture may comprise a binning classification network configured to predict a heatmap of the most accurate values of each ellipse parameter of the ellipse description.
[0029]
[0031] In another embodiment, the binning classification network may comprise a feature encoder configured to abstract an input video frame into a small spatial resolution, and the abstraction may be employed to generate a heatmap.
[0030]
[0032] In another embodiment, the neural network architecture may comprise a regression network configured to extract features from an input video frame and directly predict values characterizing each of the ellipse parameters.
[0031]
[0033] In another embodiment, the animal may be a rodent.
[0034] In one embodiment, a system is provided that may include an arena and an acquisition system. The arena may include a frame and a housing attached to the frame. The housing may be dimensioned to accommodate an animal and may include a door configured to allow access to the interior. The acquisition system may include a camera, at least two sets of light sources, a controller, and a data storage device. Each set of light sources may be configured to emit light incident on the housing at a different wavelength from the others. The camera may be configured to acquire video data of at least a portion of the housing when irradiated by at least one of the multiple sets of light sources. The controller may be in electrical communication with the camera and the multiple sets of light sources. The controller may be configured to generate control signals that operate to control the acquisition of video data by the camera and the emission of light by the multiple sets of light sources, and to receive the video data acquired by the camera. The data storage device may be in electrical communication with the controller and may be configured to store the video data received from the controller.
[0032]
[0035] In another embodiment of this system, at least a portion of the housing may be substantially opaque to visible light.
[0036] In another embodiment of this system, at least a portion of the housing may be formed of a material that is substantially opaque to visible light wavelengths.
[0033]
[0037] In another embodiment of this system, at least a portion of the housing may be formed of a material that is substantially non-reflective to infrared light wavelengths.
[0038] In another embodiment of this system, at least a portion of the housing may be formed of a sheet of polyvinyl chloride (PVC) or polyoxymethylene (POM).
[0034]
[0039] In another embodiment of this system, the first set of light sources may include one or more first illuminations configured to emit light at one or more visible light wavelengths, and the second set of light sources may include one or more second illuminations configured to emit light at one or more infrared (IR) light wavelengths.
[0035]
[0040] In another embodiment of this system, the wavelength of the infrared light may be about 940 nm.
[0041] In another embodiment of this system, the camera may be configured to acquire video data at a resolution of at least 480 × 480 pixels.
[0036]
[0042] In another embodiment of this system, the camera may be configured to acquire video data at a frame rate higher than the frequency of the movement of the mouse.
[0043] In another embodiment of this system, the camera may be configured to acquire video data at a frame rate of at least 29 frames per second (fps).
[0037]
[0044] In another embodiment of this system, the camera may be configured to acquire video data having a depth of at least 8 bits.
[0045] In another embodiment of this system, the camera may be configured to acquire video data at an infrared wavelength.
[0038]
[0046] In another embodiment of this system, the controller may be configured to compress the video data received from the camera.
[0047] In another embodiment of this system, the controller may be configured to compress the video data received from the camera using an MPEG4 codec including a filter employing distributed-based background subtraction.
[0039]
[0048] In another embodiment of this system, as the filter of the MPEG codec, Q0 HQDN3D is possible.
[0049] In another embodiment of this system, the controller may be configured to request the first light source to irradiate the housing according to a schedule that simulates the light and dark cycle.
[0040]
[0050] In another embodiment of this system, the controller may be configured to request the first light source to irradiate the housing with visible light having an intensity of approximately 50 lux to approximately 800 lux in the bright part of the light and dark cycle.
[0041]
[0051] In another embodiment of this system, the controller may be configured to request the second light source to irradiate the housing with infrared light so that the temperature rise of the housing due to infrared irradiation is less than 5°C.
[0042]
[0052] In another embodiment of this system, the controller may be configured to request the first light source to irradiate the housing according to logarithmically scaled 1024-level illumination.
[0043]
[0053] In one embodiment, a method is provided that may include irradiating a housing configured to house an animal with at least one set of light sources. Each set of light sources may be configured to emit light of different wavelengths. The method may also include obtaining video data of at least a part of the housing irradiated by at least one of a plurality of sets of light sources by a camera. The method may also include generating a control signal that operates to control the acquisition of video data by the camera and the emission of light by the plurality of sets of light sources by a controller in electrical communication with the camera and the plurality of sets of light sources. Further, the method may include receiving, by the controller, the video data obtained by the camera.
[0044]
[0054] In another embodiment of this method, at least a part of the housing may be substantially opaque to visible light.
[0055] In another embodiment of this method, at least a part of the housing may be formed of a material that is substantially opaque to visible light wavelengths.
[0045]
[0056] In another embodiment of this method, at least a part of the housing may be formed of a material that is substantially non-reflective to infrared light wavelengths.
[0057] In another embodiment of this method, at least a part of the housing may be formed of a sheet of polyvinyl chloride (PVC) or polyoxymethylene (POM).
[0046]
[0058] In another embodiment of this method, the first set of light sources may include one or more first illuminations configured to emit light at one or more visible light wavelengths, and the second set of light sources may include one or more second illuminations configured to emit light at one or more infrared (IR) light wavelengths.
[0047]
[0059] In another embodiment of this method, the wavelength of the infrared light may be about 940 nm.
[0060] In another embodiment of this method, the camera may be configured to acquire video data at a resolution of at least 480×480 pixels and resolution.
[0048]
[0061] In another embodiment of this method, the camera may be configured to acquire video data at a frame rate higher than the frequency of movement of the mouse.
[0062] In another embodiment of this method, the camera may be configured to acquire video data at a frame rate of at least 29 frames per second (fps).
[0049]
[0063] In another embodiment of this method, the camera may be configured to acquire video data having a depth of at least 8 bits.
[0064] In another embodiment of this method, the camera may be configured to acquire video data at an infrared wavelength.
[0050]
[0065] In another embodiment of this method, the controller may be configured to compress the video data received from the camera.
[0066] In another embodiment of this method, the controller may be configured to compress the video data received from the camera using an MPEG4 codec that includes a filter employing distributed-based background subtraction.
[0051]
[0067] In another embodiment of this method, as the filter of the MPEG codec, Q0 HQDN3D is possible.
[0068] In another embodiment of this method, the controller may be configured to request the first light source to irradiate the housing according to a schedule that simulates a light and dark cycle.
[0052]
[0069] In another embodiment of this method, the controller may be configured to request the first light source to irradiate the housing with visible light having an intensity of approximately 50 lux to approximately 800 lux in the bright part of the light and dark cycle.
[0053]
[0070] In another embodiment of this method, the controller may be configured to request the second light source to irradiate the housing with infrared light so that the temperature rise of the housing due to infrared irradiation is less than 5°C.
[0054]
[0071] In another embodiment of this method, the controller may be configured to request the first light source to irradiate the housing according to logarithmically scaled 1024-level illumination.
[0055]
[0072] The above and other features will be easily understood from the following detailed description in conjunction with the accompanying drawings.
Brief Description of the Drawings
[0056]
Figure 1
[0073] FIG. 1 is a flowchart showing an exemplary embodiment of an operating environment for animal tracking.
Figure 2
[0074] FIG. 2 is a schematic diagram of an embodiment of a system for animal behavior monitoring.
Figure 3
[0075] FIGS. 3A - 3F are images showing sample frames acquired by the system of FIG. 2; (A - C) visible light; (D - F) infrared (IR) light.
Figure 4
[0076] FIGS. 4A - 4B are plots of quantum efficiency as a function of wavelength for two camera models; (A) relative response for Sentech STC - MC33USB; (B) quantum efficiency of Basler acA1300 - 60gm - NIR.
Figure 5
[0077] FIG. is a plot showing the transparency - wavelength profile of an IR long - pass filter.
Figure 6
[0078] FIGS. 6A - 6D are images showing exemplary embodiments of video frames to which different compression techniques are applied; (A) uncompressed; (B) MPEG4 Q0, (C) MPEG4 Q5; (D) MPEG4 Q0 HQDN3D;
Figure 7
[0079] FIG. 7 is a diagram showing an embodiment of components of an acquisition system suitable for use with the system of FIG. 2.
Figure 8A
[0080] FIG. 8A is a schematic diagram of an exemplary embodiment of an observation environment analyzed in accordance with the present disclosure, including a black mouse, a gray mouse, an albino mouse, and a mouse with a spotted pattern.
Figure 8B
[0081] FIG. 8B is a schematic diagram of a state where animal tracking is insufficient.
Figure 8C
[0082] FIG. 8C is a schematic diagram of an exemplary embodiment of mouse tracking including object tracking in an elliptical form.
Figure 9
[0083] Figure 9 is a schematic diagram of an exemplary embodiment of a segmentation network architecture.
Figure 10
[0084] Figure 10 is a schematic diagram of an exemplary embodiment of a binning classification network architecture.
Figure 11
[0085] Figure 11 is a schematic diagram of an exemplary embodiment of a regression classification network architecture.
Figure 12A
[0086] Figure 12A shows an exemplary embodiment of a graphical user interface showing the arrangement of two marks for foreground (F) and background (B).
Figure 12B
[0087] Figure 12B shows an exemplary embodiment of a graphical user interface showing the segmentation as a result of the marking in Figure 12A.
Figure 13A
[0088] Figure 13A shows a plot of the training curves of the embodiments of the segmentation, regression, and binning classification networks of Figures 9-11.
Figure 13B
[0089] Figure 13B shows a plot of the validation curves of the embodiments of the segmentation, regression, and binning classification networks of Figures 9-11.
Figure 13C
[0090] Figure 13C shows a plot of the training and validation performance of the segmentation network architecture of Figure 9.
Figure 13D
[0091] Figure 13D shows a plot of the training and validation performance of the regression network architecture of Figure 11.
Figure 13E
[0092] Figure 13E shows a plot of the training and validation performance of the binning classification network architecture of Figure 10.
Figure 14A
[0093] FIG. 14A is a diagram showing a plot of training error as a function of the step of training a plurality of sets of different sizes according to an embodiment of the present disclosure.
Figure 14B
[0094] FIG. 14B is a diagram showing a plot of verification error as a function of the step of training a plurality of sets of different sizes according to an embodiment of the present disclosure.
Figure 14C
[0095] FIG. 14C is a diagram showing a plot of training and verification errors as a function of the step of the full training set of training samples.
Figure 14D
[0096] FIG. 14D is a diagram showing a plot of training and verification errors as a function of the step of a training set including 10,000 (10k) training samples.
Figure 14E
[0097] FIG. 14E is a diagram showing a plot of training and verification errors as a function of the step of a training set including 5,000 (5k) training samples.
Figure 14F
[0098] FIG. 14F is a diagram showing a plot of training and verification errors as a function of the step of a training set including 2,500 (2.5k) training samples.
Figure 14G
[0099] FIG. 14G is a diagram showing a plot of training and verification errors as a function of the step of a training set including 1,000 (1k) training samples.
Figure 14H
[0100] FIG. 14H is a diagram showing a plot of training and verification errors as a function of the step of a training set including 500 training samples.
Figure 15
[0101] FIGS. 15A - 15D are frames of captured video data in which color indicators for distinguishing each mouse from each other are superimposed; (A - B) visible light irradiation; (C - D) infrared light irradiation.
Figure 16
[0102] Figure 16 is a plot comparing the performance of the segmentation network architecture of Figure 9 with the beam break system.
Figure 17
[0103] Figure 17A is a plot of a prediction by an embodiment of the present disclosure and Ctrax.
[0057]
[0104] Figure 17B is a plot of the relative standard deviation of the minor axis prediction determined by the segmentation network architecture of Figure 9.
Figure 18A
[0105] Figure 18A is a plot of the total distance tracked for large strain surveys of genetically different animals determined by the segmentation network architecture of Figure 9.
Figure 18B
[0106] Figure 18B is a plot of the circadian locomotor patterns observed in six animals continuously tracked over four days in a dynamic environment determined by the segmentation network architecture of Figure 9.
DETAILED DESCRIPTION OF THE INVENTION
[0058]
[0107] Note that the drawings are not necessarily to scale. The drawings are intended to show only representative aspects of the subject matter disclosed herein and should not be considered as limiting the scope of the present disclosure.
[0059]
[0108] For clarity, exemplary embodiments of systems and corresponding methods that facilitate behavioral monitoring by video capture of one or more animals and tracking of one or more animals are discussed herein with respect to small rodents such as mice. However, the disclosed embodiments can be employed and / or configured to monitor other animals without limitation.
[0060]
[0109] FIG. 1 illustrates an arena 200, an acquisition system 700, and a neural network. FIG. 7 is a schematic diagram illustrating an example embodiment of an operating environment 100 including a tracking system configured to implement a mouse tracker. As discussed in more detail below, one or more mice may be housed in an arena 200. Video data of at least one animal (such as a mouse) is acquired. The video data may be acquired alone or in combination with other data related to animal monitoring, such as audio and environmental parameters (e.g., temperature, humidity, light intensity). The process of acquiring this data, such as controlling cameras, microphones, lighting, other environmental sensors, data storage, and data compression, may be performed by an acquisition system 700. The acquired video data may be input to a tracking system capable of implementing a convolutional neural network (CNN) to track one or more animals based on the video data.
[0061] I. Video Data Acquisition
[0110] In one embodiment, a system for capturing video data including animal movements and As discussed below, video data may be acquired continuously over a predetermined period of time (e.g., one or more minutes, one or more hours, one or more days, one or more weeks, one or more months, one or more years, etc.). The characteristics of the video data may be sufficient to facilitate subsequent analysis for the extraction of behavioral patterns, including, but not limited to, one or more of the following: resolution, frame rate, and bit depth. A practical solution is provided, which appears to be more robust and of higher quality than existing video capture systems. Embodiments of the present disclosure are tested with multiple methods of visually marking mice. A practical example of synchronous acquisition of video and ultrasonic vocalization data is also presented.
[0062]
[0111] In one embodiment, the animals are monitored for a period of approximately 4 to 6 weeks. A video monitoring system for [purpose] can be deployed. The deployment may include one or more of image capture and arena design, fine-tuning of chamber design, development of video acquisition software, acquisition of audio data, load testing of cameras, chambers, and software, and determination of chamber production in the deployment stage. Each of these will be described in detail below. The aforementioned 4-6 week observation period is provided for illustrative purposes, and it can be understood that embodiments of the present disclosure can also be employed for longer or shorter periods as necessary.
[0063] a. Arena design
[0112] Proper arena design can be important for obtaining high-quality behavioral data. This arena is the "dwelling" of the animal and can be configured to provide one or more of separation from environmental disturbances, appropriate circadian lighting, food, water, and bedding. Also, it is generally an environment without stress.
[0064]
[0113] From a behavioral perspective, the arena should desirably minimize stress and environmental disturbances while also being able to exhibit natural behavior.
[0114] From a breeding perspective, the arena should desirably facilitate cleaning, addition or removal, removal of mice, and addition and removal of food and water.
[0065]
[0115] From a veterinary perspective, the arena should desirably facilitate monitoring of environmental conditions (such as temperature, humidity, light, etc.) in addition to providing health diagnosis and treatment without substantially inhibiting the behavior of interest.
[0066]
[0116] From a computer vision perspective, the arena should desirably facilitate acquisition of high-quality video and audio without substantial occlusion, distortion, reflection, and / or noise pollution, and without substantially interfering with the expression of the behavior of interest.
[0067]
[0117] From a facility perspective, the arena should desirably provide relatively easy storage with a substantially minimized floor area and without the need for disassembly or reassembly.
[0068]
[0118] Thus, the arena can be configured to provide a balance of behavior, breeding, computing, and facilities. An exemplary embodiment of arena 200 is shown in FIG. 2. Arena 200 may include a frame 202 on which a housing 204 is mounted. The housing 204 may include a door 206 configured to allow access to the interior. One or more cameras 210 and / or lighting 212 can be attached adjacent to the frame 202 (e.g., above the housing 204) or can be attached directly to the frame 202.
[0069]
[0119] As discussed in more detail below, in certain embodiments, the lighting 212 can include at least two sets of light sources. Each set of light sources can include one or more lights configured to emit light incident on the housing 204 at a wavelength different from the other set. As an example, a first set of light sources can be configured to emit light at one or more visible wavelengths (e.g., approximately 390 nm to approximately 700 nm), and a second set of light sources can be configured to emit light at one or more infrared (IR) wavelengths (e.g., greater than approximately 700 nm to approximately 1 mm).
[0070]
[0120] The camera 210 and / or lighting 212 are connected to a user interface 214 It can be electrically connected. As the user interface 214, a display configured to display video data acquired by the camera 210 is possible. In a specific embodiment, as the user interface 214, a touch screen display configured to display one or more user interfaces for controlling the camera 210 and / or the lighting 212 is possible.
[0071]
[0121] As an alternative or addition to the above, the camera 110, the lighting 212, and the user interface 214 can be electrically connected to the controller 216. The controller 216 can be configured to generate a control signal that operates to control the acquisition of video data by the camera 210, the emission of light by the lighting 212, and / or the display of the acquired video data by the user interface 214. In a specific embodiment, the user interface can be optionally omitted.
[0072]
[0122] Also, the controller 216 can communicate with the data storage device 220. The controller 216 can be configured to receive the video data acquired by the camera 210 and transmit and store the acquired video data in the data storage device 220. Communication between one or more of the camera 210, the lighting 212, the user interface 214, the controller 216, and the data storage device 220 can be performed using a wired communication link, a wireless communication link, and combinations thereof.
[0073]
[0123] As discussed below, the arena 200 can have an open field design configured to achieve a desired balance of behavior, breeding, operation, and equipment while enabling completion in a predetermined period (e.g., approximately five months). It can have an open field design configured to achieve a desired balance of behavior, breeding, operation, and equipment while enabling completion in a predetermined period (e.g., approximately five months).
[0074] Materials
[0124] In a specific embodiment, the housing 204 (e.g., the lower part of the housing 204) is constructed At least a portion of the material making up the housing 204 may be substantially opaque to visible light wavelengths. In this manner, visible light emitted by sources other than the light 212, as well as visual cues observable by an animal within the housing 204 (e.g., movement of objects and / or a user) may be reduced and / or substantially eliminated. In additional embodiments, the material making up the housing 204 may be substantially non-reflective to infrared wavelengths to facilitate acquisition of video data. The wall thickness of the housing 204 may be selected within a range suitable to provide mechanical support (e.g., approximately 1 / 8 inch to approximately 1 / 4 inch).
[0075]
[0125] In one embodiment, the housing 204 is made of polyvinyl chloride (PVC) or polyvinyl chloride (PVC). The arena 200 may be constructed using foam sheets formed of polyoxymethylene (POM). One example of POM is Delrin® (DuPont, Wilmington, DE, USA). Such foam sheets are beneficial because they may provide sufficient versatility and durability for long-term animal monitoring of the arena 200.
[0076]
[0126] In one embodiment, the frame 202 includes a plurality of legs 202a and a and one or more shelves 202b extending vertically (e.g., horizontally). As an example, the frame 202 can be a commercially available shelving system of a predetermined size with fixed wheels for transportation to a storage area. In one embodiment, the predetermined size is approximately 61 cm. m (2 ft) x 61 cm (2 ft) x 183 cm (6 ft) (e.g., Super Erecta Metroseal 3™, InterMetro Industries Corporation, Wilkes-Barre, PA, USA), although in other embodiments arenas of different sizes may be employed without limitation.
[0077] b. Data Acquisition
[0127] The video acquisition system may include a camera 210, a lighting 212, a user interface 214, a controller 216, and a data storage device 220. The video acquisition system may be adopted to have a predetermined balance of performance characteristics. The performance characteristics may include, but are not limited to, the frame rate of video acquisition, the bit depth, the resolution of each frame, and the spectral sensitivity in the infrared region, as well as one or more of video compression and storage. As discussed below, these parameters may be optimized to maximize the quality of the data and minimize the quantity.
[0078]
[0128] In one embodiment, the camera 210 may acquire video data having at least one of a resolution of approximately 640×480 pixels, a frame rate of approximately 29 fps, and a bit depth of approximately 8 bits. By using these video acquisition parameters, uncompressed video data of approximately 33 GB / hour may be generated. As an example, the camera 210 may be a Sentech USB2 (Sensor Technologies America, Inc., Carrollton, TX, USA). FIGS. 3A-3F show sample frames obtained from one embodiment of the video acquisition system using visible light (FIGS. 3A-3C) and infrared (IR) light (FIGS. 3D-3F).
[0079]
[0129] As discussed below, the collected video data may be compressed by the camera 210 and / or the controller 216.
[0130] In another embodiment, the video acquisition system may be configured to double the resolution of the acquired video data (for example, approximately 960×960 pixels). As shown below, four other cameras having a higher resolution than the Sentech USB were investigated. As shown below, four other cameras having a higher resolution than the Sentech USB were investigated.
[0080]
Table 1
[0081]
[0131] These cameras can differ in terms of cost, resolution, maximum frame rate, bit depth, and quantum efficiency.
[0132] Embodiments of the video acquisition system can be configured to collect monochrome, approximately 30 fps, and approximately 8-bit depth video data. According to the Shannon-Nyquist theorem, the frame rate should be at least twice the frequency of the event of interest (see, e.g., Shannon (1994)). Mouse behavior can vary from a few hertz in the case of grooming to 20 hertz in the case of rapid movement (see, e.g., Deschenes et al. (2012), Kalueff et al. (2010), Wiltschko et al. (2015)). Since grooming is observed to occur at a maximum of approximately 7 Hz, it is considered appropriate to record video at a frame rate higher than the frequency of mouse movement (e.g., approximately 29 fps) to observe most mouse behavior. However, cameras can rapidly lose sensitivity in the IR region. This loss of contrast can be overcome by increasing the level of IR light, but increasing the intensity of IR light can raise the environmental temperature.
[0082] Illumination
[0133] As described above, the illumination 212 can be one or more types, such as visible white light and infrared light It can be configured to emit light of a certain class. Visible light can be employed for illumination and programmed (e.g., by the controller 216) to provide a light-dark cycle and adjustable intensity. The ability to adjust the lighting cycle enables the simulation of the light from the sun that animals are exposed to in the wild. The length of the light and dark periods is adjusted to simulate the seasons, and a lighting shift can be performed to simulate jet lag (circadian advance and retreat) experiments. Also, high-intensity lighting can be employed to cause anxiety in certain animals, and low-intensity lighting can be employed to elicit different exploratory behaviors. Thus, the ability to temporally control the length of light and dark and the light intensity is essential for proper behavioral experiments.
[0083]
[0134] In certain embodiments, the controller 216 can be configured to request irradiation of the housing 204 with visible light having an intensity of approximately 50 lux to approximately 800 lux in the bright part of the light-dark cycle. The light intensity selected can vary according to the type of movement of the observation target. In one aspect, a relatively low intensity (e.g., approximately 200 lux to approximately 300 lux) can be employed to promote and observe the exploratory movement by mice. In the bright part of the light-dark cycle, it can be configured to request irradiation of the housing 204 with visible light having an intensity of approximately 50 lux to approximately 800 lux. The light intensity selected can vary according to the type of movement of the observation target. In one aspect, a relatively low intensity (e.g., approximately 200 lux to approximately 300 lux) can be employed to promote and observe the exploratory movement by mice.
[0084]
[0135] In certain embodiments, by using an IR long-pass filter, in the IR region, substantially all video data can be acquired by the camera 210. The IR long-pass filter can remove substantially all visible light input to the camera 210. IR light is beneficial for enabling uniform illumination of the housing 104 regardless of day or night. In the IR region, substantially all video data can be acquired by the camera 210 by using an IR long-pass filter. The IR long-pass filter can remove substantially all visible light input to the camera 210. IR light is beneficial for enabling uniform illumination of the housing 104 regardless of day or night.
[0085]
[0136] Two wavelengths of IR light (850 nm and 940 nm LEDs) were evaluated. 850 nm light exhibits a distinct red hue visible to the naked eye and can result in low-intensity exposure to animals. However, such dim light can cause mood fluctuations in mice. Therefore, 940 nm light is selected for recording.
[0086]
[0137] Recording at a wavelength of 940 nm can result in a very low quantum yield in the camera and may appear as a coarse-looking image due to high gain. Therefore, different cameras were used to evaluate various infrared illumination levels to identify the maximum light level that can be obtained without substantially increasing the temperature of the housing 204 due to infrared irradiation. In certain embodiments, the temperature of the housing 204 can be increased by approximately 5 °C or less (e.g., approximately 3 °C or less).
[0087]
[0138] Also, the Basler acA1300-60gm-NIR camera was evaluated . This camera has a spectral sensitivity at 940 nm that is approximately 3 - 4 times that of the other cameras listed in Table 1, as shown in FIGS. 4A and 4B. FIG. 4A shows the spectral sensitivity of the Sentech camera as a representative example from the perspective of relative response, and FIG. 4B shows the spectral sensitivity of the Basler camera from the perspective of quantum efficiency. Quantum efficiency is a measure of the electrons emitted in response to photons impinging on the sensor. Relative response is the quantum efficiency expressed on a scale of 0 - 1. In FIGS. 4A and 4B, the wavelength of 940 nm is further shown as a vertical line for reference.
[0088]
[0139] The visible light cycle provided by the illumination 212 can be controlled by the controller 216 or another device in communication with the illumination 212. In certain embodiments, the controller 216 can comprise an illumination control panel (Phenome Technologies, Skokie, IL). The control panel is logarithmically scaled, controllable via an RS485 interface, and has 1024 levels of illumination capable of performing dawn / dusk events. As discussed in more detail below, the control of visible light can be incorporated into the control software executed by the controller 216.
[0089] Filter
[0140] As described above, optionally, during video data acquisition, substantially all visible light is skipped by the camera To prevent reaching λ210, an IR long-pass filter can be employed. As an example, a physical IR long-pass filter can be employed with the camera 110. This configuration can provide substantially uniform illumination regardless of the light and dark phases of the arena 200.
[0090]
[0141] Filters potentially suitable for use in embodiments of the disclosed systems and methods Profiles are shown in FIG. 5 (e.g., IR pass filters 092 and 093). An IR cut filter 486 that blocks IR light is shown for comparison. Additional profiles for RG-850 (glass, Edmunds Optics) and 43-949 (plastic, laser curable, Edmunds Optics) are also considered suitable.
[0091] Lens
[0142] In one embodiment, as the camera lens, 0.847 cm (1 / 3”), 3.5 - 8 mm, f1.4 (CS mount) is possible. This lens can generate the images seen in FIGS. 3A and 3B. Similar lenses for C-mount can also be employed.
[0092] Video compression
[0143] Ignoring compression, the camera 210 can generate raw video data at a rate of approximately 1 MB / frame, approximately 30 MB / second, approximately 108 GB / hour, approximately 2.6 TB / day. When choosing a storage method, various purposes can be considered. Depending on the video situation, removing specific elements of the video before long-term storage can be a beneficial option. Also, when considering long-term storage, the application of filters or other forms of processing (e.g., by the controller 216) should be desirable. However, if the processing method is changed later, saving the original video data, i.e., the raw video data, can be a beneficial solution. An example of a video compression test is described below.
[0093]
[0144] Pixel resolution approximately 480×480, approximately 29fps, and approximately 8 bits / Regarding video data collected at approximately 100 minutes at approximately 100 pixels, multiple compression standards were evaluated. Two lossless formats tested from the raw video are Dirac and H264. H264 has a slightly smaller file size but a slightly longer encoding time. Dirac can be more widely supported by subsequent encoding into another format.
[0094]
[0145] The MPEG4 lossy format was also evaluated. Since it is closely related to H264 and is known to be able to control the bitrate well. There are two ways to set the bitrate. The first is to set a constant fixed bitrate throughout the encoded video, and the second is to set a variable bitrate based on the deviation from the original video. In ffmpeg using the MPEG4 encoder, setting the variable bitrate can be easily achieved by selecting a quality value (0 - 31 (0 is almost lossless)).
[0095]
[0146] In FIGS. 6A - 6D, three different image compression methods are compared for the original (raw) captured video frame. The original image is shown in FIG 6 A. The other three methods are shown in FIG 6 B - FIG 6 D by the difference in pixels from the original image, showing only the effect of compression. That is, the compressed image is not very different from the original image. Therefore, less difference is better, and a higher compression ratio is better. As shown in FIG 6 B, the compression performed according to the MPEG4 codec with a Q0 filter shows a compression ratio of 1 / 17. As shown in FIG 6 C, the compression performed according to the MPEG4 codec with a Q5 filter shows a compression ratio of 1 / 237. As shown in FIG 6 D, the compression performed according to the MPEG4 codec with an HQDN3D filter shows a compression ratio of 1 / 97.
[0096]
[0147] When the video data collected according to the disclosed embodiments uses the quality 0 parameter (Q0 filter (Figure 6 B), Q0 HQDN3D filter (Figure 6 D)), approximately 0.01% of the pixels have changed from the original image (the intensity has increased or decreased by a maximum of 4%). This occupies approximately 25 pixels per frame. Most of these pixels are located at the boundaries of the shadows. Naturally, this small change in the image follows the scale of the noise interfering with the camera 210 itself. At higher quality values (e.g., Q5 (Figure 6 C)), artifacts may be introduced in order to compress the video data better. These often lead to artifacts with block noise that appear when attention is not paid during compression.
[0097]
[0148] In addition to these formats, other suitable lossless formats can be generated to accommodate individual user datasets. Two of these are the FMF codec (fly movie format) and the UFMF codec (micro fly movie format). The purpose of these formats is to minimize irrelevant information and optimize readability for tracking. Since these formats are lossless and function on a fixed background model, substantial data compression was impossible due to unfiltered sensor noise. The results of this compression evaluation are shown in Table 2.
[0098]
[0098]
Table 2
[0099]
[0149] In addition to the selection of the codec for data compression, it is desirable to reduce the background noise of the image as well. Background noise is inherent in all cameras and is often referred to as dark noise, representing the reference noise within the image.
[0100]
[0150] To eliminate this noise, you can increase the exposure time, increase the aperture, and reduce the gain. However, these methods are not viable options if they directly affect the experiment. Therefore, the HQDN3D filter in ffmpeg, which incorporates spatiotemporal information and removes small fluctuations, can be adopted.
[0101]
[0151] As shown in FIG. 6B to FIG. 6D, the HQDN3D filter It is observed that the file size of the MPE with HQDN3D filter is significantly reduced (e.g., about 100 times smaller compared to the file size of the original video data). After compression with the G4 codec, the resulting average bitrate is approximately 0.34 GB / hour for the compressed video. Moreover, it has been experimentally verified that virtually all information loss is several orders of magnitude less than the artifacts from sensor noise (video acquired without a mouse). This kind of noise removal significantly improves compressibility.
[0102]
[0152] Unexpectedly, the HQDN3D filter uses convolutional neural networks, which are discussed in detail below. It has been found that the HQDN3D filter significantly improves the performance of neural network (CNN) tracking. Without being bound by theoretical constraints, we believe that this improvement is achieved because the HQDN3D filter is a variance-based background subtraction method. With low variance, it is easier to identify the foreground, resulting in high-quality tracking.
[0103] Ultrasound Audio Acquisition
[0153] Mice use ultrasonic vocalizations for social communication, mating, and It is possible to perform attacks and breeding (see, for example, Grimsley et al. (2011)). Along with olfactory and tactile stimuli, this vocalization can become one of the most prominent forms of mouse communication. Although not tested in mice, in humans, changes in voice and vocalization (aging) can define transitions such as puberty and aging (see, for example, Decoster and Debruyne (1997), Martins et al. (2014), Mueller (1997)).
[0104]
[0154] Therefore, as will be discussed in detail below, embodiments of the arena 200 may further comprise one or more microphones 222. The microphone 222 may be attached to the frame 202 and configured to acquire audio data from an animal placed in the housing 204. By using the microphone 222 in the form of a microphone array, synchronous data collection can be led. This configuration of the microphone 222 enables the identification of the vocalizing mouse. The ability to further determine the vocalizing mouse among a group of mice has been demonstrated in recent years using a microphone array (see, for example, Heckman et al. (2017), Neunuebel et al. (2015)).
[0105]
[0155] A data collection setup can be provided similar to Neunuebel et al. Four microphones can be positioned on the side of the arena capable of capturing sound. When integrated with video data, the maximum likelihood method can be used to identify the vocalizing mouse (see, for example, Zhang et al. (2008)).
[0106] Environmental sensor
[0156] In one embodiment, the arena 200 is for temperature, humidity, and / or light intensity One or more environmental sensors 224 configured to measure one or more environmental parameters such as (e.g., visible and / or IR) may be further provided. In certain embodiments, the environmental sensors 224 may be integrated and configured to measure two or more environmental parameters (see, e.g., Phenome Technologies, Skokie, IL). The environmental sensors 224 are in electrical communication with the controller 216 and can collect daily temperature and humidity data along with the light level. The collected environmental data can be output for display in a user interface indicating the lighting state, in addition to the minimum and maximum temperatures (see the description regarding the control software below).
[0107] Software control system
[0157] Software control by the controller 216 for data acquisition and light control The system can be executed. The software control system can be configured to independently collect video, audio / ultrasonic, and environmental data along with corresponding timestamps. In this way, data can be collected without interruption over any predetermined period (e.g., 1 second or more, 1 minute or more, 1 hour or more, 1 day or more days, 1 year or more years, etc.). This can enable subsequent editing or synchronous analysis or presentation of the acquired video, audio / ultrasonic, and / or environmental data.
[0108] Operating system
[0158] The selection of the operating system can be driven by the availability of drivers for various sensors For example, only the Avisoft Ultrasonic microphone driver is compatible with the Windows operating system. However, this selection may affect the following.
[0109] Inter - process communication: The options for inter - process communication are affected by the underlying operating system. Similarly, the operating system influences the selection of communication between threads. However, development on a cross - platform framework such as QT can serve as a bridge.
[0110] Access to the system clock: The method of accessing a high - resolution system clock varies for each operating system, as will be discussed in more detail below. Hardware options
[0159] In certain embodiments, the control system can be implemented in the form of a single - board computer by the controller 216. Multiple options are available, such as military - grade / industrial computers that are highly robust for continuous operation.
[0111] External clock vs. system clock
[0160] Without introducing an external clock into the system, appropriate real - time clock values can be utilized from the system clock. In the POSIX system, the clock_gettime(CLOCK_MONOTONIC,...) function can return seconds and nanoseconds. The resolution of the clock can be queried with the clock_getres() function. The clock resolution of the control system embodiment should desirably be less than approximately 33 milliseconds frame period. In one embodiment, the system clock is a Unix system.
[0112]
[0161] The GetTickCount64() system function, which is used to obtain the number of milliseconds since the system was started, has been developed. The expected resolution of this timer is approximately 10 to approximately 16 milliseconds. Although this can serve the same purpose as the clock_gettime() system call, it can be beneficial to check and account for value wrapping.
[0113]
[0162] On a Macintosh computer, access to the system clock is similarly It is possible. The following code snippet is evaluated and sub-microsecond resolution has been observed.
[0114] clock_serv_t cclock; mach_timespec_t mts; host_get_clock_service(mach_host_self(), SYSTEM_CLOCK, &cclock); clock_get_time(cclock, &mts);
[0163] In any operating system, the system call that returns the time may be adjusted periodically and may move backward. In one embodiment, a monotonically increasing system clock may be employed. GetTickCount64(), clock_gettime(), and clock_get_time() may all meet this criterion.
[0115]
[0115] Video file size
[0164] It is unlikely that the software of the camera provider saves an output file with appropriate time stamps automatically split into a proper size. In an embodiment of the controller 116, it is desirable to collect video data without interruption, read each frame from the camera 110, and provide the collected video data in a simple form. For example, the controller 116 may be configured to provide video frames of approximately 10 minutes per file in a raw format, together with a time stamp header or time stamps between frames, to the data storage device 120. And each file will be less than 2GB.
[0116]
[0116] Control system architecture
[0165] FIG. 7 is a block diagram showing the components of the acquisition system 700. In a particular embodiment, the acquisition system 700 may be executed by the controller 216. Each block represents a separate process or thread of execution. Controller process Controller process
[0166] The control process can be configured to start and stop other processes or threads. Also, the control process can be configured to provide a user interface for the acquisition system 700. The control process is configured to save a log of activities and can record errors that occur during acquisition (e.g., in the log). Also, the control process can be configured to resume a paused process or thread.
[0117]
[0167] The method of communication between components can be determined after the selection of the system OS. As a user interface for the control process, a command line interface or a graphical interface is possible. The graphical interface can be built on a portable framework such as QT that provides independence from the OS.
[0118] Video acquisition process
[0168] The video acquisition process can be configured to communicate directly with the camera 210 and save timestamped frames to the data storage device 220. The video acquisition process can minimize the possibility of frame drops by operating with high priority. The video acquisition process can be kept relatively simple by minimizing the processing between frames. Also, the video acquisition process can be configured to guarantee proper exposure at a minimum effective shutter speed by controlling the IR illumination emitted by the lighting 212.
[0119] Audio acquisition process
[0169] A separate audio acquisition process can acquire ultrasonic audio with an appropriate timestamp The audio system may be configured to acquire audio data. In one embodiment, the audio system may include an array of microphones 222 disposed in audio communication with the housing 204. In certain embodiments, one or more of the microphones 222 may be positioned within the housing 204. Each microphone of the microphone array may have one or more of the following capabilities: a sampling frequency of approximately 500 kHz, an ADC resolution of approximately 16 bits, a frequency range of approximately 10 kHz to approximately 20 kHz, and an 8th order, 210 kHz anti-aliasing filter. As an example, each microphone of the microphone array may include a Pettersson M500 microphone (Pettersson Elektronik AB, Uppsala, Sweden) or a functional equivalent thereof. As described above, audio data captured by the microphones 222 may be time-stamped and provided to the controller 216 for analysis and / or provided to the data storage device 220 for storage.
[0120] Environmental Data Acquisition Process
[0170] A separate environmental data acquisition process collects environmental data such as temperature, humidity, and light levels. The environmental data may be collected at a low frequency (e.g., approximately 0.01 Hz to 0.1 Hz). The environmental data may be stored by the data storage device 220 (e.g., as one or more CSV files) with a timestamp for each record.
[0121] Lighting Control Process
[0171] The lighting control process controls the lighting 212 to provide a day-night cycle for the mice. In one embodiment, as described above, the camera 210 is configured to filter out substantially all visible light and respond only to IR, and this process can be avoided by affecting video capture since the visible light can be filtered so that no IR occurs.
[0122] Video Editing Process
[0172] The video editing process can be configured to repackage the acquired video data into a predetermined compression and a predetermined format. This process can be separated from video acquisition to minimize the chance of frame drops. The video editing process can operate as a low-priority background task or after data acquisition is complete.
[0123] Watchdog process
[0173] The watchdog process can be configured to monitor the soundness of the data acquisition process. As an example, it can record problems (e.g., in a log) and bring about a restart if necessary. Also, the watchdog process can listen for a "heartbeat" from the monitored component. Generally, as a heartbeat, a signal can be sent to the controller 216 to confirm that the components of the system 700 are operating normally. As an example, if a component of the system 700 stops functioning, it can be detected that no heartbeat is sent from this component by the controller 216. After this detection, the controller 216 can record the event and issue an alarm. Such alarms include, but are not limited to, audio alarms and visual alarms (e.g., light, alphanumeric display, etc.). As an alternative or addition to such alarms, the controller 216 can attempt to restart the operation of the component, such as sending a re-initialization signal or switching off the power. The method of communication between the components of the system 700 and the controller 216 can vary depending on the selection of the OS.
[0124] Mouse marking
[0174] In certain embodiments, the mouse can be marked to facilitate tracking. However, as will be discussed in more detail below, the marking can be omitted and tracking can be facilitated by other techniques.
[0125]
[0175] Mouse marking for visual identification involves multiple non-trivial parameters There exists. In one embodiment, by making it invisible to the mouse itself, long-term (several weeks) marking that minimizes the impact on the mouse's communication and behavior can be performed on the mouse. As an example, long-term IR-sensitive markers that are not visible within the normal mouse's field of view can be employed.
[0126]
[0176] In an alternative embodiment, human hair color and hair bleach can be used to mark the mouse's fur. With this method, the mouse can be clearly distinguishable over several weeks and can be successfully used in behavioral experiments (see, for example, Ohayon et al. (2013)). However, the process of marking the fur requires anesthesia of the mouse, which is a process not acceptable for this mouse monitoring system. al.(2013)). However, the process of marking the fur requires anesthesia of the mouse, which is a process not acceptable for this mouse monitoring system. Anesthesia can change physiological functions, and the hair dye itself can often be a stimulant that changes the mouse's behavior. Since each DO mouse is unique, this can result in a pigment / anesthesia × genotype effect, introducing unknown variables. Anesthesia can change physiological functions, and the hair dye itself can often be a stimulant that changes the mouse's behavior. Since each DO mouse is unique, this can result in a pigment / anesthesia × genotype effect, introducing unknown variables.
[0127]
[0177] Also, yet another method using IR dye-based markers and tattoos can be adopted and optimized. Also, yet another method using IR dye-based markers and tattoos can be adopted and optimized.
[0178] In another embodiment, shaving can be employed to create a pattern on the mouse's back as a form of marking. In another embodiment, shaving can be employed to create a pattern on the mouse's back as a form of marking.
[0128] Data Storage
[0179] During the development stage, a total of less than 2 TB of data may be required. These data can include raw and compressed videos of samples by various cameras and compression methods. Thus, in addition to integrating USV·video data, data transfer of video data for as long as 7 - 10 days during the load test can be achieved. The size of the video can be reduced according to the selected compression standard. Estimated values of sample data storage are given below. During the development stage, a total of less than 2 TB of data may be required. These data can include raw and compressed videos of samples by various cameras and compression methods. Thus, in addition to integrating USV·video data, data transfer of video data for as long as 7 - 10 days during the load test can be achieved. The size of the video can be reduced according to the selected compression standard. Estimated values of sample data storage are given below. Test: One arena Up to five cameras Video duration: Approximately 1 - 2 hours each Total approximately 10GB (upper limit) Load test: One arena One camera Video duration: 14 days Resolution: Twice the current (960×960) Total approximately 2TB Production: Total 120 runs (12 - 16 arenas, 80 animals per group run, alternating experiments) Duration (each): 7 days Resolution: Twice the current (960×960) 32.25TB II. Animal Tracking
[0180] Video tracking of animals such as mice is complex and impossible to perform on genetically heterogeneous animals in existing animal monitoring systems, even in a dynamic environment, without a high level of user involvement, making large - scale experiments infeasible. As will be described below, if one attempts to track multiple different mouse strains in multiple environments using existing systems and methods, it becomes clear that these systems and methods are inappropriate for large - scale experimental datasets.
[0129]
[0129]
[0181] An exemplary dataset including mice of different coat colors such as black, agouti, albino, gray, brown, nude, and mottled patterns was used for analysis. All animals were tested according to the JAX - IACUC procedure outlined below. The mice were tested at 8 - 14 weeks of age. The dataset included 1857 videos of 59 strains, for a total of 1702 hours.
[0130]
[0130]
[0182] All animals were procured from the Jackson Laboratory's production colony. Jackson The behavior of adult mice, 8 - 14 weeks old, was tested according to the certification procedures of the Institutional Animal Care and Use Committee guidelines of the institute. As described by Kumar (2011), an open - field behavior assay was performed. Briefly, the weight of group - housed mice was measured and they were acclimated to the test room for 30 - 45 minutes before the start of video recording. In this specification, the data of the first 55 - minute locomotion is presented. When available, 8 males and 8 females were tested from each inbred strain and F1 congenic strain.
[0131]
[0183] In one aspect, it should be desirable to track multiple animals in the same open - field apparatus (e.g., Arena 200) on a white background. Examples of full - frame and cropped video images obtained by the video acquisition system are shown in the first column (full - frame) and the second column (crop) of FIG. 8A. Examples of ideal tracking frames and actual tracking frames are shown in each environment with various genetic backgrounds (the third column (ideal tracking) and the fourth column (actual tracking) of FIG. 8A).
[0132]
[0184] In another aspect, it was desirable to perform video analysis of behavior in a harsh environment, such as an embodiment of Arena 200 containing food and water containers and in the Knockout Mouse Project (KOMP2) at the Jackson Laboratory, etc. (the fifth column and the sixth column of FIG. 8A, respectively).
[0133]
[0185] In the 24 - hour apparatus, the mice were housed in Arena 200 with a white - paper bedding and food / water containers. The mice were restrained in Arena 200 and continuous recording was performed under day - night conditions by using infrared light emitted by illumination 212. The bedding and food containers were moved by the mice and the visible light emitted by illumination 212 was changed over the course of each day to simulate the day - night cycle.
[0134]
[0186] In the KOMP2 project, data was collected over a 5-year period, but for an additional analysis modality to detect the effects of walking that could not be identified by the beam break system, it was desirable to perform video-based recordings. In gait analysis, the movement of the animal is analyzed. If the animal's gait is abnormal, abnormalities in the skeleton, muscles, and / or nerves can be deduced. In the KOMP2 project, a beam break system is used in which a mouse is placed in a transparent polycarbonate box irradiated with infrared light on all sides. The matrix floor is also polycarbonate, and the underlying bench surface is dark gray. Several boxes placed at the junction of two tables enable joining, and ceiling lighting (e.g., LED lighting) can provide a unique high brightness for all boxes.
[0135]
[0187] In one aspect, the video of this dataset was attempted to be tracked using the modern open-source tracking tool Ctrax that uses background subtraction and blob detection heuristics. Ctrax abstracts the mouse for each frame with respect to five measurement criteria: the major and minor axes, the x and y positions of the center of the mouse, and the direction of the animal (Branson (2009)). Also, the MOG2 background subtraction model is utilized, in which case the software estimates both the mean and variance of the background of the video used for background subtraction. In Ctrax, the shape of the predicted foreground is used to fit an ellipse.
[0136]
[0188] In another aspect, the video of this dataset was attempted to be tracked using the commercially available tracking software LimeLight that uses its own tracking algorithm. LimeLight performs segmentation and detection using a single keyframe background model. Once the mouse is detected, LimeLight abstracts the mouse with respect to the centroid by using its own algorithm.
[0137]
[0189] This dataset poses significant challenges to these existing analysis systems . As an example, in Ctrax and LimeLight, it was difficult to handle the combination of mouse coat color and environment . Generally, environments with high contrast, such as a dark mouse (e.g., black, agouti) on a white background, yield good tracking results. However, environments with low contrast, such as a light mouse (e.g., albino, gray, or piebald mouse) on a white background, produce inadequate results. A black mouse in a white open field achieves a high foreground-background contrast, so the actual tracking closely matches the ideal. A gray mouse is often removed from the nose when it has its back to the wall because it visually resembles the arena wall. Albino mice often go undetected during tracking because they resemble the background of the arena itself. Piebald mice are split in half because their coat color is patterned. Although attempts were made to optimize and fine-tune Ctrax for each video, a significant number of poor tracking frames were still observed as shown in the fourth column (actual tracking) compared to the third column (ideal tracking) of Figure 8A. Discarding poor tracking frames is undesirable because sampling can be biased and biological interpretation can be distorted
[0138]
[0190] These errors were observed to increase significantly when the environment, such as the 24-hour environment and the KMOP2 environment, becomes less ideal for tracking . Furthermore, the distribution of errors was not random. For example, as shown in the fourth column (actual tracking) of Figure 8, it was found that tracking is extremely inaccurate when the mouse is in the corner, near the wall, or on the food dispenser, while it is less inaccurate when the mouse is in the center. Placing a food dispenser in the arena in the 24-hour environment causes tracking problems when the mouse climbs on it. Also, arenas with reflective surfaces such as KOMP2 cause errors in the tracking algorithm
[0139]
[0191] Upon further investigation of the causes of poor tracking, in most cases, improper tracking is due to the mouse It was found to be due to insufficient segmentation from the background. This included cases where the mouse was removed from the foreground or cases where the background was included in the foreground due to insufficient contrast. Conventionally, some of these hurdles have been addressed by changing the environment for optimized video data collection. For example, to track albino mice, the background color of the open field can be changed to black to increase contrast. However, such environmental changes are not suitable in this context. Since the color of the environment affects the behavior of mice and humans, such operations can potentially confound the experimental results (Valdez (1994), Kulesskaya (2014)). Also, in a 24-hour data collection system or a KOMP2 arena, such a solution may not work for mice with a mottled pattern.
[0140]
[0192] Since Ctrax uses the algorithm of a single background model, a test was performed to determine whether other background models could improve the tracking results. 26 different segmentation algorithms (Sobral (2013)) were tested and, as shown in Figure 8B, it was found that these conventional algorithms each functioned well under specific circumstances and failed in other places. Other available systems and methods for animal tracking rely on background subtraction techniques for tracking. Since all 26 background subtraction methods failed, the results of Ctrax and LimeLight are considered to represent these other technologies. These segmentation algorithms are thought to fail due to improper segmentation.
[0141]
[0193] Thus, there are many tracking solutions for the analysis of video data However, attempts to achieve high-quality mouse tracking by overcoming the basic problems related to proper mouse segmentation using representative examples of existing solutions have not been successful. Since there is no solution that appropriately addresses the basic problems related to mouse segmentation and realizes proper segmentation largely relying on environmental optimization, potential confusion occurs.
[0142]
[0194] Furthermore, the time cost of finely adjusting the parameters of the background subtraction algorithm can be exorbitant. For example, in tracking data with a 24-hour setting, if a mouse sleeps in the same posture for a long time, it becomes part of the background model and cannot be tracked. In normal monitoring, an experienced user interacts for 5 minutes every hour of the video to ensure high-quality tracking results. This level of user interaction is manageable in small and limited experiments, but in large-scale and long-term experiments, it requires a long time commitment to monitor the tracking performance.
[0143]
[0195] Embodiments of the present disclosure overcome these difficulties and construct a robust next-generation tracker suitable for the analysis of video data including animals such as mice. As will be discussed in detail below, an artificial neural network is employed that achieves high performance under complex and dynamic environmental conditions, regardless of the genetic characteristics of the coat color, and does not require continuous fine-tuning by the user.
[0144]
[0144]
[0196] Convolutional neural networks learn the representation of data with multiple levels of abstraction. It is a computing model that includes multiple processing layers to be learned. These methods have dramatically improved many other areas such as state-of-the-art speech recognition, visual object recognition, object detection, and drug discovery and genomics (LeCun (2015)). In one advantage, once an efficient network with suitable hyperparameters is developed, the neural network can be easily extended to other tasks just by adding appropriate training data. Therefore, the disclosed embodiments provide a highly generalizable solution for mouse tracking. Neural network architecture
[0197] Three main neural network architectures for solving the problem of visual tracking have been disclosed. In one embodiment, as shown in FIG. 8C, based on the segmentation mask, object tracking can take the form of an elliptical description of the mouse (see Branson (2005)). In an alternative embodiment, shapes other than an ellipse can be adopted.
[0145]
[0198] The elliptical representation can describe the position of the animal by six variables, also referred to herein as parameters. In one aspect, as one of the variables, coordinates can be defined that specify a position in a predefined coordinate system (e.g., x and y of a Cartesian coordinate system) representing the pixel position of the mouse (e.g., the average center position) in the acquired video frame. That is, a unique pixel position in the plane. Optionally, landmarks (e.g., the corners of the housing 204) in the video frame can be detected as needed to assist in determining the coordinates. In another aspect, the variables can further include the lengths of the major and minor axes of the mouse, and the sine and cosine of the vector angle of the major axis. This angle can be defined with respect to the direction of the major axis. The major axis can extend from near the tip of the animal's head (e.g., the nose) to the end of the animal's body (e.g., near the point where the animal's tail extends from the body) in the coordinate system of the video frame. For clarity herein, while a cropped frame is shown as the input to the neural network, the actual input is the unmarked full frame.
[0146]
[0199] Exemplary systems and methods for determining elliptical parameters using a neural network architecture are discussed in detail below. It can be appreciated that other parameters may be utilized and determined according to the disclosed embodiments as needed.
[0147]
[0147]
[0200] In one embodiment, the first architecture is an encoder-decoder segmentation network. As shown in FIG. 9, this network predicts a foreground-background segmented image from a given input frame and, as an output, a segmentation mask that can predict whether a mouse is present or not from the perspective of pixels.
[0148]
[0148]
[0201] This first architecture comprises a feature encoder configured to abstract the input into a set of features of a small spatial resolution (e.g., 5×5 for 480×480). Many parameters are assigned to the neural network for learning. The learning can be performed by supervised training, in which case examples are presented to the neural network and the correct predictions are produced by adjusting the parameters. The definitions of the final model and the training hyperparameters are all described in Table 3 below.
[0149]
[0149]
Table 3
[0150]
[0202] The feature encoder is followed by a feature decoder configured to return a set of features of a small spatial resolution to the same shape as the original input image. That is, the parameters learned in the neural network reverse the feature encoding operation.
[0151]
[0151]
[0203] Three fully connected layers are added to the encoded features to predict the basic direction in which the ellipse is facing It is done. The fully connected layer can represent a neural network layer where different parameters (e.g., learnable parameters) are multiplied by each number of a given layer, and the sum results in a single value for the new layer. This feature decoder can be trained to generate a foreground-background segmented image.
[0152]
[0204] The first half of the network (encoder) utilizes 2D convolutional layers and 2D max pooling layers followed by batch normalization and ReLu activation. For details, see Goodfellow (2016).
[0153]
[0205] 8 was adopted as the starting filter size that doubles after each pooling layer. The kernels used are 5×5 in the case of 2D convolutional layers and 2×2 in the case of max pooling layers. The input video has a shape of 480×480×1 (e.g., monochrome), and after repeating these layers 6 times, the resulting shape is 15×15×128 (e.g., 128 colors).
[0154]
[0206] In an alternative embodiment, other shaped pooling layers such as 3×3 can be adopted. The repeating layer represents a layer with a repeating structure. The neural network learns different parameters for each layer, and the layers are stacked. Although 6 repeating layers were described above, the number of repeating layers adopted can be more or less than this.
[0155]
[0207] After applying another 2D convolutional layer (kernel 5×5, 2× filter), A different 3×3 kernel and 2D max pooling with a stride of 3 are applied. The 15×15 spatial shape can be further reduced by the use of a factor of 3. The normal max pooling has a kernel of 2×2 and a stride of 2, but each 2×2 grid selects the maximum value and generates one value. These settings select the maximum value in a 3×3 grid.
[0156]
[0208] The final 2D convolutional layer is applied, and a feature bottleneck of the shape 5×5×512 is generated. The feature bottleneck represents the encoded feature set, and the actual matrix values are output by all these matrix operations. The learning algorithm optimizes the encoded feature set to be most significant for the task it is trained to perform such that the encoded feature set works well. This feature bottleneck is then passed on to both the segmentation decoder and the angle predictor.
[0157]
[0157]
[0209] The segmentation decoder uses a strided transpose 2D convolutional layer to reverse the encoder and bypass the pre-downsampling activations by a summation junction. Note that this decoder does not utilize ReLu activation. The pre-downsampling activations and the summation junction can also be referred to as skip connections. After the features in the layer that decodes to match the same shape as the encoder layer, the network can select either better encoding or state retention at the encoder state.
[0158]
[0158]
[0210] After the layer returns to a shape of 480×480×8, another with a kernel size of 1×1 By applying intermediate convolutions, the depth becomes two monochrome images (a background prediction and a foreground prediction). The final output is 480x480x2 (2 colors). The first color is designated to represent the background. The second color is designated to represent the foreground. For each pixel, the network considers the larger of the two as the input pixel. As discussed below, a softmax operation rescales these colors so that their cumulative probabilities sum to 1.
[0159]
[0211] Then softmax is applied over this depth. It is a form of classification into groups or binmin. Further information on softmax can be found in Goodfellow (2016).
[0160]
[0212] The feature bottleneck also generates an angle prediction, which is a function of two 2D convolutions. This is achieved by applying batch normalization and ReLu activation to the flattening layer (kernel size 5 × 5, feature depth 128 and 64). From here, one fully connected layer is flattened and used to generate a 4-neuron shape that acts to predict the quadrant the mouse's head will face. Further details on batch normalization, ReLu activation, and flattening can be found in Goodfellow (2016).
[0161]
[0213] The correct orientation (± Only one direction (+180°, 135°, 225°, 315°, 315°, 45°, 180°) needs to be selected; that is, since an ellipse is predicted, there is only one major axis. One end of the major axis is in the direction of the mouse's head. The mouse is assumed to be longer along the head-tail axis. Thus, one direction is +180° (head) and the other is -180° (tail). The four possible directions that the encoder-decoder neural network architecture can choose from are 45 to 135°, 135 to 225°, 225 to 315°, and 315 to 45° on a polar grid.
[0162]
[0214] These boundaries are selected to avoid discontinuities in the angle prediction. In particular, as described above, the angle prediction is a prediction of the sine and cosine of the vector angle of the major axis, and the atan2 function is employed. The atan2 function is discontinuous (at 180°), and the selected boundaries avoid these discontinuities.
[0163]
[0215] After the network has generated the segmentation mask, an ellipse fitting algorithm can be applied for tracking, as described in Branson ( 2009). Branson uses weighted sample means and variances for these calculations, but the segmentation neural network remains invariant to situations representing improvements. For the segmentation mask generated by the background subtraction algorithm, projected shadows may add errors. The neural network learns to exclude all these problems. Also, no significant difference is observed between the use of weighted and unweighted sample means and variances. The ellipse fitting parameters predicted by weighted and unweighted methods do not differ significantly by using the mask predicted by the disclosed neural network embodiments.
[0164]
[0216] Given the segmentation mask, the sample mean of the pixel positions is calculated to represent the center position as such.
[0165]
Number
[0166] Similarly, the sample variance of the pixel positions is calculated to represent the length of the major axis (a), the length of the minor axis (b), and the angle (θ).
[0167]
Number
[0168] To determine the axis length and angle, it is necessary to solve the eigenvalue decomposition equation.
[0169]
Math
[0170]
Math
[0171]
[0217] The second network architecture is the binning classification network. As shown in FIG. 10, the structure of the binning classification network architecture can predict the heat map of the most accurate value of each ellipse fitting parameter.
[0172]
[0218] This network architecture starts with a feature encoder that abstracts the input image to a small spatial resolution. While most of the regression predictors realize the solution with a bounding box (e.g., a square or rectangle), for an ellipse, only an additional parameter, the angle, is added. Since the angle is a repeating number that is equivalent at 360° and 0°, the angle parameter is converted to its sine and cosine components. This results in a total of six parameters regressed from the network. The first half of this network encodes a set of features related to solving the problem.
[0173]
[0219] The encoded features are flattened by converting the matrix (array) representing the features into a single vector. Then, the flattened encoded features are connected to an additional fully connected layer whose output shape is determined by the desired output resolution (e.g., by inputting the vector of features into the fully connected layer). For example, in the case of the X coordinate position of a mouse, there are 480 bins, one for each x column of a 480×480 pixel image.
[0220] When the network runs, the maximum value in each heatmap is selected as the most probable value. Each desired output parameter can be realized as a set of independent trainable fully connected layers connected to the coding features.
[0175]
[0221] Resnet V2 50, Resnet V2 101, Resnet V A wide variety of pre-built feature detectors were tested, including Resnet V2 200, Inception V3, Inception V4, VGG, and Alexnet. A feature detector represents a convolution that operates on the input image. In addition to these pre-built feature detectors, a wide variety of custom networks were also investigated. From this investigation, it was observed that Resnet V2 200 performed the best.
[0176]
[0222] The final architecture is the recurrent network shown in Figure 11. Then, the regression network takes the input video frame, extracts features by Resnet200 CNN, and directly predicts six parameters for ellipse fitting. Each value (six for ellipse fitting) is continuous and can have an infinite range. The network needs to learn the appropriate range of values. In this way, the numerical values of the ellipse describing the tracking ellipse are predicted directly from the input image. That is, instead of predicting the parameters directly, the regression network instead selects the most probable value from a selection of possible binning values.
[0177]
[0223] Other neural network architectures work differently. The encoder-decoder neural network architecture outputs the probability that each pixel is a mouse or not. The binning classification neural network architecture outputs a bin that represents the location of the mouse. The class of each parameter is pre-determined, and the network (encoder-decoder or binning) only needs to output the probability of each class.
[0178]
[0224] The regression network architecture starts with a feature encoder that abstracts the input into a small spatial resolution. In contrast to the above architecture, regression neural network training relies on a cross - entropy loss function, as opposed to the mean squared error loss function.
[0179]
[0225] Due to memory constraints, the feature dimension was reduced and only a custom VGG - like network was tested. The best - functioning network was structured with a 2D max - pooling layer after two 2D convolutional layers. The kernels used are 3×3 in the case of the 2D convolutional layer and 2×2 in the case of the 2D max - pooling layer. The initial filter depth is 16, which is doubled for each 2D max - pool layer. This sequence of two convolutional + max - pool is repeated 5 times, resulting in a shape of 15×15×256.
[0180]
[0226] This layer is flattened and connected to a fully - connected layer for each output. The shape of each output is determined by the desired resolution and range of the prediction. As an example, these encoded features were then flattened and connected to a fully - connected layer, resulting in an output shape of 6, which is the number of values the network was required to predict to fit an ellipse. For testing purposes, only the center position was observed and trained on a wide overall image (0 - 480). Additional outputs such as angle prediction can be easily added as additional output vectors. A variety of modern feature encoders were tested, but the data discussed in this document for this network is derived from a 200 - layer Resnet V2 (He(2016)) that achieved the best - functioning results for this architecture.
[0181] Training dataset
[0227] To test the network architecture, as described below, OpenCV Using the base labeling interface, a training dataset consisting of 16,234 training images and 568 separate validation images spanning multiple strains and environments was generated. This labeling interface enables foreground and background fast labeling, as well as ellipse fitting, and can be used to immediately generate training data and adapt any network to new experimental conditions by transfer learning.
[0182]
[0228] Interactive watershed-based segmentation and contour-based ellipse fitting were generated using the OpenCV library. By using this software, the user can mark points as the foreground (e.g., mouse (F)) by left-clicking and label other points as the background (B) by right-clicking, as shown in Figure 12A. The watershed algorithm is executed by keystrokes, and segmentation and ellipses are predicted, as shown in Figure 12B. If the user needs to edit the predicted segmentation and ellipses, they only need to label additional areas and run the watershed again.
[0183]
[0229] If the prediction is within the pre-determined error tolerance selected by the user of the neural network (e.g., researcher), the user selects the direction of the ellipse. The user makes the selection by choosing one of the four basic directions (up, down, left, right). Since the exact angle is selected by the ellipse fitting algorithm, the user only needs to distinguish ±90° of the direction. Once the direction is selected, all relevant data is saved and the user is presented with a new frame to label.
[0184]
[0184]
[0230] The purpose of the labeled dataset is to provide good ellipse fitting tracking for the mouse It is to identify data. During data labeling, ellipse fitting was optimized such that the center of the ellipse was the mouse's body with the end of the major axis in a state of slightly touching the mouse's nose. The tail was often removed from the segmentation mask to provide better ellipse fitting.
[0185]
[0231] To train the inference network, three labeled training sets were generated. Each dataset included a reference frame (input), a segmentation mask, and an ellipse fitting. The training sets were each generated to track mice in different environments.
[0186]
[0232] The first environment was a uniform white background open field containing 16,802 annotated frames. The first 16,000 frames were labeled by 65 separate videos from one of 24 identical setups. After the first training of the network, it was observed that the network did not function well under special circumstances not included in the labeled data. Cases of intermediate jumps, irregular postures, and urination in the arena were usually observed as unsuccessful. These unsuccessful cases were identified, correctly labeled, and incorporated into the labeled training set to further generalize and improve performance.
[0187]
[0233] The second environment was a standard open field with an α-dri bed and food dispenser under two different lighting conditions (visible irradiation during the day and infrared irradiation at night). In this dataset, a total of 2,192 frames were labeled over 4 days across 6 setups. 916 of the annotated frames were obtained from night-time irradiation and 1,276 of the annotated frames were obtained from daytime irradiation.
[0188]
[0234] The final labeled dataset is the Opto- It was generated by using an M4 open-field cage. The dataset contained 1083 labeled frames, all sampled across different videos (one frame per video labeled) and eight different setups.
[0189] Neural Network Training a) Expanding the training dataset
[0235] This training dataset is trained by applying reflections. The training dataset is expanded by a factor of 8 during training, and small random changes in contrast, brightness, and rotation are applied to make the network robust to small variations in the input data. This expansion is performed to prevent the neural network from memorizing the training dataset, which would cause it to perform poorly on examples not included in the dataset (validation). Further details can be found in Krizhevsky (2012).
[0190]
[0236] Training set expansion has been used extensively in neural networks since Alexnet. This is an important aspect of training for the ML model (Krizhevsky (2012)). A handful of training set augmentations are utilized to achieve good regularization performance. As the data comes from a bird's-eye view, it is easy to apply horizontal, vertical, and oblique reflections to instantly increase the training set size by a factor of 8. Also, at runtime, slight rotations and translations are applied to the entire frame. The rotation augmentation values are sampled from a uniform distribution. Finally, noise, brightness, and contrast augmentations may also be applied to the frames. The random values used for these augmentations are chosen from normal distributions.
[0191] b) Learning rate and batch size for training
[0237] The learning rate and batch size of the training were selected independently for each network training. Large networks such as Resnet V2 200 may fall into memory constraints of the batch size at an input size of 480×480, but good learning rates and batch sizes were experimentally identified using a grid search method. The hyperparameters selected for training these networks are shown in Table 3 above. Model construction, training, and testing were performed in TensorFlow v1.0. The presented training benchmarks were run on the NVIDIA® Tesla® P100 GPU architecture. The hyperparameters were trained through multiple training iterations. After the first training of the network, it was observed that the network was not functioning well under special circumstances where it was undervalued in the training data. Cases of intermediate jumps, irregular postures, and urination in the arena were usually observed as unsuccessful. These difficult frames were identified and incorporated into the training dataset to further improve performance. The complete description of the final model definition and all training hyperparameters are described in Table 3 above.
[0192] Model
[0238] In TensorFlow v1.0, model construction, training, and testing were performed. The presented training benchmarks were run on the NVIDIA® Tesla® P100 GPU architecture.
[0193]
[0193]
[0239] The hyperparameters were trained through multiple training iterations. After the first training of the network, it was observed that the network was not functioning well under special circumstances where it was undervalued in the training data. Cases of intermediate jumps, irregular postures, and urination in the arena were usually observed as unsuccessful. These difficult frames were identified and incorporated into the training dataset to further improve performance. The complete description of the final model definition and all training hyperparameters are described in Table 3 above. The training and validation loss curve plots shown by three full networks
[0194]
[0240] The training and validation loss curve plots shown by three full networks The results are shown in FIGS. 13A to 13E. Overall, the training and validation loss curves indicate that the three full networks are trained to achieve a performance of an average error of 1 to 2 pixels. Unexpectedly, the binning classification network exhibits an unstable loss curve, indicating overfitting during validation and insufficient generalization (FIGS. 13B and 13E). The regression architecture converged to a validation error of 1.2 pixels, which indicates better training performance than validation (FIGS. 13A, 13B, and 13D). However, the best-performing feature extractor, Resnet V2 200, is a large-scale deep network with over 200 layers and 62.7 million parameters, and the processing time per frame is substantially longer (33.6 ms). Other pre-built general-purpose networks (Zoph (2017)) can only achieve similar suboptimal performance in exchange for short computation times. Thus, the regression network is an accurate but computationally expensive solution.
[0195]
[0241] As further shown in FIGS. 13A, 13B, and 13C, the encoder-decoder segmentation architecture converged to a validation error of 0.9 pixels. The segmentation architecture not only functions well but also has good computational efficiency for GPU operations with an average processing time of 5 to 6 ms / frame. Video data can be processed at up to 200 fps (6 times real time) on an Nvidia® Tesla® P100, a server-level GPU, and 125 fps (4.2 times real time) on an Nvidia® Titan Xp, a consumer-level GPU. This high processing speed is thought to be due to the fact that the depth of the structure is only 18 layers and the number of parameters is only 10.6 million.
[0196]
[0242] Encoder-Decoder Segmentation Network Architecture To identify the relative scale of the labeled training data required for good network performance, a benchmark of the training set size was also performed. This benchmark was tested by shuffling and randomly sampling subsets of the training set (e.g., 10,000, 5,000, 2,500, 1,000, and 500). Each subsampled training set was trained and compared to the same validation set. The results of this benchmark are shown in FIGS. 14A - 14H.
[0197]
[0243] Generally, the training curves appear indistinguishable (FIG. 14A). That is , the training set size shows no performance change with respect to the error rate of the training set (FIG. 14A). Surprisingly, while the validation performance converges to the same value with more than 2,500 training samples, the error increases with less than 1,000 training samples (FIG. 14B). As further illustrated, with more than 2,500 training samples, the validation accuracy is better than the training accuracy (FIGS. 14C - 14F), while after matching the training accuracy at 1,000, signs of weak generalization are starting to show (FIG. 14G). Using only 500 training samples is clearly overfitting, as indicated by the diverging and increasing validation error rate (FIG. 14H). This suggests that the training set is no longer large enough for the network to generalize well. Thus, good results are obtained only from networks trained with only 2,500 labeled images, which takes about 3 hours to generate at the labeling interface. Therefore, while the exact number of training samples ultimately depends on the difficulty of the visual problem, the recommended starting point for the number of training samples is around 2,500.
[0198]
[0244] An exemplary video frame showing a mouse tracked according to an embodiment of the disclosure is In the case of visible light, it is shown in FIGS. 15A and 15B, and in the case of infrared light, it is shown in FIGS. 15C and 15D. As shown, the spatial range of each mouse is color-coded in pixel units.
[0199]
[0245] Given the computational efficiency, accuracy, training stability, and a small number of required training data sets, the encoder-decoder segmentation architecture was selected for predicting the position of the mouse throughout the video for comparison with other methods.
[0200]
[0246] The quality of neural network-based tracking was evaluated by inferring the entire video from mice with different coat colors and data collection environments (FIG. 8A) and visually assessing the quality of the tracking. Neural network-based tracking was also compared with the KOMP2 beam break system, an independent tracking mode (FIG. 8A, 6th column).
[0201]
[0201] Experimental arena a) Open field arena
[0247] One embodiment of arena 200 was adopted as an open field arena. The open field arena is 52 cm × 52 cm. The floor is white PVC plastic and the walls are gray PVC plastic. A 2.54 cm white surface was added to all inner edges to facilitate cleaning and maintenance. Illumination is provided by an LED lighting ring (model F&V R300). The lighting ring was calibrated to produce 600 lux of light per arena.
[0202] b) Open field arena for 24-hour monitoring
[0248] The open field arena was extended for multi-day testing. Lighting 212 It is a form of ceiling LED lighting set to a standard 12:12 LD cycle. The α-dry was placed in the arena as a bedding. A single Diet Gel 76A feeder was placed in the arena to provide food and water. This nutrient source was monitored and replaced when it ran out. Each matrix was irradiated at 250 lux during the day and at less than approximately 500 lux at night. For night-time video recording, the lighting 212 was assumed to include IR LED (940 nm) lighting.
[0203] c) KOMP open-field arena
[0249] In addition to the custom arena, embodiments of the disclosed systems and methods were benchmarked against commercially available systems. By using transparent plastic walls, an Opto-M4 open-field cage was constructed. As a result, visual tracking is very difficult due to the resulting reflections. The cage is 42 cm × 42 cm. The lighting of this arena was assumed to be performed by LED irradiation of 100 - 200 lux.
[0204]
[0204] Video acquisition
[0250] All video data was obtained using the video acquisition system discussed with respect to FIGS. 2 and 7. Obtained according to an embodiment of the item. The video data was acquired using a camera 210 in the form of a Sentech camera (model STC-MB33USB) and a computer lens (model T3Z2910CS-IR) at a resolution of 640×480 pixels, 8-bit monochrome depth, and approximately 29 fps (e.g., approximately 29.9 fps). The exposure time and gain were digitally controlled using a target luminance of 190 / 255. The aperture was adjusted to be the widest so that a low analog gain was used to achieve the target luminance. This suppresses the amplification of reference noise. The files were temporarily stored on a local hard drive using the "raw video" codec and the "pal8" pixel format. The assay ran for approximately 2 hours and generated approximately 50 GB of raw video files. The ffmpeg software was used overnight to apply a noise removal filter to a 480×480 pixel crop and perform compression using an MPEG4 codec (quality set to maximum) to generate a compressed video size of approximately 600 MB.
[0205]
[0251] To relieve projection distortion, the camera 210 was attached to the frame 202 approximately 100 cm above the shelf portion 202b. The zoom and focus were manually set to achieve a zoom of 8 pixels / cm. This resolution minimizes the unused pixels on the arena boundary and generates an area of approximately 800 pixels per mouse. Although the KOMP arena is slightly smaller, the same target zoom of 8 pixels / cm was utilized.
[0206]
[0252] Using an encoder-decoder segmentation neural network By doing so, 2,002 videos (a total of 700 hours) were tracked from the KOMP2 dataset, and the results are shown in Figure 8. These data included 232 knockout lines tested in a 20-minute open-field assay on a C57BL / 6NJ background. Since each KOMP2 arena had a slightly different background due to the transparent matrix, the tracking performance was compared for each of the eight test chambers (on average n = 250 (Figure 16)) and for all combination boxes. Over the eight full test chambers used by KOMP2, a very high correlation was observed between the two methods for the total distance traveled within the open field (R = 96.9%). From this trend (red arrow), two animals were observed with a high discrepancy. The video observations showed irregular postures present in both animals, one with a staggering gait and the other with a hunched posture. The staggering gait and the hunched gait are thought to result in abnormal beam breaks and an abnormally high total distance traveled measure from the beam break system. This example highlights one of the advantages of a neural network that is not affected by the posture of the animal.
[0207]
[0253] Regarding the performance of the trained segmentation neural network even so, it was compared with Ctrax over a wide range of videos from various test environments and over the entire coat color described above for Figure 8A. The comparison with Ctrax was motivated by many reasons. On one hand, Ctrax is considered one of the best-performing conventional trackers that allows for fine-tuning of many tracking settings. Also, Ctrax is open source and provides user support. Given the results by the BGS library (Figure 8B), similar performance as below is expected for other trackers. Twelve animals per group were tracked with both the trained segmentation neural network and Ctrax. The settings of Ctrax were fine-tuned for every 72 videos as described below.
[0208]
[0254] Ctrax includes various settings for optimizing the tracking ability (Branso n(2009)). The author of this software strongly recommends that the arena be set up under specific criteria to ensure good tracking. In most of the tests discussed in this specification (e.g., albino mice on a white background), an environment is adopted in which Ctrax is not designed to function well. Nevertheless, good performance can still be achieved by adequately adjusting the parameters. With many settings for operation, Ctrax can easily incur a high time cost to achieve good tracking performance. The setup procedure of Ctrax for tracking a mouse in the disclosed environment is as follows.
[0209]
[0255] In the first operation, a background model is generated. Since the core of Ctrax is based on background subtraction, it is functionally essential to have a robust background model. The model functions optimally when the mouse is moving. To generate the background model, the part of the video where the mouse is clearly moving is searched, and frames are sampled from that part. As a result, the mouse is not included in the background model. This method considerably improves the tracking performance of Ctrax for 24-hour data. This is because the mouse does not move much and is usually incorporated into the background model.
[0210]
[0256] The second operation is to perform the background subtraction setting. Here, the standard range is 254 A background luminance normalization method of .9 to 255.0 is used. The thresholds applied to separate the mice are adjusted based on the preliminary video, as slight changes in exposure and coat color affect performance. To adjust these thresholds, a set of good starting values is applied and the video is scrutinized to ensure generally good performance. In certain embodiments, all videos may be checked for cases of mice against a wall, as these are usually the most difficult frames to track due to shadows. Morphological filtering may also be applied to remove faint changes in the environment as well as to remove the mouse's tail for ellipse fitting. An aperture radius of 4 and an occlusion radius of 5 were adopted.
[0211]
[0257] In another move, Ctrax allows you to make the observations effectively mouse. Various tracking parameters that can be used are manually adjusted. For time considerations, these parameters are fully adjusted when and after being used for all other tracked mice. If the video was visibly underperforming, the general settings were tweaked to improve performance. For the shape parameters, a range based on two standard deviations was determined from the individual black mouse videos. The minimum was further lowered as it was expected that certain mice would not perform well in the segmentation step. This allows Ctrax to still find a good position of the mouse even though it is not possible to segment the whole mouse. This method works well since all setups have the same zoom of 8 and the mice tested have roughly the same shape. In the experimental setup, the motion settings are very loose since we only track one mouse in the arena. Under the observation parameters, the "Min Area Ignore" is mainly used, which removes large detections. Here, detections larger than 2,500 are removed. Under the Hindsight tab, the "Fix Spurious Detections" setting is used to remove detections shorter than 500 frames long.
[0212]
[0258] Since Ctrax could not generate a valid background model, videos from the 24-hour apparatus in which the animals slept continuously for a long time were manually omitted from the comparison. The cumulative relative error of the total movement distance between Ctrax and the neural network was calculated and shown in (Figure 17A). For each minute of the video, the movement distance predictions from both the neural network and Ctrax were compared. This measurement criterion measures the accuracy of the centroid tracking of each mouse. Tracking of black, gray, and spotted mice showed an error of less than 4%. However, significantly higher levels of error were seen in albino (14%), 24-hour arena (27% (orange)), and KOMP2 (10% (blue)) (Figure 17A). Therefore, without the neural network tracker, albino tracking, KOMP2, or 24-hour data could not be properly tracked.
[0213]
[0259] Also, it was observed that the elliptical fitting did not correctly represent the mouse's pose when the foreground segmentation prediction was incorrect, such as when shadows were included in the prediction. In these cases, even if centroid tracking was possible, the elliptical fitting itself had high variability.
[0214]
[0260] Modern machine learning software for behavior recognition, such as JAABA (Kabra (2013)), utilizes these features for behavior classification. The variance in elliptical tracking was quantified by the relative standard deviation of the minor axis and shown in Figure 17B. This measurement criterion shows the minimum variance across all experimental mice. This is because the width of an individual mouse does not vary and is similar through the wide range of poses represented in the behavior assay when the tracking is accurate. Even when the cumulative relative error of the total movement distance was small (Figure 17B), high tracking variance was observed in gray and spotted mice (Figure 17A). As expected, high relative standard deviations were observed for the minor axis in the cases of albino and KOMP2 tracking. Therefore, it can be seen that the neural network tracker is superior to the conventional tracker in terms of both centroid tracking and the variance of elliptical fitting. JAABA (Kabra(2013)) and other modern machine learning software for behavior recognition utilize these features for behavior classification. The variance in elliptical tracking is quantified by the relative standard deviation of the minor axis and shown in Figure 17B. This measurement criterion shows the minimum variance across all experimental mice. This is because the width of an individual mouse does not vary and is similar through the wide range of poses represented in the behavior assay when the tracking is accurate. Even when the cumulative relative error of the total movement distance was small (Figure 17B), high tracking variance was observed in gray and spotted mice (Figure 17A). As expected, high relative standard deviations were observed for the minor axis in the cases of albino and KOMP2 tracking. Therefore, it can be seen that the neural network tracker is superior to the conventional tracker in terms of both centroid tracking and the variance of elliptical fitting.
[0215]
[0261] An encoder-decoder segmentation neural network was constructed as a high-precision tracker, and its performance was further tested with two large behavioral datasets. Open-field video data was generated from 1845 mice (1691 hours) across 58 strains of mice, including all colors, patterns, nudity, and obese mice. This dataset included 47 inbred mouse strains and 11 F1 congenic mouse strains, and was the largest open-field dataset generated according to the Mouse Phenome Database of Bogue (2018).
[0216]
[0216]
[0262] Tracking results for total distance moved are shown in Fig. 18A. Each point represents an individual within a strain, and the box represents the mean ± standard deviation. All mice were tracked with high precision using a single trained network without user adjustment. Visual confirmation of tracking fidelity and excellent performance were observed in over half of the strains of mice. The observed locomotor phenotypes were consistent with the publicly available dataset of mouse open-field behavior.
[0217]
[0217]
[0263] The same neural network was employed to track 24-hour video data collected from 4 C57BL / 6J mice and 2 BTBR T + ltpr3 tf / J mice (fifth column in Fig. 8A). These mice were housed with bedding, food, and water for several days, during which the food location was changed and the lighting was set to 12:12 light-dark conditions. Video data was recorded using visible and infrared light sources. Under these conditions, the movement of all animals was tracked using the same network, and very good performance was observed under light and dark conditions.
[0218]
[0264] The results are shown in Fig. 18B, where eight light and dark dots represent the light and dark conditions respectively As expected, a locomotor rhythm (curve) was observed during the dark phase, accompanied by a high level of locomotor activity.
[0219]
[0265] In summary, video-based tracking of animals in complex environments is a promising tool for understanding animal behavior. This has been a long-standing challenge in the field (Egnor (2016)). Current state-of-the-art systems do not address the fundamental problem of animal segmentation and rely heavily on visual contrast between the foreground and background for accurate tracking. As a result, users must constrain the environment to achieve optimal results.
[0220]
[0266] Herein, we focus on modern neural networks that can function in complex and dynamic environments. A neural network-based tracker and corresponding methods of use are described. Through the use of trainable neural networks, a fundamental problem in tracking (foreground and background segmentation) is addressed. Testing three different architectures shows that an encoder-decoder segmentation network achieves a high level of accuracy and works fast (more than six times faster than real time).
[0221]
[0267] By labeling just 2,500 images (approximately 3 hours), required), a labeling interface is further provided that allows new networks to be trained for specific environments.
[0222]
[0268] The disclosed pre-trained neural network outperforms two existing solutions: Compared with Yon, it was found to be much better than them in a complex environment. Similar results are expected for any commercially available system that uses a background subtraction method. In fact, when 26 different background subtraction methods were tested, it was observed that each one was unsuccessful under specific circumstances. However, only one neural network architecture can function for mice of all coat colors in multiple environments without the need for fine-tuning or user input. This machine learning method forms the basis of the next generation of tracking architectures for behavioral research because it enables long-term tracking under dynamic environmental conditions with minimal user input.
[0223]
[0269] One or more aspects or features of the control system described herein may be implemented in digital electronic circuits, integrated circuits, application specific integrated circuits (ASICs) with special designs, field programmable gate arrays (FPGAs), computer hardware, firmware, software, and / or combinations thereof. These various aspects or features may include implementations in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor coupled to receive and transmit data and instructions to and from a storage system, at least one input device, and at least one output device, which may be dedicated or general purpose. Programmable systems or computer systems may include clients and servers. Clients and servers are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by computer programs operating on each computer and having a client-server relationship to each other.
[0224]
[0270] Program, software, software application, application A computer program, which may also be referred to as a routine, component, or code, includes machine instructions for a programmable processor and can be implemented in a high-level procedural language, an object-oriented programming language, a functional programming language, a logical programming language, and / or assembly / machine language. As used herein, the term "machine-readable medium" represents any computer program product, apparatus, and / or device, such as, for example, a magnetic disk, an optical disk, a memory, and a programmable logic device (PLD), etc. (including machine-readable media that receive machine instructions as machine-readable signals), used to provide machine instructions and / or data to a programmable processor. The term "machine-readable signal" represents any signal used to provide machine instructions and / or data to a programmable processor. A machine-readable medium can persistently store machine instructions as described above, such as, for example, non-transitory solid-state memory, a magnetic hard drive, or any equivalent storage medium. As an alternative or addition to this, a machine-readable medium can persistently store machine instructions as described above, such as, for example, a processor cache or other random access memory associated with one or more physical processor cores. nal) represents any signal used to provide machine instructions and / or data to a programmable processor. A machine-readable medium can persistently store machine instructions as described above, such as, for example, non-transitory solid-state memory, a magnetic hard drive, or any equivalent storage medium. As an alternative or addition to this, a machine-readable medium can persistently store machine instructions as described above, such as, for example, a processor cache or other random access memory associated with one or more physical processor cores.
[0225]
[0271] To enable interaction with a user, for example, a cathode for displaying information to the user On a computer having a display device such as a cathode ray tube (CRT), a liquid crystal display (LCD), or a light emitting diode (LED) monitor, as well as a keyboard and a pointing device (e.g., a mouse, a trackball, etc.) through which a user can provide input to the computer, one or more aspects or features of the subject matter described herein can be implemented. Other types of devices that enable interaction with the user can similarly be used. For example, as feedback provided to the user, any form of sensory feedback is possible, such as visual feedback, auditory feedback, or tactile feedback, and input from the user can be received in any form, including but not limited to acoustic, speech, or tactile input. Other conceivable input devices include, but are not limited to, touchscreens or other touch sensor-based devices such as single-point or multi-point resistive or capacitive trackpads, speech recognition hardware and software, optical scanners, optical pointers, digital image capture devices, and associated interpretation software.
[0226]
[0272] All references cited throughout this application (e.g., issued or registered patents or equivalents, published patent applications, and non-patent literature or other sources) are hereby incorporated by reference in their entirety as if each reference were individually incorporated by reference to the extent that each reference does not conflict at least partially with the disclosure of this application. For example, a reference that is partially conflicting is incorporated by reference except for the portion that conflicts.
[0227]
[0227]
[0273] In this specification, when a Markush group or other group is used, all individual elements of that group, as well as all combinations and subcombinations possible within that group, are intended to be individually included in the disclosure.
[0228]
[0228]
[0274] In this specification, the singular forms "a," "an," and "the" "The" includes a plurality of meanings unless there is a clear indication to the contrary in the context. For this reason, for example, the expression "a cell" includes a plurality of such cells and their equivalents known to those skilled in the art, and the same applies to other cases. Also, the terms "a" (or "an"), "one or more", and "at least one" can be used interchangeably herein.
[0229]
[0275] As used herein, the term "comprising" is synonymous with "including", "having", "containing", and "characterized by", and each can be used interchangeably. These terms are each further inclusive or open-ended and do not exclude additional elements or method steps not recited. (including)", "having", "containing", and "characterized by", and each can be used interchangeably. These terms are each further inclusive or open-ended and do not exclude additional elements or method steps not recited.
[0230]
[0276] As used herein, the term "consisting of" excludes any element, step, or component not specified in the claim elements.
[0231]
[0277] As used herein, the term "consisting essentially of" does not exclude elements or steps that do not substantially affect the basic and novel characteristics of the claim scope. In any case herein, the terms "comprising", "consisting essentially of", and "consisting of" can each be replaced by either of the other two terms.
[0232]
[0278] The embodiments exemplified herein are specifically disclosed herein One or more elements that are not present, one or more limitations can be preferably realized in a state where there are none at all.
[0233]
[0279] The expression "according to any one of claims XX to YY (of any of cl aims XX-YY)" (where XX and YY represent claim numbers) is intended to provide alternative forms of multiple dependent claims and, in some embodiments, can be used interchangeably with the expression "as in any one of claims XX-YY".
[0234]
[0280] Unless otherwise defined, all technical and scientific terms used in this specification have the same meaning as commonly understood by one of ordinary skill in the technical field to which the disclosed embodiments belong.
[0235]
[0281] In this specification, whenever ranges such as temperature ranges, time ranges, composition ranges, or concentration ranges are given, it is intended that all intermediate ranges and sub-ranges, as well as all individual values included in the given range, are included in this disclosure. In this specification, a range specifically includes the values provided as the endpoint values of the range. For example, the range of 1 to 100 specifically includes the endpoint values of 1 and 100. It is understood that any sub-range or individual value within the range or sub-range included in the description of this specification may be excluded from the claims.
[0236]
[0282] In the above and in the claims, "at least one of ~ (at l Expressions such as "one of (east one of)" or "one or more of" can appear with a list of elements or features following. Also, the term "and / or" can appear as a list of two or more elements or features. Unless there is an implicit or explicit contradiction in the context of use, such expressions are intended to mean either each of the elements or features in the list individually, or a combination of any of the listed elements or features with any of the other listed elements or features. For example, the expressions "at least one of A and B", "one or more of A and B", and "A and / or B" are each intended to mean "A alone, B alone, or a combination of A and B". The same interpretation is intended for lists containing three or more items. For example, the expressions "at least one of A, B, and C", "one or more of A, B, and C", and "A, B, and / or C" are each intended to mean "A alone, B alone, C alone, a combination of A and B, a combination of A and C, a combination of B and C, or a combination of A, B, and C". Also, in the above and in the claims, the use of the term "based on" is intended to mean "based at least in part on" so that features or elements not recited may be tolerated. one of A,B,and C)", "one or more of A,B,and C", and "A,B,and / or C" are each intended to mean "A alone, B alone, C alone, a combination of A and B, a combination of A and C, a combination of B and C, or a combination of A, B, and C".
[0237]
[0283] The terms and expressions employed herein are used as terms of explanation, Without being in any way limiting, and without any intention of excluding any equivalents of the features shown and described or parts thereof, it is recognized that various improvements are possible within the scope of the claimed embodiments. For this reason, although this application may include descriptions of preferred embodiments, exemplary embodiments, and optional features, it is understood that improvements and modifications of the concepts disclosed herein may be made by those skilled in the art. Such improvements and modifications are considered to be within the scope of the disclosed embodiments as defined by the appended claims. The specific embodiments described herein are examples of useful embodiments of the present disclosure, and it will be apparent to those skilled in the art that many variations of the devices, device components, and method steps described herein may be used. It is obvious to those skilled in the art that methods and devices useful in such methods may include many optional configurations, processing elements, and steps.
[0238]
[0284] Embodiments of the present disclosure can be embodied in other specific forms without departing from the spirit or essential characteristics thereof. Accordingly, the above embodiments are to be considered in all respects as illustrative and not restrictive of the subject matter described herein.
[0239]
[0239] References
[0285] The references listed below are hereby incorporated by reference in their entirety into this specification.
[0240]
[0240]
Chemical Formula
[0241]
Chemical Formula
[0242]
Chemical Formula
Claims
1. A system comprising: A housing dimensioned to accommodate an animal, the housing including a door configured to permit access to the interior of the housing; and An acquisition system, A camera; At least two sets of light sources, each set of light sources being configured to emit light incident on the housing at different wavelengths from each other, At least two sets of light sources configured such that when the camera is irradiated by at least one of the at least two sets of light sources, video data of at least a part of the housing is acquired; and A controller in electrical communication with the camera and the at least two sets of light sources, Generating a control signal operative to control the acquisition of video data by the camera, Generating a control signal operative to control the emission of light by the at least two sets of light sources, simulating a light and dark cycle and eliciting different behaviors from the animal, Compressing the file size of the acquired video data, and Filtering background noise of the acquired video data, wherein the compressed and filtered video data is used as input data that can improve the performance of animal tracking performed by a convolutional neural network (CNN). A controller configured to perform the above An acquisition system including
2. The system according to claim 1, wherein at least a part of the housing is formed of a material opaque to visible light wavelengths.
3. The system according to claim 1, wherein at least a part of the housing is formed of a material non-reflective to infrared light wavelengths.
4. The system according to claim 1, wherein the camera is configured to acquire video data at a resolution of at least 480×480 pixels.
5. The system according to claim 1, wherein the camera is configured to acquire video data at a frame rate of at least 29 frames per second.
6. The system according to claim 1, wherein the camera is configured to acquire video data having a depth of at least 8 bits.
7. The system according to claim 1, wherein the camera is configured to acquire video data at an infrared wavelength.
8. The system according to claim 1, comprising one or more first illuminations configured to emit light at one or more visible light wavelengths, and one or more second illuminations configured to emit light at one or more infrared (IR) light wavelengths.
9. The system according to claim 8, wherein the controller is configured to request the first set of light sources to irradiate the housing with visible light having an intensity of 50 lux to 800 lux in the bright part of the light and dark cycle.
10. The system according to claim 8, wherein the controller is configured to request the second set of light sources to irradiate the housing with infrared light, and the second set of light sources is controlled such that the temperature of the housing rises by less than 5°C due to the infrared irradiation.
11. The system according to claim 8, wherein the controller is configured to request the first set of light sources to irradiate the housing according to logarithmically scaled 1024-level illumination.
12. The system according to claim 1, wherein the housing includes an object other than the animal, the video data acquired by the camera includes the object and the animal, and the compressed and filtered video data is processed using a CNN architecture capable of identifying the object and the animal.
13. Irradiating a housing configured to accommodate an animal with at least two sets of light sources, wherein each set of light sources is configured to emit light of a different wavelength. Acquiring video data of at least a part of the housing irradiated by at least one set of the at least two sets of light sources by a camera. Generating a control signal that operates to control the acquisition of video data by the camera by a controller electrically communicating with the camera. Generating a control signal that operates to control the emission of light by the at least two sets of light sources by a controller electrically communicating with the at least two sets of light sources, simulating a light and dark cycle, and eliciting different behaviors from the animal. Receiving the video data acquired by the camera by the controller. Compressing the file size of the acquired video data by the controller. Filtering the background noise of the acquired video data by the controller, wherein the compressed and filtered video data is used as input data capable of improving the performance of animal tracking performed by a convolutional neural network (CNN). A method comprising the steps.
14. The method according to claim 13, wherein the camera is configured to acquire video data at an infrared wavelength.
15. The method according to claim 13, wherein the camera is configured to acquire video data at a resolution of at least 480×480 pixels.
16. The method according to claim 13, comprising one or more first illuminations configured such that a first set of light sources emits light at one or more visible light wavelengths, and one or more second illuminations configured such that a second set of light sources emits light at one or more infrared (IR) light wavelengths.
17. Generating at least a first control signal that requests the first set of light sources to irradiate the housing with visible light having an intensity of 50 lux to 800 lux in the bright part of the light and dark cycle, including generating a control signal that operates to control the emission of light by the controller. The method according to claim 16.
18. Generating a control signal that operates to control the emission of light by the controller includes generating a first control signal that requests the second set of light sources to irradiate the housing with infrared light, and the temperature of the housing is less than 5 ° C by infrared irradiation. The method according to claim 16, wherein the second set of light sources is controlled so as to rise.
19. Generating a control signal that operates to control the emission of light by the controller includes generating a first control signal that requests the first set of light sources to irradiate the housing according to logarithmically scaled 1024-level illumination. The method according to claim 16.
20. The housing includes an object other than the animal, the video data acquired by the camera includes the object and the animal, and the method further includes using a CNN architecture capable of identifying the object and the animal. Processing the compressed and filtered video data. The method according to claim 13.
21. The system according to claim 1, wherein the controller uses a Dirac, H264, or MPEG4 codec to compress the acquired video data.
22. The system according to claim 1, wherein the controller uses an HQDN3D filter for the acquired video data.
23. The method according to claim 13, wherein the step of compressing the acquired video data by the controller includes using a Dirac, H264, or MPEG4 codec.
24. The method according to claim 13, wherein the step of filtering the acquired video data by the controller includes using an HQDN3D filter.
Citation Information
Patent Citations
Automation method for behavior observation of experimental animal
JP1999296651A
Systems and methods for monitoring behavioral informatics
JP2005502937A
Vehicle periphery monitoring device
JP2014106685A
Automated Monitoring of Animal Nutriment Ingestion
US20160050888A1
Systems and methods of video monitoring for vivarium cages
US20160150758A1