System and method for generating a three-dimensional motion simulation of an animal subject of interest

The system enhances markerless motion capture by maintaining animal subjects near the frame center and adjusting camera parameters to generate a three-dimensional simulation, addressing data limitations and enabling more natural motion capture with fewer cameras.

GB2643519APending Publication Date: 2026-02-25UNIVERSITY OF CAPE TOWN
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
GB2024012171
Authority / Receiving Office
GB · GB
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-19
Publication Date
2026-02-25

AI Technical Summary

Technical Problem

Markerless motion capture systems for animals are limited by the amount of data that can be captured, which restricts the length or duration of the three-dimensional simulation, and are often confined to linear movements along a lure track, lacking practicality and natural motion capture.

Method used

A system and method that detects and tracks an animal subject in successive image frames using a camera positioning system, maintaining the subject near the frame center, and generates a three-dimensional motion simulation by controlling camera parameters and mapping position measurements to create an image frame data set, compensating for camera movements and adjusting zoom and focus.

Benefits of technology

Increases the number of usable image frames and reduces the need for multiple cameras, allowing for more natural and prolonged motion capture with improved accuracy and reduced setup time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A system and method for generating a three-dimensional motion simulation of an animal subject of interest are provided. The method includes, over successive image frames captured by a camera mounted o
Need to check novelty before this filing date? Find Prior Art

Description

FIELD This disclosure relates to motion capture generally and markerless three-dimensional motion capture more specifically. In particular, the disclosure relates to a system and method for generating a three-dimensional motion simulation of an animal subject of interest. BACKGROUND Animal motion capture is an active field of research within robotics and computer vision and has a variety of applications, such as in biomechanics and biomimetics. Motion capture research is useful for understanding the kinematics of animals in motion, and can be used to design better performing robots. An output from a motion capture system may be a three-dimensional simulation of the motion of the animal, which may also be termed a three-dimensional trajectory. The three-dimensional simulation of the motion of the animal may capture the movement of various “key points” (such as the eyes, nose, neck base, spine, tail base, tail tip, tail mid, shoulder, front knee, front ankle, shoulder, etc.) in relation to one another. The relative movement of these key points captured by the three-dimensional simulation may represent the movement of the animal in three-dimensional space. Animal motion capture can be performed with or without markers. So-called “markerless” motion capture can be advantageous in some settings. For example, such systems can be better suited to capturing motion of an animal in the wild. Some implementations of such markerless motion capture systems may for example make use of a number of static red-green-blue (RGB) cameras positioned around a “lure track”. An animal can be encouraged to traverse the lure track while the cameras are running and the resultant video data can be processed to generate a three-dimensional simulation of the motion of the animal. Generally, the more cameras there are recording the animal subject, the better the resultant three-dimensional motion simulation of the animal subject. However, a limitation associated with such an implementation of markerless motion capture is the limited amount of data that can be captured. This may for example limit the length or duration of the resultant three-dimensional simulation. While more cameras could be added, this can add cost and may therefore be impractical. Further, motion capture may be limited to the movement of the animal along the “lure track,” whereas more natural and potentially less linear movement of the animal may also be of interest. There is accordingly scope for improvement. The preceding discussion of the background is intended only to facilitate an understanding of the present disclosure. It should be appreciated that the discussion is not an acknowledgment or admission that any of the material referred to was part of the common general knowledge in the art as at the priority date of the application. SUMMARY In accordance with an aspect of the disclosure there is provided a computer-implemented method comprising: over successive image frames captured by a camera mounted on a camera positioning system, detecting and tracking a location of an animal subject of interest in each image frame, including controlling parameters of the camera positioning system so as to maintain the animal subject of interest near a centre of each image frame; obtaining position measurements from the camera positioning system while capturing the successive image frames and mapping each image frame to a position measurement representing the position of the camera at the time the image frame is captured to generate an image frame data set including the successive image frames mapped to corresponding position measurements; and, outputting the image frame data set for generating a three-dimensional motion simulation of the animal subject of interest. Detecting the location of the animal subject of interest may include using an object detection algorithm which identifies the animal subject of interest. Detecting the animal subject of interest may include fitting a bounding box around the detected animal subject of interest. Tracking the location of the animal subject of interest may include using a tracking algorithm to estimate a position control parameter for controlling the camera positioning system to reposition the camera to a next position from which the animal subject of interest is estimated to be near a centre of a next image frame to be captured by the camera. The position control parameter may be an armature current value for controlling armature current of a motor of the camera positioning system. Tracking the location of the animal subject of interest may include tracking a location of a or part of a bounding box fitted around the animal subject of interest. The camera positioning system may be configured to rotate the camera about an axis. The position measurement may be an encoder measurement obtained from an encoder that encodes an orientation of the camera about an axis. The axis may be a vertical axis for tracking the location of the animal subject of interest along the horizontal. Tracking the location of the animal subject of interest in each image frame may include controlling optical parameters of the camera so as to maintain a size of the animal in the image frame. Tracking the location of the animal subject of interest may include using an algorithm to estimate an optical parameter for controlling the camera to adjust an optical setting so as to maintain the size of the animal in from one image frame to the next. Maintaining the size of the animal may include maximising the size of the animal in the successive image frames. The method may include obtaining optical measurements from the camera while capturing the successive image frames and mapping each image frame to an optical measurement representing the optical settings of the camera at the time the image frame is captured to generate an image frame data set including the successive image frames mapped to corresponding optical and position measurements for generating the three-dimensional motion simulation of the animal subject of interest. The optical measurements may include one or both of zoom and focus measurements. The image frame data set may be an optical-position-mapped data set. The method may include: obtaining the image frame data set; and, processing the image frame data set to generate the three-dimensional motion simulation of the animal subject of interest. Processing the image frame data set may include fitting a skeletal model of the animal subject of interest to a set of two-dimensional pose measurements obtained from the image frame data set. Processing the image frame data set may include obtaining, from each image frame, a set of two-dimensional pose measurements. Obtaining the set of two-dimensional pose measurements from each image frame may include labelling, in each image frame, the animal subject of interest with key points corresponding to key points of the skeletal model of the animal subject of interest. The set of two-dimensional pose measurements may include values for each key point in a coordinate system based on the pose of the animal subject of interest in the image frame. The coordinate system may be a two-dimensional position coordinate system (e.g. being a coordinate system of the image frame). Processing the image frame data set may include compensating for changes in camera position from one image frame to the next. Compensating for changes in camera position may include updating extrinsic camera parameters for each image frame. Updating extrinsic camera parameters for each image frame may include: generating a transformation from the camera into the camera positioning system; generating a rotation of the transformation about an axis, wherein a magnitude of the rotation is determined based on the position measurement mapped to the image frame; and, generating an inverse transformation from the camera positioning system to the camera. Compensating for changes in camera position may include applying the inverse transformation to a skeletal model of the animal subject of interest to generate a position-compensated skeletal model of the animal subject of interest. Processing the image frame data set may include fitting the position-compensated skeletal model of the animal subject of interest to the two-dimensional pose measurements. In accordance with a further aspect of the disclosure there is provided a system comprising a camera positioning system including a control module and which mounts a camera configured to capture image frames, wherein the control module is configured: over successive image frames captured by the camera, detect and track a location of an animal subject of interest in each image frame, including controlling parameters of the camera positioning system so as to maintain the animal subject of interest near a centre of each image frame; to obtain position measurements from the camera positioning system while capturing the successive image frames and to map each image frame to a position measurement representing the position of the camera at the time the image frame is captured to generate an image frame data set including the successive image frames mapped to corresponding position measurements; and, to output the image frame data set for generating a three-dimensional motion simulation of the animal subject of interest. The control module may be configured to track the location of the animal subject of interest using a tracking algorithm which estimates a position control parameter for controlling the camera positioning system to reposition the camera to a next position from which the animal subject of interest is estimated to be near a centre of a next image frame to be captured by the camera. The camera positioning system may include a motor arranged to position the camera. The position control parameter may be an armature current value for controlling armature current of the motor. The camera positioning system may be configured to rotate the camera about an axis. The position measurement may be an encoder measurement obtained from an encoder that encodes an orientation of the camera about the axis. The axis may be a vertical axis for tracking the location of the animal subject of interest along the horizontal. The system may include a computing device configured: to obtain the image frame data set; and, to process the image frame data set to generate the three-dimensional motion simulation of the animal subject of interest. The system may include a plurality of camera positioning systems, each of which mounts a camera and outputs an image frame data set. The computing device may be configured to obtain a plurality of image frame data sets and to process the plurality of image frame data sets to generate the three-dimensional motion simulation of the animal subject of interest. In accordance with a further aspect of the disclosure there is provided a system comprising: a non-transitory computer-readable storage medium; and one or more processors coupled to the non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium comprises program instructions that, when executed on the one or more processors, cause the system to perform operations comprising: over successive image frames captured by a camera mounted on a camera positioning system, detecting and tracking a location of an animal subject of interest in each image frame, including controlling parameters of the camera positioning system so as to maintain the animal subject of interest near a centre of each image frame; obtaining position measurements from the camera positioning system while capturing the successive image frames and mapping each image frame to a position measurement representing the position of the camera at the time the image frame is captured to generate an image frame data set including the successive image frames mapped to corresponding position measurements; and, outputting the image frame data set for generating a three-dimensional motion simulation of the animal subject of interest. In accordance with a further aspect of the disclosure there is provided a system including a memory for storing computer-readable program code and a processor for executing the computer-readable program code, the system comprising: a detecting component for, over successive image frames captured by a camera mounted on a camera positioning system, detecting an animal subject of interest; a tracking and control component for, over the successive image frames, tracking a location of the animal subject of interest in each image frame and controlling parameters of the camera positioning system so as to maintain the animal subject of interest near a centre of each image frame; a mapping component for obtaining position measurements from the camera positioning system while capturing the successive image frames and mapping each image frame to a position measurement representing the position of the camera at the time the image frame is captured to generate an image frame data set including the successive image frames mapped to corresponding position measurements; and, an output component for outputting the image frame data set for generating a three-dimensional motion simulation of the animal subject of interest. In accordance with a further aspect of the disclosure there is provided a computer program product comprising a computer-readable medium having stored computer-readable program code for performing the steps of: over successive image frames captured by a camera mounted on a camera positioning system, detecting and tracking a location of an animal subject of interest in each image frame, including controlling parameters of the camera positioning system so as to maintain the animal subject of interest near a centre of each image frame; obtaining position measurements from the camera positioning system while capturing the successive image frames and mapping each image frame to a position measurement representing the position of the camera at the time the image frame is captured to generate an image frame data set including the successive image frames mapped to corresponding position measurements; and, outputting the image frame data set for generating a three-dimensional motion simulation of the animal subject of interest. Further features provide for the computer-readable medium to be a non-transitory computer-readable medium and for the computer-readable program code to be executable by a processing circuit. Embodiments of the technology will now be described, by way of example only, with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS In the drawings: Figure 1 is a schematic diagram which illustrates aspects of a system and method for generating a three-dimensional motion simulation of an animal subject of interest according to aspects of the present disclosure; Figure 2A is a schematic diagram which illustrates an exemplary system for generating a three-dimensional motion simulation of an animal subject of interest according to aspects of the present disclosure; Figure 2B is a three-dimensional view of an example camera positioning system pedestal according to aspects of the present disclosure; Figure 3A is a flow diagram which illustrates example initialisation and calibration operations that may be performed in a method for generating a three-dimensional motion simulation of an animal subject of interest; Figure 3B is a flow diagram which illustrates example operations for generating an image frame data set in a method for generating a three-dimensional motion simulation of an animal subject of interest; Figure 3C is a flow diagram which illustrates an example method for generating a three-dimensional motion simulation of an animal subject of interest according to aspects of the present disclosure; Figure 4 is a schematic diagram which illustrates coordinate axes and states for the motor control and extended Kalman filter according to aspects of the present disclosure; Figure 5A is a schematic diagram which illustrates an example three-dimensional motion simulation and the corresponding image frame from which it is generated according to aspects of the present disclosure; Figure 5B is a schematic diagram which illustrates another example three-dimensional motion simulation and the corresponding image frame from which it is generated according to aspects of the present disclosure; Figure 6A is a histogram showing distribution of subject location in an image frame (10 bins) based on an experimental implementation of the described system and method for a first type of animal subject of interest; Figure 6B is a histogram showing distribution of subject location in an image frame (10 bins) based on an experimental implementation of the described system and method for a second type of animal subject of interest; Figure 7 is a schematic diagram which illustrates re-projected points and a 3D cheetah skeleton with respect to the inertial frame for selected frames based on an experimental implementation of the described system and method; and, Figure 8 illustrates an example of a computing device in which various aspects of the disclosure may be implemented. DETAILED DESCRIPTION WITH REFERENCE TO THE DRAWINGS A system and method for generating a three-dimensional motion simulation of an animal subject of interest are provided. In the described system and method, a location of an animal subject of interest in each image frame of a series of successive image frames is detected and tracked so as to maintain the animal subject near a centre of each of the image frames. This may include controlling parameters of a camera positioning system mounting a camera which captures the image frames so as to maintain the animal subject near the centre of each image frame. Position measurements may be obtained from the camera positioning system while capturing the successive image frames. The position measurements may be mapped to corresponding image frames to generate an image frame data set including the successive image frames mapped to corresponding position measurements. The image frame data set may be processed to generate a three-dimensional motion simulation of the animal subject of interest. In some cases, a plurality of camera positioning systems and cameras are provided to obtain a plurality of image frame data sets. The parameters of the camera positioning system that are controlled may include zoom and / or position parameters. In this manner, the camera may be repositioned and / or a zoom level may be adjusted to maximize the number of image frames containing the animal subject and / or the relative size of the animal subject in each of the image frames. Thus, the number of usable image frames may be increased and / or the number of cameras required may be decreased. Reducing the number of cameras may further reduce the setup time required. Referring to Figure 1, an example pipeline for detecting and tracking an animal subject of interest is illustrated. A camera captures an image frame (2) which is input into an object detection algorithm (4). The object detection algorithm may output the image frame with a graphical indication, such as a bounding box (6), superimposed thereon and which indicates the location of the animal subject in the image frame. The location of the animal subject in the image frame is determined (e.g., based on the graphical indication) and fed into a tracking algorithm (8), together with pseudo (9A) and / or a position (9B) measurements, the output of which feeds into a full state feedback (FSF) controller (10). The tracking algorithm and FSF controller estimate a position control parameter for controlling the camera positioning system to reposition the camera to a next position from which the animal subject of interest is estimated to be near a centre of a next image frame to be captured by the camera. The position control parameter is input into a motor or servo controller (12) which controls operation of a motor (14) so as to reposition the camera to the next position from which the animal subject of interest is expected to be near the centre of the next image frame. The system and method described herein may use noisy measurements relating to the camera’s angular position to compensate for the camera's movement and variable zoom. For example, in tracking the animal subject, noisy measurements of the subject's position in the frame (acquired using an object detection model or algorithm) may be filtered using a tracking algorithm, such as an extended Kalman filter (EKF) with a constant acceleration model. This provides state estimates of the motor's required angular position at each timestep. The EKF may estimate the true angular position of the camera by filtering out noise from the measurements. This may be required in some implementations, for example where the object detection algorithm runs at a lower frequency (e.g., 8-12 Hz) than the motor controller (e.g., 50 Hz). The EKF may predict the animal subject's position in the image frame at a higher rate, which may allow the motor controller to adjust the camera's position smoothly. In some examples, a feedback (such as a full state feedback) control loop drives a motor controller to position the motor, and by extension, the camera. An encoder or other position sensor may measure the angular position of the motor and / or camera. With calibrated cameras whose relative pose is known before they begin moving, new camera pose can be determined by applying a transformation to the initial pose based on the measured angular position. In some examples, motorized lenses with position feedback control the camera zoom and focus levels. The positions of both motors (zoom and focus) may be recorded at each timestep. These positions may be used to update the "camera intrinsic parameters" at each timestep, which describe the camera's physical properties and convert 2D image coordinates to 3D coordinates in the camera's frame. This allows for variable zoom and focus compensation during 3D reconstruction using the data gathered by the camera positioning system. The system and method described herein may create a dynamic mapping from 3D points to pixel measurements that can adapt to changes caused by camera motion and focal length adjustments. This dynamic mapping from 3D points to pixel measurements may for example be implemented using real-time adjustments based on the camera’s intrinsic and extrinsic parameters. As the camera moves or its focal length changes, the camera’s intrinsic parameters, orientation, and position may be updated. These parameters may be used to project 3D points into 2D image coordinates for each camera. The mapping may adapt dynamically by recalculating the projection matrix for each frame, ensuring that changes in camera motion and focal length are accurately reflected in the pixel measurements. Such mappings may be integrated into a 3D reconstruction algorithm. For example, the dynamic mappings may be integrated into the 3D reconstruction algorithm by continuously updating the projection matrices as the camera moves and adjusts its zoom. A key point labelling algorithm may be configured to extract 2D key points from image frames. These key points may be used with the updated projection matrices to triangulate the 3D positions of the key points. Using the so-called full trajectory estimation (FTE) algorithm, these triangulated 3D key points may be compared to the 3D positions of a kinematic model of the subject, allowing an optimal trajectory to be determined which fuses the measured 3D key point positions from the cameras with the predicted key point positions from the kinematic model. The algorithm combines data from multiple cameras, accounting for their respective positions and orientations, to perform accurate 3D reconstruction. The described system and method may be arranged to process data from multiple dynamic camera positioning systems (being “image frame data sets”) to generate three-dimensional (3D) motion simulation of the subject of interest. In this manner, the subject’s motion may be reconstructed in 3D. Processing the data sets may include: data collection, synchronization, key point extraction, dynamic mapping update, 3D point triangulation, and trajectory optimisation. Data collection may include collecting video (i.e., a series of image frames) and encoder data from multiple cameras tracking the subject. Synchronization may include time-synchronizing videos from multiple cameras to acquire matching frames. Key point extraction may be performed using a key point labelling algorithm (such as DeepLabCut) to extract 2D key points from the image frames of each camera. Dynamic mapping updates may include using the measured angular position of the motors controlling the camera positions and the positions of the zoom and focus lenses to update projection matrices of each camera per frame to account for camera movement and focal length changes. 3D point triangulation may include using the updated projection matrices and extracted key points to triangulate the 3D positions of the key points. Trajectory optimisation may include using 3D positions in a kinematic model-based trajectory optimization method to estimate the full trajectory of the subject. This method may be arranged to minimize the difference between the reprojected model points and the observed key points, ensuring robust and accurate 3D motion capture. Figure 2A is a schematic diagram which illustrates an exemplary system (100) for generating a three-dimensional motion simulation of an animal subject of interest. The system may include one or more camera sub-systems (102) and a computing device (104). Each camera sub-system may include a camera positioning system (106) arranged to position one or more cameras (108) associated therewith. Although the description which follows refers to camera in the singular, it should be appreciated that, in some examples, each of the one or more camera positioning systems may mount and position multiple cameras. The camera may be a red-green-blue (RGB), thermal and / or infrared camera. In examples where the camera positioning system includes a plurality of cameras, there may be a combination of RGB, thermal and / or infrared cameras. The camera captures and outputs image frames (109) at an operating frame rate. In some examples, the camera includes a controllable zoom lens (e.g. a motorized zoom lens) and outputs optical measurements for or in association with each of the image frames. The optical measurements may include zoom and / or focal length measurements. In some examples, the optical measurements are output in metadata attached to each image frame. In other embodiments, the optical measurements are output separately. In this manner, intrinsic values of the camera (relating to zoom and / or focus levels) can be calculated later on a per frame basis. The camera positioning system may mount the one or more cameras. The camera positioning system may include a control module (110) which interfaces with a controller (112) configured to control operation of a motor (114). In some examples, the control module interfaces with the camera and / or controllable zoom lens to control optical parameters, such as parameters which control zoom and / or focal length settings of the camera. The control module may be in the form of a computing device, one or more microcontroller or the like. The control module may include a processor for executing the functions of components described below, which may be provided by hardware or by software units executing on the control module. The software units may be stored in memory and instructions may be provided to the processor to carry out the functionality of the described components. The control module may include a detection component (110A) arranged to detect the object of interest. The detection component may use an object detection algorithm to detect and graphically indicate the animal subject in each of the successive image frames. The control module may include a tracking and control component (110B) arranged to track the object of interest through successive image frames and control parameters of the camera and / or camera positioning system so as to maintain the animal subject of interest near a centre of each image frame. The tracking and control component may implement a tracking algorithm. The tracking and control component may implement full state feedback (FSF) control using outputs from the tracking algorithm. The control module may include a mapping component (110C) which may be configured to map each image frame to a position measurement representing the position of the camera at the time the image frame is captured to generate an image frame data set including the successive image frames mapped to corresponding position measurements. The control module may include an output component (110D) arranged to output the image frame data set for generating a three-dimensional motion simulation of the animal subject of interest. The controller may be a servo or motor controller. The controller may be configured to control the motor by controlling an armature current or other parameters. The motor may be a brushless direct current (DC) motor. The motor may be configured to move (e.g. rotate) a platform or support on which the camera is mounted, so as to move or change position of the camera. In some examples, the motor, platform / support and camera are arranged such that the motor rotates the camera about an axis. The axis about which the camera is rotated may be the vertical (z-) axis such that the camera can be controlled to pan along the horizontal. Thus, in some examples, the camera positioning system may be configured to swivel the camera horizontally from a fixed position. In other examples, the camera positioning system may be configured to move or reposition the camera in other ways, for example by rotating the camera about one or more other axes, moving the camera along a track, or the like. In other examples, the camera positioning system may be configured to position the camera in all six degrees of freedom. The camera positioning system may include a position sensor (116) configured to measure the position of the camera and output corresponding position measurements (118). In some examples, the position sensor is an encoder that measures an angular position of the shaft of the motor. In some examples, therefore, the position sensor may measure angular position of the camera about an axis, such as the vertical (z-) axis. Other examples may make use of other forms of position sensor, such as an inertial measurement unit (IMU) or the like, instead of or as well as an encoder. In some examples, the camera positioning system includes one or more intrinsic parameter sensors for measuring intrinsic parameters of the camera, such as one or more of zoom, focal length or the like. In some examples, the intrinsic parameter sensors include an encoder which encodes a position of a motor controlling camera zoom. In some examples, the intrinsic parameter sensors form a part of the camera and / or the controllable zoom lens of the camara. The control module may be configured to receive the image frames from the camera, the position measurements from the position sensor and the optical measurements. The control module may be configured, over successive image frames, to detect and track a location of an animal subject of interest in each image frame. This may include controlling parameters of the camera and / or camera positioning system so as to maintain the animal subject of interest near a centre of each image frame. The control module may for example be configured to control optical and / or positioning parameters. The control module may further be configured to obtain the optical and / or position measurements while capturing the successive image frames and to map each image frame to optical and / or position measurements representing optical settings of the camera and / or the position of the camera at the time the image frame is captured to generate an image frame data set (120) including the successive image frames mapped to corresponding optical and / or position measurements. The control module may be configured to output the image frame data set for generating a three-dimensional motion simulation of the animal subject of interest. The image frame data set may for example be output to the computing device and / or stored in a data store (122). In some examples, the camera positioning system is fixed relative to a world frame and provided by a camera positioning housing or pedestal, such as the example pedestal illustrated in Figure 2B. The example pedestal (150) includes a rotatable platform (152) which mounts the camera (108) and optionally other sensors. The pedestal includes a frame (154) which supports the rotatable platform and houses the control module, controller and a power supply (156). In the illustrated example, the rotatable platform is rotatable around a vertical (z-) axis. Rotation of the rotatable platform is driven by a motor and the position sensor measures the angular position of the motor shaft and thus of the camera. In other examples, the camera positioning system is mobile (e.g., drone-mounted) and the motor or other actuator may be controlled to move the camera by moving the camera positioning system itself. Returning to Figure 2A, the computing device (104) may be configured to receive or retrieve one or more image frame data sets from the camera positioning system and / or the data store. The computing device may be configured to process the one or more image frame data sets to generate and output a three-dimensional motion simulation (160) of the animal subject of interest. The three-dimensional motion simulation may be output as a three-dimensional motion simulation file for storage (e.g., in the data store or elsewhere) and / or may be displayed on a display connected to the computing device. The computing device may include a processor for executing the functions of components described below, which may be provided by hardware or by software units executing on the computing device. The software units may be stored in memory and instructions may be provided to the processor to carry out the functionality of the described components. The computing device may include a data set obtaining component (162) arranged to obtain one or more image frame data sets. The data set obtaining component may be configured to obtain the data sets from the data store or from the one or more camera positioning systems directly. The computing device may include a data set processing component (164) arranged to process the one or more image frame data sets to generate the three-dimensional motion simulation of the animal subject of interest. The computing device may include an output component (166) arranged to output the three-dimensional motion simulation of the animal subject to the data store and / or a display. In the system described above, rotating the camera affects the “extrinsic” measurements of the camera, which describe the position of the camera relative to the world coordinate frame. Changing the zoom / focus of a camera changes the “intrinsic” measurements, which describe how the camera maps a 3D point in the real world to a 2D point in the image frame which captures the point. The system (100) described above may implement a method for generating a three-dimensional motion simulation of an animal subject of interest. An exemplary method for generating a three-dimensional motion simulation of an animal subject of interest is illustrated in the flow diagrams of Figure 3A to 3C. Figure 3A illustrates example initialisation and calibration operations that may be performed prior to motion capture. The initialisation and calibration operations may be performed for each camera positioning system and its associated camera. The method may include an initialisation stage of receiving (200) input of a specific animal subject of interest and configuring or instructing an object detection algorithm to identify and mark or tag a location, in an image frame, of the specific animal subject of interest. In other words, if the specific animal subject of interest is a cheetah, the object detection algorithm may be configured or instructed to identify and tag only cheetahs in the image frames. The method may include calibrating (202) the camera at different zoom and / or focus levels. This may include capturing images of a calibration object, such as a checkerboard, at various zoom and focus settings. Intrinsic parameters of the camera (such as focal length, principal point, and distortion coefficients) are calculated for each setting. This multi-level calibration allows the system to adapt to different zoom and focus settings, maintaining accuracy in 3D reconstruction. This multi-level calibration is then used to determine (204) a first measurement function for extrapolating zoom and focus levels as a continuous function, rather than relying solely on discrete values obtained from the manual calibration process. The first measurement function may for example map a motor position for zoom / focus to a camera intrinsic parameter. The function may be obtained by calibrating the camera using distinct settings, then fitting a function to the data to allow for extrapolation at different zoom / focus settings. Calibrating the camera at different zoom and / or focus levels may be referred to as intrinsic calibration to determine the intrinsic camera parameters for each setting. The intrinsic calibration may be used for development of a measurement function (which may be non-linear) and which allows for the determination of the intrinsic parameters of the camera at all zoom / focus levels. Since the zoom and focus levels can be continuous, and the intrinsic calibration is done at discrete steps, this step may be needed so that the system is not only limited to a few different zoom and focus settings. Referring now to Figure 3B, the method may include operations which may be performed for generating an image frame data set. The operations described with reference to Figure 3B may be performed after the initialisation and / or calibration operations described above with reference to Figure 3A. The method may include receiving (212) image frames captured by a camera mounted on a camera positioning system. Image frames may be captured and output by the camera at an operating frame rate, such as 120 frames per second, or the like. The method may include detecting (214) and tracking (216) a location of an animal subject of interest in each image frame over successive image frames captured by the camera. Detecting the animal subject of interest may include using an object detection algorithm which identifies the animal subject of interest in each image frame in real-time. Detecting the animal subject of interest may include tagging or marking the location of the animal subject of interest in the image frame, for example by fitting a bounding box around the identified animal subject of interest. An example real-time object detection algorithm includes the YOLO range of object detection algorithms (such as YOLOv4-Tiny or the like). In some examples, detecting the animal subject of interest includes inputting the image frame into a real-time object detection algorithm which processes the image frame and receiving, as an output from the object detection algorithm, a tagged image frame in which the location of the animal subject of interest in the image frame is delimited with a bounding box or other graphical indication. Tracking the location of the animal subject of interest may include controlling (220) parameters of the camera positioning system so as to maintain the animal subject of interest near a centre of each image frame. The parameters may be optical and / or positioning parameters. Tracking the location of the animal subject of interest may include using a tracking algorithm to estimate a position control parameter for controlling the camera positioning system to reposition the camera to a next position from which the animal subject of interest is estimated to be near a centre of a next frame. An example tracking algorithm is the EKF tracking algorithm. In some examples, tracking the location of the animal subject of interest may include controlling optical parameters of the camera so as to maintain a size of the animal in the image frame. The optical parameters may include zoom and / or focal length of the camera. For example, tracking the location of the animal subject of interest may include using an algorithm to estimate one or more optical parameters (such as parameters for controlling zoom and / or focal length) for controlling the camera to adjust an optical setting (such as zoom and / or focal length) so as to maintain the size of the animal in from one image frame to the next. Maintaining the size of the animal may include maximising (within a margin) the size of the animal in the successive image frames. In some examples, controlling the optical parameters may include setting zoom and focus levels of the camera based on the size of a bounding box in each frame. If the bounding box is small, the zoom level is increased so that the subject appears larger in the frame and vice versa. In some examples, tracking the location of the animal subject of interest includes tracking a location of a or part of a bounding box fitted around the animal subject of interest. For example, a centre of the bounding box along the horizontal (x-) axis may be tracked and a position control parameter may be determined to maintain the centre of the bounding box near the centre of the horizontal axis. In some examples, the camera positioning system is configured to rotate the camera about a vertical (z-) axis for tracking the location of the animal subject of interest along the horizontal. The position control parameter may be an armature current value for controlling armature current of a motor of the camera positioning system. In some examples, a state vector is defined as: x = [0 <p 0 <p 0 <p], where 0 and <p are the motor shaft and subject’s angular positions in the world frame respectively. An example world frame and relative camera and subject positions are illustrated in Figure 4. The world frame includes xw (302), yw (304) and zw (not shown) coordinates. The camera frame similarly includes xc (306), yc (308) and zc (not shown) coordinates, in which the zc axis is coaxial with zw. The camera frame is rotatable relative to the world frame about the zc axis, which rotation is represented by 0. The angular position of the animal subject (310) relative to the z-axes (zc and zw) is denoted by (p. The subject’s angular position <p is used as the target position and is sent to the control module of the camera positioning system. The motor shaft’s angular position 0 may be measured directly using a position sensor, such as an encoder. The subject’s angular position <p may for example be calculated using the centre of the bounding box and a non-linear measurement function that converts from pixels to radians through imaging geometry, for example as follows: h(x) = ftan(<p - 0), where f is the camera’s focal length in pixels and % is the centre of the bounding box along the horizontal. This non-linear measurement function may for example map image locations of subject to an angular measurement and may for example be obtained based on the systems geometry. In some examples, the motor may be modelled in state space (with an augmented integrator state) as follows: where 9 and 9 are the motor shaft position and velocity, e is an added state for the integral of the error, N is the gear ratio, KT is the motor torque constant, / is the combined motor and load inertia, and i is the motor armature current. Thus, the motor’s shaft position may be controlled as a function of the armature current based on the next expected location of the animal subject of interest in the next image frame. The method may include obtaining (222) position measurements from the camera positioning system while capturing the successive image frames and mapping (224) each image frame to the corresponding position measurement to generate an image frame data set. The “corresponding position measurement” may be the position measurement that represents the position of the camera at the time the image frame is captured. In some examples, the position measurement is an encoder measurement obtained from an encoder that encodes an orientation of the camera about an axis (such as the vertical axis). In some examples, each image frame includes optical measurements from the camera (e.g. in image frame metadata) and the image frame data set implicitly includes a mapping of each image frame to its corresponding optical measurement(s) representing the optical settings of the camera at the time the image frame is captured to each image frame. Optical settings may be zoom and / or focus (such as focal length) settings and optical measurements may be the corresponding values representing these settings. In other examples, the method includes obtaining the optical measurements from the camera while capturing the successive image frames and mapping each image frame to an optical measurement representing the optical settings of the camera at the time the image frame is captured to generate an image frame data set including the successive image frames mapped to corresponding optical and position measurements for generating the three-dimensional motion simulation of the animal subject of interest. In either case, the image frame data set is an optical-position-mapped data set in that it includes the successive image frames mapped to corresponding optical and position measurements. The method may include outputting (226) the image frame data set including the successive image frames mapped to corresponding optical and / or position measurements. The image frame data set may be output for generating a three-dimensional motion simulation of the animal subject of interest. Outputting the image frame data set may include storing the image frame data set in a data set storage. Operations (212) to (226) may be conducted by a camera positioning control module, which may form part of the camera positioning system. In some examples, there may be a plurality of cameras, each having their own camera positioning system and each detecting and tracking the animal subject of interest from another perspective. In this manner, a plurality of image frame data sets may be generated, each of which is time-synchronized and each capturing the animal subject of interest from a different angle or perspective. Referring now to Figure 3C, the method may include obtaining or retrieving (230) the image frame data set. In examples which include a plurality of cameras, this may include obtaining or retrieving the plurality of image frame data sets. The image frame data set(s) may be retrieved from a data set storage. The method may include processing (232) the image frame data set to generate the three-dimensional motion simulation of the animal subject of interest. Processing the image frame data set(s) to generate the three-dimensional motion simulation of the animal subject of interest may include obtaining (234), from each image frame, a set of two-dimensional pose measurements. Obtaining the set of two-dimensional pose measurements from each image frame may include labelling, in each image frame, key points on the animal subject of interest. Labelling may be performed by a key point labelling algorithm, such as DeepLabCut or the like. The key points labelled on the animal subject of interest may correspond to key points of a skeletal model of the animal subject of interest used in a subsequent operation of the method. Obtaining the set of two-dimensional pose measurements from each image frame may include determining, for each key point, two-dimensional pose measurements. The two-dimensional pose measurements for each key point may for example be in the form of values indicating the position and / or orientation of the key point in a coordinate system based on the particular pose of the animal subject of interest in the image frame. The coordinate system may be a two-dimensional position coordinate system (e.g. being a coordinate system of the image frame). The number of key points and degrees of freedom may vary from subject to subject, but the generalised coordinates follow a common structure. For example, the set of two-dimensional pose measurements for each image frame may include: • x, y, z: The position of the first or base key point in the world frame; • 01...0L: The roll angles of each subsequent key point, where L is the total number of joints; • 01 ...6L: The pitch angles of each subsequent key point; • ip'L.ApL: The yaw angles of each subsequent key point. In some examples, the set of two-dimensional pose measurements for each image frame is updated to include corresponding position measurements for that image frame, being: ar... an (which may for example be the position measurement (such as encoder angles) for each frame, where n is the number of image frames in the set of image frame data sets). The set of two-dimensional pose measurements may thus represent the pose parameters of a kinematic model and may be optimised over a given trajectory to obtain a set of optimised states for each frame. Processing the image frame data set to generate the three-dimensional motion simulation of the animal subject of interest may include compensating (236) for changes in camera position from one image frame to the next. This may include updating extrinsic camera parameters for each image frame. For example, updating extrinsic camera parameters for each image frame may include: generating a transformation from the camera into the camera positioning system; generating a rotation of the transformation about an axis, wherein a magnitude of the rotation is determined based on the position measurement mapped to the image frame; and, generating an inverse transformation from the camera positioning system to the camera. Compensating for changes in camera position may include applying the inverse transformation to the skeletal model of the animal subject of interest to generate a position-compensated skeletal model of the animal subject of interest. In examples where a plurality of image frame data sets is retrieved, the operations of obtaining (234) two-dimensional pose measurements and compensating (236) for changes in camera optics and / or position may be performed serially for each image frame data set. Processing the image frame data set(s) to generate the three-dimensional motion simulation of the animal subject of interest may include fitting (238) the skeletal model of the animal subject of interest to the set of two-dimensional pose measurements obtained from the image frame data set(s). This may include, in some examples, fitting the skeletal model to each set of the two-dimensional pose measurements obtained from each of the respective image frame data sets. The skeletal model of the animal subject of interest may be the position-compensated skeletal model of the animal subject of interest. In this manner, compensation may be performed on the three-dimensional skeletal model of the animal subject of interest. The skeletal model may be defined relative to the world coordinate frame. Then, for each camera, the position of the skeletal model may be expressed in the camera coordinate frame. This may then be reprojected into the image frame and compared to the two-dimensional pose measurements (e.g., being the labelled (measured) points, which may be labelled by hand or a machine learning labelling tool, as described above). Thus, the skeletal model may be “reprojected” into the image frame for each camera. This may include using the extrinsic and intrinsic parameters for each camera for each image frame to estimate where a particular 3D point would be mapped to in the image frame. These reprojected points may be compared to the actual labelled points, and the distance between them may be minimized (the measurement error). The skeletal model may be fitted to the two-dimensional pose measurements over a given trajectory. Fitting the skeletal model may minimize both the model error and measurement error. The model error may be based on physical limitations and assumptions made about the motion of the subject. The limitations may be used to limit the optimisation computation to ensure that the results produced are physically possible. The assumptions made about the motion of the subject are based on analysis of the motion of the subject. These assumptions are used to inform the predicted motion of the subject, such as the position, velocity and acceleration of each point on the skeletal model. The predicted motion is compared to the motion determined (estimated) by the trajectory optimisation, and the difference between the predicted and estimated motion is the model error. Processing the image frame data set(s) to generate the three-dimensional motion simulation of the animal subject of interest may include using a trajectory optimisation method, such as so-called “full trajectory estimation” described by D. Joska, L. Clark, N. Muramatsu, R. Jericevich, F. Nicolls, A. Mathis, M. W. Mathis, and A. Patel in “Acinoset: a 3d pose estimation dataset and baseline models for cheetahs in the wild,” in 2021 IEEE international conference on robotics and automation (ICRA). IEEE, 2021, pp. 13 901-13 908. The method may include outputting (240) the three-dimensional motion simulation of the animal subject of interest. Outputting the three-dimensional motion simulation of the animal subject of interest may include outputting the three-dimensional motion simulation of the animal subject of interest to a display for viewing and / or manipulation. Outputting the three-dimensional motion simulation of the animal subject of interest may include storing the three-dimensional motion simulation. Example three-dimensional motion simulations (250, 252) and the corresponding image frames (254, 256) are illustrated in Figures 5A and 5B. The histograms illustrated in Figures 6A and 6B illustrate the number of image frames having the subject in view, comparing an example static camera implementation against a dynamic camera implementation in accordance with aspects of the present disclosure and are discussed in greater detail below. A system and method for three-dimensional (3D) markerless motion capture of animals in the wild using autonomous tracking cameras is provided. The system and method described herein provide a vision-based tracking system which can autonomously track a moving animal in the wild. An example implementation of the described system was tested with a stereo camera pair mounted to the rotating sensor platform with to demonstrate 3D markerless motion capture. An object detection neural network is used to locate the subject in the video stream from a camera, being a webcam mounted to a motor, and is combined with an EKF and an FSF controller to control the position of the motor and keep the subject centred in the frame. FTE is used for reconstruction of 3D trajectories. Experimental implementation The described system and method may implement a tracking system which combines an object detection neural network with an EKF and FSF control module to accurately track a moving subject. In an example implementation of the described system and method, a Maxon 18V 70W brushless DC motor with 500 CPT quadrature encoder and 51 / 1 gearbox was used along with a Maxon 50 / 5 servo controller to actuate the sensor pedestal. A Teensy 4.0 micro-controller was used to interface with the servo controller via PWM. A Jetson Nano 4GB Developer Kit and Logitech C920 HD USB Webcam was used for object detection. A pedestal was mounted to the motor shaft for attaching various sensors. Object detection was done using YOLOv4-Tiny, for its accuracy and speed on the Jetson Nano. The model was trained on a custom data set and then optimised for deployment on the Jetson Nano. The data set contained 500 photos from the Cheetah class from OIDv6 and 10000 hand-labelled images of cheetahs. When multiple objects are initially detected, the one with the highest confidence score is chosen. Outliers are excluded using a geometric method that selects the object whose bounding box centre is nearest to the tracked object’s bounding box centre from the previous frame. Object detection operates at a 10Hz sampling rate. To enable motor control with sampling rate of 50Hz, a constant acceleration model EKF was used and a state vector was defined as introduced above. The measurement co-variance of YOLOv4-Tiny was estimated to be ~52 pixels2 by comparing detected bounding box centres to hand-labelled bounding box centres. The measurement covariance of the high resolution encoder was set to a low value of 0.012 rad2. To determine process noise, we scaled the maximum expected angular jerk (calculated based on a cheetah running at max speed past a stationary observer) between samples for each set of states, with the least noise for acceleration states and the most for position states. The discrepancy in sampling rates between the EKF at 50 Hz and object detection at 10 Hz necessitates a different approach for EKF updates. For four out of every five samples, the EKF is updated using encoder readings alone. On the fifth sample, the EKF both predicts and updates states using the new measurement, but only if an object has been detected. To compensate for unsuccessful object detection we used a noisy-pseudo measurement {(p = 0°, a = rr4 0 whenever the EKF went more than 10 samples without receiving a new measurement. The motor state was modelled in the manner defined above. In the experimental implementation, MATLAB Control System Toolbox was used with gains chosen to be [6.0918 0.9208 14.6307]. Sensor interfacing and motor control were done using a Teensy 4.0 microcontroller. FTE was used for 3D reconstruction. This method fits a skeletal model of the subject species to a set of 2D pose measurements over a given trajectory to minimise both the model error and measurement error. In terms of object detection, for the Cheetah network, a mean average precision (mAP) of 34.46% was achieved with an Intersection of Union (loU) threshold of 0.5. The average inferencing time was 0.043 seconds or 23.3 FPS on the Jetson Nano. This was limited to 10 FPS to conserve resources for additional sensors. With a confidence threshold of 0.5 the subject was missed by YOLOv4-Tiny in ~2% of the frames where the subject was within a range of 15m. Frames with no detections were saved to augment the training data set in future. The tracking system was tested on animal subjects of interest including three free-running cheetahs and a human. The cheetahs ran perpendicular to the system at approximate distances of 8m and 12m. To validate the tracker’s performance in 3D reconstruction, two GoPro HERO5 Session cameras (S Cams) were attached to the static base of the tracker and two more were attached to the rotating sensor pedestal (R Cams). The image data collected by the GoPros during these experiments is summarised in the table below and the histograms in Figures 6A and 6B. Table 1 Total number of Frames with Subject in View Animal subject R Cam 1 R Cam 2 S Cam 1 S Cam 2 Cheetah 1206 1213 519 512 Human 3306 3287 2188 2145 Total frames obtained for the human was 6593 via rotating cameras and 4333 via static cameras. Thus, the rotating cameras implemented in accordance with aspects of the present disclosure resulted in a 52.16% increase in data collected using the system. Referring to the histogram of Figure 6A, the distribution also shows that the human subject’s position in the image frame was evenly distributed for the static cameras, but was mostly distributed around the centre of the image frame using the rotating camera system. Further, the centre of the approximate Gaussian distribution of human subject locations coincides with the centre of the image frame, which indicates that the tracking system performed well and was able to keep the human subject very close to the centre of the image frame. For the cheetahs, across all three experiments, there was a 134.63 % increase in data collected where the subject was seen by the camera. The histogram in Figure 6B shows the distribution of where the subject was in the camera frame for each frame. With the static cameras the subject is evenly located throughout the image frame, while for the rotating cameras the subject’s centre is within the centre 20% of pixels in the majority of the captured frames. The cheetahs’ positions in the images are skewed slightly to the right of the image, indicating that the tracker was lagging slightly behind the target position. This can be remedied by adjusting the gains of the FSF control module. Similarly, the centre of the approximate Gaussian distribution of cheetah subject locations is situated slightly to the right of the centre of the image frame which indicates that the tracking system was able to track the cheetah, but the response speed can be improved to achieve the same tracking accuracy as for slower moving subjects. In terms of pose estimation, the final output of experimental implementation for 3D motion capture is a three-dimensional motion simulation including or made up of a set of states describing the absolute positions of each key point on the subject in 3D space over a period of time. From these, coordinates x, y, and z of each key point in the world frame are obtained. For the purposes of quantitative evaluation, a 2D reprojection error is used. This was done by reversing the rotational compensation to obtain 3D key point positions relative to the rotating cameras, and using the previously obtained camera extrinsic and intrinsics to reproject each 3D key point to the 2D camera plane. Here, hand-labelled 2D ground truth data for the pose of each subject was compared with the reprojected FTE pose estimates. For the cheetah trajectory, n = 380 points were used. Shown below are the normalised RMSE and Percentage Correct Key points (PCK): Table 2 RMSE, Normalised RMSE and PCK RMSE NRMSE PCK 8.33 0.06 0.94 The normalised RMSE is obtained by dividing the RMSE by * w, where h and w are the height and width of the subject bounding box. The bounding box is also used to calculate the PCK. A keypoint is considered "correct" when the error is less than a*^h * w, where a = 0.1 and was chosen according to the average subject size and shape in the image plane. The 3D poses obtained using the rotating motion capture system were compared to ground truth obtained from static cameras filming the same scene. A sparse bundle adjustment and FTE were used to estimate the ground truth poses from the static cameras, having previously been shown to be robust for cheetah pose estimation. For both SBA and FTE, mean average error (MAE), the RMSE and the mean relative error (MRE) are compared. Table 3 below summarises the results. (S) and (R) denote the use of the static or rotating camera pairs respectively. Table 3 3D Errors for Full Body Reconstruction Using Static and Rotating Cameras Comparison MAE (m) RMSE (m) MRE FTE (S) vs SBA (S) 0.497 0.860 0.307 FTE (R) vs FTE (S) 0.715 0.976 1.900 FTE (R) vs SBA (S) 1.129 1.723 1.749 As shown in Table 3, the RMSE for both FTE (S) and FTE (R) when compared to SBA (S) are significantly high. From visual analysis, the FTE method was determined to introduce a depth error by causing the subject to slide inwards towards the cameras. This is primarily due to missing key-points from the occlusion on the side of the cheetah opposite the camera. To verify this, the 3D positions of a single point (the tail tip) which was seen by all the cameras across the trajectory were compared. The results are shown in Table 4, and show a significant improvement over the results obtained for the entire body. Table 4 3D Errors for Tail Tip Position Estimation Using Static and Rotating Cameras Comparison MAE (m) RMSE (m) MRE Tail Tip: FTE (R) vs SBA (S) 0.444 0.538 1.295 In terms of qualitative results, to provide a visual indication of the performance of the FTE reconstruction, the estimated 3D key points are plotted in the world frame. Figure 7 illustrates reprojected points and a series (702, 704, 605, 708) of three-dimensional motion simulations of an animal subject of interest with respect to the inertial frame for selected frames based on an experimental implementation of the described system and method. Figure 7 shows that the experimental implementation of the system and method described herein yielded plausible reconstructions of the entire cheetah despite sliding caused by the occluded points. Effectiveness of the described system and method on cheetahs in the wild is thus demonstrated. The system and method described herein may be extended to any species with similar results, and RGB cameras may be substituted (or augmented) with thermal or infrared imaging for low-light applications. Thus, the described system and method is versatile and can be effective in various settings. The wide-angle lenses of the GoPros and 180 degree range of rotation of the camera positioning system allow data to be captured within a field of view of up to 320 degrees, greatly increasing the capture volume and quality of research data obtained through the system. This, coupled with the autonomous nature of the described system and method, minimises the need for human labour in setup and operation when deployed in the wild. A system and method for three-dimensional (3D) markerless motion capture of animals in the wild using autonomous tracking cameras is provided. The system may be low-cost and configured to autonomously track moving animals in the wild. The described system and method may overcome limitations of static camera systems with a hardware-constrained field of view. The described system and method may use a camera and sensor platform mounted to a motor to increase the capture volume (i.e. the number of usable image frames) of directional sensors. The position of the subject within the camera frame may be determined, in some examples, using a custom YOLO object detection model. An FSF control module along with an EKF may be implemented to enable precise tracking by positioning the motor shaft such that the subject is always in the centre of the camera’s frame. Utilising markerless 2D pose estimation, two-dimensional (2D) key points may be extracted and a trajectory optimisation-based 3D pose estimation method (e.g., FTE) may be implemented to reconstruct the subject’s 3D trajectories. Performance of an example implementation of the system and method described herein is validated using quantitative and qualitative experiments, showcasing successful tracking of cheetahs in the wild. The described system and method not only address the limitations of conventional camera traps but also presents a cost-effective solution for comprehensive motion capture in diverse environmental conditions. A system and method for generating a three-dimensional motion simulation of an animal subject of interest are thus provided. The described system and method may perform autonomous 3D skeletal pose reconstruction system with adaptive camera tracking &zooming. A motion capture system configured to autonomously track and film subjects (animals or humans) in motion is described herein which maintains the subject of interest within the camera's frame and at an adequate size. This system leverages cameras mounted on motorized pedestals, employing an object detection algorithms (such as YOLO for real-time tracking) and control system (including using EKFs for precise steering). By automatically adjusting camera angles and zoom-level, the system significantly expands the usable capture area, enabling detailed 3D skeletal pose reconstruction without the need for physical markers. The described system and method may find application in various use-cases, from wildlife research and conservation efforts to enhancing animation and virtual reality experiences, offering a scalable, cost-effective solution. The described system and method use motorized pedestals and tracking algorithms allowing cameras to pan, tilt, and / or zoom autonomously. The cameras may be used to capture multiple angles of the subject to facilitate 3D pose reconstruction. Tracking and steering may use realtime object detection and tracking algorithms, coupled with control mechanisms such as an EKF, for precise camera steering based on the subject's movement and zooming based on the size of the detected subject. The system and method described herein identify a subject within a camera’s field of view using the detection algorithms. Once identified, the camera automatically adjusts its orientation to keep the subject centrally framed. As the subject moves, the system dynamically recalibrates, ensuring continuous and comprehensive coverage. Simultaneously, the captured footage is processed to reconstruct the subject's 3D skeletal pose in real-time or near real-time, depending on the processing power available. The system and method described herein may expand capture area and flexibility, allowing for motion capture in larger and more dynamically varying environments. The ability of the cameras to autonomously track and adjust their positioning may provide for the use of fewer and / or lower-cost cameras while still maintaining comprehensive coverage of the subject. By increasing the efficiency of the capture area and allowing for the use of lower-cost cameras, the described system and method may increase accessibility of motion capture technology for a broader range of applications, from research and development to consumer-level products. An experimental implementation of the system and method described herein was found to reconstruct 3D skeletal poses accurately using lower-cost, lower-resolution cameras, which is contrary to conventional wisdom which would suggest that high-resolution cameras were necessary to capture the fine details required for precise pose estimation. However, by leveraging and maximising usable image frames from multiple viewpoints, the described system and method compensate for lower camera resolutions, delivering high-fidelity 3D reconstructions. In this way, overall cost of motion capture setups may be reduced and accessibility of the technology may be extended to a broader range of applications and users. The combination of autonomous tracking and efficient data processing as described herein means that the described system and method can be easily scaled up or down based on the requirements of the capture area or the complexity of the subject's movements. Unlike prior art technologies that require extensive manual setup and adjustment, the autonomous camera steering and subject tracking functionality described herein may increase operational flexibility and efficiency, enabling capture in diverse environments without additional setup time or personnel. The ability to use lower-cost cameras without sacrificing capture quality may open up motion capture technology to sectors previously constrained by budget limitations, including independent filmmakers, small-scale game developers, academic researchers, and the like. By eliminating the need for physical markers or special suits, the system and method described herein may reduce the preparation time for subjects and may simplify the capture process, making it less intrusive and more suited to naturalistic motion capture. In some examples, the described system and method may be integrated with virtual reality and / or augmented reality technologies to immersive experiences in gaming, training simulations, and remote collaboration. This could involve real-time rendering of the captured motion within virtual environments, allowing for interactive scenarios that respond dynamically to user movements. In some examples, the described system and method may be provided on portable hardware, for example using mobile applications and / or compact, battery-powered camera units to provide for on-the-go motion capture in, e.g., outdoorsports analysis, field research, guerrilla filmmaking and the like. Figure 8 illustrates an example of a computing device (800) in which various aspects of the disclosure may be implemented. The computing device (800) may be embodied as any form of data processing device including a personal computing device (e.g. laptop or desktop computer), a server computer (which may be self-contained, physically distributed over a number of locations), a client computer, or a communication device, such as a mobile phone (e.g. cellular telephone), satellite phone, tablet computer, personal digital assistant or the like. Different embodiments of the computing device may dictate the inclusion or exclusion of various components or subsystems described below. The computing device (800) may be suitable for storing and executing computer program code. The various participants and elements in the previously described system diagrams may use any suitable number of subsystems or components of the computing device (800) to facilitate the functions described herein. The computing device (800) may include subsystems or components interconnected via a communication infrastructure (805) (for example, a communications bus, a network, etc.). The computing device (800) may include one or more processors (810) and at least one memory component in the form of computer-readable media. The one or more processors (810) may include one or more of: CPUs, graphical processing units (GPUs), microprocessors, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs) and the like. In some configurations, a number of processors may be provided and may be arranged to carry out calculations simultaneously. In some implementations various subsystems or components of the computing device (800) may be distributed over a number of physical locations (e.g. in a distributed, cluster or cloud-based computing configuration) and appropriate software units may be arranged to manage and / or process data on behalf of remote devices. The memory components may include system memory (815), which may include read only memory (ROM) and random access memory (RAM). A basic input / output system (BIOS) may be stored in ROM. System software may be stored in the system memory (815) including operating system software. The memory components may also include secondary memory (820). The secondary memory (820) may include a fixed disk (821), such as a hard disk drive, and, optionally, one or more storage interfaces (822) for interfacing with storage components (823), such as removable storage components (e.g. magnetic tape, optical disk, flash memory drive, external hard drive, removable memory chip, etc.), network attached storage components (e.g. NAS drives), remote storage components (e.g. cloud-based storage) or the like. The computing device (800) may include an external communications interface (830) for operation of the computing device (800) in a networked environment enabling transfer of data between multiple computing devices (800) and / or the Internet. Data transferred via the external communications interface (830) may be in the form of signals, which may be electronic, electromagnetic, optical, radio, or other types of signal. The external communications interface (830) may enable communication of data between the computing device (800) and other computing devices including servers and external storage facilities. Web services may be accessible by and / or from the computing device (800) via the communications interface (830). The external communications interface (830) may be configured for connection to wireless communication channels (e.g., a cellular telephone network, wireless local area network (e.g. using Wi-Fi™), satellite-phone network, Satellite Internet Network, etc.) and may include an associated wireless transfer element, such as an antenna and associated circuitry. The computer-readable media in the form of the various memory components may provide storage of computer-executable instructions, data structures, program modules, software units and other data. A computer program product may be provided by a computer-readable medium having stored computer-readable program code executable by the central processor (810). A computer program product may be provided by a non-transient or non-transitory computer-readable medium, or may be provided via a signal or other transient or transitory means via the communications interface (830). Interconnection via the communication infrastructure (805) allows the one or more processors (810) to communicate with each subsystem or component and to control the execution of instructions from the memory components, as well as the exchange of information between subsystems or components. Peripherals (such as printers, scanners, cameras, or the like) and input / output (I / O) devices (such as a mouse, touchpad, keyboard, microphone, touch-sensitive display, input buttons, speakers and the like) may couple to or be integrally formed with the computing device (800) either directly or via an I / O controller (835). One or more displays (845) (which may be touch-sensitive displays) may be coupled to or integrally formed with the computing device (800) via a display or video adapter (840). The foregoing description has been presented for the purpose of illustration; it is not intended to be exhaustive or to limit the technology to the precise forms disclosed. Persons skilled in the relevant art can appreciate that many modifications and variations are possible in light of the above disclosure. Any of the steps, operations, components or processes described herein may be performed or implemented with one or more hardware or software units, alone or in combination with other devices. Components or devices configured or arranged to perform described functions or operations may be so arranged or configured through computer-implemented instructions which implement or carry out the described functions, algorithms, or methods. The computer-implemented instructions may be provided by hardware or software units. In one embodiment, a software unit is implemented with a computer program product comprising a non-transient or non-transitory computer-readable medium containing computer program code, which can be executed by a processor for performing any or all of the steps, operations, or processes described. Software units or functions described in this application may be implemented as computer program code using any suitable computer language such as, for example, Java™, C++, or Perl™ using, for example, conventional or object-oriented techniques. The computer program code may be stored as a series of instructions, or commands on a non-transitory computer-readable medium, such as a random access memory (RAM), a read-only memory (ROM), a magnetic medium such as a hard-drive, or an optical medium such as a CD-ROM. Any such computer-readable medium may also reside on or within a single computational apparatus, and may be present on or within different computational apparatuses within a system or network. Flowchart illustrations and block diagrams of methods, systems, and computer program products according to embodiments are used herein. Each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, may provide functions which may be implemented by computer readable program instructions. In some alternative implementations, the functions identified by the blocks may take place in a different order to that shown in the flowchart illustrations. Some portions of this description describe the examples in terms of algorithms and symbolic representations of operations on information. These algorithmic descriptions and representations, such as accompanying flow diagrams, are commonly used by those skilled in the data processing arts to convey the substance of their work effectively to others skilled in the art. These operations, while described functionally, computationally, or logically, are understood to be implemented by computer programs or equivalent electrical circuits, microcode, or the like. The described operations may be embodied in software, firmware, hardware, or any combinations thereof. The language used in the specification has been principally selected for readability and instructional purposes, and it may not have been selected to delineate or circumscribe the 5 inventive subject matter. It is therefore intended that the scope of the present disclosure be limited not by this detailed description, but rather by any claims that issue on an application based hereon. Accordingly, the present disclosure is intended to be illustrative, but not limiting, of the scope of any accompanying claims. 10 Finally, throughout the specification and any accompanying claims, unless the context requires otherwise, the word ‘comprise’ or variations such as ‘comprises’ or ‘comprising’ will be understood to imply the inclusion of a stated integer or group of integers but not the exclusion of any other integer or group of integers.

Claims

1. A computer-implemented method comprising:over successive image frames captured by a camera mounted on a camera positioning system, detecting and tracking a location of an animal subject of interest in each image frame, including controlling parameters of the camera positioning system so as to maintain the animal subject of interest near a centre of each image frame;obtaining position measurements from the camera positioning system while capturing the successive image frames and mapping each image frame to a position measurement representing the position of the camera at the time the image frame is captured to generate an image frame data set including the successive image frames mapped to corresponding position measurements; and,outputting the image frame data set for generating a three-dimensional motion simulation of the animal subject of interest.

2. The method as claimed in claim 1, wherein detecting the location of the animal subject of interest includes using an object detection algorithm which identifies the animal subject of interest.

3. The method as claimed in claim 1 or claim 2, wherein detecting the animal subject of interest includes fitting a bounding box around the detected animal subject of interest.

4. The method as claimed in any one of the preceding claims, wherein tracking the location of the animal subject of interest includes using a tracking algorithm to estimate a position control parameter for controlling the camera positioning system to reposition the camera to a next position from which the animal subject of interest is estimated to be near a centre of a next image frame to be captured by the camera.

5. The method as claimed in claim 4, wherein the position control parameter is an armature current value for controlling armature current of a motor of the camera positioning system.

6. The method as claimed in any one of the preceding claims, wherein tracking the location of the animal subject of interest includes tracking a location of a or part of a bounding box fitted around the animal subject of interest.

7. The method as claimed in any one of the preceding claims, wherein the camera positioning system is configured to rotate the camera about an axis.

8. The method as claimed in any one of the preceding claims, wherein the position measurement is an encoder measurement obtained from an encoder that encodes an orientation of the camera about an axis.

9. The method as claimed in claim 7 or claim 8, wherein the axis is a vertical axis for tracking the location of the animal subject of interest along the horizontal.

10. The method as claimed in any one of the preceding claims, including: obtaining the image frame data set; and, processing the image frame data set to generate the three-dimensional motion simulation of the animal subject of interest.

11. The method as claimed in claim 10, wherein processing the image frame data set includes fitting a skeletal model of the animal subject of interest to a set of two-dimensional pose measurements obtained from the image frame data set.

12. The method as claimed in claim 10 or claim 11, wherein processing the image frame data set includes obtaining, from each image frame, a set of two-dimensional pose measurements.

13. The method as claimed in claim 12, wherein obtaining the set of two-dimensional pose measurements from each image frame includes labelling, in each image frame, the animal subject of interest with key points corresponding to key points of the skeletal model of the animal subject of interest.

14. The method as claimed in claim 13, wherein the set of two-dimensional pose measurements include values for each key point in a coordinate system based on the pose of the animal subject of interest in the image frame, wherein the coordinate system is a two-dimensional position coordinate system, wherein the coordinate system is a coordinate system of the image frame.

15. The method as claimed in any one of claims 10 to 14, wherein processing the image frame data set includes compensating for changes in camera position from one image frame to the next.

16. The method as claimed in claim 15, wherein compensating for changes in camera position includes updating extrinsic camera parameters for each image frame.

17. The method as claimed in claim 16, wherein updating extrinsic camera parameters for each image frame includes:generating a transformation from the camera into the camera positioning system;generating a rotation of the transformation about an axis, wherein a magnitude of the rotation is determined based on the position measurement mapped to the image frame; and, generating an inverse transformation from the camera positioning system to the camera.

18. The method as claimed in claim 17, wherein compensating for changes in camera position includes applying the inverse transformation to a skeletal model of the animal subject of interest generate a position-compensated skeletal model of the animal subject of interest, and wherein processing the image frame data set includes fitting the position-compensated skeletal model of the animal subject of interest to the two-dimensional pose measurements.

19. A system comprising a camera positioning system including a control module and which mounts a camera configured to capture image frames, wherein the control module is configured: over successive image frames captured by the camera, to detect and track a location of an animal subject of interest in each image frame, including controlling parameters of the camera positioning system so as to maintain the animal subject of interest near a centre of each image frame; to obtain position measurements from the camera positioning system while capturing the successive image frames and to map each image frame to a position measurement representing the position of the camera at the time the image frame is captured to generate an image frame data set including the successive image frames mapped to corresponding position measurements; and, to output the image frame data set for generating a three-dimensional motion simulation of the animal subject of interest.

20. The system as claimed in claim 19, wherein the control module is configured to track the location of the animal subject of interest using a tracking algorithm which estimates a position control parameter for controlling the camera positioning system to reposition the camera to a next position from which the animal subject of interest is estimated to be near a centre of a next image frame to be captured by the camera.

21. The system as claimed in 20, wherein the camera positioning system includes a motor arranged to position the camera and wherein the position control parameter is an armature current value for controlling armature current of the motor.

22. The system as claimed in any one of claims 19 to 21, wherein the camera positioning system is configured to rotate the camera about an axis, wherein the position measurement is an encoder measurement obtained from an encoder that encodes an orientation of the camera about the axis, and wherein the axis is a vertical axis for tracking the location of the animal subject of interest along the horizontal.

23. The system as claimed in any one of claims 19 to 22, including a computing device configured: to obtain the image frame data set; and, to process the image frame data set to generate the three-dimensional motion simulation of the animal subject of interest.

24. The system as claimed in claim 23, including a plurality of camera positioning systems, each of which mounts a camera and outputs an image frame data set, wherein the computing device is configured to obtain a plurality of image frame data sets and to process the plurality of image frame data sets to generate the three-dimensional motion simulation of the animal subject of interest.

25. A computer program product comprising a computer-readable medium having stored computer-readable program code for performing the steps of:over successive image frames captured by a camera mounted on a camera positioning system, detecting and tracking a location of an animal subject of interest in each image frame, including controlling parameters of the camera positioning system so as to maintain the animal subject of interest near a centre of each image frame;obtaining position measurements from the camera positioning system while capturing the successive image frames and mapping each image frame to a position measurement representing the position of the camera at the time the image frame is captured to generate an image frame data set including the successive image frames mapped to corresponding position measurements; and,outputting the image frame data set for generating a three-dimensional motion simulation of the animal subject of interest.

Citation Information

Patent Citations

  • Wide Surveillance Camera using Panoramic View

    KR1020030073490A

  • Array-camera motion picture device, and methods to produce new visual and aural effects

    US20010028399A1

  • Method and apparatus for streamlined wireless data transfer

    US20090201380A1

  • Systems and methods for creating target motion, capturing motion, analyzing motion, and improving motion

    US20180357472A1

  • Method and system for image centering during zooming

    WO2013096331A1