Motion capture system and method for generating synchronized scene images and marker position data
By integrating a marker tracking and removal subsystem into a motion capture camera, high-fidelity synchronized marker data and scene data are generated, solving the problem of inaccurate correspondence between marker position data and scene images in existing technologies, and improving motion capture accuracy and training data quality for machine learning systems.
Patent Information
- Application Number
- CN202510553887.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-03-14
- Filing Date
- 2025-04-29
- Publication Date
- 2026-03-03
AI Technical Summary
Existing motion capture systems are unable to achieve precise synchronization between scene data and marker data in some applications, resulting in inaccurate correspondence between marker position data and scene images, which affects the accuracy of motion capture and the quality of training data for machine learning systems.
A motion capture camera is used to integrate a marker tracking subsystem and a marker removal subsystem. Digital image data is generated through an image sensor, the marker tracking subsystem processes the marker position, and the marker removal subsystem removes the markers and encodes the video data to generate high-fidelity synchronized marker data and scene data for training a machine learning system.
It achieves precise synchronization between marker data and scene data, improves the accuracy of motion capture, and provides high-quality training data for markerless object tracking, thereby enhancing the accuracy and efficiency of machine learning systems.
Smart Images

Figure CN121603646A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to motion capture systems, and more particularly to motion capture cameras / cameras and methods for collecting digital video data and synchronized position data about subjects in a scene, and related methods for generating high-fidelity training data for machine learning systems for markerless object tracking. Background Technology
[0002] Motion capture systems are used to track the movement of one or more real-world objects. Computer models can then be mapped onto these objects to create animations and film effects that accurately mimic real-world movement. Furthermore, motion capture allows for more efficient animation and effects production compared to frame-by-frame generation techniques. Motion capture systems also allow animation or visual effects directors to experiment with different movements or perspectives before mapping movement onto computer models, leading to more flexible content creation.
[0003] Typical motion capture equipment includes multiple cameras that detect one or more objects in a scene by identifying the location of markers mounted on the object (e.g., a human body). The markers can be active markers that emit light (e.g., light of a selected wavelength) or passive markers, such as reflectors or white dots that only reflect incident light generated by an external source (e.g., infrared illumination). In many cases, motion capture cameras are equipped with filters to improve the signal-to-noise ratio of the image detected by the camera, making marker identification easier. Alternatively, motion capture equipment may include one or more cameras without filters to record a normal view of the scene in the visible spectrum.
[0004] U.S. Patent No. 9,019,349, owned by the assignee of this application, discloses a motion capture camera system including a marker-tracking optical filter that relatively amplifies light from a marker on a moving object in a scene and is selectively interchangeable with scene-view optics. These motion capture cameras are remotely controllable to selectively switch between marker-tracking mode and scene mode by turning the marker-tracking optical filter on or off. This remote switching allows the same camera to capture object position data in marker-tracking mode and capture a reference scene in scene mode, but not simultaneously.
[0005] The inventors have recognized that asynchronous capture of scene data and marker data may not be the optimal choice for certain applications where accurate correspondence between marker location data and scene images is crucial. Summary of the Invention
[0006] A motion capture system includes one or more motion capture cameras, each having an image sensor operable to generate a series of frames representing digital image data of a scene visible to the motion capture camera. In some embodiments, the motion capture system may include a set of motion capture cameras arranged around a capture space to capture different aspects of the scene, and the motion capture cameras may be interconnected with each other via a local area network and / or connected to a host computer system for synchronization and / or calibration. Each motion capture camera includes a marker tracking subsystem configured to access the digital image data generated by the image sensor and process at least a portion of the digital image data to determine the current position of each of a plurality of reflective or luminescent markers attached to a moving subject in the scene for each of at least some of the frames in the series. Thus, the marker tracking subsystem generates a series of marker datasets, each marker dataset corresponding to one frame in the series of tracked image frames. Each motion capture camera may also include an encoder configured to access the digital image data and encode at least some of the frames in the series into compressed video data, including at least some of the frames in the series of tracked frames processed by the marker tracking subsystem. The data communication device of a motion capture camera can be configured to transmit compressed video data and a series of marker datasets. For example, the frame rate of a motion capture camera can be between 10 and 1000 frames per second. The series of frames of tracked digital image data can include an entire series of frames, or essentially consist of one of the following: a series of digital image data frames acquired at a frame rate (e.g., a subset of adjacent frames), a series of non-adjacent frames, or a series of adjacent and non-adjacent frames.
[0007] Each motion capture camera may also include a marker removal subsystem configured to modify each frame of digital image data to remove markers in the scene, after which an encoder encodes the series of such modified digital image data frames. The encoded and compressed video data, along with the corresponding dataset of markers, can be received by the host computer system of the motion capture system for subsequent use and processing, and the host computer system may optionally store this data to maintain synchronization and / or correspondence between each frame of compressed video data and its corresponding dataset for some or all motion capture cameras.
[0008] To improve efficiency and reduce processing overhead, the marker tracking subsystem and / or marker removal subsystem may process only a subset of the digital image data for each frame, including one or more regions of interest (ROIs) identified for marker acquisition. In some embodiments, a series of marker datasets and / or modified scenes (the scene after marker removal) may be generated at or near the frame rate of the image sensor, and the marker datasets and compressed video data may be transmitted at or near that frame rate. The marker tracking subsystem, marker removal subsystem, and encoder may all be implemented in a digital data processor, such as one or more field-programmable gate arrays (FPGAs) and / or one or more application-specific integrated circuits (ASICs), each communicating with the image sensor and communication devices. In one embodiment, the image sensor and digital data processor may be implemented in a single ASIC.
[0009] According to another aspect of this disclosure, a method for generating motion capture data and image data may include the following steps: (1) generating a series of adjacent and / or non-adjacent digital image data frames representing a scene visible to a motion capture camera via an image sensor; (2) processing at least a portion (e.g., ROI) of the digital image data via a marker tracking subsystem to determine the current position of each of a plurality of reflective or luminescent markers attached to a moving object in the scene for each of at least some of the frames in the series of frames, the marker tracking subsystem generating a series of marker datasets, each marker dataset corresponding to one frame in the series of frames being tracked and including the current position of the markers in the scene; (3) encoding at least some of the series of digital image data frames via an encoder to generate compressed video data, wherein the compressed video data includes at least some of the tracked frames processed by the marker tracking subsystem; and (4) transmitting the compressed video data and the series of marker datasets from the motion capture camera via a communication device. Before encoding the series of digital image data frames (or a portion thereof), each frame (or its ROI) may be modified by removing markers from the scene to generate modified digital image data frames for encoding by the encoder.
[0010] The compressed-coded video data generated by the system and method according to this disclosure, along with a corresponding dataset of markers, can be used to train machine learning systems or other AI systems for markerless motion capture. Removing markers from an image scene provides a markerless modified video that corresponds precisely frame-by-frame to the marker dataset, thereby enabling the generation of accurate object data regarding the position and orientation of a subject or object in the markerless scene to aid in training the machine learning system (e.g., by informing or validating the training). The compressed-coded video data (with or without marker removal) and the corresponding marker dataset can also be transferred to the trained machine learning system. For example, the marker data can be used for high-precision tracking of some objects or elements in a scene, while markerless AI-based tracking can be used for other elements where lower precision requirements are needed or where markers are difficult to attach.
[0011] Other aspects and advantages will become apparent from the following detailed description of exemplary embodiments, which will be made with reference to the accompanying drawings. Attached Figure Description
[0012] Figure 1 A motion capture system according to one embodiment is shown.
[0013] Figure 2 yes Figure 1 A schematic block diagram of a networked camera and host computer system for a motion capture system.
[0014] Figure 3 An image showing a scene in the normal visible spectrum, including figures with markers attached.
[0015] Figure 4 Showing from Figure 3 The data captured in the scene shows the location of the markers in the scene.
[0016] Figure 5 Show Figure 3 The character's object model, including the main skeletal joints, has been based on... Figure 4 It is constructed from marker tracking data.
[0017] Figure 6 A method for generating training data for a machine learning system for markerless motion capture is shown according to one embodiment.
[0018] Figure 7 According to one embodiment Figure 1 An isometric view of one of the cameras in a motion capture system.
[0019] Figure 8 Schematic illustration according to one embodiment Figure 7 The components of the camera. Detailed Implementation
[0020] For ease of identification in the discussion of any particular element or action, the most prominent one or more digits in the reference numerals appearing in the accompanying drawings and the following detailed description refer to the drawing number described when the element was first introduced. The same reference numeral appearing in multiple drawings refers to the same element throughout the document.
[0021] Figure 1 An embodiment of a motion capture system 100 according to the present disclosure is shown. The motion capture system 100 includes a plurality of motion capture cameras 102 configured to receive light from a scene 104. In the illustrated embodiment, six cameras 102 are distributed around and pointed at a capture space 106 to capture different aspects of the scene 104. In other embodiments, more or fewer cameras 102 may be used. For example, in some embodiments, a single motion capture camera may be used, while in other embodiments, two (2) to one thousand (1000) or more motion capture cameras may be used to capture the scene of a single capture space from different perspectives.
[0022] Multiple markers 108 may be attached to various locations on the subject 110 (such as a person or animal) and / or on other objects in the capture space 106. In some embodiments, the markers 108 are passive markers that reflect incident light to enhance the brightness of the markers 108 relative to the surrounding scene 104 so that they can be detected by the multiple cameras 102. In other embodiments, the markers 108 are active markers that emit their own light rather than just reflect it, so they are brighter than other elements of the subject 110 or scene 104, making these active markers easily detectable by the cameras 102. As an example, each active marker may include one or more light-emitting diodes (LEDs) housed within a spherical diffuser housing having a predetermined diameter. Passive markers may include various reflective objects or materials, such as white spheres, reflective paint dots, circles or spheres made of reflective material, retroreflective corner prisms, or retroreflective material having a pattern of multiple corner prism reflectors. Markers may be implemented in any of a variety of shapes and sizes. In some embodiments, the cameras 102 may include one or more illumination sources 702 ( Figure 7 These illumination sources can be positioned substantially along the optical axis 802 of the camera 102. Figure 8The light source 702 emits light to illuminate scene 104 and marker 108. In some embodiments, the light source 702 may emit light substantially coaxially aligned with the optical axis 802. In some embodiments, the light source 702 for each camera 102 includes an LED emitting broad-spectrum visible light, and in some cases also includes infrared (IR) illumination. In other embodiments, the light source 702 is a narrowband emitter, such as an IR LED or other IR illumination device that emits only wavelengths in the near-infrared spectrum. This IR illumination is reflected by marker 108 without affecting the visible appearance of scene 104. In alternative embodiments, various other broadband or narrowband illumination wavelengths (different from visible light or IR) may also be used.
[0023] The position of marker 108 in scene 104 can be identified by a marker tracking subsystem of camera 102. In some embodiments, the marker tracking subsystem can identify the size and shape of the marker, providing additional information about the marker's extent (distance from the camera) and orientation. The marker positions and sizes detected by multiple cameras 102 can be correlated, triangulated, and mapped to a three-dimensional (3D) object model to determine the 3D spatial position and movement of subject 110 or other objects in capture space 106. Host computer system 120 can communicate with cameras 102 and is configured to receive marker position data from multiple cameras 102 via a wired or wireless local area network, and perform marker data correlation, triangulation, and mapping to a 3D object model to record the movement of subject 110. Subject 110 can include any suitable subject or object or set of subjects or objects whose movement can be tracked using markers 108 fixed to or relative to the moving subject or object. For example, the subject to be tracked can include facial features, animals, people, etc. Furthermore, any suitable number of markers can be deployed on the object to appropriately track its movement. For example, one to dozens or hundreds of markers can be attached to a single moving subject. In some cases, one or more markers 108 can be attached to a stationary object or other subject, such as a reference square 122 with three markers defining a plane, which is tracked as a reference reference in scene 104.
[0024] Cameras 102 can also be interconnected via a wired or wireless local area network so that marker data output by one camera 102 can be received by other cameras, thereby providing marker position feedback. For example, such marker position feedback can improve the operation and fidelity of the marker tracking subsystem of each camera. The motion capture system 100 can be configured such that each of the multiple cameras 102 has a different position and orientation relative to the capture space 106 to capture the scene 104 from different vantage points, so that marker data from multiple cameras can be used to accurately triangulate the position of marker 108. Cameras 102 can be jointly synchronized and calibrated, which may involve determining and recording the relative timing and position of cameras 102 through synchronization and calibration between cameras via a host computer system 120 and / or without using a host computer. During calibration, one or more reference markers (such as a set of markers on calibration rod 124) can be moved within the field of view of camera 102 to create a set of marker position and timestamp data, which is organized into a calibration dataset from which the relative position offset and viewpoint offset of camera 102 can be derived. The capture space 106 can be defined based on the camera calibration procedure or as a result of the camera calibration procedure, wherein locations outside the capture space 106 cannot be seen by all or a sufficient number of cameras 102, and therefore the motion capture system 100 may not be able to accurately track objects outside the capture space 106 in 3D space. Further aspects and features of the calibration procedure are well known, and much of it is described in U.S. Patent No. 9,019,349.
[0025] Figure 2 This is a schematic block diagram of the network connection between the motion capture system 100, the camera 102, and the host computer system 120. (Reference) Figure 2 The camera 102 can be directly connected to the host computer system 120 via a suitable data connection (such as USB, Ethernet, wireless network (e.g., Wi-Fi 802.11), etc.) as shown in the figure. In some embodiments, the camera 102 can be connected (e.g., via Ethernet) to one or more network switches (not shown), which are then connected to the host computer system 120 via further network connections. The host computer system 120 may include a display subsystem 202 and a data processing subsystem 204, which communicate with a memory 206 storing a motion capture application 208. The roles of these components of the host computer system 120 will be described below. Figures 6-8 This becomes clear in the description of the components and operation of camera 102.
[0026] Figure 3The image shows a raw visual image of scene 104 in the normal visible spectrum, captured by one of the cameras 102, which includes a subject 110 (individual) with a marker 108 attached, for example by attaching the marker through a motion capture suit worn by the subject 110.
[0027] Figure 4 Showing from Figure 3 The scene 104 captures marker tracking data, which shows the position of marker 108 in scene 104, but the scene image is omitted.
[0028] Figure 5 Show Figure 3 The animated rendering of the object model 502 of the subject 110, which includes major skeletal joints 504, was performed by the host computer system 120 based on... Figure 4 The marker tracking data shown is used to generate the marker 108. Marker tracking data can be acquired from multiple cameras 102 to obtain the accurate 3D position of the marker 108. The object model 502 can represent or apply motion constraints to the joints 504. The position of the marker 108 can be displayed relative to the object model 502, or as part of the object model 502.
[0029] The inventors have noted recent efforts to develop artificial intelligence (AI) systems for markerless motion tracking, utilizing video from one or more conventional video cameras. These markerless motion tracking systems operate as their name suggests, where subjects and objects in a scene are presented without attached markers. Instead of marker position data, AI-based markerless systems utilize software architectures such as neural networks and other machine learning systems to directly determine object models from image data. This AI-based image processing technique can largely derive object models (e.g., the position of joint 504) from the edges and shapes of objects appearing in the video. To date, such AI-based systems have not proven reliable or accurate, frequently producing artifacts and errors in the object model. One reason for the poor performance of existing AI-based markerless motion capture systems may be the lack of high-quality training data. For example, most machine learning-based AI systems may only be trained on scene images and possibly some user corrections or other supervised feedback. Therefore, the inventors have identified an opportunity to acquire and utilize large amounts of accurate, high-fidelity training data, including scene data and synchronized marker position data. However, known conventional camera systems cannot generate such high-fidelity synchronized data.
[0030] refer to Figure 6The method 600 for generating high-fidelity synchronized image data and marker data according to this disclosure includes the following steps: providing one or more motion capture cameras, such as camera 102 with enhanced image capture and marker data capture capabilities (hereinafter referred to in detail). Figure 7 and Figure 8 (Further description). According to method 600, the image sensor 804 of camera 102 ( Figure 8 A series of frames representing digital image data of a scene visible to the motion capture camera are generated at a certain frame rate, wherein the scene includes a moving subject with multiple passive or active markers attached. Image sensor 804 can operate at frame rates from 10 frames per second (fps) (10 Hz) to approximately 500 fps (500 Hz), 1000 fps (1000 Hz), or higher, but more typical frame rates are 30 fps to 120 fps (20 Hz to 120 Hz), 30 fps to 100 fps (30 Hz to 100 Hz), or 30 fps to 60 fps (30 Hz to 60 Hz) to produce relatively smooth video images. After frames of digital image data are generated by image sensor 804 in step 602, the image data is processed on camera 102 in steps 604-608, and then transferred in step 610 to host computer system 120 or a data repository for storage and subsequent use, such as as training data for a machine learning system.
[0031] In step 604 of method 600, the marker tracking subsystem 806 via camera 102 ( Figure 8The marker tracking subsystem 806 processes each frame of at least some digital image data to determine the current position of each marker 108 in the scene. The marker position data generated in this way is metadata about the original image frames, which can be used to annotate the image frames. In some examples, an entire series of frames of digital image data acquired by the image sensor 804 is processed by the marker tracking subsystem 806 in step 604 to generate marker position data for adjacent frames in the series. In other examples, only non-adjacent frames in this series are processed in step 604 as the frames to be tracked. And in a further example, a series consisting of adjacent and non-adjacent frames is processed by the marker tracking subsystem 806 in step 604 as the frames to be tracked. Because the marker tracking subsystem 806 operates on the raw image data on the camera 102, the accuracy of marker tracking is improved compared to image data that has been compressed and transmitted from the camera 102 to, for example, a host computer system. It is worth noting that bandwidth limitations make it impossible or infeasible to transmit raw image data at the full frame rate of the image sensor 804 for processing outside the camera, especially when multiple cameras are used. Therefore, transmitting video image data from a camera at frame rate typically requires compression of the video image data on the camera before transmission. Compared to systems that use different cameras to capture video and capture marker tracking data, implementing the marker tracking subsystem 806 in the same camera 102 for both video acquisition and transmission allows marker data and video images to be spatially and temporally aligned, at least for the tracked frames processed by the marker tracking subsystem 806. This "duplex" method of capturing video and marker position data enables the generation of marker position data with higher fidelity using half the number of cameras.
[0032] In optional step 606, the optional marker removal subsystem 810 of camera 102 ( Figure 8 This involves selectively "removing" markers from image frames in image data. Removing markers from images allows video image data to simulate marker-free scenes, resulting in improved, more realistic training data, while marker location data (metadata) provides "realistic" feedback for machine learning.
[0033] In step 608, the encoder 812 on the camera 102 ( Figure 8The encoder 812 encodes digital image data (which may optionally be modified digital video image data with markers removed) to generate compressed video data at the frame rate. A suitable encoder may use an intra-frame-only compression scheme (such as M-JPEG) to compress the digital image data. In other embodiments, the digital image data may be encoded using an inter-frame video compression scheme (such as H.264). Encoder 812 may include multiple encoding engines that process different portions of frames of image data or different frames in a series of frames, for example, in parallel. Thus, even if each of the multiple encoding engines of encoder 812 is operable to compress the digital image data or a portion thereof at a rate much lower than the frame rate, encoder 812 is still able to compress the digital image data (e.g., modified digital image data) at the frame rate. In some embodiments, where markers are not removed or are not removed before compression, the compression of a video image data frame in step 608 may be performed simultaneously with the processing of the same video image data frame in step 604 to determine the location of markers in that frame, for example, in the same digital data processor 808 of camera 102. Figure 8 Parallel processing is performed on the image sensor. In some embodiments, only a portion of the digital image data is compressed for transmission. For example, when extremely high-precision marker location data is required but lower-precision video data is needed, marker location data can be acquired at a high frame rate (e.g., 1000 fps), but only a portion of the frames acquired at this high frame rate are encoded as compressed video; for example, one frame out of every 10 frames of marker location data acquired is encoded at 100 fps. In other embodiments, only a subset of a series of frames of digital image data acquired by the image sensor at a frame rate is processed to generate marker location data (i.e., the tracked frames are a subset of a series of frames of the acquired digital image data), but the entire series of frames is encoded as compressed video, for example, when lower precision is required or for objects that are not moving or moving slowly.
[0034] In step 610, the communication device 814 of camera 102 ( Figure 8Encoded compressed video data is transmitted from the motion capture camera. Corresponding synchronization marker data for at least some frames encoded as compressed video data may also be transmitted. In some embodiments, the compressed video data and synchronization marker data are transmitted at or near frame rate. Alternatively, synchronization marker data for a series of frames or a subset thereof may be accumulated in memory 816 (such as DRAM memory on camera 102), and the accumulated series of marker datasets may then be transmitted periodically, or read periodically or on demand by the host computer system. Subsequently, steps 602 to 610 may be repeated for each consecutive frame captured by the image sensor. Therefore, the marker tracking subsystem preferably generates a series of marker datasets at frame rate, wherein each marker dataset corresponds to the visible content in a corresponding frame of a series of frames of an image generated by the image sensor, including the current position of the marker in each frame. In some embodiments, the marker datasets are generated at frame rate, while only some frames are encoded as compressed video and transmitted to reduce bandwidth while acquiring high-speed marker data. In other embodiments, the entire series of video image data acquired at the frame rate is encoded and compressed, but only a subset of these frames is used to acquire marker location data. Thus, in some embodiments, marker location data can be acquired from adjacent frames, while in others it is acquired only from non-adjacent frames, and in still others, it can be acquired from a combination of adjacent and non-adjacent frames. In yet another embodiment, a series of adjacent frames of digital image data can be encoded into compressed video data. And in still others, only a portion of a series of adjacent frames of digital image data is encoded, such that the compressed video consists substantially of non-adjacent frames, or substantially of a combination of adjacent and non-adjacent frames. In any case, at least some frames of the generated marker location data are synchronized with corresponding frames of the video image data, because each frame of the marker location data and its corresponding video image frame (if transmitted) are generated from the same frame of image data acquired by the image sensor.
[0035] In some embodiments, the training data generated by method 600 may involve acquiring training data from a single camera 102. Alternatively, by applying the above method to multiple synchronized cameras 102, different advantageous perspectives of scene 104 and subject 110 can be acquired to generate training data for training a machine learning system to perform markerless motion capture using a multi-camera setup, achieving higher accuracy and fidelity compared to a single-camera system. In some embodiments, object models can be used to train the machine learning system. For example, labeled data can be mapped to corresponding object models, and then the mapped labeled data (object model data) can be used to train the machine learning system. For example, labeled data of subject 110 (which is a person) can be mapped to a skeletal object model to deduce the location and orientation of major bones in a human skeleton. Labeled data of different subjects 110 (such as rigid subjects or other types of objects (not a person)) can be mapped to different object models (different from the human skeleton model). In some cases, multiple object models of the same type or various types may correspond to multiple subjects and / or objects in a single video scene; and the scene video and real data provided by multiple object models can be used to train the machine learning system.
[0036] Similar to the training methods described above for machine learning systems that derive marker position data from markerless videos, machine learning systems trained using object models can be configured to derive skeletal positions or other object model data from modified scene images with markers removed, and to improve their training by comparing their results with skeletal positions or other object model data derived by motion capture system 100.
[0037] The motion capture system 100 and method 600 can also be used to generate marker tracking data and video data (with or without marker removal), which are then sent to a previously trained machine learning system, where both the marker data and video data can be used for tracking by the machine learning system. In a further example, the camera 102 of the motion capture system 100 can perform some aspects of AI processing (preprocessing) on camera 102, after which the video data, the output of the AI preprocessing, and optional marker location data are sent to a central host system or network for further AI processing.
[0038] Figure 7 Details of an exemplary motion capture camera 102 for practicing the systems and methods of the present invention according to this disclosure are shown. References Figure 7 The camera 102 includes a lens 704 and an illumination source 702 surrounding the lens 704 on the front portion 706 of the camera 102. The electronic components of the camera 102 (see below) Figure 8(As described) can be mainly housed in the main body 710 of the camera 102 behind the lens 704.
[0039] Turn now Figure 8 The motion capture camera 102 includes an image sensor 804 that generates frames of digital image data using light focused onto it by the lens 704 of the camera 102. The image sensor 804 may include a CMOS image sensor or other type of sensor with a global shutter and may operate at frame rates from approximately 10 frames per second (fps) (10Hz) to approximately 1000 fps (1000Hz) or higher, but more typically from approximately 30 fps to approximately 120 fps (30Hz to 120Hz) or 30 fps to 60 fps (30Hz to 60Hz) to produce relatively smooth video images. In some embodiments, the image sensor 804 may operate in servo mode, with its shutter triggered by the digital data processor 808 of the camera 102, to allow the digital data processor 808 to maintain shutter synchronization with other cameras 102. Camera 102 includes a marker tracking subsystem 806, which can be implemented in a digital data processor 808 of camera 102 that communicates with image sensor 804. The marker tracking subsystem 806 is configured to receive, read, or access digital image data generated by image sensor 804, and process one or more frames of the digital image data to determine the position of markers in each specific image frame of a series of frames. The marker tracking data generated in this manner is synchronized with the frames of the video image data thus processed. The marker tracking subsystem 806 can generate a set of marker tracking data for images containing 1 to 100 markers, or up to 1000 markers, or more preferably up to 10000 or more markers, at a rate equal to or higher than the frame rate.
[0040] The marker tracking subsystem 806 of camera 120 can use any of a variety of image processing techniques to determine marker location data. For example, determining the XY position of each marker 108 in an image frame may involve a first step of scanning pixel rows to identify regions of interest (ROIs) in the image that meet certain minimum criteria, such as a group of two or more adjacent pixels having a predetermined minimum brightness. In some embodiments, the marker tracking subsystem 806 may utilize marker location data previously determined for a previous frame or several previous frames of video image data to assist in quickly finding the XY position of the same marker in the current frame. For example, the XY marker location data of marker 108 in a previous frame may be stored in memory 816 and used in subsequent "current" frames of video image data to determine an ROI window in which the same marker 108 in the current frame is analyzed. As a further example, marker position data of a specific marker 108 in a series of previous frames can be used to approximate or represent the trajectory of marker 108, which can be stored in memory 816 and used by marker tracking subsystem 806 in subsequent "current" frames to determine the ROI window for processing the current frame.
[0041] Camera 102 may optionally include a marker removal subsystem 810 configured to modify each frame of digital image data to remove or otherwise exclude markers 108 from the image data, thereby creating a modified image. The modified image data is then compressed by encoder 812 and transmitted from the camera via communication device 814. During the optional step 604 of determining the marker location, markers 108 can be conveniently and efficiently removed from each frame of the original digital image data using the ROI already stored in memory 816, and then the modified (marker-removed) digital image data is recombined and encoded, rather than removing the markers from the entire frame of the digital image data, or removing them after encoding and compressing the digital image data. Removing markers from the video image before encoding the video image allows for removal using the original background data closely surrounding markers 108, which is more accurate than using encoded data from the same area (which may be corrupted during compression). Removing markers before encoding also allows the removed portions of the modified digital image data to be smoothed and / or blurred during the encoding and compression processes, thereby reducing imperfections in the removed areas. In one embodiment, the marker removal subsystem 810 can conveniently and efficiently operate on pixel data in each ROI stored in memory 816 shortly after determining the XY position of the marker 108 in the ROI. From a data processing perspective, removing markers in the ROI data is more convenient and efficient than removing markers in the entire image frame. In some embodiments, the encoder 812 (or its multiple encoding engines) can cooperate with the marker removal subsystem 810 so that encoding and compression begin only after the portion of the image frame being processed by the encoder 812 has been marker-removed and the removed areas have been reassembled.
[0042] The digital data processor 808 may include, for example, a CPU, GPU, field-programmable gate array (FPGA), or application-specific integrated circuit (ASIC), and the marker tracking subsystem 806 and / or marker removal subsystem 810 may be programmed into the digital data processor 808, and / or embodied in software stored in memory 816, or embodied in another machine-readable medium. In other embodiments, the marker tracking subsystem 806 and marker removal subsystem 810 may be embodied in a separate processor (such as a separate ASIC). In one embodiment, the image sensor 804 and the digital data processor 808 may be implemented in a single ASIC, which may optionally include on-board memory 816.
[0043] The digital data processor 808 can communicate with a memory 816 for storing software programs and / or temporarily storing image data and / or marker tracking data. In some embodiments, the encoder 812 can be included in or implemented as part of a codec. The encoder 812 can be implemented in, for example, a stand-alone hardware encoder or hardware codec that communicates with the digital data processor 808, or it can be implemented in a software program running on the digital data processor 808. A data communication device 814 (such as a wireless data transceiver or an Ethernet transceiver) communicates with the encoder 812.
[0044] Software instructions for implementing method 600 and other methods disclosed herein, or for implementing marker tracking subsystem 806, optional marker removal subsystem 810, and optional encoder 812, may be stored in a non-transitory computer-readable medium (such as memory 206 or memory 816).
[0045] It will be apparent to those skilled in the art that many changes can be made to the details of the above embodiments without departing from the basic principles of the invention. Therefore, the scope of the invention should be determined only by the appended claims.
Claims
1. A motion capture system comprising at least one motion capture camera, each motion capture camera comprising: An image sensor that operates at a frame rate of 10 to 1000 frames per second generates a series of frames representing digital image data of the scene visible to the motion capture camera. A marker tracking subsystem is configured to access the digital image data generated by the image sensor and process at least a portion of the digital image data to determine the current position of each of a plurality of reflective or luminescent markers attached to a moving subject in the scene for each of at least some of the frames in the series of frames. The marker tracking subsystem generates a series of marker datasets, each marker dataset corresponding to one frame in the series of frames being tracked and including the current position of the marker in the scene. An encoder configured to access the digital image data and encode at least some of the frames in the series of frames into compressed video data, the encoded frames including at least some of the frames in the series of tracked frames processed by the marker tracking subsystem; as well as A communication device configured to transmit the compressed video data and the dataset of markers.
2. The motion capture system of claim 1, wherein the marker tracking subsystem generates the series of marker datasets at or near the frame rate of the image sensor.
3. The motion capture system of claim 1, wherein the communication device transmits the series of marker datasets at the frame rate.
4. The motion capture system of claim 1, wherein the marker tracking subsystem and the encoder are implemented in a digital data processor that communicates with the image sensor and the communication device.
5. The motion capture system according to claim 4, wherein the digital data processor comprises a field-programmable gate array and / or an application-specific integrated circuit.
6. The motion capture system according to any one of the preceding claims further includes a marker removal subsystem configured to modify each frame of the digital image data to remove the markers in the scene, and then the encoder encodes the series of frames of the thus modified digital image data.
7. The motion capture system of claim 6, wherein both the marker tracking subsystem and the marker removal subsystem process a subset of the digital image data containing the region of interest.
8. The motion capture system according to any one of claims 1 to 5, wherein the motion capture camera further includes an illumination source.
9. The motion capture system according to claim 8, wherein the illumination source comprises an infrared illumination device.
10. The motion capture system according to any one of claims 1 to 5, further comprising a group of motion capture cameras arranged around the capture space for capturing different aspects of the scene, the group of motion capture cameras being interconnected via a local area network and synchronizing and calibrating together.
11. The motion capture system of claim 10, further comprising a host computer system communicating with the motion capture cameras via the local area network, the host computer system being configured to receive the compressed video data and the corresponding set of marker datasets from each of the motion capture cameras, and to store such compressed video data and the set of marker datasets from the motion capture cameras to maintain a synchronization or correspondence between each frame of the compressed video data and its corresponding set of marker datasets.
12. The motion capture system according to any one of claims 1 to 5, wherein the series of frames of the digital image data being tracked consists substantially of one of the following: a series of adjacent frames, a series of non-adjacent frames, or a series of adjacent and non-adjacent frames.
13. The motion capture system according to any one of claims 1 to 5, wherein the series of frames being tracked comprises the entire series of frames.
14. A method for generating motion capture data and image data, the method comprising the following steps: A motion capture camera is provided, the motion capture camera including an image sensor operating at a frame rate of 10 to 1000 frames per second, the motion capture camera being configured to perform the following steps: The image sensor generates a series of frames representing digital image data of the scene visible to the motion capture camera; Processing at least a portion of the digital image data to determine the current position of each of a plurality of reflective or luminous markers attached to a moving object in the scene for each of at least some of the frames in the series of frames, the processing including generating a series of marker datasets, each marker dataset corresponding to one frame in the series of frames being tracked and including the current position of the marker in the scene; At least some of the frames in the series of digital image data are encoded to generate compressed video data, wherein the encoded frames include at least some of the frames in the series of tracks. as well as The compressed video data and the dataset of markers are transmitted from the motion capture camera.
15. The method of claim 14, further comprising storing the compressed video data together with the corresponding set of marker datasets.
16. The method of claim 14, wherein the marker dataset is generated at or near the frame rate of the image sensor.
17. The method of any one of claims 14 to 16, wherein the step of transmitting the compressed video data and the corresponding series of marker datasets comprises transmitting the series of marker datasets at the frame rate.
18. The method according to any one of claims 14 to 16, further comprising: Prior to the step of encoding the series of frames of digital image data, for each frame of the digital image data, the digital image data is modified to remove the markers in the scene, thereby generating a frame of modified digital image data. and The step of encoding the series of frames of digital image data includes encoding the frames of the modified digital image data.
19. The method of claim 18, wherein the step of processing at least a portion of the digital image data to determine the current position of each of the markers includes identifying and processing a region of interest in the digital image data, and wherein the step of modifying the digital image data to remove the markers is performed on the region of interest.
20. The method of claim 18, wherein the following steps are performed by a digital data processor of the motion capture camera: (a) processing the digital image data to generate the dataset of the series of markers, (b) modifying the digital image data to remove the markers, and (c) encoding the modified digital image data.
21. The method according to any one of claims 14 to 16, further comprising receiving, at a host computer system, the compressed video data and the corresponding set of marker datasets from each of the motion capture cameras, and storing such compressed video data and the set of marker datasets from the motion capture cameras to preserve the synchronization or correspondence between each frame of the compressed video data and its corresponding set of marker datasets.
22. The method according to claim 21 further includes interconnecting the group of motion capture cameras and the host computer system via a local area network, and jointly synchronizing and calibrating the group of motion capture cameras.
23. The method according to any one of claims 14 to 16, wherein the series of frames of the tracked digital image data consists substantially of one of the following: a series of adjacent frames, a series of non-adjacent frames, or a series of adjacent and non-adjacent frames.
24. The method according to any one of claims 14 to 16, wherein the series of frames being tracked comprises the entire series of frames.
25. A non-transitory computer-readable medium for storing a software program for implementing the method according to claim 14.
26. A method for training a machine learning system to perform markerless motion capture using compressed video data generated by the method according to claim 14 and a corresponding set of marker datasets.
Citation Information
Patent Citations
Automated collective camera calibration for motion capture
US9019349B2