A real-time system for generating 4D spatiotemporal models of real-world environments

The method uses neural networks and holonomic constraints to derive 2D and 3D skeletons from image data, addressing accuracy and efficiency challenges in real-time 4D model generation for dynamic environments.

JP7754511B2Active Publication Date: 2025-10-15MOVE AI LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2022554966
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-03-13
Filing Date
2020-11-18
Publication Date
2025-10-15
Estimated Expiration
2040-11-18

AI Technical Summary

Technical Problem

Existing methods for deriving 3D and 4D models of real-world environments face challenges in accuracy and computational efficiency, particularly in real-time applications, due to the complexity of processing large amounts of image data and handling deformable objects.

Method used

A method involving neural networks to derive 2D and 3D skeletons from image data, applying holonomic constraints and recursive estimators to reduce computational complexity, and integrating these with 3D environment models to generate accurate 4D spatiotemporal models.

Benefits of technology

Enables highly accurate, real-time generation of 4D spatiotemporal models of environments and objects within, reducing processing power requirements and handling deformable objects effectively.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007754511000002
    Figure 0007754511000002
  • Figure 0007754511000003
    Figure 0007754511000003
  • Figure 0007754511000004
    Figure 0007754511000004
Patent Text Reader

Abstract

The present invention relates to a method for deriving 3D data from image data, the method comprising: receiving image data representing an environment from at least one camera; detecting at least one object in the environment from the image data; classifying the at least one detected object; and, for each classified object of the at least one classified object, implementing a neural network to identify features of the classified object in the image data that correspond to the classified object, determining a 2D skeleton of the classified object by identifying the classified object's features in the image data that correspond to the classified object; and constructing a 3D skeleton of the classified object, the method comprising mapping the determined 2D skeleton to a 3D skeleton.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a real-time system configured to derive 4D spatiotemporal data describing a real-world environment and generate an accurate 4D spatiotemporal model from the data. Specifically, the system is configured to process image data to detect, classify, and identify objects in the real-world environment, derive three-dimensional descriptor information that accurately describes the objects, and use the descriptor information to generate a 4D model that accurately depicts the environment. In this manner, the system is configured to model the objects by extracting the semantics of the objects represented by the descriptor data and reconstructing the model using the semantic description. [Background technology]

[0002] Three-dimensional digital representations of environments are employed in a wide variety of fields. One use is the visualization of virtual environments, for example in video games to create an immersive experience for the player, or in design work to help designers understand an environment before building it. Another, and particularly advantageous, use is for 3D models to visualize and represent existing real-world environments.

[0003] Observing a real-world environment through direct observation (either in person or via video recording) presents a limited perspective of the environment. The observer is limited to their own perspective or the perspective of the camera selected for observation. If one wants to observe the same real-world environment from a different perspective (e.g., location / angle / viewpoint), the observer must physically reposition themselves or set up an alternative camera to provide an alternative / additional feed. However, such repositioning has drawbacks. For example, setting up multiple cameras and the associated hardware and structured cabling is time-consuming and costly. There is an upper limit to the number of cameras that can realistically be present in an environment, there are some points in the environment where cameras cannot be placed (e.g., the middle of a soccer pitch during a match), and setting up new cameras may not be useful for providing a new observational perspective on an event that has already occurred. Furthermore, direct observation is limited to the information that can realistically be derived about the environment.

[0004] For example, sports matches (such as soccer, tennis, or rugby) are intensively analyzed by observers (humans and cameras) for compliance with the rules of the sport. While the rules of the game vary from sport to sport, a common component is that compliance with a rule violation depends on the determination of the location of the user or sports equipment within the sport's environment. One example is the offside rule in soccer. Under this rule, a player is "offside" if any part of the player's body, excluding the hands and arms, is within the opponent's half of the pitch and is closer to the opponent's goal line than both the ball and the opponent's penultimate player. While the exact definition of the offside rule may evolve or change over time, compliance is always determined by taking into account the relative placement of objects on the pitch (including players). Therefore, it is important to accurately observe the locations of multiple players and their body parts in relation to the pitch.

[0005] Referees and other individuals are employed to conduct such observations, but they cannot observe all aspects of play at once. Existing technologies have been developed to improve the observation of sporting environments and their components, but these have also attracted criticism regarding the accuracy of their observations. For example, in soccer, "video assistant referees" (or VARs) are often employed. VARs are additional observers who view video footage collected from multiple angles to determine whether rules are being followed (e.g., viewing players from different angles to confirm compliance with the offside rule). However, judging pitch conditions even from additional perspectives can be difficult, and accuracy can be limited if individual components cannot be accurately located.

[0006] A solution to the limitations of direct observation is instead to construct a 3D representation of the environment in question, which can then be manipulated and observed to determine physical environmental properties. It is therefore desirable to implement techniques that construct highly accurate 3D versions of real-world environments, which can then be manipulated to observe the pitch from multiple angles not present in the original data. The derived 3D models can also be further analyzed and manipulated to infer additional information not obtainable through traditional observation. Thus, previously unseen (but nevertheless accurate) viewpoints can be seen and inferences made about the data without the need to employ separate observers or additional cameras.

[0007] One method for 3D modeling known in the art is to implement a motion capture system in which multiple trackers can be positioned relative to the observed object, allowing an imaging system to track specific surface points on the object in three dimensions to implement 3D tracking. However, such systems are limited to monitored objects that have been specially prepared with the necessary trackers for the purpose. This may not be practical for the environment being modeled (e.g., a sports match). Therefore, it is preferable to be able to perform 3D modeling on data that is already available (e.g., camera / video data of the environment).

[0008] Existing techniques that enable the derivation of 3D information from camera / video data involve collecting image data from multiple viewpoints and rendering a 3D version of the environment by calculating the positions of objects in that environment, pixel by pixel. However, image data contains millions of pixels, and comparing pixel-by-pixel positions requires finding correspondences between trillions of combinations. This not only represents a significant computational challenge, but also requires processing large amounts of data to provide any degree of accuracy in any derived 3D model. Therefore, using such techniques to generate sufficiently accurate models is not practical or feasible for real-time execution, and an accurate 3D representation of the environment can only be realistically rendered after the completion of any event the system is attempting to model (e.g., a sports match).

[0009] Furthermore, real-world environments are not static quantities. Many objects in an environment may move relative to the environment and may themselves be deformable objects (e.g., human limb movement or the bending of non-rigid objects). Existing techniques for 3D data extraction may operate by separating and extracting 3D information from individual frames of image data, then merging them together to form a time-dependent sequence (i.e., 4D image data). This can result in frame-to-frame anomalies (e.g., motion "stutter"), which translates to inaccuracies in the final 4D image data. Furthermore, deformable objects across multiple frames may be perceived differently, creating inconsistent and inaccurate development of the fitted 4D data.

[0010] It is therefore desirable to provide a method for describing a 3D environment that processes image data to derive 4D spatiotemporal information that accurately describes the 3D environment in space and time. The present invention seeks to address the aforementioned disadvantages of the prior art by providing a system that derives accurate real-time 4D descriptor data and uses that data to generate a 4D spatiotemporal model that accurately represents the environment. Summary of the Invention

[0011] According to a first aspect of the present invention, there is provided a method for deriving 3D data from image data, the method comprising: receiving image data representing an environment from at least one camera; detecting at least one object in the environment from the image data; and classifying the at least one detected object, the method comprising: for each classified object, implementing a neural network to identify features of the classified object in the image data corresponding to the classified object, thereby determining a 2D skeleton of the classified object; and constructing a 3D skeleton of the classified object, the 3D skeleton comprising mapping the determined 2D skeleton to a 3D skeleton. This method thus provides a means for generating 3D data descriptor data through the intermediate step of generating 2D data. Computing the 2D descriptor data in the form of a 2D skeleton and the 3D descriptor data in the form of a 3D skeleton reduces the number of elements describing the object by several orders of magnitude from the millions of pixels that describe the object in the image data. This method is therefore significantly more computationally efficient and faster, allowing it to be performed in real time.

[0012] Advantageously, each of the at least one object comprises a plurality of associated sub-objects, and detecting at least one object for each of the at least one object comprises detecting each of the plurality of associated sub-objects, classifying the at least one object comprises classifying each of the plurality of associated sub-objects, and determining a 2D skeleton of the classified object comprises identifying features of each of the plurality of classified sub-objects in the image data corresponding to the classified sub-objects. In this way, the object from which 2D descriptor information can be derived is segmented, and a 2D skeleton is constructed by identifying the positions of the sub-objects relative to each other. This simplifies the task of deriving semantic information for each sub-object, and in particular reduces the complexity of the neural network for performing the determination of the 2D descriptor information.

[0013] Advantageously, for each classified object, mapping the determined 2D skeleton to a 3D representation can involve implementing a neural network and / or applying statistical and / or probabilistic methods to the determined 2D skeleton and applying holonomic constraints to the statistical and / or probabilistic methods. The 3D skeleton describing the object can be composed of many individual points, with a correspondingly large parameter space. By applying holonomic constraints to the degrees of freedom available when fitting these multiple points, the parameter space can be significantly reduced without sacrificing the accuracy of the final model. Therefore, the processing speed of the estimation process can be significantly reduced.

[0014] Further advantageously, classifying the at least one detected object includes classifying a first detected object of the at least one detected object as a human object, and the holonomic constraints include human anatomical holonomic constraints when statistical and / or probabilistic methods are applied to the determined 2D skeleton. Human objects are understood to be deformable in that their shape and configuration can change regardless of the environment in which they are placed. Such a high degree of freedom represents a challenging computational problem that is solved by appropriate application of trained neural networks and / or reduction of the problem dimension through application of specific holonomic constraints.

[0015] Advantageously, the 3D skeleton comprises an anchor point and a plurality of child points, each child point being defined by applying a transformation to the anchor point or another child point of the plurality of child points, and holonomic constraints defining the degrees of freedom of the transformation that define each child point and / or the range of possible values ​​of each degree of freedom of the transformation. Defining each point relative to the previous point allows the skeleton to be segmented into subgroups or networks, which can be adapted independently and to which specific holonomic constraints can be applied. This reduces the dimensionality of the problem, thereby speeding up processing and reducing the load on the processor.

[0016] Advantageously, classifying the at least one detected object includes classifying a second detected object of the at least one detected object as a non-human object. The present invention demonstrates great flexibility in being able to accurately model a wide variety of objects.

[0017] Advantageously, the at least one camera includes multiple cameras, detecting the at least one object includes detecting the at least one object in image data collected from each camera, classifying the at least one object includes classifying the at least one object detected in the image data from each camera, determining a 2D skeleton for each classified object includes determining multiple 2D skeletons, each of the multiple 2D skeletons being determined from identifying features of the classified object in image data collected from a different camera among the multiple cameras, and constructing a 3D skeleton includes combining the multiple determined 2D skeletons. Providing simultaneous feeds of the same object from different cameras improves the accuracy of the final constructed 3D skeleton. Furthermore, the intermediate steps of calculating 2D descriptor data in the form of a 2D skeleton and 3D descriptor data in the form of a 3D skeleton reduce the number of elements describing the object to those modeled from millions of pixels describing the object in the image data. Therefore, the complexity of finding correspondences between images from different cameras is significantly reduced compared to performing pixel-by-pixel comparisons of millions of pixels (which may require performing trillions of comparisons).

[0018] Advantageously, the image data comprises a time-ordered sequence of frames, the time-ordered sequence of frames comprising frames collected from each of the at least one camera, said detecting of at least one object in the environment comprises detecting said at least one object in each frame of the time-ordered sequence of frames, said classifying of the at least one detected object comprises classifying the at least one detected object in each frame of the sequence of frames, and each classified object of the classified at least one object is tracked throughout the time-ordered sequence of frames. Thus, by taking into account time-dependent input data, the method is able to provide a spatiotemporal output model.

[0019] Further advantageously, each classified object is tracked throughout the sequence of time-ordered frames by implementing a recursive estimator, and for each classified object, mapping the determined 2D skeleton to 3D includes determining a plurality of 3D skeletons to form a time-varying 3D skeleton, and applying the recursive estimator throughout the sequence of time-ordered frames to determine the time-varying 3D skeleton of the classified object. Thus, the method can refine the accuracy of the determined 3D descriptor information based on historical determinations.

[0020] Further advantageously, applying the recursive estimator includes applying time-dependent holonomic constraints. As described above, the 3D skeleton of any one object can include many 3D points and a corresponding large parameter space. This parameter space further increases as the object moves over time because the time dependence of each 3D point of the 3D object is determined to determine the time evolution of the 3D skeleton. By applying time-dependent holonomic constraints, the parameter space is reduced without sacrificing the accuracy of the derived 3D information. This can significantly slow down the processing speed of the estimation process. Optionally, the 3D skeleton includes an anchor point and multiple child points, each child point defined by applying a transformation to the anchor point or another child point among the multiple child points. The time-varying 3D skeleton is defined by a time-variable transformation of each child point, and the time-dependent holonomic constraints define, for each time point, the degrees of freedom of the transformation defining each child point and / or the range of possible values ​​for each degree of freedom of the transformation. Defining each point relative to the previous point allows the skeleton to be segmented into subgroups or networks that can be adapted independently and subject to certain time-dependent holonomic constraints, which reduces the dimensionality of the problem and can therefore speed up processing and reduce processor load.

[0021] Further advantageously, the at least one camera includes at least one first type camera and at least one second type camera, where each of the first type cameras captures image data at a first frame rate and each second type camera captures image data at a second frame rate different from the first frame rate. The additional image data from the additional cameras may improve the accuracy of the adapted 4D descriptor data and further provide additional time points for measuring and determining the 4D descriptor data, providing smoother and more accurate determination of the 4D skeleton.

[0022] Advantageously, the method further comprises constructing a 3D model of the environment from the image data. The same data used to generate the 3D avatar is used to construct the 3D environment model, providing consistency across the output of the method. The use of an intermediate stage of an environment map avoids the need to perform a time-consuming and processor-intensive derivation of the 3D environment in real time.

[0023] Further advantageously, the 3D environment model is constructed by applying positional mapping techniques to the image data to estimate the position and orientation of each of the at least one camera and constructing an environment map, and by mapping the environment map to a predetermined 3D model of the environment to construct the 3D model. SLAM techniques can use the image data even when the positions of the cameras are unknown, such as when the cameras are moving through the environment. Thus, this method can incorporate a wider variety of input image data sources.

[0024] More advantageously, the method further includes integrating the time-varying 3D skeleton with a 3D environment model to construct a time-varying integrated environment. Thus, an accurate 3D representation of the environment may be output without fully rendering the 3D model. Thus, the environment may be described using minimal data. In this method, the integrated model may be stored or transmitted to another computing device for later reconstruction of a fully rendered 3D environment model. This is advantageous because a fully rendered 3D model is described with a large amount of data, but is not necessary for a completely accurate representation of the scene. Thus, the integrated model may be stored and transmitted with minimal data requirements, and the 3D pre-generated model is mapped to the integrated environment only as needed to create the final model.

[0025] Advantageously, the method further comprises, for each classified object of the at least one classified object, constructing a 3D avatar of the classified object comprising integrating the constructed 3D skeleton with a 3D model corresponding to the classified object. The method thus provides a means for generating a 3D model of an object from the generated 3D descriptor data and for mapping pre-rendered objects to that descriptor data. Thus, generation of the 3D rendered model becomes significantly more computationally efficient and faster, and can be performed in real time.

[0026] Advantageously, the method further comprises capturing a 3D model corresponding to the classified object at a resolution higher than that of the image data. In conventional 3D image capture techniques, the resolution of the ultimately generated 3D model is limited by the resolution of the image data used to generate the model. In the present invention, low-resolution image data can be used to generate accurate descriptor information, to which a high-resolution model can then be mapped. Thus, the method is not limited by the resolution of the input image data of the environment being modeled.

[0027] Advantageously, the method further comprises: constructing a 3D avatar for each classified object of at least one classified object, the 3D avatar comprising integrating the constructed 3D skeleton with a 3D model corresponding to the classified object; and integrating the determined 3D environment model with the constructed 3D avatar for each classified object to construct a 3D model of the integrated environment, the integrated environment comprising the environment and the objects within the environment. Thus, the system provides a final output integrating all determined outputs to completely and accurately recreate the observed environment.

[0028] Advantageously, the method further includes refining the construction of the integrated environment by applying time-dependent smoothing to the constructed 3D avatars using a filtering method. Even more advantageously, the method further includes refining the construction of the integrated environment by constraining interactions between each constructed 3D avatar and the 3D environment model. Additionally, even more advantageously, constructing a 3D avatar for each classified object includes constructing multiple 3D avatars, and the method further includes refining the construction of the integrated environment by constraining interactions between the multiple 3D avatars. Such refinement of the output integrated environment reduces errors or inconsistencies that may result from the integration, thereby providing a more accurate output model.

[0029] In another aspect of the invention, there is provided a computer comprising a processor configured to perform any of the above methods.

[0030] In yet another aspect of the present invention, there is provided a computer program comprising instructions which, when executed by a processor, cause the processor to carry out any of the methods described above.

[0031] The present invention can provide highly accurate spatiotemporal representations of a wide variety of environments and the objects therein in real time. Environments that can be modeled include, but are not necessarily limited to, soccer pitches (and constituent players and elements), race tracks (the race tracks can be for human or animal racing, and the components can include humans and / or animals accordingly), Olympic sporting events (cycling, javelin throwing, archery, long jump, etc.), tennis matches, baseball games, etc. [Brief explanation of the drawings]

[0032] Embodiments of the present invention will now be described, by way of example, with reference to the accompanying drawings, in which:

[0033] [Figure 1]1 is a flow chart illustrating the general architecture of the method according to the invention; [Figure 2] 1 is a flowchart illustrating an exemplary method for constructing a 3D avatar in accordance with the present invention. [Figure 3A] 1 is a diagram of an exemplary object for constructing a 3D avatar as part of the present invention. [Figure 3B] FIG. 3B is a diagram of an exemplary 2D skeleton of the object of FIG. 3A. [Figure 3C] FIG. 3B is a diagram of another exemplary 2D skeleton of the object of FIG. 3A. [Figure 4A] FIG. 1 is a diagram of a human object for constructing a 3D avatar as part of the present invention. [Figure 4B] FIG. 3B is a diagram of an exemplary 2D skeleton of the human object of FIG. 3A. [Figure 4C] FIG. 3B is another exemplary 2D skeleton diagram of the human object of FIG. 3A. [Figure 5] 1 is a flowchart illustrating an exemplary method for constructing a 3D avatar based on image data including a time-ordered series, according to the present invention. [Figure 6] 1 is a diagram of an exemplary environment depicted in received image data. [Figure 7A] 1 is an exemplary system architecture suitable for implementing the present invention; [Figure 7B] 1 is an exemplary processor architecture suitable for implementing the present invention; DETAILED DESCRIPTION OF THE INVENTION

[0034] The present invention seeks to address the above-identified shortcomings of the prior art through an improved means of generating accurate 3D models of objects in an environment and the environment itself. The 3D models accurately represent objects in the environment as they move in space and, advantageously, over time. Specifically, instead of attempting to recreate the 3D environment pixel-by-pixel from image data, the image data is instead processed to derive underlying descriptor data that accurately describes the environment and objects within the environment. Pre-rendered objects can then be mapped to the descriptor data to reconstruct the 3D environment. This allows the system to significantly reduce the processing power required to construct the 3D environment without sacrificing accuracy.

[0035] A general architecture 100 of the present invention that sets out the steps by which a three-dimensional (3D) model or a four-dimensional (4D) spatiotemporal model can be derived from input image data is shown in FIG.

[0036] Input image data is provided at 102. The input image data may include, but is not limited to, a single image, multiple images, video, depth camera data, and laser camera data. The input image data is processed at 104 to derive 3D descriptor data that describes the environment and objects within the environment. 3D descriptor data in the context of the present invention is understood as data that describes the underlying features of an environment that describe that environment, and the underlying features of objects within the environment that describe those objects. In this sense, the 3D descriptor data is understood to describe the underlying "semantics" of the objects and / or environment. For example, in the context of a soccer match, the positioning of a person on the pitch can be described by identifying the locations of anatomical body parts of the human body (e.g., elbows, neck, heels, wrists) relative to each other and to the pitch, and the positioning of a ball in the environment can be described by the location of the ball's center of mass relative to the pitch. For example, the 3D descriptor data may include the 3D coordinates of body joints or the 3D coordinates of sub-objects such as the head and arms. The derived 3D descriptor data may also be processed and analyzed to perform measurements of elements of the modeled environment (e.g., ball or player speed, average distance between players or other items in the environment). As will be described in more detail below, the derivation of 3D descriptor data may be performed over multiple time points to create 3D descriptor data that evolves over time (i.e., 4D descriptor data).

[0037] After identifying the 3D descriptor data, the system then performs an integration step 106, in which the 3D descriptor data for an object in the image data is mapped to a corresponding 3D model of that object, resulting in a 3D model that accurately represents the orientation of the object as described by the descriptor data. For example, in the case of a 3D model of a human, the 3D coordinates of points in the model match the corresponding 3D coordinates derived from the image data. Thus, the resulting 3D model is a representation of the human object as shown in the image data. In the case of a modeled soccer, the center of mass of the 3D model of the soccer can be mapped to the center of mass derived from the image data so that the 3D soccer model is a representation of the soccer as shown in the image data. Integration step 106 is performed separately for multiple objects in the environment, which are then integrated with a 3D model of the environment itself to provide an integrated environment at output 108.

[0038] The final integrated environment is provided as a 3D model in a form that can then be flexibly used by a user of the graphics rendering software. For example, the model can be used to generate new images of the environment from viewpoints that are not present in the original image data. Building the 3D model with intermediate derivation of descriptor data allows for accurate modeling of the environment in real time. Further details of processing step 104 and integration step 106 are provided below.

[0039] FIG. 2 illustrates an exemplary image data processing methodology for use in processing step 104. The methodology is a method 200 for deriving 3D data from image data. The method 200 includes receiving 203 image data 204a representing an environment from at least one camera 202a. The camera 202a may be any vision system suitable for collecting image data, including, but not limited to, a 2D camera, a 2D video camera, a 3D depth camera, a 3D depth video camera, or a light detection and ranging (LIDAR). The camera 202a thereby generates appropriate image data 204a (e.g., 2D image, 2D video, 3D image, 3D video, LIDAR data). The method described below with respect to FIG. 2 is provided in the context of still images. 3D object modeling of video image data (i.e., image data including a sequence of time-ordered frames) is discussed below with respect to FIG. 5. The image data may be, for example, a live video feed from a television camera positioned within the environment of a soccer stadium.

[0040] Image data 204a represents a view of an environment. The view will include a number of objects that represent the environment, as represented by the image data. These objects will include static objects in the environment (e.g., goalposts, plots) and dynamic objects in the environment (e.g., players, balls).

[0041] The method 200 includes detecting at least one object in an environment from the image data 204a. Any suitable methodology can be used to detect objects in the environment from the image data 204a. For example, an autonomous object detection system can be used to detect objects in the image. The autonomous object detection system uses one or more neural networks to define the location and bounding box of the object. The input of the autonomous object detection system is the raw image data, and the output is the location and bounding box of the object. Any suitable neural network may be employed, including a multi-layer deep neural network (e.g., a convolutional neural network).

[0042] One or more neural networks are configured to recognize a wide variety of objects through training the neural network using image repositories such as ImageNet and Coco, which contain over 1.5 million stock images. For example, the stock images include many images of people of different ages, sizes, genders, and ethnicities, wearing a variety of clothing, presented in many different backgrounds, and standing in a variety of poses and positions. A similar variety of images of other items (e.g., sports equipment, seats) are provided for training the neural network. By training the system on this repository of images, the system can recognize, to a required level of confidence, which parts of the image data are objects and which parts of the image are not objects.

[0043] Objects detected in the environment may be of many different classes and types. Accordingly, method 200 further includes classifying at least one detected object to identify a classified object 205a. Any suitable methodology may be used to classify objects in the environment detected from the image data. For example, an autonomous object classifier may be used to classify the objects. A confidence value may be calculated for all possible classes from a predefined set of classes, and the object with the highest confidence value is selected as the class of the object. The autonomous object classification system may implement one or more neural networks. Any suitable neural network may be employed, including a multi-layer deep neural network (e.g., a convolutional neural network).

[0044] In a neural network for an autonomous object classification system, the input is the object to be classified (in the form of image data contained within the bounding box of the detected object), and the output is the object class to which the detected object belongs. The neural network is configured to classify a wide variety of objects through suitable training of the neural network. For example, the neural network is trained using input images tagged with appropriate classes. The training input images may again be sourced from visual databases (including open datasets such as ImageNet and Coco) where objects have been manually classified by human users prior to training the neural network.

[0045] Optionally, the method can include an object identification step (not shown), in which the classified object can be identified as belonging to a subclass or as being a particular instance of a class of objects (e.g., a particular player or a particular type of tennis racket). For example, the system can incorporate OCR technology to read any text that may be present in image data corresponding to the classified object (e.g., athlete names in an Olympic event or player numbers in a soccer event). Additionally or in combination, the method can implement a neural network for object identification. The neural network for object identification can be trained and applied in a manner substantially similar to those neural networks described above for object classification.

[0046] The above method steps relate to identifying objects within the environmental data. The following method 200 steps relate to deriving semantic information for each object in the form of deriving 3D descriptor information, with each subsequent method step being applicable to one or more of the classified objects.

[0047] As described above, the object classes classified and from which 3D descriptor data is derived may be of different classes and types. One such class of object is a "non-deformable" object. A non-deformable object is an object whose shape and pose do not change (e.g., a cricket bat or a goal post). Another such class of object is a "deformable" object. A deformable object is an object whose shape and pose are variable (e.g., a human's arms, legs, and head may change orientation). Modeling identified deformable objects is further complicated because they will have poses that change not only relative to the camera collecting the image data, but also relative to the pose the object is in at the time the image is collected. While a human is an example of a deformable object, it will be readily apparent that many other types of deformable objects exist, including other living creatures (e.g., dogs, horses, birds) and various inanimate objects (e.g., archery bows, javelin throwers, high jump poles, etc.). Given the above complexities, deriving semantic / descriptor information for deformable objects is technically more difficult than deriving semantic / descriptor information for non-deformable objects, and deriving semantic / descriptor information for some non-deformable objects is more difficult than other non-deformable objects.

[0048] For example, for a non-deformable, fully 3D symmetrical object (such as a rigid ball), the 3D semantic / descriptor information may be a single point in 3D space, whereas for a non-deformable object with partial or no symmetry, the 3D semantic / descriptor information may include two or more points in 3D space. For example, a javelin or rigid pole may be described by two points in 3D space. The more complex the non-deformable object, the more points in 3D space may be used to form the object's 3D semantic / descriptor information. However, it is understood that in the non-deformable case, the arrangement of 3D points in the descriptor information may change depending on the environment, but the relative positioning of each point in 3D space does not change. For any deformable object, given the variable shape and pose of the overall surface, the 3D descriptor information may include multiple 3D points with variable positions relative to the environment and each other.

[0049] Regardless of the object from which 3D semantic / descriptor information is derived and the level of complexity involved, the same challenge exists, namely, deriving three-dimensional descriptor data from input image data alone. This challenge is addressed by the steps described below.

[0050] Specifically, for each classified object of the at least one classified object, the method 200 further includes determining 205 a 2D skeleton 206 a of the classified object 205 a by implementing a neural network to identify features of the classified object 205 a in the image data 204 a corresponding to the classified object 205 a. In the context of the present invention, a “skeleton” of an object is understood to be a reduction of the object to a number of its representative elements, and a 2D skeleton is understood to mean those representative elements projected into a two-dimensional plane, thus representing the 2D semantic / descriptor data of the classified object. As noted above, the number of 3D elements required to fully describe each object may vary in number. Accordingly, the number of 2D elements in the 2D skeleton may also vary as needed to fully describe the dimensions and pose of the classified object.

[0051] The process of deriving a 2D skeleton is described below with reference to the examples shown in Figures 3A-3C and 4A-4C.

[0052] FIG. 3A illustrates a 2D image of an object 300 classified as a tennis racket from which semantic / descriptor data is derived. Such an image may be, for example, a portion of image data captured from camera 202a. FIG. 3B illustrates a 2D skeleton describing tennis racket object 300. As illustrated, the object class "tennis racket" can be described by a 2D arrangement of vertices 302a, 302b, 302c, etc., and lines 303a, 303b, etc., connecting the vertices. In this example, the combined arrangement of lines and vertices forms 2D skeleton 301, although a 2D skeleton can be described with reference to only the vertices or only the connecting lines. In such an example, each vertex is described by an X,Y coordinate in 2D space, and each line is described by a vector in 2D space. The arrangement of vertices and lines, as well as the length of the lines, will vary depending on the object's orientation in the image data, such that the 2D skeleton describes the object's pose in two dimensions.

[0053] An alternative example of 2D skeleton derivation is described below with reference to FIG. 3C. As shown in FIG. 3C, a tennis racket object 304 may be described by a combination of constituent sub-objects 304a, 304b, and 304c, each representing a component part of the tennis racket. In the illustrated example, the sub-objects include objects representing the handle 304a, throat 304b, and head 304c of the tennis racket 300. Depending on the relative placement and arrangement of these constituent sub-objects, alternative 2D skeletons may be described. For example, one may identify a "handle" object 304a at a first point in 2D space, a "throat" object 304b at a second point in 2D space, and a "head" object 304c at a third point in 2D space. The placement of sub-objects will vary depending on the orientation of the object in the image data, just as a 2D skeleton describes the pose of an object in two dimensions and therefore represents 2D descriptor information (which itself describes the underlying 2D "semantics" of the object - i.e., the placement of its constituent sub-objects relative to each other).

[0054] Optionally, each sub-object 304a, 304b, and 304c may also be described in 2D by multiple objects, forming a 2D sub-skeleton for each sub-object. These sub-skeleton may then be combined to form a larger 2D skeleton, such as the 2D skeleton shown in FIG. 3B. This may be done by aligning aspects of one sub-skeleton with another sub-skeleton based on known alignments of the sub-objects. For example, the method may include aligning the determined throat sub-skeleton with the determined handle sub-skeleton by mapping the 2D locations of vertices of the throat sub-skeleton to the 2D locations of other vertices of the handle sub-skeleton so that both vertices share the same 2D locations.

[0055] FIG. 4A illustrates a 2D image of an object 400 classified as a human from which semantic / descriptor data is derived. Such an image may be, for example, a portion of image data captured from camera 202a. FIG. 4B illustrates a 2D skeleton describing the human object 400. As illustrated, the object class "human" can be described by a 2D arrangement of vertices 402a, 402b, 402c, etc., and lines 403a, 403b, etc., connecting the vertices. In this example, the combined arrangement of lines and vertices forms 2D skeleton 401, although a 2D skeleton can be described with reference to only the vertices or only the connecting lines. In such an example, each vertex is described by an X,Y coordinate in 2D space, and each line is described by a vector in 2D space. The arrangement of vertices and lines, as well as the length of the lines, will vary depending on the object's orientation in the image data, such that the 2D skeleton describes the object's pose in two dimensions.

[0056] An alternative example of deriving a 2D skeleton for a human object is described below with reference to FIG. 4C. As shown in FIG. 4C, a human object 304 may be described by a combination of constituent sub-objects 404a, 404b, and 404c, each representing a component part of the human. In the illustrated example, the sub-objects include objects representing a head 404a, a neck 404b, and a torso 404c of a human 400. The relative placement and arrangement of these constituent sub-objects may describe alternative 2D skeletons. For example, for a human, a "head" object 404a may be identified at a first point in 2D space, a "neck" object 404b at a second point in 2D space, and a "torso" object 404c at a third point in 2D space. The placement of the sub-objects will vary depending on the object's orientation in the image data, such that a 2D skeleton describes the object's pose in two dimensions and, therefore, represents 2D descriptor information.

[0057] Optionally, each sub-object 404a, 404b, and 404c may also be described in 2D by multiple objects, forming a 2D sub-skeleton for each sub-object. These sub-skeleton may then be combined to form a larger 2D skeleton, such as the 2D skeleton shown in FIG. 4B. This may be done by aligning aspects of one sub-skeleton with another sub-skeleton based on known alignments of the sub-objects. For example, the method may include aligning the determined head sub-skeleton with the determined neck sub-skeleton by mapping the 2D locations of vertices of the head sub-skeleton to the 2D locations of other vertices of the neck sub-skeleton so that both vertices share the same 2D locations.

[0058] To identify sub-objects of an object, the method may include performing object detection and classification techniques on image data representing the object from which the sub-objects are to be identified. These techniques may be the same as those previously described in connection with FIG. 2 for detecting objects. For example, a suitably trained neural network may be employed to detect features of the object in the image data to detect the sub-objects, and a suitably trained neural network may be employed to classify each detected sub-object.

[0059] Alternatively, or in addition, detecting and classifying sub-objects of object 205a may form part of detecting and classifying object 205a as in method 200 described above, where each of the at least one object includes multiple associated sub-objects, and for each of the at least one object, detecting the at least one object includes detecting each of the multiple associated sub-objects, and classifying the at least one object includes classifying each of the multiple associated sub-objects. Thus, determining a 2D skeleton of the classified object includes identifying features of each of the multiple classified sub-objects in the image data corresponding to the classified sub-object. The additional step of detecting sub-objects may be performed by an additional neural network or as one or more additional layers in the neural network implemented to detect and classify object 205a. It will be appreciated that identifying and classifying all sub-objects of a larger object is not necessary in all instances of image data. For example, a complete 3D skeleton can be inferred from multiple smaller sub-objects. In the case of a tennis racket, the position of the head is known relative to the throat, so the larger object can be inferred by identifying and classifying only the handle and throat. Furthermore, if a sub-object is occluded in some instances of the image data, other cameras with different viewpoints may nevertheless capture the object in the image data captured by those cameras. In this manner, the method can perform successfully even when parts of the object are occluded (e.g., occluded by other objects or cut off at the edge of the field of view of the camera collecting the image data).

[0060] The above example 2D skeleton is provided for illustrative purposes, but it will be understood that any arrangement of 2D elements can be used to construct the 2D skeleton as long as those elements can be arranged to describe the size and pose of the object in two dimensions.

[0061] As described above, a neural network is used to determine the appropriate positioning of two-dimensional features of the 2D skeleton. The input of the neural network is image data 204a corresponding to the classified object 205a (e.g., image data contained within the bounding box of the classified and detected object), and the output is a 2D skeleton 206a describing the classified object 205a by appropriately positioning the 2D elements on the image data in the manner described above with reference to Figures 3A-3C and 4A-4C. Any suitable neural network, including a convolutional neural network, can be used for the determination. The neural network is trained using a training set of data containing multiple images in which the appropriate elements of the 2D skeleton have been manually identified. For example, the images in the training set could be images of a person with vertices and lines identified in the manner shown in Figure 4B, or images of a person with the relative positioning of sub-objects identified in the manner shown in Figure 4C. The training set includes many such images of various types of objects from many different viewpoints.

[0062] The method 200 further includes constructing a 3D skeleton 208 of the classified object, which includes mapping 207 the determined 2D skeleton 206a to 3D. To map the 2D skeleton to 3D, 2D elements of the 2D skeleton 206a can be mapped to corresponding points on a three-dimensional skeleton that includes 3D elements that correspond to the 2D elements in the 2D skeleton. For example, if the 2D skeleton is described by X,Y coordinates for the placement of the neck, shoulders, elbows, hips, knees, feet, and hands, the 3D skeleton is described by X,Y,Z coordinates for the placement of the neck, shoulders, elbows, hips, knees, feet, and hands, and the mapping includes mapping each point from 2D to 3D. Thus, the 3D skeleton is a coordinate representation of the underlying semantics of the object being modeled.

[0063] Advantageously, a point on the 3D skeleton may be defined by a transformation from another point on the 3D skeleton. For example, if two points are represented by respective values ​​in the coordinate systems (X1,Y1,...) and (X2,Y2,...), the second point is defined by (X2,Y2,...)=T 12 (X1,Y1,...), where T 12 is the appropriate transformation from the first point to the second point (e.g., defined by x, y, or z transformations in Cartesian coordinates, or by changes in length, azimuth, and tilt in polar coordinates). The third point (X3,Y3,...) is (X3,Y3,...)=T 13 Apply a transformation T to the first point so that (X1,Y1,...) 13 by applying (X3,Y3,...)=T 23 Apply a transformation T to the second point so that (X2,Y2,...) 23 Thus, the third point can be defined by two successive transformations on the first point, i.e., (X3,Y3,...)=T 23 T 12(X1, Y1, ...). Advantageously, the 3D skeleton thus includes an anchor point and a number of child points, each of which is defined by applying a transformation to the anchor point or to another of the number of child points. Thus, for each child point, the transformation can be a single transformation from the root point, or a series of successively applied transformations that are reapplied to the root point. Thus, the 3D skeleton can be represented by one or more "chains" or networks of points defined relative to the placement of the root point. Thus, by accurately placing the root point in 3D space, the remaining points of the 3D skeleton can also be placed in 3D space. The transformation between the two coordinates is realized by changing the values ​​that define each point in the coordinate system.

[0064] Additionally, each point of the 3D skeleton may be defined in a coordinate system having six values ​​to define the point in three-dimensional space: three translation values ​​(e.g., up / down, left / right, forward / backward, or similar lengths, azimuth, tilt) that define the absolute position of the point in 3D space, and three rotation values ​​(pitch, yaw, and roll) that define the orientation of the point in 3D space.

[0065] In such a coordinate system, the transformation from one point on the skeleton to a second point on the skeleton can be defined by a change in six values. For example, the second point can be defined by a translation in one, two, or three dimensions and / or a reorientation of the point around one, two, or three axes. Thus, upon transformation, each child point of the 3D skeleton can be described not only by a shift in its absolute position in three-dimensional space relative to another point, but also by a change in orientation relative to another point. Thus, each child point is defined by its associated transformation, and when a child point is defined by a transformation relative to another child point, the result is equivalent to a continuous transformation of position and orientation applied to the root point. The root point can be defined by specifying values ​​for each parameter relative to the environment or another object. The location of the root point can be specified by the user or determined from image data according to the methods described above.

[0066] Any suitable method can be used to map the 2D skeleton 206a to 3D, where mapping the determined 2D skeleton to 3D includes applying statistical and / or probabilistic methods to the determined 2D skeleton and applying holonomic constraints to the statistical methods, resulting in a model of 3D points. Applying statistical and / or probabilistic methods can include, for example, applying nonlinear estimators such as an extended Kalman filter, particle filter, unscented Kalman filter, or lattice partitioning filter to the determined points of the 2D skeleton to provide estimates for mapping to three dimensions to determine corresponding points on the 3D skeleton. These estimation techniques can be constrained by applying holonomic constraints to the model.

[0067] Holonomic constraints are understood to be known relationships between the spatial (and optionally temporal) coordinates of a system that constrain the possible values ​​of those spatial (and optionally temporal) coordinates within the system. For example, consider a rigid ball of radius a, the x, y, z coordinates of the ball's surface are given by the equation x 2 +y 2 +z 2 =a 2 The present invention can apply suitable holonomic constraints to each classified object to constrain the estimation method. It will be appreciated that in addition to the ball example above, other devices can have appropriate rules (e.g., points along the length of a rigid pole are constrained by equations describing an elongated rigid body, the head of a tennis racket is located in a specific location relative to the handle of the racket, and, as explained in more detail below, the anatomical limits of the human body can be described by equations for forming holonomic constraints). In each case, the holonomic constraints are employed as equations to reduce the parameter space of possible 3D skeletons that can be fitted to the points of the 2D skeleton.

[0068] Specifically, given an example where a 3D skeleton is described by a root point and multiple child points, holonomic constraints can define the limits within which values ​​can vary between points in the 3D skeleton and the amount by which they can vary. Specifically, holonomic constraints can constrain the transformations that define each child point by defining the number of values ​​that can change when performing the transformation (i.e., the degrees of freedom of the transformation) and the range within each degree of freedom for any given change. (For example, the yaw of a transformation may be fixed within a range of angles, or the absolute length between two points may be fixed.) Thus, holonomic constraints can be thought of as defining the degrees of freedom of the transformations that define each child point and / or the range of possible values ​​for each degree of freedom of the transformation. Equations and / or parameters describing holonomic constraints can be obtained from a database suitable for use in real-time determination of the 3D skeleton.

[0069] Alternatively, or in combination with the statistical and / or probabilistic methods described above, other suitable methods for mapping the 2D skeleton 206a to 3D may be used, including the use of neural networks. The input to the neural network is the 2D skeleton 206a of the classified object 205a, and the output is a 3D skeleton 208 that describes the classified object 205a through the appropriate arrangement of 3D elements. Any suitable neural network, including a convolutional neural network, may be used for the determination. The neural network is trained using a training set of data containing multiple examples of mapping from 2D elements to 3D elements. This data may be identified manually, or the 3D and 2D skeletons may be determined by simultaneous 2D and 3D observation of the object (the 3D observation may be implemented via appropriate motion capture technology). The training set includes many different 2D and 3D skeletons of various types of objects from many different perspectives.

[0070] As described above, humans can be classified as "deformable" objects. Detecting, classifying, and constructing 2D and 3D skeletons of deformable objects can be performed in the same manner as described above, i.e., by implementing one or more neural networks and through appropriate training of the neural networks. Specifically, the neural networks are trained not only with different types of objects at different viewpoints, but also with a wide variety of different poses that represent the extent to which each object can deform (e.g., for humans, many examples are used that demonstrate different positions within the full range of human body motion).

[0071] Although modeling deformable objects is more difficult than modeling non-deformable objects, humans represent a particularly difficult object to model. This is because human anatomy has a large number of degrees of freedom (over 200) due to the range of possible positions of each element of the human body (compared to simpler deformable objects, such as bows used in archery). Specifically, mapping from a 2D skeleton to a 3D skeleton can be difficult because 2D elements may be very close together at many angles, and there may be multiple possible mappings from 2D points to the 3D skeleton. Thus, there may be a significant number of degrees of freedom in the 2D to 3D mapping, increasing the complexity of the mapping.

[0072] Nevertheless, as described above with reference to examples, method 200 may be implemented such that classifying at least one detected object 205a includes classifying a first detected object of the at least one detected object as a human object. The above problem can be addressed by a suitable methodology for mapping a determined 2D human skeleton to 3D. Specifically, mapping the determined 2D skeleton to 3D includes applying statistical and / or probabilistic methods to the determined 2D skeleton and applying human anatomical holonomic constraints to the statistical and / or probabilistic methods to provide a model of 3D points. As described above, applying statistical and / or probabilistic methods may include, for example, applying a nonlinear estimator, such as an extended Kalman filter, a particle filter, an unscented Kalman filter, or a lattice partitioning filter, to the 2D skeleton to provide an estimate for mapping to three dimensions. These estimation techniques can be constrained by applying holonomic constraints to the model of 3D points.

[0073] As described above, holonomic constraints include equations that constrain the possible parameters when fitting a 2D skeleton to a 3D skeleton. Holonomic constraints for human anatomy can include constraints on the entire system of points (e.g., simultaneously specifying the placement of all points relative to one another based on known human spatial proportions), or holonomic constraints can be applied to specific pairs or groups of points on a 3D skeleton. For example, holonomic constraints for a given human anatomy can impose restrictions on the dimensions of 3D skeletal joints, such as the angles between points or the 3D locations of some points relative to other points (e.g., the knee joint can only bend in one direction, or the knee joint must be located between the foot and hip), or holonomic constraints can set upper or lower limits on the lengths between points (e.g., eliminating anatomically impossible distances between the ankle and knee). Thus, applying holonomic constraints objectively reduces the dimensionality of the fitting problem, improving the speed and accuracy of fitting. Although the human body has over 200 degrees of freedom, this method can be implemented with fewer degrees of freedom (e.g., 58) to accurately estimate the pose and position of the human body. For example, the feet have over 20 degrees of freedom (e.g., toe movement), and the spine has over 60 degrees of freedom. The degrees of freedom of the feet may not be considered for purposes of the present invention because they are blocked by shoes. In the case of the spine, the precision of all the spine's degrees of freedom can be reduced to a much lower number (e.g., 3), because accurate placement of human objects relative to the environment, or accurate placement of sub-objects relative to the spine, does not require precision of all the spine's degrees of freedom.

[0074] As an example application of holonomic constraints, consider again the example where a 3D skeleton's points include an anchor point and multiple child points. As further specified above, holonomic constraints define, for each child point, the number of degrees of freedom and / or the range of values ​​for each degree of freedom. Thus, in the context of a human skeleton, holonomic constraints can constrain the 3D skeleton within human anatomical parameters. For example, holonomic constraints can constrain the conformity of 3D points to account for the limited range of motion of the joints in the human body and to account for the fixed lengths of the body's bones.

[0075] For example, consider the definition of the elbow and wrist joints of a human arm. The wrist joint is separated from the elbow joint by a defined distance in the forearm (e.g., referencing the length of the radius or ulna). However, the wrist joint can also rotate relative to the elbow joint around the wrist joint's axis (pronation and supination, leading to hand rotation), anterior-posterior (flexion and extension, leading to a hand flapping back and forth), and lateral (radial and ulnar deviation, leading to a hand waving motion). Each of these rotations is limited to a range of possible values ​​(e.g., radial deviation can be in the range of 25–30 degrees, and ulnar deviation can be in the range of 30–40 degrees). Thus, to define the wrist joint, the equations and parameters defining the transformations from the elbow joint are constrained to allow only translational and rotational transformation values ​​within a set range. Other joints can be reduced in dimension in various ways, such as the elbow joint, which does not rotate relative to the shoulder and therefore does not require consideration of the relative orientation of the elbow joint with respect to the shoulder.

[0076] Applying such holonomic constraints significantly reduces the number of degrees of freedom that need to be considered when fitting a 3D skeleton from points in a 2D skeleton by taking into account the objective, real-world constraints of the body that the 3D skeleton describes. Furthermore, as identified above, the complexity of the problem is further reduced by defining each point in the 3D skeleton by applying a transformation to another point in the skeleton. Specifically, the fitting of each 3D point is fitted and constrained only in relation to the 3D points to which it is connected within the 3D skeleton, rather than the more complex problem of attempting to simultaneously fit all parameters in the skeleton or applying holonomic constraints to 3D skeleton points that are not directly connected. Furthermore, as noted above, the application of transformations and consideration of holonomic constraints within the application of transformations allows sections of the 3D skeleton to be separated into subsets of points in a chain or network, each emanating from a root point. Each of these sections of the 3D skeleton can be fitted and constrained independently of the other sections of the skeleton without loss of accuracy.

[0077] While holonomic constraints for statistical and / or probabilistic methods are discussed above in the context of human objects, it will be appreciated that similar methods can be applied to mapping the 2D skeleton to a 3D skeleton of other deformable objects through the use of appropriate holonomic constraints. For example, in the case of other animals, the holonomic constraints used can be the anatomical holonomic constraints of the animal (e.g., for dogs or horses, etc.). As another example, suitable structural holonomic constraints can be applied to deformable inanimate objects. For example, an archery bow will be significantly deformable in one plane (to model the tension on the string). As explained above with respect to human objects, holonomic constraints can impose limits on the dimensionality of aspects of a 3D skeleton, such as the angles between points of the 3D skeleton or the 3D locations of some points relative to other points in the skeleton (e.g., in an archery bow, the tip of the bow can elastically compress inward but not outward, or the grip of the bow must be located between the feet and hips), or holonomic constraints can set upper or lower limits on the lengths between points (e.g., precluding physically impossible distances between the grip and tip of the bow). Thus, the application of holonomic constraints objectively reduces the dimensionality of the fitting problem, improving the speed and accuracy of fitting.

[0078] While the holonomic constraints discussed above are provided to define limits on points in space, as noted earlier, holonomic constraints can equally specify time-dependent equations that constrain how elements of an object during known kinematic behavior (e.g., running, walking) vary over time. Their application to kinematic models will be discussed in more detail below.

[0079] As an alternative to using statistical and / or probabilistic methods, in the case of classified human objects, mapping the determined 2D skeleton to 3D involves the implementation of a neural network. As noted above, a suitable neural network (e.g., a convolutional neural network) can be trained on a sufficiently large data set that includes many examples of humans in a wide variety of poses.

[0080] As a result of the above method, the system can generate 3D descriptor information from input image data that accurately describes objects detected in the environment without the need for additional specialized equipment or computationally intensive pixel-by-pixel processing. The system may then utilize this 3D descriptor information to derive a wide variety of information about the objects detected in the image data. For example, the 3D descriptor information of an object can be used to derive detailed placement of the object with respect to other objects in the environment and with respect to the 3D environment itself. For example, in the context of a soccer match, the 3D descriptor information of multiple players and soccer players can be compared to the 3D environment to accurately determine whether certain placement criteria are being adhered to (e.g., compliance with the offside rule). While placement criteria may be determined by the rules of the match, accurately determining whether real-world objects meet such placement criteria remains a significant technical challenge that is solved by the present invention.

[0081] Additionally, the 3D descriptor information may be used to determine other characteristics and / or properties of objects detected in the environment. These characteristics may also be determined from the 4D descriptor data using temporal information from the time-ordered data (as described in more detail below). For example, the velocity and / or acceleration of any one detected object may be accurately determined (e.g., for comparison with existing records of objects in the same context). As another example, the movement of a deformable object may be determined (e.g., maximum joint angles, limb orientations when a player assumes a particular pose and posture). Furthermore, the proximity and interaction between objects may be accurately recorded (e.g., jump height, number of times a player contacts the ball, length of time a player possesses the ball). The characteristics and / or properties may be determined automatically in real time (and simultaneously with determining the 3D data) by any suitable pre-programmed equation, or may be determined manually by an operator after calculating the model. Accordingly, optionally, the method may further include, for each classified object of at least one classified object, deriving properties of the classified object using the constructed 3D skeleton of the classified object. The properties may include, for example, three-dimensional location information regarding the position of the classified object relative to the camera, or may include 3D location data of sub-objects or sub-objects of the classified object relative to other sub-objects of the same classified object.

[0082] Advantageously, the method also includes constructing a 3D avatar 210 of the classified object 205a, which includes integrating the constructed 3D skeleton 208 with a 3D model corresponding to the classified object. The 3D model is a representation of the object being modeled and may include high-resolution surface textures representing the object's surfaces. The 3D model may include 3D elements corresponding to the constructed 3D skeleton 208. Thus, a simple mapping can be implemented such that the 3D model maps directly onto the constructed 3D skeleton to create a 3D representation of the object 205a, including a representation of the object's size and pose information in the image. Because the 3D model is pre-rendered, it does not need to be generated "on the fly." As described above, the only real-time component data generated are 2D and 3D descriptor data. Due to the large amount of data used to describe the 3D model, deriving such intermediate descriptor data is computationally significantly faster and easier than direct pixel-by-pixel reconstruction of the 3D scene. Thus, the above-described method 200 achieves the goal of deriving an accurate 3D representation of an object from image data in a simpler and more computationally efficient manner.

[0083] Furthermore, this method can generate high-resolution 3D images from low-resolution input image data (which is not possible using existing technologies). That is, in the case of pixel-by-pixel analysis and reconstruction of a 3D scene, the image resolution of the final output model may be limited by the resolution of the input image data. For example, if the image data captured at a soccer match consists of low-resolution images or video from a television feed, the final output model of the scene may be limited accordingly (i.e., the 3D / 4D model cannot be created at a higher resolution than the input data). In contrast, the above-described method of the present invention may not build a final model directly from low-resolution television images. The low-resolution television images are used only to derive 3D descriptor data onto which a pre-generated high-resolution 3D model can be mapped. The high-resolution pre-generated model may be generated, for example, by 3D scanning of the object using photogrammetry techniques.

[0084] The pre-generated model may be a generic model maintained in a database, or advantageously, the method of the present invention itself may also include a process of pre-generating a unique 3D model of the object for mapping to the 3D skeleton. That is, the method described herein may further include, for each classified object of the at least one classified object, capturing a 3D model of the classified object using a photogrammetry system. Specifically, the method may further include capturing a 3D model corresponding to the classified object at a resolution higher than the resolution of the image data.

[0085] The method need not be limited to modeling one class of object, and may include classifying at least one detected object as a non-human object. After the object is classified, an appropriate neural network is identified for appropriate analysis of the classified object.

[0086] Method 200 was described above with reference to image data of an environment collected by one camera 202a. This method can be improved by collecting image data from multiple cameras. That is, additional image data 204b, 204c...204z of the same environment can be collected. Each instance of the additional image data is collected from a respective additional camera 202b, 202c...202z. As described above for camera 202a, each additional camera 202b, 202c...202z can be any vision system suitable for collecting image data, including, but not limited to, a 2D camera, a 2D video camera, a 3D depth camera, a 3D depth video camera, or a light detection and ranging (LIDAR). This causes camera 202a to generate appropriate image data 204a (e.g., 2D image, 2D video, 3D image, 3D video, LIDAR data). Each camera is positioned at a different location relative to the environment, and each instance of the image data accordingly provides an alternative representation of the environment (e.g., different position, different angle, different field of view, etc.) Thus, the objects detected, classified, and modeled may be depicted in multiple alternative instances of image data representing the scene.

[0087] Each step of the method 200 described above can be applied to each instance of further image data 204b, 204c...204z to identify a particular classified object 205b, 205c...205z in each instance of further image data 204b, 204c...204z. The same object is classified in each instance of image data such that the method includes providing multiple instances 205a...205z of the same object classified based on image data provided from different viewpoints. Thus, the method is improved in that detecting the at least one object includes detecting the at least one object in image data collected from each camera, and classifying the at least one object includes classifying the at least one object detected in image data from each camera.

[0088] Further, determining a 2D skeleton for each classified object includes determining multiple 2D skeletons, each determined from identifying features of the classified object in image data collected from a different one of the multiple cameras. Thus, by classifying the object within each instance of image data 204a, 204b...204z, the method can determine multiple versions 206a, 206b...206z of an object's 2D skeleton, each describing a different respective instance 205a, 205b...205z of the classified object, thereby enabling refinement of the descriptor data determined from the image data. Information describing the position and orientation of each camera 202a...202z is known to the method, thus enabling each instance of image data to be associated with information about its source. This information may be included in the received image data, or the information may be predetermined prior to commencement of the method.

[0089] The method further includes combining the determined 2D skeletons to construct a 3D skeleton, providing multiple views of the same 2D skeleton, but from different perspectives, to triangulate each element of the classified object in 3D space, thereby constraining the degrees of freedom available in mapping the 2D skeleton to the 3D skeleton. Correspondences between objects and / or sub-objects identified with different cameras are identified, as well as relationships between those cameras that are utilized to triangulate the 3D locations of the objects or sub-objects relative to the locations of each camera.

[0090] For example, for a particular point on the determined 2D skeleton 205a determined from an instance of image data 204a, a line can be drawn from the X, Y, and Z coordinates of the location of the camera 202a used to capture that instance of image data 204a through the determined location of the particular point (i.e., a vector from the camera location extending through the 2D location on the image plane). A similar process is performed for each camera 202b, 202c, ..., 202z, resulting in multiple 3D vectors. The point in 3D space with the smallest distance to all lines (e.g., Euclidean distance or Mahalanobis distance) is determined as the 3D coordinate of the particular point on the 2D skeleton. This process is repeated for all points on the 2D skeleton to construct multiple 3D points that make up the 3D skeleton. It will be appreciated that the above methods are merely exemplary, and any suitable statistical and / or probabilistic method can be applied to simultaneously map multiple corresponding 2D points collected from cameras in known positions to points in 3D space, such as mapping multiple 2D skeletons to the same 3D skeleton. As described in more detail below, this methodology can be refined through the collection of time-varying image data and the determination of 2D skeletons across time.

[0091] The above-identified method can be performed with significantly improved computational efficiency. Specifically, in pixel-by-pixel comparison approaches, finding correspondence between multiple cameras requires attempting to find correspondence between millions of pixels in various instances of image data (which requires trillions of comparisons). In contrast, the present method involves identifying correspondence between elements of a determined 2D skeleton (which could be anywhere from 1 element to 1000 elements, depending on the object). Thus, significantly fewer comparisons are required to determine the final elements of the 3D skeleton.

[0092] It will be appreciated that any suitable estimator model can be implemented to achieve triangulation of points from multiple 2D skeletons, for example, a non-linear statistical and / or probabilistic model can be implemented and / or a suitably trained neural network (e.g., trained to triangulate 3D points from multiple lines as described above) can be implemented.

[0093] Typically, triangulation requires the use of at least two cameras to accurately estimate a single point in 3D space, but the more cameras used, the more accurate the final determination. The above method is possible using a single camera without triangulation techniques. While triangulation cannot be performed from a single image, the system can still accurately identify and classify each object from a single instance of image data. Furthermore, each individual element can accurately identify each 2D element of a 2D skeleton because each individual element exists in combination with other elements of the skeleton, thereby constraining the degrees of freedom of any single point within the particular skeleton to which it is mapped. The accuracy of fitting a 2D skeleton to a single instance of image data can be supplemented by any number of auxiliary methods, including dense mapping techniques. These same auxiliary methods may also be used to improve the accuracy of 2D skeleton determination from multiple instances of image data from different cameras. The above method describes the construction of a 3D model representing an object depicted by still image data from a single camera or multiple cameras. Nevertheless, the above method can also be applied to video image data including a sequence of time-ordered frames by applying each step of the method to each individual frame of the sequence of time-ordered frames. Thus, the final derived descriptor data includes multiple time-ordered sets of 3D descriptor data that describe any spatially classified object in time. The 3D descriptor data can then be used to construct a 4D model (i.e., a temporal 3D model) of the object depicted in the image data. For example, the sequence of time-ordered frames could be a live video feed from a camera installed in a soccer stadium. The method steps described herein can be applied in real time to frames of the live video feed to enable real-time generation of a 4D model.

[0094] A refined method 500 is shown in Figure 5. In this figure, the steps of method 200 described above are applied to image data 502 collected from each camera of at least one camera, the image data 502 comprising a time-ordered sequence of frames 502a, 502b, 502c...502z. In the same manner as described above with respect to Figure 2, detection and classification of objects in the environment is applied to an individual frame (e.g., 502a or 502b) of the time-ordered sequence of frames 502a, 502b, 502c...502z. Thus, the detection of at least one object in the environment includes detecting the at least one object in each frame of the time-ordered sequence of frames 504a, 504b, 504c...504z, and the classification of the at least one detected object includes classifying the at least one detected object in each frame of the sequence of frames 504a, 504b, 504c...504z. As a result, the method provides a time-ordered sequence of classified objects 504a, 504b, 504c...504z. In the illustrated example, a single object is discussed, and each of the classified objects 504a, 504b, 504c...504z is the same classified object, although identified at different times. It will be appreciated that the method 500 can also identify multiple objects in a time-ordered sequence of frames, with each classified object of at least one classified object being tracked throughout the time-ordered sequence of frames. Because each object is classified by the method 500, the method can identify each object in the image data and therefore track the object individually across multiple frames.

[0095] Advantageously, each classified object is tracked throughout the sequence of time-ordered frames by implementing a recursive estimator. In this way, the classification of object 504a in a first frame 502a is used by a method to assist in the classification of object 504b in a second frame 502b. For example, in the case of a process for detecting object 504b, the input to the detection method includes both the individual frames of image data 502b and the output 505a of the detection and classification step for object 504a. Thus, the detection and classification of object 504b is refined by the output of the detection and classification of the object at a previous time instance. Similarly, the output 505b of the detection and classification of object 504b can be used to refine the detection and classification of object 504c. Any form of statistical and / or probabilistic method may be used as the recursive estimator, and data from one time point can be used as input to derive information at the next time point. One such recursive estimator may be a neural network.

[0096] Such tracking is known as "object" tracking. Object tracking can be achieved in any of a variety of ways, including motion model tracking (including Kalman filters, Kanade-Lucas-Tomashi (KLT), mean shift tracking, and optical flow calculations), and vision-based tracking (including the application of deep neural networks such as online and offline convolutional neural networks (CNNs), convolutional long short-term memory (ConvLSTMs), etc.).

[0097] Implementation of such a tracking method is advantageous in instances where multiple objects of the same class are present within a single video stream. Typical detection and classification methods may be able to accurately identify the presence of two objects of the same class and recognize the two objects as unique instances of the same class. However, when presented with a second image containing further examples of the same two objects, the method may be unable to identify which object is which. Method 500 can uniquely identify each classified object within each frame of image data by refining the object detection and classification of each frame within the context of the preceding frame.

[0098] However, in instances where a group of objects within the same class are in close proximity at a particular moment, after the objects separate after the particular moment, the objects may no longer be uniquely tracked. Thus, simply tracking the objects may not be sufficient to uniquely distinguish between various objects that frequently come into close proximity. Method 500 may be refined to address this possibility by applying the object identification steps described above to improve the accuracy of unique identification of individual objects even during and after instances of proximity. Alternatively, or in addition, image data from an additional camera located at a different position (and having a different viewpoint) may be incorporated into the method as an additional camera of the multiple cameras described above.

[0099] Further alternatively, or in addition, the method may include receiving input from an observer of real-time processing of the image data, who may manually tag each classified object with a subclass in a period immediately following the moment of proximity. Tracking of the object then becomes capable of uniquely identifying each object again at each time instance. Manual tagging may occur at the beginning of the method, and may uniquely identify each classified object upon initial identification of the classified object to ensure tracking from instantiation.

[0100] In the same manner as described above with respect to Figure 2, the step of determining a 2D skeleton is applied to each classified object (e.g., 504a or 504b) of the time-ordered sequence of classified objects 504a, 504b, 504c...504z to generate a time-ordered sequence of 2D skeletons 506a, 506b, 506c...506z. Similarly, and as described above with respect to Figure 2, the step of constructing a 3D skeleton is applied to each determined 2D skeleton (e.g., 504a or 504b) of the time-ordered sequence of classified objects 506a, 506b, 506c...506z to generate a time-ordered sequence of 3D skeletons 508a, 508b, 508c...508z.

[0101] Advantageously, for each classified object, mapping the determined 2D skeleton to 3D involves applying a recursive estimator across the sequence of time-ordered frames. In this way, a 3D skeleton 508a derived at a first instance is used as a further input 509a to the derivation of a second 3D skeleton 508b for the same object at a later time instance. Using such temporal information provides greater precision and accuracy for the complex problem of deriving 3D skeleton placement on a moving object. This not only refines the method by ruling out certain placements of 3D features (e.g., it may be physically impossible for a human body to execute the leg movements required to kick a ball in less than a second), but also helps map detected movements to well-recognized movements (e.g., kicking a ball has a well-observed movement pattern that can be matched to the time-dependent placement of 3D features). Any form of statistical and / or probabilistic method may be used as the recursive estimator, and data from one time point can be used as input to the derivation information at the next time point. An example of a recursive estimator is a neural network (such as a convolutional neural network). Such a neural network is trained using time-dependent 3D input data collected by observing an object undergoing a variety of different motions. For example, a human can be recoded using motion capture techniques (e.g., placing optical markers on an individual and observing that individual perform various actions).

[0102] As discussed above, holonomic constraints can be applied in statistical and / or probabilistic ways to constrain the coordinates of a system in space. Advantageously, recursive estimators can also implement time-dependent holonomic constraints for fitting 3D skeletons across time (i.e., in 4D). Time-dependent holonomic constraints are implemented by a kinematic model, which constrains the points of the 3D skeleton across a sequence of time-ordered frames. The kinematic model represents the 3D coordinates of a point or group of points of the 3D skeleton over time. Thus, when attempting to fit these points of the 3D skeleton at time t, the available values ​​of these points are constrained not only by the physical holonomic constraints (detailed above), but also by the determined values ​​of these points at the previous time t-1 and the kinematic model that defines the relationship between the points at time t and time t-1. An example of a kinematic model may be equations that describe the motion of various points on a human body over time as the human engages in various activities (e.g., running, walking, serving a ball with a tennis racket, etc.). Other kinematic models may include equations that describe the elastic deformation of a non-rigid body under the application of a force (e.g., a pole in high jumping or a bow in archery). Additionally, kinematic equations may describe the interaction of an object with an external environment (e.g., the movement of an object such as a ball through 3D space under the influence of gravity, or the movement of a ball bouncing or rolling on a plane).

[0103] One example of implementing time-dependent holonomic constraints on 3D skeletons is through the use of hidden Markov models, considering Eq. X t =f h (X t-1 ,u t ) (1) Y t =g c (X t ) (2) In the formula, X tis the vector of states of the elements of the 3D skeleton of object h (e.g., for a human object, X t are the joint vectors of the human skeleton in 3D space, and X t =[x leftwrist ,y leftwrist ,z leftwrist ,x head... ]).

[0104] X t represents the "hidden state" of the hidden Markov model and is not itself directly measured. Instead, X t is estimated from measurements of the object. In this case, Y t is the measurement vector of the observed object h. If we observe object h through a set of cameras c, we have

number

[0105] The goal of the predictive model is to predict the measured value Y t-1 Then, we refine it to obtain the measured value Y t From X t The goal is to find the best estimate of the measured Y t The value is the measurement function g c From the application of X tY is the measured value t-1 The value is the measurement function g c From the application of X t-1 value, and further determined value X t-1 is a function f h Using X t is used to derive an estimate of f h (X t-1 ,u t ) to find the most likely value X t The system then determines the measured Y t Value and measurement function g c From X t Next, determine the state X t The final value of X t-1 and Y t Based on the value determined from X t It is refined by optimizing the value of X, for example using a Kalman filter. t-1 X based on the value inferred from t and the value of Y t to minimize the residual error between the value determined from

[0106] Function f h describes a kinematic model of object h, which incorporates the spatial holonomic constraints described above, as well as how these points in space fluctuate over time for various types of kinematic motion, thus defining time-dependent holonomic constraints for the kinematic model. As described above, spatial holonomic constraints constrain the number and / or range of available degrees of freedom for each point in the 3D skeleton. Time-dependent holonomic constraints further constrain the degrees of freedom over time.

[0107] Consider an example where a 3D skeleton comprises an anchor point and multiple child points, each child point defined by applying a transformation to the anchor point or to another child point of the multiple child points, and the time-varying 3D skeleton is defined by a time-variant transformation for each child point. For example, the difference in value between two points of the 3D skeleton will vary over time, and these value differences are represented by a different transformation at each time point. Time-dependent holonomic constraints define, for each time point, the degrees of freedom of the transformations that define each child point and / or the range of possible values ​​for each degree of freedom of the transformation. The amount of time that can vary depends on the time interval between measurements, the selected kinematic model, and the u t (described below). Time-dependent holonomic constraints are advantageously applied in combination with space-holonomic constraints.

[0108] For example, f h The model includes holonomic constraints that constrain the relative orientations of the knee, hip, and ankle, as well as how the knee, hip, and ankle move relative to one another over time for certain types of kinematic motion (e.g., running). Holonomic constraints may limit the degrees of freedom available in defining the 3D points describing the leg to the point at which a human is running. For example, during running, the ankle joint is not expected to demonstrate significant supination or eversion, so this degree of freedom for the ankle may be removed entirely or may be limited to a small value characteristic of a running human.

[0109] u t describes the control inputs for a kinematic model (e.g., the rotational velocity of a joint such as the left elbow for human anatomical constraints, or for use in a kinematic model of an object moving under the influence of gravity, which takes approximately 9.807 ms -2The Earth's gravitational constant, g, is the gravitational constant of the Earth, or the coefficients of friction between the object and various types of materials. These values ​​can be measured or calculated prior to applying the method, such as through motion capture measurements of an individual or object, using methods known in the art. t In state X t-1 It is also possible to include variable inputs to the kinematic model, including values ​​of t-1 (e.g., leg velocity or trunk angular velocity measured at time t-1). Optionally, u t These variable values ​​of u can be measured via suitable physical sensors in the environment (e.g., LIDAR or 3D depth cameras) along with the camera observations Y. Additionally or alternatively, u t One or more values ​​of u can form part of the hidden state inferred from the measurements Y. t The value of is mainly X t is used to estimate the value of Y t-1 It only needs to be an approximate measurement, as it will be refined by the measurement of

[0110] Function f h is the kinematic model of a particular object h, and various f h Functions can be applied to describe the kinematic models of other types of objects (e.g., a single f for a human object). h There is an equation, and the ball has another f h equations, etc.), but all objects of the same class have the same kinematic model f h You can use the same f for all the balls you are modeling (for example, h Different kinematic models can be applied to different subclasses, or the same kinematic model can be applied but with different parameters (e.g., the same f h The equations are applied to different people, but with different parameters because different people have different body dimensions and human gait will vary when running.

[0111] Therefore, for any one object h, the measurement Y t , Y t-1 , ..., u t , u t-1 , ... and f h Knowing X t We can find the best estimate of Y in these cases. t Vectors describe only the measurement data of individual objects, and X t The system is further advantageous because it describes the 3D skeleton information of the object, meaning that the number of parameters that need to be simultaneously fitted by the system is significantly less than pixel-by-pixel comparisons, which may involve time-dependent kinematic modeling of millions of pixels (resulting in trillions of comparisons).The hidden Markov model described above is just an example of how time-dependent holonomic constraints can be implemented in a nonlinear estimator process to refine a time-dependent 3D skeleton determined from time-evolving image data.

[0112] The time-ordered sequence of frames described above can include frames captured by multiple cameras in the environment, each operating at the same frame rate and synchronized so that image data is available from two or more viewpoints at each point in time. This allows the fit of the 3D model to be constrained at each point in time by triangulation as described above, but also by using a recursive estimator as described above. This improves the overall accuracy of the determined 4D descriptor information.

[0113] Furthermore, by using a recursive estimator, the above method can be performed when the at least one camera includes at least one first type camera and at least one second type camera, where each of the first type cameras captures image data at a first frame rate and each of the second type cameras captures image data at a second frame rate different from the first frame rate. Any number of cameras may be operating at the first frame rate or the second frame rate, and there may be additional cameras operating at different frame rates. Commonly used frame rates are 24 frames per second (fps) and 60 fps, but data can be collected at higher rates of 120, 240, and 300 fps. Thus, a time-ordered sequence of frames including frames collected from multiple cameras may include a variable number of frames at each instant in time, including some instants in which only one frame is captured.

[0114] Using a recursive estimator like the one above, similar constraints on the 3D skeleton are possible even for 3D skeletons determined at a single frame in the combination of a time-ordered sequence of frames. For such frames, the 3D skeleton determined from the 2D image can be constrained and refined by a 3D skeleton adapted from an image at a previous time point. For example, in the Hidden Markov Model above, Y t The data contains a single frame, from which X t is determined, and the previous time state X t-1 X inferred from tSuch methods and systems are advantageous over methods that do not employ a recursive estimator and instead individually calculate 3D skeletons for each time point without knowledge of past history. Such systems that use image data collected from cameras that are not precisely synchronized may produce time-dependent models of varying accuracy due to inconsistent numbers of cameras used at each time point. Advantageously, the above methods using recursive estimators can be used to determine 4D descriptor data from multiple cameras with different frame rates, as well as cameras with the same frame rate but that are asynchronous.

[0115] Further advantageously, the use of a recursive estimator can enable the determination and tracking of a 3D skeleton from data collected by only a single camera, where each 3D skeleton determined from an individual frame in a time-ordered sequence of frames is refined from the 3D skeleton estimate from the previous frame by applying a recursive estimator. As long as the 3D skeleton is accurately determined relative to the initial state, it can still provide accuracy without triangulation from multiple cameras. Such applications may be advantageous, for example, in situations where a person or object may temporarily move from an area monitored by multiple cameras to an area monitored by only one camera. The object will be accurately tracked until it moves out of the area monitored by the multiple cameras, providing an initial accurate input for the 3D skeleton at the start of the single-camera tracking region. The 3D skeleton can continue to be determined and tracked over time by a single camera until the person or object returns to the area monitored by the multiple cameras. For example, in the context of a soccer match, a human player may be running on the pitch past a structure that obscures all but one camera, thus enabling continuous and accurate tracking in the event that systems known in the art may fail.

[0116] Thus, the above method allows for continued accuracy even during periods of low input image data, but also allows for a more flexible system because it can incorporate additional image data from cameras not specifically configured to synchronize with existing cameras. For example, a temporary camera might be quickly configured to cover an area with poor visibility, or camera data from spectators in the audience at a sports match could be incorporated into the 3D skeleton fitting process for objects in that sports match.

[0117] The output of method 500 is a 4D skeleton 510 (i.e., a temporal representation of the 3D object information) of the classified object. Each frame of this 4D skeleton can be mapped to a 3D avatar in the manner described above in connection with FIG. 2, thereby providing a 4D avatar. While FIG. 5 illustrates a single camera, it is understood that the method described with respect to FIG. 5 can also be applied to instances of image data received from multiple cameras simultaneously. Thus, each instance of a determined 2D skeleton 506a, 506b, 506c...506z shown in FIG. 5 can be understood to include multiple 2D skeletons in the manner described above in connection with FIG. 2, with the multiple 2D skeletons determined from each instance of image data being used to determine an individual 3D skeleton, as shown at 508a, 508b, 508c...508z. Thus, in multiple camera examples of method 500, the output is a single, time-ordered sequence of 3D skeletons 508a...508z, which can be refined in the same manner as described above.

[0118] Advantageously, the method may further comprise, for each classified object of the at least one classified object, using the 4D skeleton of the classified object to derive properties of the classified object, as described above. The properties may include, for example, time-varying three-dimensional location information relating to the time-dependent position of the classified object relative to the camera, or may include time-varying three-dimensional location data of a sub-object or sub-objects of the classified object relative to other sub-objects of the same classified object.

[0119] For example, if the classified object is a person (e.g., a sports player), the method may use the 3D skeleton to derive one or more of the person's position (either in a two-dimensional sense as a position on the ground, or in a three-dimensional sense, where the person should jump), the person's speed (or velocity), and the person's acceleration from the sequence of images.

[0120] In a further example, the collection of classified objects may represent a player of a sport and the additional object may represent an item of equipment for the sport (e.g., a ball), and the method may use the 3D skeleton of the player to identify interactions between the player and the item of equipment. The method may derive timing of the interactions. The interactions may include contact between the player and the item of equipment.

[0121] In a further example, the collection of classified objects may represent players of a sport, and the classified objects may represent items of equipment for that sport (e.g., bats or rackets) and / or balls carried by the players, and the method may use the 3D skeletons of the players and the carried items of equipment to identify interactions between any players and the items of equipment or between the items of equipment. The interactions may include contact between the classified objects. In the soccer game example, such interactions may include a player touching the ball with any part of his / her body, including a foot or hand. As described below, this interaction can be used as part of determining compliance with the offside rule.

[0122] For example, in a soccer match, the method may use multiple 3D skeletons representing players and an object representing the ball to derive from a sequence of images whether a rule, such as the offside rule, has been violated during the sequence of images.

[0123] As discussed above, processing speed advantages are obtained by mapping existing textures to the 3D skeleton rather than generating 3D textures "on the fly." However, the method can nevertheless be refined by generating several 3D textures to combine with the 3D model applied to the 3D skeleton. Specifically, the method may include generating a 3D surface mesh from the received image data and combining the derived 3D surface mesh with the 3D model and mapping it to the 3D skeleton. Examples of surface meshes may be meshes with surfaces describing clothing features or meshes with surfaces describing facial features. The mesh may be derived from image data of the object or from image data corresponding to sub-objects of the object.

[0124] For example, if the object is a soccer ball, the 3D mesh describes the surface of the soccer ball and is derived from image data of the soccer ball. As another example, the sub-object could be the face of a human object, and the 3D mesh could describe the surface of the face and be derived from image data corresponding to the face. The image data could be collected from a single camera or from multiple cameras of the same object, with the image data from each camera merged. Any suitable image processing technique, including but not limited to a neural network (such as a convolutional neural network), may be used to derive the 3D mesh from the image data. Such a single neural network (or multiple neural networks) could be trained on training data with multiple 3D surface outputs for various objects, with inputs being multiple flattened 2D perspectives corresponding to each 3D surface output. The 3D mesh could be mapped to a 3D skeleton in a similar manner to how a 3D model can be mapped to a 3D skeleton. That is, the 3D mesh could be mapped to corresponding points on the 3D skeleton (e.g., a face 3D mesh could be mapped to a point on the head of the 3D skeleton). Using custom 3D meshes for parts of an object in combination with pre-rendered 3D models means that elements unique to a particular object (e.g., a face) can be rendered more accurately, while pre-determined 3D models can be used for the rest of the model, thus providing a balance that provides accuracy when needed while reducing processing speed to allow for real-time processing.

[0125] 6 illustrates an environment 600 in which the present invention can be implemented. The environment includes multiple cameras 601a...601d. The cameras can include stationary cameras 601a and 601b. Such stationary cameras may have a fixed field of view and angle, or may be capable of rotating and / or zooming. The cameras can include movable camera 601c, which can translate relative to the environment and may also have a fixed field of view and angle, or may be capable of rotating and / or zooming. Additionally, the cameras can include handheld camera 601d, for example, held by an observer in the audience. Each camera can be a different type of vision system, as described above.

[0126] The environment includes multiple objects 602a...602d. The above method may be employed to receive image data collected from multiple cameras 601a...601d, identify each of the multiple objects, and construct a 3D / 4D avatar for each object. The method may not necessarily construct 3D avatars for all identified objects 602a...602d, but may construct 3D avatars only for certain types of object classes. For example, object 602e is a goal post and object 602f is a corner flag. These objects are static objects that do not change their position or shape. This information is already known as a result of identifying the object classes and locations in the environment. Therefore, there is no need to re-derive the 2D and 3D descriptor information.

[0127] The above method allows for the generation of multiple 3D / 4D avatars for objects located within an environment. Given that the avatars are generated by a camera with a known position and orientation, the avatars are generated at known 4D locations relative to the environment. It is desirable to also generate a 3D environment with an orientation for the 3D avatar that matches the orientation of the real-world environment in which the objects are located. The 3D avatars can then be placed within the 3D environment to more perfectly replicate the data captured by the camera.

[0128] The method may further include generating a 3D model of the environment. This model may be generated prior to performing any method in the absence of image data from calibration of cameras used to capture the image data. For example, multiple cameras may be set up with known locations and orientations relative to the environment. Knowing the locations and orientations allows the pre-generated environment to be mapped to the corresponding locations.

[0129] In a further advantageous refinement of the above method, the method further includes constructing a 3D model of the environment from the image data. Specifically, the same image data used to generate 3D / 4D avatars of objects in environment 600 is also used to generate environment 600 itself. Such a method is advantageous in situations where the camera used to collect the image data is not stationary, or where the position / orientation of the stationary camera relative to the environment is unknown.

[0130] As mentioned above, existing examples in the art of pixel-by-pixel rendering of a complete 3D environment require the application of significant computational resources. Thus, these existing methods are impractical for real-time reconstruction of 3D models. In a manner similar to the generation of 3D / 4D avatars of objects as described above, the present invention instead seeks to generate a 3D environment model by using 3D descriptor data of that environment.

[0131] In one advantageous embodiment, a 3D environment model may be constructed by receiving image data from multiple cameras, each positioned at a different known point in the environment, performing image processing techniques to identify stationary objects and / or features of the environment in the image data of each camera, and triangulating the stationary features and / or objects detected in the image data collected from each camera to identify the location of each stationary object and / or feature. Once the location of each stationary object and / or feature is determined, a predetermined 3D model of the environment may be mapped to the determined stationary objects and / or features. Suitable image processing techniques may also be used to identify stationary objects and / or features of the environment in the image data, including implementations of convolutional neural networks. Examples of determined features include goal posts, corner posts, ground line markings, etc. The same techniques for identifying stationary objects may also be applied to additional stationary cameras whose locations are unknown to identify multiple stationary objects in the image data collected by those cameras. Because the same stationary objects are identified in the image data of cameras with known locations, correspondences can be made between different stationary objects in different sets of camera data, allowing for location inference of additional stationary cameras with unknown locations.

[0132] In addition to, or as an alternative to, the above-described feature extraction for constructing the 3D environment model, in one further advantageous refinement, the 3D environment model may be constructed by applying simultaneous localization and mapping (SLAM) techniques to the image data to estimate the position and orientation of each of the at least one camera to construct an environment map, and by mapping the environment map to a predetermined 3D model of the environment to construct the 3D model. SLAM techniques are understood as computational methods used by an agent to build and / or update a map of the environment while simultaneously keeping track of the agent's location within the environment in real time.

[0133] In the present invention, a SLAM methodology can be applied to image data of environment 600 collected from at least one non-stationary camera (e.g., camera 601c) in environment 600 to simultaneously derive the location of each camera within the environment and map the environment including the cameras. Applying the SLAM methodology to the image data results in the generation of a low-resolution 3D model of the environment and the estimation of the position and pose of the camera within the low-resolution 3D model. The low-resolution 3D model can then be mapped to a high-resolution version stored in a database. SLAM methods can be used simultaneously with multiple cameras in the same environment to simultaneously locate the cameras within the same environment.

[0134] Forming part of the SLAM methodology can be a process of image detection of image data, in which stationary objects are detected in the environment. The object detection methodology used may be similar to that described above for detecting non-stationary objects in image data. For example, a convolutional neural network may be implemented, where the convolutional neural network is trained on a dataset containing images of stationary objects to be detected.

[0135] The output of the implemented SLAM methodology is an environment map that identifies feature points within the environment 600, as well as their locations. The map can be topological in nature, describing only the feature points of the environment, rather than a complete geometric reproduction. A topological map can be as few as 100 pixel coordinates, with each pixel describing the location of the feature point relative to the camera used to determine the pixel. For example, SLAM techniques can identify many static features of the environment, such as goal posts, corners of the pitch, and seating stands on the sides of the pitch. The environment map is then mapped to a predetermined 3D model of the environment. For example, each pixel in the environment map can have a corresponding point in the predetermined 3D model.

[0136] Any suitable SLAM methodology can be applied to simultaneously determine the location, position, and orientation of the camera, for example, a Hidden Markov Model can be applied where the measured state is image data and the hidden state is the location of elements of the environment.

[0137] Examples of fast SLAM techniques suitable for use in the present invention are also discussed in "FastSLAM: A factored solution to the simultaneous localization and mapping problem" by Montemerlo, Michael, et al., Aaai / iaai 593598 (2002), and "6-DoF Low Dimensionality SLAM (L-SLAM)" by Zikos, Nikos; Petridis, Vassilios (2014), Journal of Intelligent & Robotic Systems, Springer.

[0138] After constructing the 3D model of the environment, the method may further include integrating the determined 3D environment model with the constructed 3D or 4D skeletons for each classified object to construct a 3D or 4D model of the integrated environment, the integrated environment including the environment and the 3D or 4D skeletons of the classified objects within the environment. This integration step may be performed independently of the step of integrating the 3D model with the 3D skeleton to create a 3D avatar. Thus, the 3D or 4D integrated environment model accurately represents the environment and its associated objects while minimizing the data used to describe the environment. This 3D or 4D integrated model may be exported and saved in a data file for later integration with a previously generated 3D model, allowing the scene to be fully recreated at a later date, either on the same computer system that generated the model or on a different computer system.

[0139] The integrated 3D or 4D environment model may also be used to derive characteristics related to the environment. Advantageously, the method may further include, for each classified object of the at least one classified object, deriving properties of the classified object using the integrated 3D or 4D environment. When a 3D environment is used, the properties may include, for example, three-dimensional location information regarding the position of the classified object (and any points on the 3D skeleton of the classified object) relative to the environment and other objects in the environment (and any points on the 3D skeletons of the other objects). These properties may be used, for example, to determine whether players on the pitch are complying with the rules of a sports game, such as the offside rule in soccer. In this example, portions of a player's 3D skeleton may be analyzed relative to the environment and other objects within the environment to determine whether any of the player's body parts, excluding his or her hands and arms, are within the opponent's half-pitch and closer to the opponent's goal line than both the ball and the penultimate opponent's player (the positions of the ball and the penultimate opponent's player relative to the environment may be derived from the respective 3D skeletons of the ball and the penultimate opponent's player). For example, a two-dimensional vertical plane may be generated in the 3D environment that runs the width of the pitch and intersects behind the point where a player's body parts may be considered offside. The determined locations of elements of the player's constructed 3D skeleton are then compared to this 2D plane to determine whether they are behind the plane and therefore in an area that would be considered offside. For example, dimensions of the 3D skeleton elements (e.g., legs) may be derived from the 3D skeleton, from which it is determined whether any parts of the body parts represented by the 3D skeleton elements are behind the 2D plane. These elements may be predefined elements that include the legs, feet, knees, or any other part of the human body, but not including the hands and arms of the human body.

[0140] As the definition of the offside rule differs (e.g., because the rules change over time, or when considering other sporting environments such as rugby), the point beyond which a player may be considered offside may change (and therefore the location of the 2D plane may change). However, it is easy to see that the same principle as above can still be applied: the point beyond which a player may be considered offside is determined by considering the location of objects or elements of objects in the sporting environment, and determining whether any element of the player is behind this point. The specification of the offside rule may be pre-programmed into the system, or may be specified by a user at the time compliance is determined.

[0141] When using a 4D environment model, properties may include, for example, time-varying three-dimensional location information regarding the time-dependent position of a classified object relative to the environment and relative to other objects in the environment. Thus, the integrated 3D or 4D environment model may be used to derive quantities such as the speed of one or more players, changes in the team's formation and location on the pitch, and the speed and movement of a player's individual joints. Furthermore, properties derived from the 4D environment model may also include determining compliance with the offside rule at different times, in the manner described above for each moment in time. This determination may be made at or after a given time, either at the request of a user of the system, or at or after a specific event is detected. In the case of the offside rule, the integrated 4D environment model may also be used to identify the exact time the ball was kicked. It is at or after that time that the above-described analysis comparing the position of a player's body parts to a 2D plane may be performed. Subsequent events after the ball is kicked may also be determined, including further interactions between the player and another player and / or the ball (e.g., the player touching another player or touching the ball), or the relative positioning of classified objects in the sports environment (e.g., whether a player in an offside position blocks an opposing player's line of sight to the ball). For example, contact of a player by another player or contact of the ball by a player may be determined by the relative positioning of the 3D skeletons of the respective objects. Thus, methods and systems can determine whether an offside violation occurred after the ball is kicked.

[0142] Alternatively, or in addition, after constructing the 3D model of the environment, the method may further include integrating the determined 3D environment model with the constructed 3D avatars for each classified object to construct a 3D model of the integrated environment, the integrated environment including the environment and the objects within the environment. As described above, each of the 3D / 4D avatars constructed by the foregoing method includes location information for placing each 3D avatar within the environment. Thus, by integrating the 3D / 4D avatar with the environment, an integrated 4D model is created that accurately depicts in 3D the environment and all objects within that environment, as well as the time-based progression of non-static elements within the environment.

[0143] It will be appreciated that the integrated 3D or 4D environment including the constructed avatar can be used to derive characteristics related to the environment in a similar manner as described above for the integrated 3D or 4D environment including the 3D or 4D skeleton. Advantageously, the method may further include, for each classified object of the at least one classified object, deriving properties of the classified object using a model of the integrated 3D or 4D environment. When a 3D model is used, the properties may include, for example, three-dimensional location information regarding the position of the classified object (and any points on the surface 3D avatar of the classified object) relative to the environment and other objects in the environment (and any points on the surface 3D avatar of the other objects). Following the above example, these properties may be used, for example, to determine whether a player on the pitch is complying with the rules of a sports game, such as the offside rule in soccer. In this example, portions of the player's 3D avatar may be analyzed relative to the environment and other objects within the environment to determine whether any of the player's body parts, excluding the hands and arms, are within the opponent's half pitch and closer to the opponent's goal line than both the ball and the penultimate opponent player (the positions of the ball and the penultimate opponent player relative to the environment may be derived from the 3D avatars or 3D skeletons of the ball and the penultimate opponent player, respectively). For example, a two-dimensional vertical plane may be generated in the 3D environment that runs the width of the pitch and intersects behind the point where a player's body parts may be considered offside. Portions of the player's 3D avatar (including the surface of the 3D avatar) are then compared with this 2D plane to determine whether any part of the 3D avatar (including the surface of the 3D avatar) is behind the plane, i.e., in an area that is considered offside. The parts compared may relate to predefined parts of the player, including the legs, feet, knees, or any other part of the human body, excluding the hands and arms.

[0144] As the definition of the offside rule differs (e.g., because the rules change over time, or when considering other sporting environments such as rugby), the point beyond which a player may be considered offside may change (and therefore the location of the 2D plane may change). However, it is easy to see that the same principle as above can still be applied: the point beyond which a player may be considered offside is determined by considering the location of objects or elements of objects in the sporting environment, and determining whether any element of the player is behind this point. The specification of the offside rule may be pre-programmed into the system, or may be specified by a user at the time compliance is determined.

[0145] When using a 4D environment model including a constructed 4D avatar, properties can include, for example, time-varying three-dimensional location information regarding the time-dependent position of a classified object relative to the environment and relative to other objects in the environment. Thus, the integrated 3D or 4D environment model can be used to derive quantities such as the speed of one or more players, changes in the team's formation and location on the pitch, and the speed and movement of a player's individual joints. Furthermore, properties derived from the 4D environment model can also include determining compliance with the offside rule at different times, in the manner described above for each moment in time. This determination can be made at or after a given time, either at the request of a user of the system or at or after a specific event is detected. In the case of the offside rule, the integrated 4D environment model can also be used to identify the exact time the ball was kicked. It is at or after that time that the above-described analysis comparing the position of a player's body parts to a 2D plane can be performed. Subsequent events after the ball is kicked may also be determined, including further interactions between the player and another player and / or the ball (e.g., the player touching another player or touching the ball), or the relative positioning of classified objects in the sports environment (e.g., whether a player in an offside position blocks an opposing player's line of sight to the ball). For example, contact of a player by another player or contact of the ball by a player may be determined by the relative positioning of the 3D skeletons of the respective objects, or by whether the surfaces of the 3D avatars of the respective objects are touching. Thus, methods and systems can determine whether an offside violation occurred after the ball is kicked.

[0146] It is understood that the separate derivation and construction of components of an integrated environment may result in mismatch errors when combined. For example, a human 3D avatar may be slightly misaligned relative to the ground of the pitch, resulting in the 3D avatar "floating" above the pitch or "sinking" within the pitch. The method of the present invention may reduce such mismatch errors by refining the construction of the integrated environment by constraining the interaction between each constructed 3D avatar and the 3D environment model. For example, the integration of the 3D avatar and the 3D environment may be constrained so that elements of the texture of the mapped 3D model cannot overlap with the texture of the environment (e.g., "pinning" a standing human feed to the ground of the environment).

[0147] Such methods can be reapplied using appropriately trained neural networks. For example, motion capture techniques can be used to observe objects for which 3D avatars are constructed interacting with an environment. The training data can include objects performing a wide variety of actions (e.g., running, jumping, walking) to identify a corresponding wide variety of configurations of the object's features with respect to the environment. For example, when a human runs, their feet will only contact the floor of the environment at specific points in their gait. By training a neural network with such examples, the method of the present invention can recognize appropriate interactions between the classified objects and the environment to constrain the interactions between the corresponding constructed avatars and the 3D environment model when constructing an integrated environment model.

[0148] The generated integrated environment model can be further improved. Specifically, it will be appreciated that the image data collected by the at least one camera may be a comprehensive observation of each object for which an avatar is generated. The time interval between captured instances of an object may be greater than the time interval over which a 4D avatar is constructed. Thus, attempts to generate animations from 3D avatars will demonstrate "jitter" as the animation progresses.

[0149] This jitter can be reduced or eliminated by a method that further includes applying time-dependent smoothing to each constructed 3D avatar or to each constructed 3D skeleton from which the avatar was created. Smoothing can be performed using various filtering methods, such as temporal neural networks, temporal estimators, surface fitting, etc. The temporal neural network can be trained using training data with fine temporal resolution (e.g., frame rates higher than the 60 fps used in typical video transmission). Additionally or alternatively, smoothing of each individual 4D avatar can be based on the time-dependent holonomic constraints described above. Specifically, additional points of the 3D skeleton representing the object at additional time points can be calculated by interpolating between two previously determined sets of values ​​at the measurement time points. These interpolated values ​​can then be constrained based on the same kinematic model used to determine the previously determined values. For example, if a finite number of 3D skeletons for a 4D avatar, such as a bowstring or a javelin thrower, are determined from image data using holonomic constraints based on elastic physics, additional 3D skeletons can be derived from the existing 3D skeletons by applying the same equations used for the holonomic constraints in the 3D skeleton determination. Thus, additional 3D skeletons for the 4D avatar may be derived, and new surface textures may be applied to the refined 4D avatar to create a smoothed 4D avatar. Furthermore, anatomical constraints on the degree of human body movement may constrain the modeling of human avatar movement.

[0150] Smoothing can be applied to each 4D descriptor data or each 4D avatar, regardless of any integration with the environment, or it can be performed as part of the integration step to refine the construction of the integrated environment, where the interaction between the 3D environment and each 4D avatar is modeled as part of the smoothing technique. Specifically, the interaction between the 3D environment and the 4D avatar may include further refinement based on established laws of physics. For example, the trajectory of a ball can be modeled using Newtonian physics and the laws of gravity to limit the degrees of freedom of the ball's trajectory. Optionally, additional features of the environment that are not detected as objects in the image data can be generated on the fly based on the interaction between the environment and the 4D avatar. One example is shadows cast on the ground. These can be generated based on the position of each avatar and appropriately programmed parameters (such as solar radiation levels, time of day, etc.) and integrated with the 3D model.

[0151] The generated integrated environment model can be further improved. Specifically, in situations where an environment includes at least two interacting objects for which 4D avatars are generated, it can be understood that combining two such avatars in a single environment may result in mismatch errors between the 4D avatars and the environment similar to those described above. For example, a 3D avatar of a person may be slightly misaligned with respect to the tennis racket held by the person, resulting in the racket 3D avatar "floating" above or "sinking" within the person's hand. Similarly, when modeling multiple non-human objects, such as billiard balls in a billiards game, the modeled 4D avatar balls may overlap in physically impossible ways.

[0152] The method of the present invention can reduce such mismatch errors. Specifically, in an improved method, constructing a 3D avatar for each classified object includes constructing multiple 3D avatars, and the method further includes refining the construction of the integrated environment by constraining interactions between the multiple 3D avatars to refine an estimation of the placement of one or more 4D objects relative to the environment. The constraints may be implemented by applying a neural network, including a temporal quadratic estimator, or another suitable statistical and / or probabilistic method. Additionally or alternatively, heuristically developed rules may be implemented. For example, a system operator can specify a set of equations that effectively model specific object-to-object interactions, and these equations are applied to refine the positioning of the objects relative to each other and the environment. Furthermore, the constraints may include refining possible interactions between the multiple 3D avatars based on known equations from the laws of physics.

[0153] Heuristically developed rules or known equations from the laws of physics may describe the time-dependent locations of multiple objects and may be applied simultaneously to multiple objects from which a 4D model has been derived. Thus, if the time-dependent evolution of each object is known in the form of a 4D vector for that object, the system may simultaneously position / orient these 4D vectors relative to each other and the environment by constraining that each 4D vector must simultaneously satisfy constraints set by heuristic equations, equations from the laws of physics, or other constraints.

[0154] For example, a game of billiards involves multiple solid spheres interacting on a plane, and the interactions of the multiple solid spheres can be modeled as objects at least partially predictably based on Newtonian collision dynamics. Thus, the 4D vectors of each modeled ball on a billiard table can be simultaneously oriented / positioned relative to each other and the environment by constraining that each 4D vector must simultaneously satisfy equations describing the collision dynamics in space and time.

[0155] The integrated 4D model represents the output of the method of the present invention. As described above, the output integrated 4D model may include only the 3D environment and a 4D skeleton (to which the 3D model may later be mapped), or it may include the 3D environment and a 4D avatar already mapped to the 4D skeleton. This integrated 4D model can then be manipulated and implemented for various purposes in line with the advantages of the present invention. For example, the model can be used to render new 2D images or videos of the integrated environment that were not displayed in the image data. For example, the rendered image or video may represent the environment from a new perspective (e.g., the pitch of the soccer match itself, or any requested point in the stands of the crowd). The generated data may also be VAR data for use in a virtual augmented reality (VAR) headset. An operator can also manually refine the determined 4D model to increase precision (e.g., to smooth out jitter or inconsistencies not already corrected by the above methods). The output data (i.e., the 4D environment model and each of its components) may be configured for use in known engines used to generate visual output (e.g., 2D output, 3D output). For example, a 4D model may be suitable for use in the Unity video game engine.

[0156] As described above, the output data may include an environment-independent determined 4D skeleton, or the 4D skeleton may be separate from the data describing the 4D integrated environment model. Advantageously, the 4D skeleton included in the output data may be used to generate a 3D or 4D model of one or more objects and integrated with additional image data. Specifically, the method may further include capturing additional image data including an additional time-ordered sequence of frames, generating a 4D avatar of the object, and overlaying a two-dimensional representation of the 4D avatar on the additional image data to augment the additional time-ordered sequence of frames with a depiction of the object. For example, a video feed of a live soccer match may be augmented with additional players to depict those players. Alternatively, the 4D model may augment a video feed of a studio environment to display images of soccer players and their actions alongside a studio presenter.

[0157] As mentioned above, the intermediate 3D and 4D descriptor information may be used to determine other characteristics and / or properties of objects detected in the environment. These characteristics and / or properties may also be inferred / determined from the final integrated 4D model. Thus, characteristics and / or properties regarding the relationship of movable objects to the environment may be inferred from the model (e.g., the amount of time a player spends in a particular area of ​​the pitch, their relative speed to other players / objects, etc.).

[0158] The characteristics and / or properties may be determined automatically in real time (simultaneously with the determination of the 3D data) by pre-programmed equations, or may be determined manually by an operator after the model has been calculated.

[0159] While the above examples relate to methods, it will be understood that any or all of the above methods can be implemented by a suitably configured computer. Specifically, forming part of the present invention is a computer having a processor configured to execute any of the above methods. In the above description, reference has been made to specific steps that may be performed by a human operator to refine the method. However, it will be understood that such human-based input is optional, and the method can be performed entirely in an automated manner by a computer system to automatically create and output a 4D spatiotemporal model of an environment (and the results of automated analysis of that 4D spatiotemporal model) based on input image data. For example, a user of the system may upload image data (in the form of a video feed), identify the context of the image data (e.g., the type of sport), and then instruct the computer to analyze the data. The computer generates output data in the described manner in an automated manner, suitable for export to a suitable engine for rendering a visualizeable model of the environment.

[0160] An exemplary system 700 for implementing the present invention is shown in FIG. 7. System 700 includes a computer 702. Computer 702 includes a processor 704 configured to execute a method according to any of the methods described above, including receiving image data 708. Image data 708 is received from one or more of several vision systems 710. Each vision system is configured in any suitable manner to transmit data from the camera to computer 702, which may include a wired camera 710a that communicates with the computer via a wired connection 712a and a remote camera 710b that communicates with computer 702 via a network 712b. The network may be any type of network suitable for communicating data (e.g., WLAN, LAN, 3G, 4G, 5C, etc.). The computer includes storage 714 that may store any suitable data (e.g., image data 708 received from vision system 710, or the predetermined 3D model described above).

[0161] Also forming part of the invention is a computer program containing instructions that, when executed by a processor (such as processor 704), cause the processor to perform any of the methods described above. The computer program may be maintained on a separate computer storage medium, or the computer program may be stored in storage 714. The computer program includes the instructions and coding necessary to implement all of the steps of the methods, including the implementation of statistical and / or probabilistic methods and neural networks, as described above. The computer program may reference other programs and / or data maintained or stored remotely from computer 702.

[0162] A processor 704 configured to perform the present invention can be understood as a series of appropriately configured modules, as shown in Figure 7B. For example, the processor can include a detection module 720 that performs all of the above-described object detection steps, a classification module 722 that performs all of the above-described object classification steps, a 2D skeleton determination module 724 that performs all of the above-described 2D skeleton determination steps, a 3D skeleton construction module 726 that performs all of the above-described 3D skeleton construction steps, a tracking module 728 that performs all of the above-described tracking steps, a 3D environment construction module 730 that performs all of the above-described 3D environment construction steps, and an integration module 732 that performs all of the above-described integration steps. The described modules are not intended to be limiting, and it will be understood that many suitable configurations of a processor are possible to enable the processor to perform the above-described method.

[0163] While preferred embodiments of the present invention have been described, it is to be understood that these are by way of example only and that various modifications can be contemplated.

[0164] The following numbered paragraphs (NP) explain various examples.

[0165] NP1. A method for deriving 3D data from image data of a sports environment for use in determining compliance with offside rules in a sports match, the method comprising: receiving image data representative of an environment from at least one camera; Detecting at least one object in the environment from the image data; classifying the at least one detected object, wherein the method comprises: determining a 2D skeleton of the classified object by implementing a neural network to identify features of the classified object in the image data corresponding to the classified object; and constructing a 3D skeleton of the classified object, comprising mapping the determined 2D skeleton to 3D.

[0166] NP2. Each of the at least one object includes a plurality of associated sub-objects, and for each of the at least one object: detecting the at least one object includes detecting each of a plurality of associated sub-objects; classifying the at least one object includes classifying each of a plurality of related sub-objects; The method of claim NP1, wherein determining the 2D skeleton of the classified object includes identifying features of each of a plurality of classified sub-objects in the image data that correspond to the classified sub-object.

[0167] NP3. For each classified object, map the determined 2D skeleton to 3D. Neural network implementations, and / or A method according to NP1 or NP2, comprising applying statistical and / or probabilistic methods to the determined 2D skeleton and applying holonomic constraints to the statistical and / or probabilistic methods.

[0168] NP4. The method of NP3, wherein classifying the at least one detected object includes classifying a first detected object of the at least one detected object as a human object, and wherein the holonomic constraints include human anatomical holonomic constraints when statistical and / or probabilistic methods are applied to the determined 2D skeleton.

[0169] NP5. The method of any preceding NP, wherein classifying the at least one detected object includes classifying a second detected object of the at least one detected object as a non-human object.

[0170] NP6. At least one camera includes multiple cameras; detecting the at least one object includes detecting the at least one object in image data collected from each camera; classifying the at least one object includes classifying the at least one object in the image data collected from each camera; determining a 2D skeleton for each classified object comprises determining a plurality of 2D skeletons, each of the plurality of 2D skeletons determined from identifying features of the classified object in image data collected from a different one of the plurality of cameras; The method of any preceding NP, wherein constructing the 3D skeleton comprises combining a plurality of determined 2D skeletons.

[0171] NP7. Constructing a 3D model of the environment from image data; constructing a 3D integrated environment, including, for each of the at least one classified object, integrating a 3D skeleton of the classified object with a 3D environment model; The method of any preceding NP, further comprising: determining characteristics of the integrated 3D environment, including, for each classified object, deriving placement information of a 3D skeleton of at least one classified object with respect to the 3D environment to determine whether a player in the 3D environment is complying with an offside rule.

[0172] NP8. Constructing a 3D model of the environment from image data; constructing a 3D avatar of the classified object, including, for each classified object of the at least one classified object, integrating the constructed 3D skeleton with a 3D model corresponding to the classified object; integrating the determined 3D environment model with the constructed 3D avatar for each of the at least one classified object to construct a 3D model of the integrated environment, the integrated environment including the environment and the object within the environment; The method of any one of NP1 to 6, further comprising: determining characteristics of the 3D model of the integrated environment, including, for each classified object, deriving placement information of a 3D avatar of the classified object with respect to the 3D environment to determine whether a player in the 3D environment complies with the offside rule.

[0173] NP9. the image data collected from each camera of the at least one camera comprises a sequence of time-ordered frames; wherein said detecting at least one object in the environment includes detecting the at least one object in each frame of a time-ordered sequence of frames; said classifying the at least one detected object includes classifying the at least one detected object in each frame of the sequence of frames; The method of NP1-6, wherein each classified object of the at least one classified object is tracked throughout the sequence of time-ordered frames.

[0174] NP10. Each classified object is tracked throughout a sequence of time-ordered frames by implementing a recursive estimator, and for each classified object, the determined 2D skeleton is mapped to 3D. determining a plurality of 3D skeletons to form a time-varying 3D skeleton; and applying a recursive estimator over the sequence of time-ordered frames to determine a time-varying 3D skeleton of the classified object.

[0175] NP11. The method of NP10, wherein applying the recursive estimator includes applying time-dependent holonomic constraints.

[0176] NP12. The method of any one of NP9 to 11, further comprising constructing a 3D model of the environment from the image data.

[0177] NP13.3D environment model, applying a positioning and mapping technique to the image data to estimate a position and orientation of each of the at least one camera and construct an environment map; constructing a 3D model by mapping the environment map onto a predetermined 3D model of the environment.

[0178] NP14. Integrating the time-varying 3D skeleton and the 3D environment model to construct a time-varying integrated environment; The method of NP12 or NP13, further comprising: determining characteristics of the integrated time-varying 3D environment, including, for each classified object, deriving time-varying positioning information of a 3D skeleton of the classified object relative to the 3D environment to determine whether a player in the 3D environment is complying with the offside rule.

[0179] NP15. The method comprises: for each classified object of at least one classified object: The method of any preceding NP, further comprising constructing a 3D avatar of the classified object, comprising integrating the constructed 3D skeleton with a 3D model corresponding to the classified object.

[0180] NP16. The method of NP15, wherein the method further includes capturing a 3D model corresponding to the classified object at a resolution higher than the resolution of the image data.

[0181] NP17. constructing a 3D avatar of the classified object, including, for each classified object of the at least one classified object, integrating the constructed 3D skeleton with a 3D model corresponding to the classified object; integrating the determined 3D environment model with the constructed 3D avatars for each classified object to construct a 3D model of an integrated environment, the integrated environment including the environment and the objects within the environment; The method of NP12 or NP13, further comprising: determining characteristics of the 3D environment model, including, for each classified object, deriving placement information of a 3D avatar of the classified object with respect to the 3D environment to determine whether a player in the 3D environment complies with the offside rule.

[0182] NP18. The method for determining whether a player in a 3D environment is in compliance with the offside rule is to: generating a 2D vertical plane that intersects behind the point where any part of a player's body may be considered offside; and determining on which side of a two-dimensional plane a predefined portion of the player may be found.

[0183] NP19. The method of NP17, further comprising refining the construction of the integrated environment by applying time-dependent smoothing to the constructed 3D avatar using a filtering method.

[0184] NP20. The method of NP17 or NP19, further comprising refining the construction of the integrated environment by constraining the interaction between each constructed 3D avatar and the 3D environment model.

[0185] NP21. A method according to any one of NP17 to NP20, wherein constructing a 3D avatar for each classified object includes constructing multiple 3D avatars, and the method further includes refining the construction of the integrated environment by constraining interactions between the multiple 3D avatars.

[0186] NP22. A computer comprising a processor configured to perform the method of any of the preceding NPs.

[0187] NP23. A computer program comprising instructions which, when executed by a processor, cause the processor to carry out the method of any of NP1 to NP21.

Claims

1. 1. A method for deriving 3D data from image data, comprising: receiving image data representative of an environment from at least one camera; detecting at least one object in the environment from the image data; classifying the at least one detected object, wherein the method comprises, for each classified object of the at least one classified object: determining a 2D skeleton of the classified object by implementing a neural network to identify features of the classified object in the image data that correspond to the classified object; constructing a 3D skeleton of the classified object, including mapping the determined 2D skeleton to 3D; the image data includes a time-ordered sequence of frames, the time-ordered sequence of frames including a frame collected from each camera of the at least one camera; wherein said detecting at least one object in said environment comprises detecting said at least one object in each frame of said time-ordered sequence of frames; wherein the classifying the at least one detected object includes classifying the at least one detected object in each frame of the sequence of frames; each classified object of the at least one classified object is tracked throughout the time-ordered sequence of frames; each classified object is tracked throughout the sequence of time-ordered frames by implementing a recursive estimator, and for each classified object, the determined 2D skeleton is mapped to 3D; determining a plurality of 3D skeletons to form a time-varying 3D skeleton; applying a recursive estimator over the sequence of time-ordered frames to determine a time-varying 3D skeleton of the classified object; wherein in applying the recursive estimator throughout the sequence of time-ordered frames, a 3D skeleton derived at a first time instance in the time order from the plurality of 3D skeletons forming the time-varying 3D skeleton is used as a further input to the derivation of a second 3D skeleton at a second time instance in the time order from the plurality of 3D skeletons forming the time-varying 3D skeleton.

2. each of said at least one object includes a plurality of associated sub-objects, and for each of said at least one object: detecting at least one object includes detecting each of the plurality of associated sub-objects; classifying at least one object includes classifying each of the plurality of related sub-objects; 2. The method of claim 1 , wherein determining the 2D skeleton of the classified object comprises identifying features of each of a plurality of classified sub-objects in the image data that correspond to the classified sub-object.

3. mapping the determined 2D skeleton to 3D for each classified object; Neural network implementations, and / or The method of claim 1 or 2, comprising applying statistical and / or probabilistic methods to the determined 2D skeleton and applying holonomic constraints to the statistical and / or probabilistic methods.

4. 4. The method of claim 3, wherein classifying the at least one detected object comprises classifying a first detected object of the at least one detected object as a human object, and wherein the holonomic constraints comprise human anatomical holonomic constraints when statistical and / or probabilistic methods are applied to the determined 2D skeleton.

5. 5. The method of claim 3 or 4, wherein the 3D skeleton comprises an anchor point and a plurality of child points, each child point being defined by applying a transformation to the anchor point or to another child point of the plurality of child points, and wherein the holonomic constraints define the degrees of freedom of the transformation that define each child point and / or the range of possible values ​​of each degree of freedom of the transformation.

6. 6. The method of claim 1, wherein classifying the at least one detected object comprises classifying a second detected object of the at least one detected object as a non-human object.

7. the at least one camera includes a plurality of cameras; detecting the at least one object includes detecting the at least one object in image data collected from each camera; classifying the at least one object includes classifying the at least one object detected in the image data from each camera; determining the 2D skeleton for each classified object comprises determining a plurality of 2D skeletons, each of the plurality of 2D skeletons determined from identifying features of the classified object in image data collected from a different one of the plurality of cameras; The method of any one of claims 1 to 6, wherein constructing the 3D skeleton comprises combining the determined plurality of 2D skeletons.

8. The method of any one of claims 1 to 7, wherein applying the recursive estimator comprises applying time-dependent holonomic constraints.

9. the 3D skeleton includes an anchor point and a plurality of child points, each child point being defined by applying a transformation to the anchor point or another child point of the plurality of child points; 9. The method of claim 8, wherein the time-varying 3D skeleton is defined by a transformation of a time variable for each child point, and the time-dependent holonomic constraints define, for each time point, the degrees of freedom of the transformations that define each child point and / or the range of possible values ​​of each degree of freedom of the transformations.

10. 10. The method of claim 1, wherein the at least one camera includes at least one camera of a first type and at least one camera of a second type, each of the cameras of the first type capturing image data at a first frame rate and each camera of the second type capturing image data at a second frame rate different from the first frame rate.

11. The method of any one of claims 1 to 10, further comprising constructing a 3D environment model of the environment from the image data.

12. The 3D environment model applying a simultaneous positioning and mapping technique to the image data to estimate the position and orientation of each of the at least one camera and construct an environment map; and constructing the 3D model by mapping the environment map onto a predetermined 3D model of the environment.

13. The method of claim 11 or 12, further comprising integrating the time-varying 3D skeleton with the 3D environment model to construct a time-varying integrated environment.

14. The method further comprises: for each classified object of the at least one classified object:

14. The method of claim 1, further comprising constructing a 3D avatar of the classified object, comprising integrating the constructed 3D skeleton with a 3D model corresponding to the classified object.

15. The method of claim 14 , wherein the method further comprises capturing the 3D model corresponding to the classified object at a resolution higher than a resolution of the image data.

16. constructing a 3D avatar of the classified object, including, for each classified object of the at least one classified object, integrating the constructed 3D skeleton with a 3D model corresponding to the classified object; 13. The method of claim 11 or 12, further comprising: integrating the determined 3D environment model with the constructed 3D avatars for each of the classified objects to construct a 3D model of an integrated environment, the integrated environment including the environment and the objects within the environment.

17. 17. The method of claim 16, further comprising refining the construction of the integrated environment by applying time-dependent smoothing to the constructed 3D avatar using a filtering method.

18. The method of claim 16 or 17, further comprising refining the construction of the integrated environment by constraining interactions between each constructed 3D avatar and the 3D environment model.

19. 19. The method of any one of claims 16 to 18, wherein constructing a 3D avatar for each classified object comprises constructing a plurality of 3D avatars, the method further comprising refining the construction of the integrated environment by constraining interactions between the plurality of 3D avatars.

20. 1. A method for deriving 3D data from image data of a sports environment for use in determining compliance with offside rules in a sports match, the method comprising: receiving image data representative of the environment from at least one camera; detecting a plurality of objects in the environment from the image data; classifying a plurality of said detected objects as players; The method comprises, for each player: determining a 2D skeleton of the player by implementing a neural network to identify player features in the image data corresponding to the player; constructing a 3D skeleton of the player, including mapping the determined 2D skeleton to 3D; and constructing a 3D avatar of the player by integrating the constructed 3D skeleton with 3D models corresponding to the classified objects; The method, in response to a user-detected or automatically detected event, defining a two-dimensional plane within said environment; The method further includes determining on which side of the two-dimensional plane a predefined portion of the player's avatar can be found.

21. 1. A method for deriving 3D data from image data of a sports environment for use in determining compliance with offside rules in a sports match, the method comprising: receiving image data representative of the environment from at least one camera; detecting a plurality of objects in the environment from the image data; classifying a plurality of said detected objects as players; The method comprises, for each player: determining a 2D skeleton of the player by implementing a neural network to identify player features in the image data corresponding to the player; constructing a 3D skeleton of the player, including mapping the determined 2D skeleton to 3D; The method, in response to a user-detected or automatically detected event, defining a two-dimensional plane within said environment; The method further includes determining on which side of the two-dimensional plane a predefined portion of the 3D skeleton of the player can be found.

22. A computer comprising a processor configured to carry out the method of any one of claims 1 to 21.

23. A computer program comprising instructions which, when executed by a processor, cause the processor to carry out the method of any one of claims 1 to 21.

Citation Information

Patent Citations

  • Imaging system and imaging method

    JP2019032660A

  • JPP6534499B

  • Method circuit and system for human to machine interfacing by hand gestures

    KR1020130140772A

  • Depth projector system with integrated vcsel array

    US20120051588A1

  • Monitoring system

    WO2018061616A1