3D object tracking using unvalidated detections registered by one or more sensors

The system effectively tracks moving objects in 3D space by leveraging multiple sensors to manage high false positive rates, applying filters and physical models to validate object movement hypotheses, ensuring reliable detection and tracking of objects like golf balls.

JP7699208B2Active Publication Date: 2025-06-26TOPGOLF SWEDEN AB
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023526086
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-11-03
Filing Date
2021-10-28
Publication Date
2025-06-26
Estimated Expiration
2041-10-28

AI Technical Summary

Technical Problem

Existing systems for tracking moving objects, such as a golf ball in flight, face challenges in accurately detecting objects in 3D space due to a high number of false positives from sensors, which can overload the tracking system and lead to missed detections.

Method used

The system employs multiple sensors, including cameras and radar devices, to generate 3D data points. It uses unvalidated object detections from these sensors, allowing more false positives to be used as input, which minimizes false negatives. The system applies filters and physical models to validate hypotheses of object movement, forming accurate 3D tracks despite high false positive rates.

Benefits of technology

This approach enables reliable detection and tracking of small objects like golf balls in 3D space, even in conditions where sensors face challenges in data accuracy, by effectively managing false positives and negatives, thus maintaining system performance and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007699208000010
    Figure 0007699208000010
  • Figure 0007699208000011
    Figure 0007699208000011
  • Figure 0007699208000012
    Figure 0007699208000012
Patent Text Reader

Abstract

In at least one aspect, methods, systems, and devices, including media coding computer program products, for three-dimensional object tracking include a method that includes obtaining three-dimensional positions of registered objects by a detection system configured to tolerate more false positives and minimize false negatives; forming hypotheses using a filter that enables associations between registered objects when estimated three-dimensional velocity vectors approximately correspond to objects moving in three-dimensional space; eliminating a suitable subset of hypotheses that are not further expanded during the forming; specifying at least one three-dimensional track of at least one ball moving in three-dimensional space by applying a complete three-dimensional physical model to data about the three-dimensional positions used in forming the at least one hypothesis that remains after the eliminating; and outputting for display the at least one three-dimensional track of the at least one ball moving in three-dimensional space.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to tracking moving objects such as a golf ball in flight using data obtained from different sensors that can employ different sensor technologies.

Background Art

[0002] Systems and methods for tracking the flight of a golf shot using sensors include launch monitors, full-flight two-dimensional (2D) tracking, and full-flight three-dimensional (3D) tracking. Commonly used types of sensors are cameras, Doppler radars, and phased array radars. The launch monitor method is based on the measurement of a set of parameters that can be observed during the swing of a golf club and during the first few inches of ball flight after the club hits the ball. Next, the measured parameters are used to estimate the expected ball flight using mathematical and physical modeling.

[0003] In contrast, full-flight 3D tracking systems are characterized by designs that attempt to track the full flight of a golf shot rather than estimating from launch parameters. Further, full-flight 2D tracking systems track the shape of a golf shot as seen from a particular angle but do not generate 3D information and generally cannot be used to determine important parameters such as the distance the ball has traveled. Full-flight 3D tracking using a combination of camera and Doppler radar data is described in U.S. Patent No. 10,596,416. Finally, full-flight 3D tracking using a stereo camera having image frame acquisition units synchronized with each other is described as potentially usable in some situations for 3D tracking of objects.

Summary of the Invention

[0004] This specification describes techniques related to tracking an object moving in a three-dimensional (3D) space, such as a golf ball in flight, using data obtained from two or more sensors, such as a camera, a radar device, or a combination thereof. 3D object tracking uses unvalidated object detections registered by two or more sensors. By enabling more false positives of the detected objects to be used as input to 3D tracking in 3D space, the number of false negatives of the detected objects can be minimized, and thus, even in cases where the distance from the sensor (and / or the relative velocity with respect to the sensor) makes it difficult for the sensor to obtain sufficient data to accurately detect the object, small objects (e.g., golf balls) can be reliably detected.

[0005] However, this large number of false positives means that most (e.g., at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) of the 3D data input to the 3D tracking system is not actually an accurate detection of an object, which, without the systems and techniques described in this document for performing motion analysis in a 3D cloud of unvalidated detections of tracked objects, could overload the ability of a 3D tracking system to discover actual objects moving in 3D space in real time. False positives can include incorrect sensor detections that are not actually tracked objects (e.g., more than 90% of the data from each camera identifying a potential golf ball is not actually a golf ball, e.g., mere noise in the signal or other objects that are not golf balls), as well as incorrect pairings of data that are not actually one object (e.g., more than 50% of the pairings of object detections from different cameras, or pairings of object detections from a camera sensor and a radar device or other sensor, are incorrect, e.g., only the data from one of the sensors is actually a golf ball, or the data from both sensors is a golf ball but not the same golf ball). Nevertheless, the systems and techniques described herein can accurately identify objects moving within 3D space despite the input 3D data containing many false detections of objects and many incorrect pairings of data from the sensors used to detect the objects.

[0006] In general, one or more aspects of the subject matter described in this specification are one or more methods (and one or more non-transitory computer-readable media tangibly encoding instructions to cause a data processing apparatus associated with an object detection system comprising two or more sensors to perform operations), the object detection system being configured to allow more false positives and minimize false negatives for registered objects of interest, obtaining the three-dimensional position of a registered object of interest; forming a hypothesis of an object moving in three-dimensional space using a filter applied to the three-dimensional position of the registered object of interest, the filter enabling an association between specific objects of interest among the registered objects of interest when an estimated three-dimensional velocity vector for a specific object of interest substantially corresponds to an object moving in three-dimensional space over time; using at least one additional object of interest registered by the object detection system to remove an appropriate subset of hypotheses that are not further extended by the associations made during the forming; applying a complete three-dimensional physical model to data about the three-dimensional positions of the registered objects of interest used in forming at least one hypothesis remaining after the removing to specify at least one three-dimensional track of at least one ball moving in three-dimensional space; and outputting to display at least one three-dimensional track of at least one ball moving in three-dimensional space.

[0007] Other embodiments of this aspect include corresponding systems, apparatuses, and computer program products recorded on one or more computer storage devices, each configured to perform the operations of the method. Accordingly, one or more aspects of the subject matter described herein can be embodied in one or more systems that include two or more sensors (e.g., a hybrid camera / radar sensor, at least one camera and at least one radar device, a stereo camera, a plurality of cameras that can form various stereo camera pairs, etc.), one or more first computers of an object detection system configured to register a target object and to allow more false positives and minimize false negatives for the registered target object, and one or more second computers configured to perform operations according to one or more methods. The foregoing and other embodiments can optionally include one or more of the following features, alone or in combination.

[0008] Forming a hypothesis of an object moving in three-dimensional space using a filter can involve using, for a given hypothesis, a first motion model in three-dimensional space, a first model having a first level of complexity, when the number of registered detections for the given hypothesis is less than a threshold, and using, for a given hypothesis, a second motion model in three-dimensional space, the second model having a second level of complexity higher than the first level of complexity of the first model, when the number of registered detections for the given hypothesis is greater than or equal to the threshold. The first motion model in three-dimensional space can be a linear model that associates a first three-dimensional position with a given hypothesis when the first three-dimensional position of a registered target object extends the given hypothesis in substantially the same direction predicted by the given hypothesis, and the second motion model in three-dimensional space can be a recurrent neural network model trained to perform short-term prediction of motion in three-dimensional space.

[0009] Forming a hypothesis of an object moving in three-dimensional space may include using a first model or a second model to predict the next three-dimensional position of a registered detection for a given hypothesis according to the number of registered detections for the given hypothesis and a threshold, and searching a spatial data structure that includes all of the three-dimensional positions of the objects of interest registered by an object detection system for a next time slice to find a set of three-dimensional positions within a defined distance of the predicted next three-dimensional position; if the set of three-dimensional positions is an empty set, using the predicted next three-dimensional position for the given hypothesis; if the set of three-dimensional positions includes only one three-dimensional position, using the one three-dimensional position for the given hypothesis; if the set of three-dimensional positions includes two or more three-dimensional positions, sorting the two or more three-dimensional positions based on proximity to the predicted next three-dimensional position to form a sorted set; removing any non-proximate three-dimensional positions in the sorted set that exceed a defined threshold number greater than two; and using the two or more three-dimensional positions remaining in the sorted set to branch the given hypothesis into two or more hypotheses within a group of hypotheses.

[0010] Forming a hypothesis of an object moving in three-dimensional space may include, if the number of registered detections for a given hypothesis is greater than or equal to a threshold, identifying a data point in the given hypothesis that has a minimum vertical position, checking the respective vertical components of the estimated three-dimensional velocity vectors before and after the data point in the given hypothesis, and designating the data point as a ground collision if a first vertical component of the respective vertical components is negative and a second vertical component of the respective vertical components is positive.

[0011] Each vertical component can be obtained from the average of the estimated 3D velocity vectors before and after the data point in a given hypothesis, and the method may include selecting at least one time window based on the noise level associated with the estimated 3D velocity vector, the shortest predicted flight time on one or both sides of the data point, or both the noise level and the shortest predicted flight time, and calculating the average of the estimated 3D velocity vectors before and after the data points falling within at least one time window.

[0012] Forming a hypothesis of an object moving in 3D space involves, when the number of registered detections for a given hypothesis is above a threshold, dividing the given hypothesis into distinct segments including a flight segment, one or more bound segments, and a roll segment, including identifying, checking, and designating, dividing, classifying the first segment of the given hypothesis before the first designated ground collision as a flight segment, and when the angle between the first estimated velocity vector before the next designated ground collision and the second estimated velocity vector after the next designated ground collision is greater than a threshold angle, classifying the next segment of the given hypothesis after each next designated ground collision as one of the one or more bound segments, and when the angle between the first estimated velocity vector before the next designated ground collision and the second estimated velocity vector after the next designated ground collision is less than or equal to the threshold angle, classifying the next segment of the given hypothesis after the next designated ground collision as a roll segment.

[0013] A given hypothesis can be at least one hypothesis that remains after deletion, and specifying at least one 3D track of at least one ball moving in 3D space can form at least one hypothesis by triangulating a more accurate 3D path using sensor observations registered by an object detection system in at least a flight segment, generating data about the 3D position of a registered object used in forming the at least one hypothesis, and verifying at least one moving ball by applying a complete 3D physical model to the generated data for at least the flight segment.

[0014] Obtaining can include receiving or generating the 3D position of a registered object of interest, but most of the registered objects of interest are false positives including false detections by individual sensors and incorrect combinations of detections from each of the individual sensors. The individual sensors can include three or more sensors, and obtaining can include generating the 3D position by creating each combination of detections from each pair of three or more sensors for the current time slice. The three or more sensors can include two cameras and a radar device, and creating can include combining detections of a single object of interest by the two cameras using stereo pairing of the two cameras to generate a first 3D position of the single object of interest, and combining detections of a single object of interest by at least one of the radar device and the two cameras to generate a second 3D position of the single object of interest. Further, the three or more sensors can include additional cameras, and creating can include combining detections of a single object of interest by the additional cameras and the two cameras using stereo pairing of the additional camera with each of the two cameras to generate a third 3D position and a fourth 3D position of the single object of interest.

[0015] Details of one or more embodiments of the subject matter described in this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the invention will become apparent from the description, the drawings, and the claims.

Brief Description of the Drawings

[0016]

Figure 1A

Figure 1B

Figure 1C

Figure 1D

Figure 2

Figure 3

Figure 4A

Figure 4B

Figure 4C

Figure 4D

Figure 5

Figure 6A

Figure 6B

Figure 7

Figure 8

Figure 9

Figure 10A

Figure 10B

DETAILED DESCRIPTION OF THE INVENTION

[0017] Like reference numerals and names in the various drawings indicate like elements.

[0018] Figure 1A shows an example of a system 100 that performs motion-based preprocessing of two-dimensional (2D) image data and then performs three-dimensional (3D) object tracking of an object moving through a 3D space 110. The tracked object can be a golf ball or another type of object that is struck, kicked, or thrown (e.g., a baseball, a soccer ball, or a football / rugby ball). In some implementations, the 3D space 110 is a golf driving range, a grassland, or another open area where an object can be launched. For example, the 3D space 110 can include a golf entertainment facility that includes one or more targets 114, a building that includes a golf bay each having at least one tee area 112 (more generally, a launch area 112), and potentially other entertainment and dining options.

[0019] In some implementations, the 3D space 110 is a competition area for sports such as a golf course, the launch area 112 can be a golf tee for a particular hole on the golf course or an intermediate landing point for a golf ball in play on the course, and the target 114 can be the cup at the end of a particular hole on the golf course or an intermediate landing point for a golf ball in play on the course. Other implementations are possible, such as a launch area 112 that is one of a plurality of designated tee areas along a tee line where a golfer can hit a golf ball into the open field 110, or a launch area 112 that is one of a plurality of designated tee areas within the stands of a sports stadium where a golfer can hit a golf ball over and onto the playing field 110 of the sports stadium.

[0020] System 100 includes two or more sensors 130 including at least one camera 120 and its associated computer 125. One or more of the sensors 130 (including at least one camera 120 and its associated computer 125) may be placed near the launch area 112 for tracking an object, but this is not necessary. In some implementations, one or more sensors 130 (including cameras 120 and computers 125) may be placed along one or both sides of the 3D space 110 and / or on the other side of the 3D space 110 on the opposite side of the launch area 112. For example, in a golf tournament, assuming a shot is hit towards the green, the camera 120 and computer 125 may be placed behind the green facing the golfer. Thus, in various implementations, the sensors 130 can observe and track objects moving away from, towards, and / or through the field of view of the sensors 130 (note that each set of three dots leaders in the figure may also include one or more additional instances such as sensors, computers, communication channels, etc.).

[0021] The sensors 130 can include cameras (e.g., a stereo camera pair), radar devices (e.g., a single antenna Doppler radar device), or a combination thereof, potentially including a hybrid camera / radar sensor unit as described in U.S. Patent No. 10,596,416. However, at least one of the sensors 130 is the camera 120 and its associated computer 125, which are connected by a communication channel. FIGS. 1B - 1D are diagrams showing examples of different sensor and computer configurations that may be used in the system of FIG. 1A.

[0022] FIG. 1B shows an example of a pair of cameras 152, 156 connected to a first computer 150 via first communication channels 154, 158 having a first data bandwidth wider than the bandwidth of at least one other communication channel used in the system. For example, the first communication channels 154, 158 may employ one or more high-bandwidth short-distance data communication technologies such as USB (Universal Serial Bus) 3.0, MIPI (Mobile Industry Processor Interface), PCIx (Peripheral Component Interconnect Extended). As will be described in more detail below, preprocessing of the image data from the cameras 152, 156 may be performed near the cameras in one or more first computers 150, and if the preprocessing in the first computer(s) 150 reduces the data bandwidth, the output of this preprocessing may be transmitted via a second communication channel 162 having a second data bandwidth narrower than the first data bandwidth. Thus, the second communication channel 162 may employ one or more narrower-bandwidth, longer-distance data communication technologies such as copper wire Ethernet or wireless data connections (e.g., using WiFi and / or one or more mobile phone communication technologies).

[0023] This is important to enable the system to be implemented with higher resolution cameras 120, 152, 156 and computers 125, 150 that operate on raw (uncompressed) image data from these cameras 120, 152, 156. Whether using stereo camera tracking or hybrid camera / radar tracking, using higher resolution cameras with higher frame rates enables higher quality 3D tracking, but note that this is only the case if the data can be processed efficiently and effectively. Further, if object tracking is intended to function on very small objects (e.g., if the object appears as only a single pixel even in a high resolution camera image), using conventional irreversible video compression techniques (such as MPEG) may remove valuable information about small objects from the image, so object detection may need to access raw (uncompressed) image data.

[0024] To address these issues, the first computer(s) 150 may perform preprocessing (including object detection and optionally 2D tracking) on image data close to cameras 152, 156 to reduce the bandwidth requirements for transmitting sensor data to one or more second computers 160 via a second communication channel 162. Additionally, the preprocessing (as described in this document) enables virtual time synchronization of measured object positions downstream (after image capture) in time and space, and enables 3D tracking to be performed at the second computer(s) 160 using data received via one or more second communication channels 162. This makes it possible to easily perform downstream processing on a remote server since after preprocessing, it is trivial to transmit data over long distances due to the very narrow data bandwidth.

[0025] Note that this can provide a significant advantage when setting up system 100 due to the flexibility provided by system 100. For example, in the case of a golf competition television (TV) broadcast where system 100 is used to track a golf ball through the 3D space of a golf course and overlay the trace of the golf ball on a TV signal generated for live transmission or recording, sensor 130 can be deployed more than one mile away from a TV production facility (where 3D tracking computer 160 can be positioned). Note that the conversion of the ball position (identified during 3D tracking) to the corresponding position within the video data acquired by a TV camera (which enables a graphical representation of the ball's flight path to be traced overlaid on the video data) can be performed using known homography techniques. As another example, in the case of a golf entertainment facility, the 3D tracking computers (e.g., server computers 140, 160) do not have to be located in the same facility, and the 3D tracking performed by these computers (e.g., to extend other data or media such as showing the path of a golf ball in a computer representation of the physical environment where a golfer is located or in a virtual environment that exists only in the computer) can be easily transferred to another computer (e.g., for failover processing).

[0026] Various sensor and computer configurations are possible. FIG. 1C shows an example where each camera 152, 156 has a dedicated first computer 150A, 150B, and computers 150A, 150B communicate their respective pre - processed data to a second computer(s) 160 via separate second communication channels 162, 164. Thus, cameras (or other sensor technologies) may or may not share first computer resources. Further, pre - processing can be split and executed on different computers.

[0027] FIG. 1D shows an example in which camera 152 is coupled to computer 150 via a first communication channel 154 having a first data bandwidth, the first computer 150 is coupled to a third computer 166 via a third communication channel 168 having a third data bandwidth, and the third computer 166 is coupled to a second computer 160 via a second communication channel 162 having a second data bandwidth, where the second data bandwidth is narrower than the first data bandwidth and the third data bandwidth is narrower than the first data bandwidth but wider than the second data bandwidth. The first computer 150 performs object detection, the third computer 166 performs 2D tracking of the object, and the second computer 160 performs virtual time synchronization and 3D tracking of the object. Further, in some implementations, the first computer 150 performs object detection and pre-tracking in 2D (using very simple / loose constraints), the third computer 166 performs more thorough 2D tracking, and the second computer 160 performs virtual time synchronization and 3D tracking of the object.

[0028] Other sensors and computer configurations consistent with the disclosure of this document are also possible. For example, instead of using an intermediate third computer 166 to perform 2D tracking of an object, the first computer 150 can perform object detection (performing pre-tracking in 2D (using very simple / loose constraints) or not performing 2D tracking of the object), and the same second computer 160 can perform 2D tracking of the object (more thorough 2D tracking or full 2D tracking after pre-tracking in 2D), virtual time synchronization, and 3D tracking of the object. Conversely, in some implementations, one or more additional intermediate computers can be used. For example, the system can use four separate computers to perform each of the four operations of object detection, 2D tracking, virtual time synchronization, and 3D tracking. As another example, the system can use five separate computers to perform each of the five operations of object detection, pre-tracking in 2D (using very simple / loose constraints), more thorough 2D tracking, virtual time synchronization, and 3D tracking. Other configurations are possible, provided that at least one of the operations is performed on a first computer communicatively coupled to at least one camera via a first communication channel, and at least one other operation of the operations is performed on a second computer communicatively coupled to the first computer via a second communication channel having a data bandwidth narrower than the data bandwidth of the first communication channel.

[0029] Various types of computers can be used in the system. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. As used herein, "computer" can include a server computer, a client computer, a personal computer, an embedded programmable circuit, or a dedicated logic circuit. FIG. 2 is a schematic diagram of a data processing system including a data processing device 200 representing one implementation of a first computer 150, a second computer 160, or a third computer 166. The data processing device 200 can be connected to one or more computers 290 via a network 280.

[0030] The data processing device 200 can include various software modules that can be distributed between the application layer and the operating system. These can include executable and / or interpretable software programs or libraries, including programs 230 that operate as an object detection program (e.g., within the first computer 150), a 2D tracking program (e.g., within the first computer 150 and / or the third computer 166), a virtual time synchronization program (e.g., within the second computer 160), and / or a 3D tracking program (e.g., within the second computer 160) as described in this document. The number of software modules used can vary for each implementation. Also, in some cases, for example, the program 230 of the 2D tracking program 230 can be implemented within embedded firmware, and in other cases, for example, the programs 230 of the time synchronization and 3D tracking programs 230 can be implemented as software modules distributed over one or more data processing devices connected by one or more computer networks or other suitable communication networks.

[0031] The data processing apparatus 200 may include a hardware or firmware device including one or more hardware processors 212, one or more additional devices 214, a non-transitory computer-readable medium 216, a communication interface 218, and one or more user interface devices 220. The processor 212 can process instructions for execution within the data processing apparatus 200, such as instructions stored on the non-transitory computer-readable medium 216, and the non-transitory computer-readable medium 216 may include a storage device such as one of the additional devices 214. In some implementations, the processor 212 is a single-core processor or a multi-core processor, or two or more central processing units (CPUs). The data processing apparatus 200 communicates with one or more computers 290 using its communication interface 218, for example via a network 280. Thus, in various implementations, the described processes may be executed in parallel or sequentially on a single-core computing machine or a multi-core computing machine, and / or on a computer cluster / cloud, etc.

[0032] Examples of the user interface device 220 include a display, a touch screen display, a speaker, a microphone, a haptic feedback device, a keyboard, and a mouse. Further, the user interface device(s) need not be local device(s) 220 and may be remote from the data processing apparatus 200, for example, a user interface device(s) 290 accessible via one or more communication network(s) 280. The data processing apparatus 200 may store instructions for implementing the operations described in this document on a non-transitory computer-readable medium 216 that may include one or more additional devices 214 such as, for example, a floppy disk device, a hard disk device, an optical disk device, a tape device, and a solid state memory device (e.g., a RAM drive). Further, the instructions for performing the operations described in this document may be downloaded from one or more computers 290 (e.g., from the cloud) to the non-transitory computer-readable medium 216 via the network 280, and in some implementations, the RAM drive is a volatile memory device in which the instructions are downloaded each time the computer is powered on.

[0033] Furthermore, the 3D object tracking of the present disclosure can employ the pre-processing of sensor data, 2D object tracking, and virtual time synchronization described in this document (and also U.S. Provisional Patent Application No. 63 / 065872, filed on August 14, 2020, and U.S. Patent Application No. 17 / 404953, filed on August 17, 2021, both entitled "Motion based Pre-Processing of Two-Dimensional Image Data Prior to Three-Dimensional Object Tracking With Virtual Time Synchronization"). Thus, the 3D object tracking systems and techniques of the present disclosure, when implemented together, can obtain the attendant advantages of the systems and techniques of data pre-processing, 2D object tracking, and virtual time synchronization. However, in some implementations, it is not necessary to use these systems and techniques to obtain the three-dimensional positions of the objects of interest registered by the object detection system. One or more other object detection systems can be used in conjunction with the 3D object tracking described in more detail in connection with FIGS. 7-10B.

[0034] FIG. 3 shows an example of a process executed on different computers to detect an object, track the object in 2D, generate a virtual 2D position for time synchronization, and construct a 3D track of the moving object. The process of FIG. 3 includes pre-processing operations 310-330 executed on one or more first computers (e.g., computers 125, 150, 166 of FIGS. 1A-1D) and additional processing operations 360-375 executed on one or more second computers (e.g., computers 140, 160 of FIGS. 1A-1D). The pre-processing operations can include object detection and 2D tracking that effectively compress the ball position data (to reduce the bandwidth requirements for transmitting the data used in 3D tracking) to enable virtual time synchronization of the measured object positions during additional processing in the second computer(s).

[0035] Accordingly, the image frame 300 is received (310) from a camera (e.g., by computers 125, 150) via a first communication channel 305 that couples the camera to the first computer(s). In this case, the first communication channel 305 used when the image frame 300 is received has a first data bandwidth. For example, the first communication channel 305 can be a USB3.0, MIPI, or PClx communication channel, e.g., communication channel(s) 154, 158. Note that the bandwidth requirements between the camera and the computer can easily exceed 1 gigabit per second (Gbps). For example, a 12 megapixel (MP) camera operating at 60 frames per second (FPS) and 12 bits per pixel requires a bandwidth exceeding 8 Gbps.

[0036] Furthermore, multiple such cameras used in combination may require a total bandwidth of 10 - 100 Gbps, which can even impose a significant burden on Ethernet communication hardware. Additionally, stereo devices (e.g., stereo cameras 152, 156 of FIG. 1B) may require a significant distance between cameras or between a camera and computer infrastructure such as a server room or cloud - based computing, which makes high - bandwidth communication even more difficult when long cables and / or communication via the Internet are required. As described above, conventional video compression techniques such as MPEG technology may not be a suitable way to reduce bandwidth because there is a risk that the tracked object is removed by conventional video compression, especially when small objects (e.g., a distant golf ball) are being tracked. Thus, by using the high - bandwidth communication channel 305 (for video frames from one or more cameras), it becomes possible to receive high - resolution, high - bit - depth, and / or uncompressed image data as input to the object detection process (310).

[0037] In the received image frame, a point of interest is identified (315) (e.g., by computers 125, 150). For example, this may involve using image difference techniques to identify each position within the image frame that has one or more image data values that change by more than a threshold amount from a previous image frame. Additionally, other techniques are possible. For example, the process can look for groups of pixels of a particular luminance or color (e.g., white for a golf ball), look for a shape that matches the shape of the object being tracked (e.g., a round shape or at least an elliptical shape to find a round golf ball), and / or use template matching to search for an object (e.g., a golf ball) within the image.

[0038] Furthermore, looking for positions that have one or more image data values that change by more than a threshold amount from one image frame to another may involve applying image difference to discover pixels or groups of pixels that change by more than the threshold amount. For example, image difference can be applied to discover pixels in each image that change by more than a particular threshold, and groups of such changing pixels that are adjacent to each other can be discovered using, for example, known connected component labeling (CCL) techniques and / or connected component analysis (CCA) techniques. Such groups of pixels (and potentially single pixels as well) that meet the object detection criteria are called "blobs", and the position and size of each such blob can be stored in a list, and the list of all blobs within each image can be sent to the 2D tracking component. Converting the image to a list of object positions (or blobs) has a bandwidth reduction effect. In some cases, the bandwidth reduction for this operation can be 10:1 or more. However, as described in this document, further bandwidth reduction can be achieved, which can result in significant benefits when small objects are being tracked.

[0039] In the case of tracking small objects, since it is difficult to identify small objects (in some cases, single pixels in the image) based on the features of the objects, there is a serious problem of false detection. Therefore, the identification (315) (to detect the object of interest at a specific position in the camera image) can be implemented using a low threshold while allowing a large number of false positives and prioritizing zero false negatives. This method is generally counterintuitive in that false positives are often not preferred in object tracking, and thus, a conflict between minimizing false positives and minimizing false negatives is set. However, since the downstream processing is designed to handle a large number of false positives, the method of the present invention for object tracking can easily handle false positives. However, since object detection is designed to allow a large number of false positives, more objects, including many "objects" that are exactly noise in the image data, are identified (315) within each image frame 300, and thus, the bandwidth reduction effect of converting the image into a list of objects is partially offset.

[0040] A sequence of positions identified in an image frame is discovered (320) (e.g., by computers 125, 150, 160, 166). The processes shown in FIG. 3 (and other figures) are shown as sequential operations for ease of understanding, but in reality, the operations can be performed in parallel or simultaneously, for example, using hardware and / or operating system-based multitasking, and / or using pipeline processing techniques. Note that pipeline processing can be used for concurrent processing. For example, object identification (315) can start processing frame n+1 if available immediately after handing off frame n to 2D tracking (320) without having to wait for the downstream component to finish first. Thus, the disclosed content presented in this document in relation to the drawings is not limited to performing the operations sequentially as shown in the drawings, except when the processes executed on each computer are described as sequential processes, i.e., when the object identification process, 2D tracking process(es), and virtual time synchronization and 3D tracking process(es) occur sequentially. Each object identification and 2D tracking processing step reduces the bandwidth of the data sent to the downstream component.

[0041] Each of the discovered (320) sequences meets the motion criteria for positions identified in at least three image frames from the camera. In some implementations, the criteria are measured with respect to four or more frames and / or one or more criteria are used (e.g., the tree start criteria described below). Generally, 2D tracking aims to discover a sequence of objects (or blobs) over three or more frames that exhibits object movement consistent with the movement of an object in Newtonian motion that is not affected by forces other than gravity, bounce, wind, air resistance, or friction.

[0042] The reference for this object movement can be defined to include displacement, velocity, and / or acceleration in each dimension (x and y within the image) within a defined range of values. This range of values is set such that the 2D movement and acceleration of a moving object (e.g., a flying golf ball) as depicted by a 2D camera are well within the specified bounds, while more spasmodic movements are rejected (unless there are known objects that might bounce the tracked object). Further, a larger system can employ a secondary tracking step in downstream processing that can filter what constitutes the actual tracked object (e.g., a golf shot) at a finer granularity, so that discovery (320) need not be a complete (or nearly complete) filter that permits only the movement of actual objects such as the movement of a golf ball after it is hit from the tee area 112.

[0043] Rather, by intentionally making the filtering done in discovery (320) an incomplete filtering, objects other than moving objects can be included in the discovered sequences (including potential sequences of noise that are misidentified as the object of interest (315) and then misdiscovered (320) to form sequences). In other words, discovery (320) can implement a loose filter that increases false positives and minimizes false negatives, such that, for example, all or nearly all moving golf balls are permitted as forming valid sequences in discovery (320).

[0044] This looser (erring on the side of caution) approach means that a much simpler tracking algorithm can be used at (320), understanding that it does not need to be perfect when distinguishing a desired object (e.g., a golf ball) from an unwanted object (e.g., a non-golf ball). The set of rules that define the tracking can be minimized, and any errors made by 2D tracking (such as when passing a non-golf ball) can be removed by downstream components and processing. Instead of emitting an entire trajectory path that has one start point and one end point each, the discovered (320) sequences can be represented as a "rooted tree", where each vertex (node in the tree) is an observation blob (x, y, and time t), and each edge is a possible movement between the positions of the object whose movement is being tracked. As will be explained in more detail in connection with FIG. 4A, each branch can further have some metadata, such as the full depth of the tree.

[0045] However, even with this looser (erring on the side of caution / lower threshold) approach, it is still possible to miss object detections. Therefore, dummy observations can be used to account for objects that should be in the image data but have not been identified. In some implementations, if a sufficiently good blob that can extend the path is not found, the 2D tracker can add a dummy observation at the predicted position. The dummy observations can be implemented with significant penalty scores, and in some implementations, the dummy observations will not be permitted unless the graph is already at a certain depth. Since there are limits on how much penalty a branch can have, there are effectively limits on how many dummy observations a path can have.

[0046] As described above, discovery (320) may involve forming a rooted tree from a point of interest, where each rooted tree is a connected acyclic graph having a root node that is the root of the tree, and all edges of the connected acyclic graph originate directly or indirectly from the root. FIG. 4A shows an example of a process for discovering a sequence of object positions that satisfy a motion criterion by forming a rooted tree. At (400), the next set of positions identified for an image frame are obtained for processing, and this processing continues while the point of interest remains within the set for the current frame (405). For example, when a blob frame is processed, all blobs in the new frame can be matched to all tree nodes added during the processing of the previous frame to see if the blob could be a possible continuation of its path, depending on how well the points within this branch appear to have the desired motion as defined by the motion criterion.

[0047] The next point of interest is searched for in the set (410), and a check is made (415) to determine whether this point of interest satisfies a tree start criterion. If the point of interest satisfies the tree start criterion, a new root node for the tree is established using this point of interest (420). For example, if the image data value at the point of interest is greater than the minimum object size, this can be used to indicate that the ball is close to the camera and a new tree should be established. FIG. 4B shows this visual example, where six blobs 460 are observed, but only four of these blobs 460 are large enough to be used to establish a new root node 465. Note that one blob observation can be added to several trees, and all observed blobs could, in theory, be the start of a new object observation. This can lead to a combinatorial explosion in a noisy environment, and thus, in some implementations, some additional constraints (such as a minimum blob size constraint) are imposed before a new tree is established.

[0048] In this example, all blobs larger than a specific minimum size are promoted to a new tree at depth 0, and the minimum size limit can be set based on the camera resolution and the noise level of the incoming video. For example, the criterion can be that a blob must be at least 2 pixels (not just a single pixel) in order to establish a new tree. Note that truly random noise affects pixels individually and it is very rare to generate larger clusters in an image. Other techniques for avoiding combinatorial explosion are also possible. For example, the noise threshold in the blob generation (i.e., position identification) process can be adjusted adaptively.

[0049] Furthermore, the minimum tree depth required for a path to be exported can be increased, and since it is rare for a random noise sequence to accidentally construct a longer sequence, the amount of data exported is reliably limited. Since the constraints are very loose, many graphs with depth 0 and depth 1 are generated, but because the predictions are better, most of these graphs never reach depth 2, and even fewer graphs reach depth 3. Thus, since a tree or branch is discarded when there are no more nodes that can be added, the statistics are valid for the benefit of the systems and techniques of the present disclosure.

[0050] Furthermore, when the camera is placed near the target with the camera facing back towards the object's launch area, i.e., when detecting an incoming object, different size thresholds can be applied to different parts of the image. For example, in the part where a distant ball is found (e.g., a rectangle), a minimum size of 1 pixel can be used, and in other parts, a minimum size of 2 pixels can be used. Thus, the criterion used to form a tree of possible object paths can be determined based on the position of the camera relative to the launch area.

[0051] For example, if a somewhat flat ground can be inferred and the camera position and aiming are known, all ball detections below the "horizon" of the ground can be "sanity" checked against an estimated value of the maximum distance (and thus minimum size) at which a ball can exist. The ray from the camera to the ball detection intersects the ground at a certain distance D. The ball must be located at distance D or closer; otherwise, the ball would be underground. For random ball detections above the "horizon", there is no simple distance heuristic since there is no intersection with the ground. However, if there is knowledge about the limits of the 3D space (e.g., the arena), some general and / or angle-dependent maximum distance / minimum object size constraints can be used, taking into account that the object must be within the 3D region of interest.

[0052] Returning to FIG. 4A, a check (425) is made to determine whether the current point of interest is within the distance threshold of the root node established for the position identified within the previous image frame. If the point of interest is within the distance threshold, the point of interest is added (430) below the root node of this tree as a first depth sub-node. Note that if the tree consists of only one vertex (depth = 0), there is no way to estimate the speed of that path yet, so it must match any blob within a reasonable distance from the vertex. FIG. 4C shows a visual example following FIG. 4B. Only two of the four observed blobs 470 are within the distance limit 472 of the root node 474 from the previous frame. Thus, only two of the blobs 470 are added to the rooted tree as first depth sub-nodes 476. In some implementations, the distance limit depends on the resolution and frame rate of the incoming sensor, as well as the maximum expected speed of the object being tracked (perpendicular to the camera's aiming direction). If the size of the object being tracked is known in advance, as in the case of a golf ball, this can be taken into account when determining the distance limit / threshold 472. Additionally, larger blobs can be estimated to be closer to the camera than smaller blobs, i.e., if there are larger blobs (either in the previous frame or the current frame, depending on the camera's placement relative to the expected direction of the object being tracked), larger movement between frames can be tolerated.

[0053] However, when matching a blob to a vertex at a depth of 1 or greater, it is possible to predict where the next point in the path should be. In some implementations, the predicted area of the next blob is calculated by assuming that the speed has been the same since the last frame. Figure 4D shows a visual example of this. In addition to the position of the root node 480 and the position of the first depth sub-node 482, a predicted position 484 is calculated based on the known frame rate of the camera. Then, a search area 486 (or region) around the predicted position 484 is determined for use in searching for the blob to add to the tree. Similar to the details shown above for determining the distance limit / threshold 472, the size of the area / region 486 around the predicted position 484 can be determined based on the resolution of the camera and the maximum predicted speed of the object being tracked (perpendicular to the aiming direction of the camera), and the maximum predicted speed of the object can be adjusted according to the known size of the object and the size of the blob in different image frames. Further, in some implementations, for example, when the degree of the graph is large enough, a Kalman filter can be used when creating the prediction.

[0054] Any blob that is close enough to the predicted position can be transformed into a vertex and added to the tree. Thus, as shown in Figure 4D, one blob is within the region 486 of the predicted position 484, and thus this blob is added as a new sub-node 488. Additionally, in some implementations, a penalty can be accumulated on the branches of the tree, and the penalty is added to the score used to rank the branches, which can be useful in determining the best branch to export. Further, in some implementations, the penalty applied to this branch can increase as the position of the blob deviates from the predicted coordinates.

[0055] In some implementations, the penalties accumulated on the branches of the tree can be used to limit the magnitude of the allowed penalties and to determine when a branch should be discarded. In the former case, if a branch already has a high penalty, it is not allowed to add new nodes such that the extended branch exceeds the limit. When the penalty is calculated from the discrepancy between the predicted and the actual positions, this is essentially a way to ensure that the acceleration is within limits; otherwise, the penalty would become too high. In the latter case, when there are several branches sharing the last three blobs, the deepest and smallest-penalty branch can be retained and the other branches discarded. Note that this type of penalty-based hypothesis tracking can be applied to 3D tracking (explained in more detail, for example, in relation to FIGS. 7 - 10B) in addition to 2D tracking.

[0056] Returning to FIG. 4A, the check (435) is made to determine whether the current focus position is within the region determined using the estimated velocity of the position used by the established parent sub-node of the tree. If the current focus position is within the region, the focus position is added to the tree (440) as a second or deeper-depth sub-node under the parent sub-node, as described above in relation to FIG. 4D. Additionally, in some implementations, as the depth of the tree increases, more advanced prediction algorithms can be used to obtain better predictions. By fitting the most recent nodes in a branch to a quadratic or cubic polynomial function, a good prediction of where the next point will be can be obtained and paths with clearly unreasonable changes in acceleration can be discarded. Further, in some implementations, observations can be provided to a Kalman filter and the model of the Kalman filter can be used to generate new predictions.

[0057] Also note that the same point of interest within a single image frame can be added as respective sub-nodes (at the first depth or a deeper depth) to two or more rooted trees that maintain a track of potential object paths through the image frames and that are potentially used to establish new root nodes for new trees. Further, the same point of interest within a single image frame can be added as a sub-node within a tree having two or more linkages back to respective parent nodes within the tree, where the tree includes the same sub-node having different depths depending on which linkage leads back to the root node.

[0058] Thus, even if many incorrect paths are generated and downstream processing can filter out these incorrect paths, there should still be a limitation as to when potential paths are output. In some implementations, until a rooted tree representing a sequence exceeds a predetermined tree depth (445), the sequence is not considered for output. Thus, the sequence of identified positions is not output until it is confirmed (450) for output based on whether the tree depth of the sequence exceeds a predetermined tree depth (445), as specified for a given implementation or determined on-the-fly based on one or more factors.

[0059] Note that only a portion of the rooted tree having a predetermined tree depth needs to be confirmed (450) for output as a sequence of identified positions. This helps prevent noise that generates incorrect paths in the rooted tree from propagating from 2D tracking in a pre-processing stage on a first computer to a post-processing stage on a second computer. Also, the minimum tree depth threshold generally depends on the environment and the acceptable latency for generating a 3D track of the object. In some implementations, the rooted tree of the sequence must have a tree depth deeper than 2 before the sequence is confirmed (450) for output. In some implementations, the tree must have a depth deeper than 3, 4, 5, 6, or 7 before the sequence is confirmed (450) for output.

[0060] For example, in some implementations, the vertices of the tree (corresponding to blobs / positions) are exported only if they are at a particular depth (D) within the tree or have children that are exported by children nodes having a particular depth (D) within the tree. A vertex with D = 5 has five edges between itself and the root. This means that the output may be delayed by the same number of frames D, as it is often not possible to know in advance whether a blob is part of a sufficiently long branch. Since only the exported blobs are the blobs that constitute the possible paths, the filtering described (by restricting the export of the nodes of the tree) can be proven to dramatically reduce the bandwidth required to communicate object candidates. The following table shows an example of the reduction of the bandwidth for golf applications based on the minimum branch depth required. [Table 1]

[0061] As can be seen from this table, most of all the notable blobs / positions are either not connected by any path or not connected by a path of sufficient depth, and thus most notable blobs / positions are rejected by the effective filtering performed in the preprocessing stage. Returning to FIG. 3, at 325, for any sequence of nodes within a tree having at least one node with a tree depth deeper than a predetermined tree depth, i.e., as described above, to determine whether the data for the sequence is ready for output, a check is made to determine whether the nodes of that sequence are ready for output. If the output is ready, the output data for one or more sequences of positions is sent (330) to one or more second computers (e.g., computers 140, 160) (e.g., by computers 125, 150, 166). However, in some implementations, the identified (315) positions are sent from the first computer (e.g., by computers 125, 150) to a second computer (e.g., computers 140, 160, 166), and the second computer discovers (320) a sequence of positions identified within the image frame and then outputs the sequence data to another process on the same or a different computer. In either case, the sequence output data includes, at least for each position within each sequence, the two-dimensional position of the position within a particular image frame having a time stamp, and the time stamp is required for virtual synchronization at the second computer.

[0062] One of the bases of 3D tracking is triangulation. However, in the triangulation of the position of an object seen by two cameras, the observations from the same time instance need to be aligned. A common way to achieve this is to use a common synchronization trigger signal (e.g., transmitted via a cable) for all cameras. However, this method is only effective when the cameras have the same capture frame rate. In some configurations, it is difficult or impossible to ensure that two cameras are synchronized (triggering image capture in a tandem manner), and thus, triangulation of observations from different cameras becomes difficult or impossible. To solve this problem, instead of trying to actually synchronize camera images with other sensor(s) at the time of data capture, timestamp information (combined with information regarding movement represented by a 2D path) is used for virtual synchronization in a second computer.

[0063] Furthermore, the systems and techniques described in this document are applicable to both global shutter cameras and rolling shutter cameras. A rolling shutter means that the timing of capture of each row within a camera image is different. Thus, when an object is imaged by a sensor, the position of the object can be determined, but the measurement time depends on where in the image the object is found. In some implementations, there can be a time difference of about 20 milliseconds between the top and bottom image rows. This also causes a triangulation problem as the requirement for simultaneous measurements in both cameras may not be met. To solve this problem, in addition to the frame timestamp, by taking into account the relative time of a specific time offset of the "position" within the frame, rolling shutter information is also taken into account, and thereby, the measured object position from a rolling shutter camera can also be made available for high-quality triangulation in a second computer during a post-processing stage. In the case of global shutter capture, this offset is always 0.

[0064] As described above in connection with Table 1, in order to adapt the tracks to the minimum path length (minimum tree depth) requirement, several data frames may be required, so there is a trade-off between introducing a delay to all blobs and implementing a more complex protocol that allows communicating those blobs as soon as they are recognized as already adapted. Thus, in some implementations, transmitting (330) involves a delayed dispatch of both frame and blob data, and in some implementations, transmitting (330) involves an incremental dispatch of frame and blob data.

[0065] In the case of delayed dispatch, transmitting (330) involves delaying the output of data for that attention position found in one or more of a given image frame and the sequence until any further attention positions identified for the given image frame are no longer included in either the sequence based on the attention positions identified in subsequent image frames. Thus, the export of a node for a frame can occur when the number of processed subsequent frames excludes the defined tree depth reached for any more blobs within the current frame. In other words, the transmission of a frame is delayed until it is determined that any further blobs identified within the frame are not included in the path within the rooted tree having the required minimum tree depth.

[0066] Furthermore, in some implementations, a dynamic minimum tree depth may be used. The minimum tree depth may initially be set low to minimize latency when there are few paths, and then may be increased dynamically when the load (total output of limited paths per unit time) exceeds a maximum load threshold. This can result in a kind of throttling effect that improves performance by adjusting the processing in response to the current input data.

[0067] Furthermore, the frame data structure may include a list of blobs detected within this frame and passing the filter criteria. For example, the frame data structure may include the data fields shown in Table 2 below. [Table 2] Furthermore, the blob data structure included in the frame data structure may include the data fields shown in Table 3 below. [Table 3] Note that these data structures can be encoded in various formats such as the JSON format as follows. [Number] In the case of transmission with limited bandwidth, a more compact encoding such as Google's Protocol Buffer format can be used.

[0068] When unfiltered blobs are passed between processing nodes, for example, when blob detection is performed on an edge device and pre-tracking and / or tracking are executed on different computers, since there is no information regarding the relationship between blobs in different frames at all, a simplified format can be used. For example, the frame data structure may include the data fields shown in Table 4 below. [Table 4] Furthermore, the SimpleBlob data structure included in the frame data structure may include the data fields shown in Table 5 below. [Table 5]

[0069] In the case of incremental frame dispatch, transmitting (330) involves outputting the data of the image frame when the identification (315) for each image frame is completed, and outputting (330) the data for each focus position only after finding (325) one or more of the sequences including the focus positions to be output. As described above, the sequence needs to have the tree depth required to export the node, and when the tree reaches the threshold depth, all the parent nodes of the tree nodes exceeding the threshold depth that have not been exported previously (by some other branch) are output retroactively (330). This is a more complex method for transmitting data that can be used when low latency is required. 2D tracking outputs a continuous stream having information about the processed frame and a list of all blobs that have passed through the ball candidate filtering (i.e., all identified (315) focus positions) as soon as the information becomes available.

[0070] For example, when a new frame is processed, the information about the frame can be transmitted using a frame data structure having the data fields shown in Table 6 below.

Table 6

Table 7

Table 8

[0071] Regardless of the data structure or output timing used, the output data 350 is received (360) from one or more first computers (e.g., from computers 125, 150) (e.g., by computers 140, 160) via a second communication channel 355 that couples the first computer(s) to the second computer(s). The second communication channel 355, which is the path through which the output data 350 is received, has a second data bandwidth that is narrower than the first data bandwidth of the first communication channel 305. For example, the second communication channel 355 can be a copper wire Ethernet or a wireless communication channel, e.g., communication channels 162, 164. Further, in an implementation where a part of the preprocessing (e.g., discovery 320) is executed on the same second computer (e.g., by computers 140, 160) as the additional processing operations 360 - 375, the received data includes the identified (315) location and associated timestamp information.

[0072] In either case, data reception (360) continues until check (365) indicates that the sequence is ready to be processed. Additionally, as described above, operations can be performed in parallel or simultaneously, for example, using hardware and / or operating system-based multitasking and / or pipeline processing techniques during construction (375). Thus, reception (360) can continue while processing (370) of one or more sequences is performed in parallel or simultaneously, and each component can start processing frame n+1 if available without waiting for the downstream component to finish first immediately after handing off frame n to the downstream component. For example, even if the 3D tracker has not finished processing the point of interest, stereo triangulation can be performed on the next frame.

[0073] The sequence(s) within the output data 350 are processed (370) by one or more second computers (e.g., by computers 140, 160) by interpolating between specified 2D positions within a particular image frame of the sequence(s) using the time stamp of the particular image frame, generating a virtual 2D position at a given point in time. The record of each point of interest within the output data 350 (e.g., each blob found within each frame) includes both the time stamp of the frame and the indication of one or more previous points of interest connected to this point of interest by the 2D tracking component (e.g., each blob is described by a pointer to one or more previous blobs to which it belongs). Thus, from this data 350, it is possible to determine the time, position, movement direction, and speed of each point of interest / blob.

[0074] Therefore, data 350 enables the use of interpolation to generate a virtual "intermediate" position / blob at any given point in the path as long as there is at least one notable position / blob within Kuchi with an earlier timestamp and at least one notable position / blob with a later timestamp. Additionally, in some implementations, a given point in time is one of a plurality of points in time at a defined constant frame rate of the 3D tracking server. For example, 3D tracking server computers 140, 160 can operate at a defined constant frame rate and use interpolation to generate virtual snapshots of all camera blobs at these points in time. Since all blob coordinates represent the same point in time, triangulation between points is possible even if the original captures are not synchronized. Further, in some implementations, a given point in time is the time specified by another sensor such as another camera sensor or a radar sensor.

[0075] Furthermore, as described above, the camera can be a rolling shutter camera, in which case the output data 350 can include time offset values for each notable position included in each sequence. Using this ongoing data, the process (370) also functions for virtual time synchronization with the rolling shutter camera. FIG. 5 is a diagram showing an example of a process of interpolating between specified 2D positions within a specific image frame obtained from a rolling shutter camera.

[0076] The first observation time for a first position having one of the specified 2D positions within a particular image frame is calculated (500) by adding a first time offset value for the first position to the timestamp of the first image frame of the particular image frame. For example, the time offset (dt) of the first blob within the first frame (as detailed above in Tables 3 and 7) can be added to the timestamp (t_sync) of that first frame (as detailed above in Tables 2, 3, 6, and 7). The second observation time for a second position having another of the specified 2D positions within a particular image frame is calculated (510) by adding a second time offset value for the second position to the timestamp of the second image frame of the particular image frame. For example, the time offset (dt) of the second blob within the second frame (as detailed above in Tables 3 and 7) can be added to the timestamp (t_sync) of that second frame (as detailed above in Tables 2, 3, 6, and 7). Next, interpolation is performed (520) using the first and second observation times calculated from the time offset and frame timestamps.

[0077] Returning to FIG. 3, when a virtual 2D position is generated, a 3D track of an object (e.g., a ball) moving in 3D space is constructed (375) (e.g., by computers 140, 160) using the generated virtual 2D position and position information obtained from at least one other sensor at a given point in time. Constructing (375) can be for the display (e.g., immediate display) of the 3D track of the object, or constructing (375) can generate a 3D track for use as an input to further processing prior to display. For example, the 3D track can be further processed so as to be effectively displayed by overlaying a trace of the golf ball on a TV signal generated for live transmission or recording. As another example, the 3D track can be further processed so as to be effectively displayed by indicating the path of the golf ball in a computer representation of the physical environment where the golfer is located, or in a virtual environment that exists only on the computer but is displayed to a user of the system, by extending other data or media.

[0078] Other types of further processing prior to display are also possible. For example, the 3D track can be further processed to determine the final resting position of the sports ball, which can be useful for a wagering application supplied to a sports website or for general statistical collection. In addition to indicating the trace, the 3D track can be further processed to measure the speed, spin, carry, and launch angle of the shot. The 3D track can be further processed to determine whether the ball has gone over the net of the site and into an adjacent building. The 3D track can be further processed to inform the range owner which bays have ongoing activity and to count the number of balls hit from each bay.

[0079] FIG. 6A is a diagram showing an example of a process for constructing a 3D track of a moving object (e.g., a ball). The 3D track construction includes (e.g., by computers 140, 160) combining the virtual 2D positions with position information obtained from at least one other sensor to form the 3D position of the object of interest (600). Generally, this involves triangulation of observations from different sensors, use of object observation data generated using those different sensors, and use of calibration data for those different sensors. Note that the generated observation data is available for triangulation because virtual time synchronization is performed by generating virtual 2D positions.

[0080] For example, the other sensor can be a second camera used as a stereo pair with the first camera for which the virtual 2D positions are generated. The object detection, 2D tracking, and virtual time synchronization techniques described above may be used with the second camera. In light of this, the output data for the second camera can generate a plurality of detected objects (having a plurality of corresponding virtual 2D positions) at the same common time point at which the virtual 2D positions are generated for the first camera. Thus, the position information for the second camera can be two or more 2D positions obtained from the second camera, and the combining (600) can include determining which of the two or more 2D positions from the second camera should be matched with the virtual 2D positions generated for the data from the first camera.

[0081] Furthermore, combining (600) may involve excluding (602) at least one but not all of the two or more 2D positions obtained from the second camera as those that cannot form a 3D point having a virtual 2D position obtained from the first camera. FIG. 6B shows an example of a process of excluding at least one 2D position obtained from the second camera to perform epipolar line filtering before triangulation (604). Using the virtual 2D position generated for the first camera, the optical center of the first camera, the optical center of the second camera, the baseline between the first camera and the second camera, and the external calibration data for the first camera and the second camera, a region around at least a portion of the epipolar line in the image plane of the second camera is determined (602A). In some implementations, the region is determined based on the error of the data being used (602A).

[0082] In some implementations, the baseline between the first camera 120 and the second camera 130 of the stereo camera pair is 30 meters or less, which, in combination with the ability to detect shots quickly enough to be able to estimate the trajectory back to where the shot was taken, provides reasonable stereo accuracy. By using lenses with a wider field of view, shots can be observed earlier, but instead, when the ball is far away, the accuracy is reduced. Higher resolution cameras can mitigate this problem to some extent. With a shorter baseline, the depth accuracy is reduced, but it becomes possible to observe shots earlier. Note that using the virtual time synchronization described in this disclosure, data from different pairs of cameras 120, 130 can be dynamically combined, for example, in real time, as needed. Thus, different pairs of cameras 120, 130 within the system can form stereo pairs, and thus, different stereo camera pairs can have different baselines and different depth accuracies.

[0083] Further, in some implementations, the determined (602A) region is reduced (or further reduced) (602A) based on known limitations on the distance to the object, e.g., a portion of the epipolar line that cannot be used due to the known distance between the camera and the tee area 112 and / or the target 114. Further, other criteria can be used when matching observations of ball candidates from different cameras used as one or more stereo pairs. For example, an additional criterion can be that the blob contrast of each blob observed by the two cameras (used as a stereo pair) is similar.

[0084] Next, the pairing of the virtual 2D position obtained from the first camera with each of the two or more 2D positions obtained from the second camera is rejected (602B) in response to each of the two or more 2D positions being outside of a region around at least a portion of the epipolar line in the image plane of the second camera. In short, a line in 3D space that appears as a single point in the 2D plane of the first camera to the first camera (since this line coincides directly with the optical center of the first camera) appears as a line in the image plane of the second camera to the second camera (which is known as the epipolar line). Given the known 2D position of an object observed by the first camera, only an object observed by the second camera that lies along (within a certain tolerance value) this epipolar line can be the same object observed by both cameras. Further, if the system is designed and set up using known limitations regarding what distance to the object should be considered (e.g., the distance based on the distance to the tee area 112, the distance to the target 114 such as the distance to the green when tracking a ball coming onto the green, and / or the distance to the currently tracked object(s)), the system can place a hard stop (boundary) on the epipolar line for use in rejecting (602B). Thus, objects that are clearly outside the boundary (e.g., airplanes, birds, light from distant traffic, etc.) can be easily ignored by the system.

[0085] Other techniques for excluding the 2D positions obtained from the second camera are also possible. For example, as described above, the output data may include a display of previous positions in the sequence for each position after the initial position in each sequence. Since this data is included (e.g., transmitting (330) includes not only the position of each blob but also the positions of previously observed blobs connected by 2D tracking within the path sequence), the possible 2D movement directions and speeds of each object of interest (blob) can be estimated, and this estimation can be used to perform motion filtering before triangulation (604). For example, a stereo triangulation algorithm can use this information to reject blob pairings with incompatible parameters, e.g., by using the velocity vectors of the left and right stereo points to exclude false positives. For example, if the positions of the objects in the next frame are predicted in both cameras (based on the current velocity), the predicted positions (when triangulated) will also be close to the next actual observations when the blob has actually observed the same object from two different angles.

[0086] Thus, the exclusion (602) can involve estimating the 2D velocity for an object at a virtual 2D position based on a specified 2D position within a particular image frame, where at least one of the specified 2D positions is identified in the output data using a display of a previous position in at least one of the sequences, estimating the 2D velocity for an object at two or more 2D positions obtained from the second camera, and rejecting the pairing of the virtual 2D position obtained from the first camera with each of the 2D positions of the two or more 2D positions obtained from the second camera based on the estimated 2D velocities for the object at the virtual 2D position and at the two or more 2D positions, as well as the internal calibration data and the external calibration.

[0087] As will be understood by those skilled in the art, it should be noted that the position / blob coordinates need to be undistorted using internal calibration data and further transformed into a compatible coordinate system using external calibration data. Internal calibration can be performed to find the optical distortion and true focal length of each camera. External calibration can be performed to find the direction of the cameras relative to each other. Both internal calibration and external calibration are referred to together as "calibration" or "calibration data" and can be used to triangulate the position of an object visible to both cameras.

[0088] Furthermore, the other sensor may include a radar device instead of a second camera. For example, in accordance with the systems and techniques described in U.S. Patent No. 10,596,416, one or more radar devices can be combined with a single camera, a stereo camera pair, or three or more cameras to form one or more hybrid camera / radar sensors to provide at least a portion of the sensors of an object detection system that registers a target object for 3D object tracking. Readings from the radar(s), which can be detection of a moving object (e.g., a golf ball), can be added to the 3D point cloud of the three-dimensional position of the target object and thus to the potential robustness and coverage area of the combined sensors.

[0089] The readings from the radar can be converted into 3D points using the following method. The distance measurements from the radar can be combined with all the ball candidates detected by one of the cameras at the same time, which can include combining the radar with the ball candidates detected by only one of the cameras of a stereo camera pair and / or each ball candidate detected by both cameras of a stereo camera pair. That is, one stereo camera and one radar device can operate as three separate sensors (stereo camera sensor, radar + left camera sensor, and radar + right camera sensor) that can generate 3D points of the ball candidates. Note that the angle for each ball candidate reported by the camera and the distance reported by the radar determine the 3D position in space. This is done for each ball candidate seen by the camera, generating an array of 3D points to add to the point cloud. Only one of these points will be the correct association between the radar data and the ball observation (i.e., they both originate from the same object and are true positives). The remaining 3D points will be incorrect associations (false positives). However, the 3D point tracking algorithm is selected to be robust against most of the 3D points that are false positives. Additionally, in some implementations, one or more radar sensors that can determine both the distance and the angle to each observed object can be used in the system.

[0090] However, in some implementations, regardless of whether position data from the second sensor is acquired and / or excluded, or how it is acquired and / or excluded, the 3D position of the object of interest can be triangulated (604) using the virtual 2D position obtained from the first camera, at least one 2D position out of two or more 2D positions obtained from the second camera, the internal calibration data for the first and second cameras (the determined optical distortion and true focal length of each camera), and the external calibration data for the first and second cameras (the determined orientation of the cameras relative to each other). Note that since all the position / blob coordinates represent the same point in time, triangulation (604) between points is possible even if the original captures are not synchronized.

[0091] Then, the 3D position of the object of interest can be added to other 3D positions of the object of interest within the cloud of 3D positions for a given point in time (610). Note that various types of data structures can be used. In some implementations, the cloud of 3D positions is stored as an octree, and this type of data structure allows any point in 3D space to be represented (up to the accuracy limits of the scalar number computer representation). Motion analysis is performed across multiple clouds of 3D positions (620) to construct a 3D track of a ball moving in 3D space, where each of the multiple clouds is for a single point in time. This motion analysis can be performed as a final "golf ball" classification that checks between hypothesis formation, hypothesis expansion, hypothesis branching, and / or whether the trajectory hypothesis adequately corresponds to the expected motion of a golf ball flying through 3D space.

[0092] Note that various types of motion analysis can be used. However, even if the number of false positives in the data 350 is greater than the number of conventional object tracking techniques, it should be understood that these false positives tend to spread in a large 3D space compared to the 2D image plane of the camera where the false positives occurred. For this reason, when a 3D tracking system identifies an object moving in 3D, it is easy for the 3D tracking system to easily discard any false positives in the 3D space that do not match the ongoing 3D track being constructed. Therefore, 3D tracking is robust enough to handle false positives included in the 3D cloud. Because (1) most of the false positives are randomly spread in 3D space and it is difficult for false positives to create a coherent trajectory (according to motion analysis), (2) when a false positive is close to the hypothesis of an actual object in flight and is added to the hypothesis of that object, 3D tracking can branch (split) the hypothesis when it receives one or more subsequent true observations of that object in flight that do not match the false positive, and the branch containing the true observation is much more likely to continue than the branch containing the false positive, which leads to the branch containing the false positive being discarded in most (but not all) cases, and (3) in some implementations, the tracking position is smoothed (e.g., after 3D tracking and / or using a Kalman filter or a low-pass filter) to increase the accuracy regarding the actual object in motion, but this also has the effect that false positives added during hypothesis formation differ by a random amount compared to actual observations and thus tend to be canceled out during smoothing, so that false positives added during hypothesis formation are effectively removed. In addition, note that hypotheses with coherent motion (typically, the trajectory of an object other than a golf ball) that include false positives can still be rejected by a "golf ball" classification (either before or after any smoothing) that checks whether the hypothesis adequately corresponds to the expected motion of a golf ball in flight in 3D space.

[0093] Furthermore, the process may include outputting a 3D track for generation of a representation of the 3D track (630), or generating and displaying a representation of the 3D track in 3D space (630). As described above, this may involve further processing the 3D track to effectively display the 3D track by overlaying a trace of the golf ball on a TV signal generated for live transmission or recording, or this may involve further processing the 3D track to augment other data or media, for example, in a computer representation of the physical environment where the golfer is located, or in a virtual environment that exists only on the computer but is displayed to a user of the system, such as a virtual environment shared among participants in a multiplayer game (locally within the same area or scattered worldwide), by showing the path of the golf ball.

[0094] FIG. 7 is a diagram illustrating an example of a process for 3D object tracking of a moving object using unvalidated detections registered by two or more sensors. A three-dimensional (3D) position of a target object registered by an object detection system is obtained (700). In some implementations, obtaining (700) involves generating the 3D position of the target object using the object detection system and techniques described above in connection with FIGS. 3 - 6B. In some implementations, obtaining (700) involves one or more other techniques for generating the 3D position of the target object. In some implementations, obtaining (700) involves receiving the 3D position of the registered target object from the object detection system and techniques described above and / or other object detection systems and techniques.

[0095] In either case, the object detection system is configured to defer a strict classification of detections into "true positives" and "false positives" until a later stage in the tracking algorithm, and thus, at this stage, tolerate a significant amount of false positives for the registered objects of interest to minimize false negatives. False positives can include both misdetections by individual sensors (e.g., detections from a single sensor when it is capable of determining the 3D position of the detection alone, such as a phased array radar having both angle measurement and distance measurement capabilities) and incorrect combinations of detections from respective sensors that can include two or more cameras, a camera and a radar device, or a stereo camera and a radar device. In some implementations, most of the registered objects of interest are false positives. In some implementations, most of the false positives can be 60% or more, 70% or more, 80% or more, 90% or more, 95% or more, or 99% or more of the registered objects of interest (generated by the object detection system).

[0096] Furthermore, the plurality of sensors 130 can be used to improve the robustness of the system for detecting a moving ball without overloading the 3D tracking algorithm. In some implementations, obtaining (700) involves generating the 3D position of a registered object of interest by creating different combinations of data from three or more sensors 130. The three or more sensors 130 can include three cameras 130, and creating each combination of detections from each pair of the three or more sensors for a current time slice can involve using two or three stereo pairings of the three cameras to combine detections of a single object of interest by the three cameras to generate two or three respective three-dimensional positions of the single object of interest. Thus, any combination of two of the three cameras can be used to form a stereo pair, and if one of the three cameras fails to detect an object in a particular time slice, the 3D tracking still has collaborative sensed 3D positions, and if all three cameras detect an object in a particular time slice, three separate 3D positions can be added to the 3D point cloud of the object at that time slice, in which case one of these three is likely to be better than the other two. Thus, this can add redundancy to the data acquisition and improve the 3D tracking.

[0097] In addition, the three or more sensors 130 may include two cameras 130 and a radar device 130, and creating each combination of detections from each pair of the three or more sensors for the current time slice may include combining detections of a single object of interest by the two cameras using stereo pairing of the two cameras to generate a first three-dimensional position of the single object of interest, and combining detections of a single object of interest by the radar device and one of the two cameras to generate a second three-dimensional position of the single object of interest. Similar to the case of three cameras, this can also add redundancy and improve 3D tracking. In some implementations, the data from the radar device is combined separately with the data from each of the two cameras to generate three separate 3D positions in 3D space. In some implementations, more sensors (e.g., cameras and radar devices) may be added to the system to further increase redundancy without overloading the 3D tracking. Further, the 3D tracking may be improved by this additional redundancy in that noise in the input 3D position data can be at least partially canceled out by the system's ability to select the best available observations to extend each hypothesis.

[0098] In either case, when the 3D position of the object of interest is obtained (700) (e.g., in real time, in batches, etc.), a working hypothesis is formed from the 3D position (710), and each hypothesis represents an object moving in 3D space. Forming a hypothesis of an object moving in 3D space (710) uses a slow filter applied to the 3D position of the registered object of interest. If the estimated 3D velocity vector for a particular object of interest approximately corresponds to an object moving in 3D space over time, the slow filter enables an association between the particular objects of interest among the registered objects of interest. Note that each velocity vector is an estimate of the direction in 3D space and the velocity (in that direction) for the object of interest (since it is not known at the time of hypothesis formation (710) whether the object is a ball or not). Further, the hypothesis may include at least some hypotheses that branch into two or more hypotheses of objects moving in 3D space, as will be described in more detail below.

[0099] The "coarseness" in this case is similar to that described above for 2D tracking in that the 3D tracker attempts to find a sequence of objects of interest in 3D space that matches the movement of the object in Newtonian motion without first considering other forces such as wind, air resistance, or friction, and in some cases even ignoring gravity, i.e., without evaluating whether the acceleration of the object exactly conforms to the physical characteristics of a golf ball in flight at this stage. Thus, verification that the apparent forces acting on the object match those of a golf ball in flight is deferred until a later stage of 3D tracking.

[0100] Furthermore, in some implementations, the 3D tracking component does not simply have cooperating 3D point clouds. For example, as described above, pre-tracking in 2D can be performed on data from one or more camera sensors 130. For example, a golf ball can be tracked independently in the 2D data from each camera as much as possible to generate several 2D tracks of the ball. Thus, the data provided to the 3D tracking process can include partial segments of the object trajectory in addition to single object detections. When attempting to match detections between different sensors, the partial trajectories (formed by pre-tracking) can be matched to each other (regardless of whether the partial trajectories completely overlap), and the 3D tracking process can compare and connect partial segments of the trajectories in 3D instead of (or in addition to) points in 3D as part of forming a hypothesis (710) of an object moving in 3D space. This approach can reduce the number of false ball detections, which can simplify one or more aspects of 3D tracking, but also requires a more complex image tracking process, which may not be desirable in some cases.

[0101] In some implementations, forming (710) involves checking (712) whether the number of object detections for a given hypothesis is greater than a threshold. If the number of registered detections for a given hypothesis is below the threshold (or simply less than the threshold and possibly equal depending on the threshold used), a first motion model in 3D space is used (714), and the first model has a first level of complexity. Also, if the number of registered detections for a given hypothesis is greater than the threshold (or, depending on the threshold used, greater than or equal to the threshold), a second motion model in 3D space is used (716), and the second model has a second level of complexity that is higher than the first level of complexity of the first model.

[0102] The first model forms the first stage of a process that uses a very loose filter on what constitutes an object in Newtonian motion. This is very fast, but allows for a wide range of hypothesis formation, since random false positives can easily form hypotheses at this first stage. However, this (combined with the number of false positives entering the hypothesis formation pipeline) makes it almost impossible to miss any actual object in 3D space that exhibits the desired motion. The second model is more complex than the first model in that some knowledge of 3D physical properties (not necessarily a complete 3D physical model) is used at the second stage to prevent the continued formation of false hypotheses.

[0103] Note that this approach is the opposite of conventional 3D tracking, which uses very strict filtering at the beginning of the tracking pipeline, burdening (e.g., overloading) the processing capabilities of systems further downstream in the tracking pipeline that have false positives. In contrast to this conventional approach, the systems and techniques of the present invention allow more information about the motion of an object to be used when filtering hypotheses that should not have been formed in the first place, for example, when combining information about movement and position in 3D space with other information, such as whether the object size information matches the triangulated distances from the sensor(s). That is, as described in this document, when filtering can be deferred to a later stage (s) of the tracking pipeline, more advanced filtering techniques can be used efficiently while minimizing the risk of missing an actual object in flight.

[0104] This method can also take on specific values when used with one or more sensors placed near the expected target of an object in flight. This is because the object decelerates within the field of view of such sensor(s) and its relative size increases as time passes. As will be explained, the system can potentially employ both a camera and a radar sensor, including data from a stereo camera which is potentially a single sensor within the system, or a combination of data from two separate (and selectable) cameras within the system. In the case of a golf shot, when one or more sensors are placed at the target (e.g., the green of a golf course), data from the radar sensor is most useful for tracking the ball in flight from the hitting position towards the target, while data from the camera sensor is not as useful for detecting this initial part of the ball trajectory when the ball is far from the sensor, depending on the selection of the lenses and baseline used for tracking with stereo vision. Thus, when not combined with a stereo vision system with two cameras having a wide baseline, radar can provide better tracking data for the flight segment, and a stereo camera system can provide better tracking data for the bounce and roll segments.

[0105] Similarly, when one or more sensors are placed at the hitting position (e.g., the tee), data from both the radar and stereo camera systems are useful for tracking the ball from the hit, even if the baseline of the stereo system is short, but the accuracy of the stereo vision system degrades faster than that of the radar system. If the baseline is short enough, the stereo vision system is likely to lose track of the 3D track of the ball when the ball cannot be detected in 3D space due to very low 3D accuracy. However, if the baseline is long enough, this can be avoided and the radar data may not be necessary for 3D object tracking.

[0106] Some system deployments can use sensors at the target (e.g., the green), the ball location (e.g., the tee), and one or more locations in between (e.g., along the fairway), and it should be noted that the type of sensor (e.g., radar vs. camera) and sensor configuration (e.g., the baseline of a stereo vision system of two cameras) selected for use at a particular location can be determined according to the expected nature of the golf shot in the vicinity of each sensor. If the bounce segment and roll segment of the ball's path in 3D space may be far from the sensors, additional cameras can be used with different optics (such as being able to zoom in more, having a longer focal length, etc.) to better view the bounce and roll.

[0107] Accordingly, a particular set of sensor types, configurations, and placements can be adjusted according to a particular system deployment, such as a combination of a hybrid camera / radar sensor unit near the tee (to provide high-quality tracking data for the flight segment) and a stereo camera unit near the target (to provide high-quality tracking data for the bounce segment and roll segment), where the stereo camera unit has two or more cameras with a spacing of at least 4 feet (about 121.91 cm) between each camera pair, for example, two cameras attached to a single tripod's horizontal beam. Each camera has a wide-angle optical system that provides a field of view of more than 90 degrees and can cover the entire golf green even in situations where space is limited. For a 4-foot baseline, the measurement error can be + / - 3 inches for a typical golf green. Additionally, the accuracy can be improved by adding one or more layers of computer vision when the ball has come to a complete stop. This triangulates the rest position based on the center of the ball observed by both cameras to increase the accuracy.

[0108] Furthermore, in some implementations, the first motion model in 3D space is a linear model that associates the first 3D position of a registered object of interest with a given hypothesis when extending the given hypothesis in substantially the same direction in which the first 3D position is predicted by the given hypothesis. For example, the first model can be a simple dot product check against a vector defined by three 3D points P1, P2, P3 representing the registered object of interest (potentially, one or both of P1 and P2 are the 3D positions predicted by the active hypothesis). The vector i→(i+1) (P2→P3) and the vector (i-1)→i (P1→P2) (or potentially, the dot product of the vector (i-1)→i (P1→P2) and the vector (i-1)→(i+1) (P1→P3)), if positive, i.e., if the angle between the previous velocity vector and the new velocity vector is less than 90 degrees, this can be considered to be close enough to Newtonian motion in 3D space to connect the specific 3D points P1, P2, P3 in the hypothesis formed. Also, in some implementations, the threshold angle between the previous velocity vector and the new velocity vector is a programmable parameter that can be experimentally adjusted for a given implementation after deployment.

[0109] Other types of first models are possible, but each such first model should be very simple as it facilitates rapid initial hypothesis formation and minimizes the risk of new objects in flight not being detected in 3D space. The system can be designed such that no hypothesis formed using only the first model is ever sent as an output, and thus it should be noted that some degree of extension of the hypothesis by the second model is required before the hypothesis can be output for verification of objects in flight. Further, forming a hypothesis (710) can include splitting the hypothesis into separate segments (718), and in some implementations, splitting (718) is only performed on a hypothesis that has been extended at least once using a second motion model in 3D space.

[0110] In some implementations, the second motion model in 3D space is a recurrent neural network (RNN) model trained to perform short-term prediction of motion in three-dimensional space. In some implementations, the second motion model in 3D space is a numerical physics model that employs numerical integration combined with numerical optimization. Other types of second models are possible, such as another machine learning model that maps observations directly to future positions, or a set of differential equations (more complex than linear models) that describe how a ball will fly in the future given the current observations. These differential equations can be solved analytically (analytical physics model), or numerically, potentially using one or more Kalman filters (viewed as low-pass filtering) to remove noise in the data as a backing for numerical optimization. Different models have different trade-offs with respect to performance, accuracy, and noise sensitivity. Thus, the particular type of second model used can depend in part on the particular deployment of the system, e.g., on the nature of the sensor(s) 130 used in the system, and whether those sensor(s) 130 are located on one, two, three, or all of the faces of the 3D space 110.

[0111] FIG. 8 is a diagram illustrating an example of a hypothesis formation process that can be used in the 3D object tracking process of FIG. 7. The 3D position of the object of interest registered by the object detection system for the next time slice is received (800). The 3D position can be stored in an octree or another data structure. The octree data structure enables a quick answer to queries such as "which points are x meters from point y?" Generally, the received 3D position data can be understood as a point cloud representing unvalidated objects in 3D space at a given time slice. Further, the time slice can be one of a plurality of time slices at a defined time increment, e.g., at a defined fixed frame rate of the 3D tracking server referred to above.

[0112] A check (805) is performed to determine whether an active hypothesis remains unprocessed using the 3D position data of the current time slice. Each hypothesis may include models of the past and future trajectories of the hypothesis. For each active hypothesis that remains unprocessed, a prediction can be made as to where in 3D space the next sensor registration of the potential moving object represented by the given hypothesis will be found, according to the number of registered detections for the given hypothesis and a threshold. Thus, a check (810) can be performed as to whether the number of object detections for a given hypothesis is greater than a threshold (similar to check 712 in FIG. 7). If the number of registered detections for a given hypothesis is below the threshold (or simply less than the threshold and possibly equal depending on the threshold used), the next 3D position of the object for the given hypothesis is predicted using a first model (815). If the number of registered detections for a given hypothesis is greater than the threshold (or greater than or equal to the threshold and possibly equal depending on the threshold used), the next 3D position of the object for the given hypothesis is predicted using a second model (820).

[0113] For example, for the first few observations of a hypothesis, a simple linear model can be used as the model for the flight of the ball. After a defined number of observations have been made, a recurrent neural network (RNN) (or an appropriate movement of quadratic complexity) can be used instead as the model for the flight of the ball. In some implementations, the defined number of observations is 10, but other numbers may be used. Generally, the defined number needs to be set to a value less than the number of sensor observations in the system corresponding to, for example, about 0.3 seconds of ball flight, which is near the edge before the ball starts to exhibit curvilinear motion, i.e., the number of observations typically required for the golf ball to start exhibiting curvilinear rather than linear motion.

[0114] In some implementations, the RNN only needs to perform short-term predictions that enable the use of less complex deep learning-based models, e.g., models architectures that are not very wide and / or deep, which leads to faster predictions. In the case of RNN implementations, the predictions can be made from the last state of the model and, even if not necessary, can predict several steps ahead. These predictions do not require updating the state of the model. To enable faster predictions, predictions for all hypotheses can be made in batches. In each batch, predictions for multiple hypotheses can be made simultaneously using vector instructions, single instruction multiple data (SIMD).

[0115] In some implementations, rather than using an RNN as the second model, for example, after 10 observations have been made, a more advanced physical model can be used. This physical model can be a conventional numerical physical model that uses numerical integration of measured aerodynamic properties such as the drag, lift, and spin decay of the ball. This physical model can be a full physical model or a reduced physical model that simplifies one or more aspects of the ball flight to reduce processing time and thus speed up the prediction of the next 3D position of the hypothesis. In some implementations, the state, velocity, and spin of the physical model are continuously updated during the hypothesis. This can be done by first selecting a subset of the observed 3D points towards the end of the hypothesis. These 3D points are low-pass filtered and 3D velocity and acceleration vectors can be calculated from the filtered data. Next, an optimization can be formulated using the selected physical model and solved for the 3D spin vector of the ball. Next, numerical integration can be applied to the estimated 3D velocity and spin vectors to reach the predicted position of the object at the current time slice.

[0116] By using the predicted location, a more rapid search of the 3D position data for the current time slice becomes possible, and any registered object of interest that can be used to expand a given hypothesis can be found. For example, the octree data structure containing all of the 3D positions of the objects of interest registered by the object detection system for the next time slice can be searched (825) to find a set of 3D positions within a defined distance from the predicted next 3D position. The octree can be queried for points close to the predicted position, and each 3D point can be determined to be "close" when it is within the defined distance of the predicted position. The use of a data structure such as an octree that enables an efficient (faster) search between points in the 3D point cloud for the current time slice (so as not to need to check all points in the 3D cloud) is an important part of achieving real-time performance even when the 3D point cloud contains many false positives. Note also that spatial data structures other than octrees, such as K-D trees, can be used to enable fast searches by these types of queries for points close to the predicted points for the current hypothesis.

[0117] Furthermore, in some implementations, each 3D point is determined to be "close" only when the 3D point is (1) within a defined distance of the predicted position and (2) indicates movement in the same general direction as the hypothesis. For example, each hypothesis may include an estimated value of the current direction in which the ball is moving (estimated velocity vector), the new direction vector can be calculated based on the same anchor point used to calculate the new potential 3D point and the ball direction, and then the 3D point is permitted only if the angle between the current estimated direction and the new direction is 90 degrees or less (i.e., a positive dot product). Note that even when a second model is used to predict the next 3D position for a given hypothesis, a simple linear motion model can be used to limit the 3D points considered for addition to a given hypothesis. This is because it provides a quick way to exclude bad 3D points from consideration for expanding the current hypothesis.

[0118] If the search (825) returns zero (0) hits (i.e., the set is an empty set), the predicted 3D positions for a given hypothesis can be used (830) to extend the given hypothesis through the current time slice, and then the process returns to checking (805) the next active hypothesis to be processed. If the search (825) returns one (1) hit or a number (#) of hits that is below a defined threshold number that controls the number of branches allowed for each hypothesis in each time slice (e.g., a branching threshold number greater than 2 if the number of branches is programmable), one or more of the discovered 3D points are used (835) to extend the given hypothesis (branch the hypothesis if necessary), and then the process returns to checking (805) the next active hypothesis to be processed.

[0119] If the set of 3D points returned by the search (825) contains a number of 3D points greater than the branching threshold number (#), the discovered 3D points are sorted (840) based on their proximity to the predicted next 3D position, and a sorted set can be formed. For example, the 3D points can be scored, e.g., according to the square of the distance L2 from the predicted position during the search (825), and then ranked (840). 3D points that are not close enough within the sorted set that exceed the defined branching threshold number can be removed (845). Then, a given hypothesis is branched (850) into two or more hypotheses within a group of hypotheses using two or more 3D points remaining in the sorted set. Then the process returns to checking (805) the next active hypothesis to be processed. Also, note that even if a set of hypotheses that are branches from the initial hypothesis can be stored as a group, each hypothesis branch can be tracked individually and returned to the branching point when that hypothesis is eliminated.

[0120] While there are hypotheses left to be checked (805), the process is repeated. However, although the flowchart is shown as sequential processing for each hypothesis, in some implementations, it should be understood that the hypotheses are checked in parallel. Thus, predictions (815, 820) can be made for all active hypotheses before the octree is searched (825) for 3D points close to the predicted positions. In any case, all 3D positions of the noted objects registered by the object detection system for the current time slice and not added to the existing hypotheses can be added to the list of inconsistent 3D points for further processing. Additionally, in some implementations, if hypotheses disappear rapidly (e.g., before the hypothesis grows to detect three or more, or four or more, or five or more objects), the 3D points (corresponding to sensor impressions) used in the disappeared hypotheses can be added to the list of inconsistent points to reduce the risk of missing new ball launches due to sensor impressions being "stolen" by the stagnant hypotheses.

[0121] Furthermore, note that regardless of the nature of the second model used (e.g., an RNN model or a numerical integration physics model), the update of each hypothesis can be delayed. When the work on the hypothesis in the current time slice is completed, the second model is updated. If a new point is added, that point is used to update the model. Otherwise, the model's position prediction is used instead. In either case, the update can be "delayed". When a new point is added, it is not necessary to immediately use that point to update the model. This is because raw 3D observations are somewhat noisy and can potentially confuse the model. This delayed work can be done as follows. The position used to update the model is delayed by a few frame numbers from the last added position. When a point is used to update the model, the point can be smoothed by the positions both before and after the point. In some implementations, a window size of 9 is used, which means that 4 points behind and 4 points in front are required before the model is updated. In such implementations, special processing is done for the first part of the trajectory.

[0122] 3D positions that do not match the existing hypothesis are used to detect (860) potential new balls and new hypothesis firings. This can involve using the octree(s) for the current time slice and the stored additional octrees for previous time slices. To reduce the number of potential firings (especially for firings in the same neighborhood), points that are closer to each other than 1 / 2 meter can be clustered, and the point with the lowest measurement error, such as the reprojection error for the stereo camera sensor, can be retained within the cluster. The remaining points are added to the octree. In some implementations, the octrees for two previous time slices are also retained.

[0123] In some implementations, clustering involves finding points that are less than a threshold distance (e.g., 2 meters or 1 meter) apart (for each point). These points are considered clusters, and if a point belongs to more than one cluster, the two clusters can be merged into one. Thus, a cluster is a set of points that are close together, and thus these points can correspond to the same object. To avoid generating more than one hit for the same object within the system, the set of points within each cluster is reduced to one point, and the metric used to do this can be the reprojection error, and the point with the lowest reprojection error is selected to be the point representing the cluster. In some implementations, a new point cloud is created using the individual points representing the identified clusters, and this new point cloud should (1) contain fewer points than the original point cloud for this time slice, and (2) can be used to detect hits (and only hits). Note that the new point cloud can be stored in a spatial data structure such as an octree to enable faster search and facilitate real-time hit detection.

[0124] Hit detection (860) can be designed to find a movement that is approximately linear (i.e., a straight line). In some implementations, this is done by iterating over all remaining 3D points and searching for 3D points that were close (e.g., within 5 meters) two time slices ago in the octree. Next, a line is calculated between these two points. Then, this line can be used to predict the position of the ball within the time slice between these two points, and can be verified by searching the octree in the previous time slice for 3D points near this predicted position. If all of these criteria are met, a new hypothesis containing these three points is created (860).

[0125] Furthermore, as described above, forming a hypothesis (710) may include dividing the hypothesis into separate segments (718). Over the duration of the movement of a ball in 3D space, for example, over the duration of a golf shot, the movement enters different phases or segments such as flight, a first bounce, a second bounce, roll, etc. The physical model may vary depending on which segment the shot is in. That is, the physical model / parameters used for the flight segment may be different from the physical model / parameters used for the roll segment, and the physical model / parameters used for the roll segment may be different from the physical model / parameters used for the skid segment. Additionally, there may be a practical advantage to switching the physical / prediction model and / or parameters (singular or plural) used depending on which segment the shot is in. For example, bounces typically have higher noise but a much less prominent Magnus effect, so the same algorithm is not used for both bounces and flight.

[0126] The process of dividing the movement of an object into such segments (718) may be performed after all extensions have been made to the currently active hypothesis for the current time slice, or after each hypothesis has been extended. For example, each time a new point is added to a hypothesis, a determination may be made as to whether this point starts a new segment. FIG. 9 is a diagram showing an example of a process (718) of dividing a hypothesis into separate segments during hypothesis formation that may be used in the 3D object tracking process of FIG. 7. Data points within a given hypothesis are identified (900), and that data point has a minimum vertical position.

[0127] For example, if a new 3D position Pt x,y,z is added to the hypothesis for time slice t, for two previous 3D positions P(t - 1) x,y,z and P(t - 2) x,y,z with respect to the Y component P(t - 1) y < P(t - 2) y and the Y component P(t - 1) y < Pt yChecks can be performed to determine whether this is the case. Here, the Y component refers to the vertical position of a 3D point in a 3D coordinate system, and it should be noted that the "Y" position refers to the coordinate axis that defines the height above / below the ground surface. Thus, a minimum value of the Y position in 3D space indicates the possibility of a collision with the ground. Other techniques can be used to find the minimum vertical position, but this technique has the advantage that the determination is very quick and easy. However, due to noise within the system, the number of false positives, and the risk that the ball may not be detected by any sensor at a given time slice (note that a 3D data point identified as the minimum vertical position in the hypothesis may be a predicted 3D point rather than a registered detection if the detection was not registered at this time slice), another check is needed to confirm that this represents a ground collision.

[0128] The respective vertical components of the estimated 3D velocity vectors before and after the data point are checked (930) for a given hypothesis. For example, the Y (vertical) component of the vector P(t - 2) x,y,z →P(t - 1) x,y,z and the Y (vertical) component of the vector P(t - 1) x,y,z →Pt x,y,z are checked to see if the first component is negative (or positive in the inverted coordinate axis, in either case, indicating the sign that the object is moving towards the ground after the data point P(t - 1) x,y,z ), and the second component is positive (or negative in the inverted coordinate axis, in either case, indicating the sign that the object is moving away from the ground after the data point P(t - 1) x,y,z ). If so, the data point is designated as a ground collision event (942), and if not, the data point is not designated as a ground collision event (944).

[0129] In some implementations, to overcome risks caused by noise within the system, two or more estimated velocity vectors are checked. Two or more estimated velocities, both before and after the minimum vertical position data point, can be averaged, and then the vertical (e.g., Y) component of each velocity vector average can be checked (930). This can delay the identification of a ground collision by a number of time slices equal to one less than the number of velocity vectors averaged after a potential ground collision, but this delay provides the corresponding advantage of more robust ground collision detection using noisy data. This facilitates the use of less sophisticated sensor technology in the system and thus can reduce the overall cost of the system.

[0130] Further, in some implementations, one or more time windows can be selected based on the noise level associated with the estimated 3D velocity vector, the shortest predicted flight time on one or both sides of the data point, or both the noise level and the shortest predicted flight time (910). Then, the average of the estimated 3D velocity vectors before and after the data point can be calculated using velocity vectors that fall within at least one time window on both sides of the data point that is a potential ground collision. In some cases, a single averaging window is used, and in some cases, two averaging windows are used, and the averaging window used before the data point can be a different length than the averaging window used after the data point.

[0131] Furthermore, although shown in the figure as being included in the segmentation process itself, the selection (910) is not to dynamically (e.g., in real time) select the time window(s) while the hypothesis is being formed, but rather the selection (910) is made periodically, e.g., outside of a processing loop or potentially only once for a given system configuration having a given set of sensors 130 disposed with respect to a given 3D space 110. For example, the window(s) can be selected (910) based on the nature of the sensor(s) 130 used in the system and whether those sensor(s) 130 are located on one, two, three, or all of the faces of the 3D space 110. Further, when more advanced sensor technology, e.g., more accurate sensors 110 and / or a higher data capture rate, is used in the system, additional checks of the vertical component of velocity can be used, such as checking whether the Y component of one or more velocity vectors is within a small range near zero to help confirm that the ball is in the roll segment, or checking whether the rate of change of the Y component of the rolling average of the velocity vector is within a small range near the Earth's gravitational constant G to help confirm that the ball is in the flight segment or the bounce segment.

[0132] In any case, if a ground collision is specified (942), a new segment is specified for the hypothesis, and a determination is made as to whether this new segment is a bound segment or a roll segment with respect to the previous segment. If the specified ground collision is the first ground collision specified for a given hypothesis (950), this ground collision divides the hypothesis into two segments, and the segment before the ground collision is classified as a flight segment (corresponding to, for example, the initial portion of a golf shot) (952). Next, a check (960) can be performed to determine the nature of the segment after the ground collision. In some implementations, the check (960) includes determining whether the angle between a first estimated velocity vector (e.g., an average velocity vector) before the specified ground collision and a second estimated velocity vector (e.g., an average velocity vector) after the specified ground collision is greater than a threshold angle (e.g., 20 degrees). If so, the segment after the ground collision is classified as a bound segment (962). If not, the segment after the ground collision is classified as a roll segment (964).

[0133] In some implementations, the threshold angle between the estimated velocity vectors is a programmable parameter that can be adjusted experimentally for a given implementation after deployment. In some implementations, if the previous segment is a roll segment, the next segment cannot be a bound segment. In this logic, a "small" bound is identified as rolling, but the purpose is not to propagate this information to the consumer of the tracking data, but rather to change the tracking behavior, as will be explained in more detail below. In some implementations, a roll segment can be a true roll or a skid depending on the rotational component of the ball as it moves along the ground. Finally, it should be noted that it is typical for a skid to be followed by a roll, and the system can track the ball until it stops (or at least until the speed of the ball is below a threshold), and can discard the detection of the stopped ball.

[0134] Returning to FIG. 7, hypotheses that are not further extended by an association made with at least one additional object of interest registered by the object detection system during formation (710) are deleted (720). For example, a hypothesis can be marked as extinct if a potential ball is missed for more time slices than x, where x is set to a different value depending on how many observations were made. In some implementations, x is set to 3 for hypotheses with fewer than 10 observations and 20 otherwise. Other values for x can be used for hypotheses with different numbers of observations, depending on the frequency of the system and sensor observations. Also, if all hypotheses within a group are extinct, the entire group is marked as extinct. In any case, over time, false detections and / or pairings (e.g., 3D positions generated from data from any two sensors, either a camera or radar, determined to be detections of potentially the same object) are randomly distributed over a relatively large 3D space, so false hypotheses are erased relatively quickly, and due to the size of the volume of the space, the likelihood of finding a pattern of false detections and / or pairings that matches the movement of a golf ball within the space is very low.

[0135] Some hypotheses will remain after deletion (720), and thus, by applying a complete three-dimensional physical model to the data on the 3D position of the registered objects of interest used in forming (710) at least one hypothesis that remains after deletion (720), at least one 3D track of at least one ball moving in three-dimensional space is specified (730). In some implementations, in each time slice, the best hypothesis (or a set of active hypotheses that meet defined criteria) is selected based on a scoring function that looks at hypothesis size and one or more characteristics of the sensor data (e.g., blob contrast), and then, using these one or more hypotheses, for example, based on the predicted average vertical component of the estimated 3D velocity vectors before and after the minimum and minimum vertical positions, after determining a more accurate 3D path according to the segmentation performed during forming (710), at least one 3D track of the moving ball is specified (730).

[0136] FIG. 10A is a diagram showing an example of a process for specifying a 3D track of a moving ball that can be used in the 3D object tracking process of FIG. 7. The data on the three-dimensional position of the registered objects used in forming at least one hypothesis can be generated by processing the sensor data to generate a more accurate 3D path using sensor observations registered by the object detection system at least in the flight segment (1000). This can include processing the sensor data to remove outliers, smooth the data, and subsequently recombine the processed sensor data to form 3D positions within a 3D path having much less noise than could be caused by the raw 3D combinations (used during hypothesis formation). For example, in the case of two cameras combined in a stereo pair, this includes performing triangulation again based on the smoothed 2D observations.

[0137] Next, the 3D path is analyzed to make a final determination as to whether it is a golf ball. A complete three-dimensional physical model can be applied to the generated data for at least the flight segment (1005). The complete 3D physical model can estimate the spin of the ball, and balls with too much spin can be discarded. If sensor observations (after processing (1000)) are found (1010) not to match the ball in flight according to the applied (1005) complete physical model, the data is rejected (1015) for output. If sensor observations (after processing (1000)) are found (1010) to match the ball in flight according to the applied (1005) complete physical model, the data is confirmed (1020) for output.

[0138] Furthermore, in some implementations, only paths that move smoothly (with small angular changes between each point) are permitted. This additional constraint is necessary when the object does not move smoothly (in the way a golf ball moves), although with low spin (low acceleration). Further, rather than permitting or rejecting individual segments within the path, the entire path can be permitted or rejected. Still, individual points can be rejected as being outliers. Additionally, in some implementations, portions of a path having an excessive number of outliers at the data points can be identified as bad and discarded.

[0139] Figure 10B shows a detailed example 1050 of a process for specifying a 3D track of a moving ball that can be used in the 3D object tracking process of FIG. 7. This example corresponds to the case of two cameras combined in a stereo pair, but the process can be more generally applied to combinations of data from any two suitable sensors. Tracking point 1055 is received by a post-hypothesis process 1060 that recombines the processed sensor data to form 3D positions within a 3D path that is much less noisy than the raw 3D combinations used during hypothesis formation. 2D outliers in the data from the left camera (for the ball track) in the stereo pair are removed (1062A), and 2D outliers in the data from the right camera (for the ball track) in the stereo pair are removed (1062B). The remaining 2D positions in the data from the left camera (for the ball track) are smoothed (1064A), and the remaining 2D positions in the data from the right camera (for the ball track) are smoothed (1064B). Note that the smoothing can involve low-pass filtering as described above.

[0140] Non-overlapping data points are removed (1066). Since outliers are removed independently for the two cameras (1062A, 1062B), it is possible that only one of the two camera observations of the ball is considered an outlier, which leads to one of the cameras having a detection while the other camera does not. Therefore, these points with only one detection are removed. Next, the remaining data points (for the ball track) from the left and right cameras are triangulated (1068). This can generate a new set of 3D points, which can then be used to form a final 3D trajectory that is a more accurate representation of the actual path of the object in 3D space (1070).

[0141] As described above, the process of FIG. 10B can be more generally applied to combinations of data from any two suitable sensors. Generally, when the tracking stage for a given ball is complete (1055), all detections for that ball can be used to refine the tracking trajectory (1060). Thus, position data coming later in the trajectory can be used to refine position data coming earlier in the trajectory. In the case of data from camera + radar, the camera data can be processed as described above, and the radar data can be processed using low-pass filtering and outlier detection to improve the quality of the radar data.

[0142] Furthermore, the final golf shot classification operation can calculate the spin and average angle change in at least the flight segment of the 3D path of the object to confirm that the object is a golf ball (1075). The characteristic that determines the trajectory of a golf ball is the effect of spin on a golf ball in free flight. Thus, estimating the spin rate of the ball from the observed (flight) segment of the ball can help determine whether the trajectory of the object is that of a golf ball, as non-golf ball objects may inappropriately generate a high spin rate. A reasonable value for the spin of a golf ball is 0 to 16,000 revolutions per minute (RPM), and some slack is included due to the inaccuracy of the estimated value.

[0143] Furthermore, if the estimated spin for an object that is not a golf ball is very low, this can be detected by an additional test of the average angular change (relating to smoothness of movement). Golf balls should have smooth movement in air, and thus objects that do not have this smooth movement in the flight segments of their trajectories should be discarded. For each point in the 3D path, a direction vector between that point and the previous point can be calculated. This direction vector can then be compared to the direction vector of the previous point in the 3D path, and the angle between the two vectors can be calculated. An arithmetic mean (average) can be calculated (1075) over two or more of these angles (e.g., all angles within a flight segment), and a large average value (e.g., exceeding 10 degrees) indicates that the object has a jagged (rapidly changing) movement that is not the expected movement of a golf ball.

[0144] In such a case, the tracked object is rejected as not being a golf ball. If the object is permitted as a golf ball (or another type of object being tracked in a given implementation), the 3D path of the ball is output (1080) for processing by another system / process and / or for display (e.g., for a video game or TV broadcast). Note that the 10-degree threshold for the average of the direction vectors between points is just an example, and different thresholds can be used in different implementations. In general, since the distance that a potential object moves between each detection point during free-flight motion is short, the higher the data capture rate, the smaller the threshold can be, as the threshold depends on the capture frequency of the sensor(s). In some implementations, the threshold for the average of the direction vectors is a programmable parameter that can be experimentally adjusted for a given implementation after deployment.

[0145] Returning to FIG. 10A, in some implementations, the processing is performed separately in each of the separate segments (1025). This processing (1025) can be part of the processing done to generate data for ball confirmation (1000). For example, based on the path segmentation performed during 3D tracking, the trajectory can be segmented, and in each segment, the observations can be smoothed and outliers removed. Then, recombination (e.g., triangulation) can be performed again based on the smoothed observations to generate a 3D path with less noise in the data. Note that the observations (e.g., 2D observations from a pair of cameras) are smoothed within the segments as opposed to smoothing over the entire trajectory. Otherwise, the bound points would be "smoothed" and their sharp edges would be softened. By performing smoothing separately on the pieces before and after the bound, the sharp edges of the bound points are preserved, and in some implementations, a more accurate 3D position for each bound point can be determined, for example, by extrapolating two smoothed segments, based on the 3D paths on both sides of the identified bound point. That is, note that since the finite sampling rate of the sensor means that the exact instant of the ball's bound may not be captured, looking for the intersection in the space between the (extrapolated) path segments before and after the bound can be the only way to estimate the actual bound position.

[0146] As another example, the processing (1025) can be performed separately in each of the separate segments after the data has been verified for output (1020). This processing (1025) can use different techniques in each segment to identify important points within the 3D path, such as the exact 3D position of each bound point and the exact 3D position of the ball's final rest point. The different techniques used in each segment can include different low-pass filtering techniques, different data smoothing techniques, and / or different ball motion modeling techniques.

[0147] In either case, the separate segments remain associated with each other as part of the same ball's trajectory, and the final path data utilized in the output may include data for each of the different identification segments, only a portion of an identification segment (e.g., only the flight segment), and / or important identification points of an identification segment (e.g., the 3D position of each bounce point and the 3D position of the ball's final rest point).

[0148] When the 3D path meets all the criteria necessary to confirm that it is the true object it is moving as, e.g., an actual golf ball shot, the final path data is transmitted as output. Referring again to FIG. 7, the detected 3D track of the ball is output (740) for generation of a representation of the 3D track or for generation and display (740) of a representation of the 3D track in 3D space. This may involve effectively displaying the 3D track by further processing the 3D track to overlay a trace of the golf ball on a TV signal generated for live transmission or recording, or this may involve effectively displaying the 3D track by further processing the 3D track to augment other data or media, e.g., in a computer representation of the physical environment where the golfer is located or in a virtual environment that exists only on a computer but is displayed to the system user, e.g., a virtual environment shared among participants in a multiplayer game (either locally in the same area or scattered around the world), by showing the path of the golf ball. Further, the processes shown in FIGS. 7 - 10B are shown as sequential operations for ease of understanding, but in reality, the operations may be performed in parallel or simultaneously, for example, using hardware and / or operating system-based multitasking, and / or using pipeline processing techniques. Thus, the present disclosure presented in this document in relation to the figures is not limited to sequentially performing the operations as shown in the figures, except when the process being executed requires sequential operation with respect to specific inputs and outputs.

[0149] Embodiments of the subject matter and the functional operations described in this specification can be implemented in digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed in this specification and their structural equivalents, or combinations of one or more of them. Embodiments of the subject matter described in this specification can be implemented using one or more modules of computer program instructions encoded on a non-transitory computer-readable medium for execution by, or to control the operation of, a data processing apparatus. The non-transitory computer-readable medium can be a product, such as a hard drive in a computer system, an optical disk sold through a retail channel, or an embedded system. The non-transitory computer-readable medium can be acquired separately and later encoded with one or more modules of computer program instructions, such as by delivery of one or more modules of computer program instructions via a wired or wireless network. The non-transitory computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, or a combination of one or more of them.

[0150] The term "data processing apparatus" encompasses, by way of example, all apparatus, devices, and machines for processing data, including programmable processors, computers, or multiple processors or computers. The apparatus can include, in addition to hardware, code for creating an execution environment for the computer program, such as code constituting processor firmware, protocol stack, database management system, operating system, runtime environment, or a combination of one or more of them. Further, the apparatus can adopt various different computing model infrastructures, such as web services, distributed computing infrastructure, and grid computing infrastructure.

[0151] A computer program (also known as a program, software, software application, script, or code) is written in any suitable form of programming language, including compiled or interpreted languages, declarative or procedural languages, and can be deployed in any suitable form, for example, as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. The program can be part of a file that holds other programs or data (e.g., one or more scripts stored within a markup language document), a single file dedicated to the program, or multiple cooperating files (e.g., files that hold one or more modules, subprograms, or portions of code). A computer program can be deployed to be executed on one computer or on multiple computers at one location or distributed across multiple locations and interconnected by a communication network.

[0152] The processes and logical flows described herein can be executed by one or more programmable processors that execute one or more computer programs to perform functions by operating on input data and generating output. Further, the processes and logical flows can be executed by dedicated logic circuitry, such as an FPGA (Field Programmable Gate Array) or ASIC (Application Specific Integrated Circuit), and the apparatus can be implemented as dedicated logic circuitry, such as an FPGA (Field Programmable Gate Array) or ASIC (Application Specific Integrated Circuit).

[0153] Processors suitable for the execution of a computer program include, by way of example, both general and special purpose microprocessors, and one or more processors of any kind of digital computer. In general, a processor receives instructions and data from a read only memory or a random access memory or both. Essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. In general, a computer may further include one or more mass storage devices (such as, for example, magnetic disks, magneto-optical disks, or optical disks) for storing data, or may be operatively coupled to a mass storage device for receiving data from or transferring data to the mass storage device, or both. However, a computer need not have such devices. Further, a computer may be embedded in another device, such as, by way of example, a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device (such as, for example, a universal serial bus (USB) flash drive). Devices suitable for storing computer program instructions and data include, by way of example, all forms of non-volatile memory, media, and memory devices, including semiconductor memory devices (such as, for example, EPROM (erasable programmable read only memory), EEPROM (electrically erasable programmable read only memory), and flash memory devices), magnetic disks (such as, for example, internal hard disks or removable disks), magneto-optical disks, CD-ROM and DVD-ROM disks, network attached storage, and various forms of cloud storage. The processor and the memory may be supplemented by, or incorporated in, special purpose logic circuitry.

[0154] To provide interaction with a user, embodiments of the subject matter described herein can be implemented on a computer having a display device (e.g., an LCD (liquid crystal display), an OLED (organic light emitting diode), or other monitor) for displaying information to the user, and a keyboard and a pointing device (e.g., a mouse or trackball) that are means for the user to input into the computer. Other types of devices can also be used to provide interaction with the user. For example, feedback to the user can be sensory feedback, such as visual feedback, auditory feedback, or tactile feedback, and input from the user can be received in various forms including acoustic input, voice input, or tactile input.

[0155] A computing system can include a client and a server. The client and server are generally remote from each other and typically interact via a communication network. The relationship of client and server arises by computer programs running on respective computers and having a client / server relationship to each other. Embodiments of the subject matter described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or middleware components (e.g., an application server), or front-end components (e.g., a client computer having a graphical user interface or a web browser through which a user can interact with an implementation of the subject matter described herein), or any suitable combination of one or more such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any suitable form or medium of digital data communication, e.g., a communication network. Examples of communication networks include local area networks (LANs) and wide area networks (WANs), the Internet (e.g., the Internet), and peer-to-peer networks (e.g., an ad hoc peer-to-peer network).

[0156] This specification includes many implementation details, but these should not be construed as limiting the scope of the invention or what may be claimed, but rather as descriptions of features specific to particular embodiments of the invention. The specific features described herein in the context of separate embodiments may be implemented in combination in one embodiment. On the other hand, the various features described in relation to a single embodiment may be implemented separately in multiple embodiments or in any suitable sub-combination. Furthermore, features are described above as acting in a particular combination and may initially be claimed as such, but one or more features from the claimed combination may, in some cases, be removable from the combination, and the claimed combination may be directed to a sub-combination or a variant of a sub-combination. Accordingly, unless explicitly stated otherwise or as would be apparent to one of ordinary skill in the art, any of the features of the above-described embodiments may be combined with any of the other features of the above-described embodiments.

[0157] Similarly, operations are shown in the drawings in a particular order, but it should not be understood that such operations must be performed in the particular order or in a sequential order shown in order to achieve the desired result, or that all of the illustrated operations must be performed. In certain situations, multitasking and / or parallel processing may be advantageous. Furthermore, the separation of the various system components in the above-described embodiments should not be understood as necessary in all embodiments, and it should be understood that the program components and systems described may generally be integrated together in a single software product or packaged into multiple software products.

[0158] The specific embodiments of the present invention have been described above. Other embodiments are within the scope of the following claims. For example, the above description has focused on tracking the movement of a golf ball, but the systems and techniques described are applicable to tracking the movement of other types of objects, such as baseballs or skeet shooting, as well as uses other than sports. Further, the tracking of a "moving" object may, in some implementations, include tracking the object when it bounces from the ground and / or rolls along the ground.

Claims

Claim 1 A method comprising: obtaining, by an object detection system configured to allow more false positives and minimize false negatives for a registered object of interest, a three-dimensional position of the registered object of interest; using a filter applied to the three-dimensional position of the registered object of interest to form, in three-dimensional space, a hypothesis of an object moving in the three-dimensional space, wherein the filter enables an association between specific objects of interest among the registered objects of interest when an estimated three-dimensional velocity vector for a specific object of interest substantially corresponds to an object moving in the three-dimensional space over time; using at least one additional object of interest registered by the object detection system to remove, during the forming, an appropriate subset of hypotheses that are not further extended by the performed association; applying a complete three-dimensional physical model to data about the three-dimensional positions of the registered objects of interest used in forming at least one hypothesis remaining after the removing to specify at least one three-dimensional track of at least one ball moving in the three-dimensional space; outputting for display at least one three-dimensional track of at least one ball moving in the three-dimensional space; A method comprising: Forming a hypothesis of an object moving in three-dimensional space using the filter includes: when the number of registered detections for a given hypothesis is less than a threshold, a first motion model in three-dimensional space, wherein the first motion model has a first complexity level; when the number of registered detections for the given hypothesis is greater than or equal to the threshold, using a second motion model in three-dimensional space, wherein the second motion model has a second complexity level higher than the first complexity level of the first motion model; A method comprising. Claim 2 The first motion model in the three-dimensional space is a linear model that associates the first three-dimensional position with the given hypothesis when the first three-dimensional position of the registered object of interest extends the given hypothesis in substantially the same direction predicted by the given hypothesis. The second motion model in the three-dimensional space is a recurrent neural network model trained to perform short-term prediction of motion in the three-dimensional space. The method according to claim 1.

3. Forming a hypothesis of an object moving in a three-dimensional space includes using the first motion model or the second motion model to predict the next three-dimensional position of the registered detection for the given hypothesis according to the number and threshold of the registered detections for the given hypothesis; searching a spatial data structure that includes all three-dimensional positions of the object of interest registered by the object detection system for the next time slice to find a set of three-dimensional positions within a defined distance of the predicted next three-dimensional position; if the set of three-dimensional positions is an empty set, using the predicted next three-dimensional position for the given hypothesis; if the set of three-dimensional positions includes only one three-dimensional position, using the one three-dimensional position for the given hypothesis; if the set of three-dimensional positions includes two or more three-dimensional positions, sorting the two or more three-dimensional positions based on their proximity to the predicted next three-dimensional position to form a sorted set; removing any non-proximate three-dimensional positions within the sorted set that exceed a defined threshold number greater than 2; and using the two or more three-dimensional positions remaining within the sorted set to branch the given hypothesis into two or more hypotheses within a group of hypotheses. The method according to claim 1.

4. Forming a hypothesis of an object moving in a three-dimensional space, when the number of registered detections for the given hypothesis is greater than or equal to the threshold, includes identifying data points in a given hypothesis having a minimum vertical position; checking the respective vertical components of the estimated three-dimensional velocity vectors before and after the data points in the given hypothesis; if the first vertical component of the respective vertical components is negative and the second vertical component of the respective vertical components is positive, designating the data point as a ground collision. The method according to claim 1, comprising

5. Each of the respective vertical components is obtained from the average of the estimated three-dimensional velocity vectors before and after the data point in the given hypothesis, and the method selects at least one time window based on the noise level associated with the estimated three-dimensional velocity vector, the shortest predicted flight time on one or both sides of the data point, or both the noise level and the shortest predicted flight time; calculates the average of the estimated three-dimensional velocity vectors before and after the data points falling within the at least one time window; The method according to claim 4, comprising

6. Forming a hypothesis of an object moving in three-dimensional space includes, when the number of registered detections for the given hypothesis is greater than or equal to the threshold, dividing the given hypothesis into separate segments including a flight segment, one or more bound segments, and a roll segment, and the dividing includes the identifying, the checking, and the designating, classifying a first segment of the given hypothesis before the first designated ground collision as the flight segment; when the angle between a first estimated velocity vector before a next designated ground collision and a second estimated velocity vector after the next designated ground collision is greater than a threshold angle, classifying a next segment of the given hypothesis after each next designated ground collision as one of the one or more bound segments; when the angle between a first estimated velocity vector before a next designated ground collision and a second estimated velocity vector after the next designated ground collision is less than or equal to the threshold angle, classifying a next segment of the given hypothesis after the next designated ground collision as the roll segment; The method according to claim 4, comprising

7. The given hypothesis is at least one hypothesis remaining after the deleting, and designating at least one three-dimensional track of at least one ball moving in the three-dimensional space generates data on the three-dimensional positions of the registered objects, which is used to form the at least one hypothesis by triangulating a more accurate 3D path using at least the sensor observations registered by the object detection system in the flight segment; Identifying at least one moving ball by applying the complete three-dimensional physical model to the generated data for at least the flight segment, and The method according to claim 6, comprising.

8. The obtaining includes receiving or generating a three-dimensional position of the registered object of interest, and most of the registered objects of interest are false positives including false detections by individual sensors and incorrect combinations of detections from each of the individual sensors. The method according to claim 1.

9. The individual sensors include three or more sensors, and the obtaining includes generating the three-dimensional position by creating each combination of detections from each pair of the three or more sensors for the current time slice. The method according to claim 8.

10. The three or more sensors include two cameras and a radar device, and the creating includes Combining detections of a single object of interest by the two cameras using stereo pairing of the two cameras to generate a first three-dimensional position of the single object of interest; and Combining detections of the single object of interest by at least one of the radar device and the two cameras to generate a second three-dimensional position of the single object of interest. The method according to claim 9, comprising.

11. The three or more sensors include an additional camera, and the creating includes using stereo pairing of the additional camera with each of the two cameras to combine detections of the single object of interest by the additional camera and the two cameras to generate a third three-dimensional position of the single object of interest and a fourth three-dimensional position of the single object of interest. The method according to claim 10, comprising.

12. A system, Two or more sensors and one or more first computers of an object detection system configured to register an object of interest and to allow more false positives and minimize false negatives for the registered object of interest; One or more second computers configured to execute operations, the operations including Obtaining a three-dimensional position of an object of interest registered by the object detection system; Using a filter applied to the three-dimensional position of the registered object of interest to form a hypothesis of an object in three-dimensional space that is moving within the three-dimensional space, wherein the filter enables an association between specific objects of interest among the registered objects of interest when an estimated three-dimensional velocity vector for a specific object of interest substantially corresponds to an object moving within the three-dimensional space over time, and Using at least one additional object of interest registered by the object detection system to delete an appropriate subset of hypotheses that are not further extended by the association made during the forming, and Applying a complete three-dimensional physical model to the data on the three-dimensional positions of the registered objects of interest used in forming at least one hypothesis remaining after the deleting to specify at least one three-dimensional track of at least one ball moving within the three-dimensional space, and Outputting for display at least one three-dimensional track of at least one ball moving within the three-dimensional space, and comprising, Forming a hypothesis of an object moving within three-dimensional space using the filter when the number of registered detections for a given hypothesis is less than a threshold, a first motion model in three-dimensional space, wherein the first motion model has a first complexity level, and when the number of registered detections for the given hypothesis is greater than or equal to the threshold, using a second motion model in three-dimensional space, wherein the second motion model has a second complexity level higher than the first complexity level of the first motion model, and a second computer including, A system comprising.

13. The system according to claim 12, wherein the two or more sensors comprise a camera and at least one radar device.

14. The system according to claim 12, wherein the two or more sensors include a hybrid camera / radar sensor.

15. The first motion model in the three-dimensional space is a linear model that associates the first three-dimensional position with the given hypothesis when the first three-dimensional position of the registered object of interest extends the given hypothesis in substantially the same direction predicted by the given hypothesis. The second motion model in the three-dimensional space is a recurrent neural network model trained to perform short-term prediction of motion in the three-dimensional space. The system according to claim 12.

16. Forming a hypothesis of an object moving in a three-dimensional space involves using the first motion model or the second motion model to predict the next three-dimensional position of the registered detection for the given hypothesis according to the number of registered detections and the threshold for the given hypothesis, searching a spatial data structure that includes all three-dimensional positions of the object of interest registered by the object detection system for the next time slice to find a set of three-dimensional positions within a defined distance of the predicted next three-dimensional position, using the predicted next three-dimensional position for the given hypothesis if the set of three-dimensional positions is an empty set, using one three-dimensional position for the given hypothesis if the set of three-dimensional positions includes only one three-dimensional position, if the set of three-dimensional positions includes two or more three-dimensional positions, sorting the two or more three-dimensional positions based on proximity to the predicted next three-dimensional position to form a sorted set, removing any non-proximate three-dimensional positions within the sorted set that exceed a defined threshold number greater than two, and using the two or more three-dimensional positions remaining within the sorted set to branch the given hypothesis into two or more hypotheses within a group of hypotheses, The system according to claim 12, comprising.

17. Forming a hypothesis of an object moving in a three-dimensional space, when the number of registered detections for the given hypothesis is greater than or equal to the threshold, identifying data points in a given hypothesis having a minimum vertical position, checking the respective vertical components of the estimated three-dimensional velocity vectors before and after the data points in the given hypothesis, When a first vertical component among the respective vertical components is negative and a second vertical component among the respective vertical components is positive, designating the data point as a ground collision; The system according to claim 12, including this.

18. The respective vertical components are obtained from the average of the estimated three-dimensional velocity vectors before and after the data point in the given hypothesis, and the operation is selecting at least one time window based on the noise level associated with the estimated three-dimensional velocity vector, the shortest predicted flight time on one or both sides of the data point, or both the noise level and the shortest predicted flight time; calculating the average of the estimated three-dimensional velocity vectors before and after the data points falling within the at least one time window; The system according to claim 17, including this.

19. Forming a hypothesis of an object moving in three-dimensional space includes, when the number of registered detections for the given hypothesis is equal to or greater than the threshold value, dividing the given hypothesis into separate segments including a flight segment, one or more bound segments, and a roll segment. The dividing includes the identifying, the checking, and the designating; classifying a first segment of the given hypothesis before the first designated ground collision as the flight segment; when the angle between a first estimated velocity vector before the next designated ground collision and a second estimated velocity vector after the next designated ground collision is greater than a threshold angle, classifying the next segment of the given hypothesis after each next designated ground collision as one of the one or more bound segments; when the angle between a first estimated velocity vector before the next designated ground collision and a second estimated velocity vector after the next designated ground collision is less than or equal to the threshold angle, classifying the next segment of the given hypothesis after the next designated ground collision as the roll segment; The system according to claim 17, including this.

20. The given hypothesis is at least one hypothesis remaining after the deleting, and designating at least one three-dimensional track of at least one ball moving in the three-dimensional space is Generating data about the three-dimensional position of the registered object, which is used in forming the at least one hypothesis by triangulating a more accurate 3D path using sensor observations registered by the object detection system at least in the flight segment; Identifying at least one moving ball by applying the complete three-dimensional physical model to the generated data at least for the flight segment; The system according to claim 19, comprising the above.

21. The obtaining includes receiving or generating the three-dimensional position of the registered object of interest, and most of the registered objects of interest are false positives including false detections by individual sensors among the two or more sensors and incorrect combinations of detections from individual sensors among the two or more sensors. The system according to claim 12.

22. The two or more sensors include three or more sensors, and the obtaining includes generating the three-dimensional position by creating each combination of detections from each pair of the three or more sensors for the current time slice. The system according to claim 21.

23. The three or more sensors include two cameras and a radar device, and the creating includes: Combining detections of a single object of interest by the two cameras using stereo pairing of the two cameras to generate a first three-dimensional position of the single object of interest; Combining detections of the single object of interest by at least one of the radar device and the two cameras to generate a second three-dimensional position of the single object of interest; The system according to claim 22, comprising the above.

24. The three or more sensors include an additional camera, and the creating includes using stereo pairing of the additional camera with each of the two cameras to combine detections of the single object of interest by the additional camera and the two cameras to generate a third three-dimensional position of the single object of interest and a fourth three-dimensional position of the single object of interest. The system according to claim 23.

25. A non-transitory computer-readable medium storing instructions for causing a data processing apparatus associated with an object detection system including two or more sensors to execute operations, the operations comprising: Obtaining a three-dimensional position of a target object registered by the object detection system, the object detection system being configured to allow more false positives and minimize false negatives for the registered target object; Using a filter applied to the three-dimensional position of the registered target object to form a hypothesis of an object moving in a three-dimensional space, the filter enabling an association between specific target objects among the registered target objects when an estimated three-dimensional velocity vector for a specific target object substantially corresponds to an object moving in the three-dimensional space over time; Using at least one additional target object registered by the object detection system to delete an appropriate subset of hypotheses that are not further extended by the performed association during the forming; Specifying at least one three-dimensional track of at least one ball moving in the three-dimensional space by applying a complete three-dimensional physical model to data about the three-dimensional positions of the registered target objects used in forming at least one hypothesis remaining after the deleting; Outputting for display at least one three-dimensional track of at least one ball moving in the three-dimensional space; Including; Forming a hypothesis of an object moving in the three-dimensional space using the filter comprises: When the number of registered detections for a given hypothesis is less than a threshold, a first motion model in a three-dimensional space, the first motion model having a first complexity level; When the number of registered detections for the given hypothesis is greater than or equal to the threshold, using a second motion model in the three-dimensional space, the second motion model having a second complexity level higher than the first complexity level of the first motion model; Including, a non-transitory computer-readable medium. Claim 26 The first motion model in the three-dimensional space is a linear model that associates the first three-dimensional position with the given hypothesis when the first three-dimensional position of the registered object of interest extends the given hypothesis in substantially the same direction predicted by the given hypothesis. The second motion model in the three-dimensional space is a recurrent neural network model trained to perform short-term prediction of motion in the three-dimensional space. The non-transitory computer-readable medium according to claim 25.

27. Forming a hypothesis of an object moving in a three-dimensional space involves using the first motion model or the second motion model to predict the next three-dimensional position of the registered detection for the given hypothesis according to the number and threshold of the registered detections for the given hypothesis, searching a spatial data structure that includes all the three-dimensional positions of the object of interest registered by the object detection system for the next time slice to find a set of three-dimensional positions within a defined distance of the predicted next three-dimensional position, if the set of three-dimensional positions is an empty set, using the predicted next three-dimensional position for the given hypothesis, if the set of three-dimensional positions includes only one three-dimensional position, using the one three-dimensional position for the given hypothesis, if the set of three-dimensional positions includes two or more three-dimensional positions, sorting the two or more three-dimensional positions based on their proximity to the predicted next three-dimensional position to form a sorted set, removing any non-proximate three-dimensional positions in the sorted set that exceed a defined threshold number greater than two, and using the two or more three-dimensional positions remaining in the sorted set to branch the given hypothesis into two or more hypotheses within a group of hypotheses, The non-transitory computer-readable medium according to claim 25.

28. Forming a hypothesis of an object moving in a three-dimensional space, when the number of registered detections for the given hypothesis is greater than or equal to the threshold, involves identifying data points in a given hypothesis having a minimum vertical position, checking the respective vertical components of the estimated three-dimensional velocity vectors before and after the data points in the given hypothesis, When a first vertical component among the respective vertical components is negative and a second vertical component among the respective vertical components is positive, designating the data point as a ground collision, The non-transitory computer-readable medium according to claim 25, including this.

29. The respective vertical components are obtained from the average of the estimated three-dimensional velocity vectors before and after the data point in the given hypothesis, and the operation is Based on the noise level associated with the estimated three-dimensional velocity vector, the shortest predicted flight time on one or both sides of the data point, or both the noise level and the shortest predicted flight time, selecting at least one time window, Calculating the average of the estimated three-dimensional velocity vectors before and after the data points falling within the at least one time window, The non-transitory computer-readable medium according to claim 28, including this.

30. Forming a hypothesis of an object moving in three-dimensional space includes, when the number of registered detections for the given hypothesis is equal to or greater than the threshold value, dividing the given hypothesis into separate segments including a flight segment, one or more bounce segments, and a roll segment, and the dividing includes the identifying, the checking, and the designating, Classifying a first segment of the given hypothesis before the first designated ground collision as the flight segment, When the angle between a first estimated velocity vector before the next designated ground collision and a second estimated velocity vector after the next designated ground collision is greater than the threshold angle, classifying the next segment of the given hypothesis after each next designated ground collision as one of the one or more bounce segments, When the angle between a first estimated velocity vector before the next designated ground collision and a second estimated velocity vector after the next designated ground collision is less than or equal to the threshold angle, classifying the next segment of the given hypothesis after the next designated ground collision as the roll segment, The non-transitory computer-readable medium according to claim 28, including this.