Method and system for optically tracking a moving object

The method addresses high memory and processing loads in object tracking by using pixel value inequalities and historical variance to filter false positives, achieving efficient and accurate object tracking with reduced computational demands.

JP2025535304APending Publication Date: 2025-10-24TOPGOLF SWEDEN AB
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2025522056
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-10-17
Filing Date
2023-10-06
Publication Date
2025-10-24

AI Technical Summary

Technical Problem

Existing object tracking methods using computer vision suffer from high memory and processing loads due to numerous false positive blob detections, especially with advanced digital cameras, leading to bottlenecks despite using high-performance hardware.

Method used

A method and system that reduces memory and processing power requirements by using a digital camera to capture images, applying an inequality based on pixel values and historical variance or standard deviation to identify true blobs, and storing data efficiently in computer memory to track moving objects.

Benefits of technology

The method significantly reduces computational power needed for object tracking by effectively filtering out false positives, allowing for accurate tracking with less hardware resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025535304000001_ABST
    Figure 2025535304000001_ABST
Patent Text Reader

Abstract

A method, system, and apparatus, including a computer program product, for tracking a moving object comprises: t ), and for two or more of the pixel values, determining an inequality that compares a first value to a second value, wherein the first value is greater than or equal to the pixel value (i x,y,t ) and predicted pixel value a second value calculated based on the product of the square of the difference between the pixel values ​​(i x,y,t storing information in computer memory indicating that the detected blob is part of the series of digital images (I) to determine the path of the moving object; t ) and correlating the detected blobs across the
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a method and system for optically tracking a moving object. [Background technology]

[0002] Known methods use computer vision to track moving objects using one or more cameras that depict the space in which the moving object resides. Tracking may be performed by first identifying the object as a single image pixel or set of adjacent pixels that deviate from the local background. Such deviating pixels are collectively referred to as a "blob." Once several blobs are detected in several image frames, possible paths of the tracked object are identified by interconnecting the identified blobs in subsequent frames.

[0003] An example of such a method is illustrated in US Patent Application Publication No. 20220051420(A1).

[0004] Blob generation in each individual frame potentially results in a large number of false positive blobs, i.e., identified blobs that do not truly correspond to any present moving objects, which may be due to noise, changing lighting conditions, and untracked objects appearing in the field of view of the camera in question.

[0005] Detection of possible tracked object paths usually results in a reduction of such false positives, for example based on filtering out physically or statistically unlikely paths. However, due to a large number of false positive blob detections, even if most of the false positives are filtered out in the tracked path detection step, blob detection itself is associated with a heavy memory and processor load and may therefore become a bottleneck for object tracking, even when high-performance hardware is used.

[0006] Furthermore, as digital cameras become more powerful, the pixel data output from such cameras correspondingly increases. To achieve accurate tracking of moving objects, it is desirable to use as accurate and precise image information as possible.

[0007] To prevent too many undetected blobs (false negatives) that could potentially lead to missed tracked object paths, it is usually preferable to allow a relatively large proportion of false positive blob detections. [Prior art documents] [Patent documents]

[0008] [Patent Document 1] U.S. Patent Application Publication No. 20220051420(A1) Summary of the Invention [Means for solving the problem]

[0009] Various embodiments described herein address one or more of the problems discussed above and provide techniques for tracking the path of a moving object using less memory and / or processing power than conventional object tracking techniques.

[0010] Accordingly, the present invention provides a method for tracking a moving object, comprising the steps of: A sequence of digital images I at successive times t t acquiring a digital image I t represents the optical input from the three-dimensional space within the field of view of the digital camera, and the digital camera captures the pixel p x,y said digital image I having a corresponding set of t and the digital image is arranged to generate corresponding pixel values ​​i x,y,t a digital camera for capturing said series of digital images (I t ) not moving relative to the three-dimensional space during generation; The pixel value i x,y,tdetermining an inequality that compares a first value with a second value for two or more of the pixel values ​​i x,y,t and predicted pixel values

[0011]

number

[0012] The second value is calculated as or based on the square of the difference between, first, the square of the number Z and, second, the pixel p in question. x,y Historical pixel values ​​of i x,y,{t-n,t-1} The estimated variance or standard deviation σ x,y,t The predicted pixel value calculated as or based on the product of

[0013]

number

[0014] But the pixel in question is p x,y Historical pixel values ​​of i x,y,{t-n,t-1} Z is calculated based on Z 2 is 10≦Z 2 a step, the step being a number selected to be an integer such that ≦20; The pixel value i for which the first value is higher than the second value x,y,t With respect to pixel value i x,y,t storing information in computer memory indicating that is part of the detected blob; and extracting the sequence of digital images I based on information stored in a computer memory to determine a path of a moving object through the three-dimensional space. t Correlating the detected blobs across the The method may be embodied as a method, including:

[0015] In some embodiments, The inequality is:

[0016]

number

[0017] and

[0018]

number

[0019] is the predicted pixel value, and σ x,y,t But the pixel in question is p x,y Historical pixel values ​​of i x,y,{t-n,t-1} is the estimated standard deviation for

[0020] In some embodiments, the method comprises: The pixel p x,y For each pixel of and for number n≦N, the sum

[0021]

number

[0022] and

[0023]

number

[0024] in said computer memory; The pixel value i x,y,t For each pixel value of

[0025]

number

[0026] determining the inequality such that Further includes:

[0027] In some embodiments, S x,y,t , Q x,y,t , or S x,y,t and Q x,y,t Both and are calculated recursively, so that pixel value i x,y,t If the calculated values ​​for the same pixel p x,y , but with respect to the previously stored calculated value S at the immediately preceding time t-1. x,y,t , Q x,y,t , or S x,y,t and Q x,y,t It is calculated using both and

[0028] In some embodiments, S x,y,t But, S x,y,t = S x,y,t-1 + i x,y,t -i x,y,t-n It is calculated as follows, and Q x,y,t but,

[0029]

number

[0030] It is calculated as follows:

[0031] In some embodiments, the method comprises: S x,y,t and Q x,y,t , pixel p x,y and combining and storing the combined data in said computer memory as a single data type containing 12 bytes or less for each data type.

[0032] In some embodiments, the method comprises: For each pixel p x,y With respect to pixel value i x,y,t A pixmap having the information indicating that the detected blobs are part of a particular digital image I t and storing in said computer memory:

[0033] In some embodiments, pixel value i x,y,t The information indicating that each pixel p is part of a detected blob is x,y is indicated in a single bit for .

[0034] In some embodiments, The pixmap also has a pixel p x,y With respect to the pixel p x,y The predicted pixel value of i x,y,t Contains a value indicating

[0035] In some embodiments, The pixel in question, p x,y The predicted pixel value of i x,y,t The value indicating the predicted pixel value, using a total of 15 bits for the integer part and the fractional part,

[0036]

number

[0037] This is achieved by storing the .times. ...

[0038] In some embodiments, predicted pixel values

[0039]

number

[0040] , the estimated variance or standard deviation σ x,y,t , or both the predicted pixel value and the estimated variance or standard deviation are x,y n historical pixel values ​​i x,y,{t-n,t-1} where 10≦n≦300.

[0041] In some embodiments, The estimated variance or standard deviation σ of the second value x,y,t The previous image I is taken into account for the estimation of t The number n is chosen to be a power of two.

[0042] In some embodiments, The pixel value i x,y,t has a depth ranging from 8 bits to 46 bits across one or several channels.

[0043] In some embodiments, predicted pixel values

[0044]

number

[0045] But the image I t Pixel p in x,y The historical pixel values ​​of the sampled set i x,y,t The estimated projected future mean pixel value μ, also determined based on x,y,t is determined based on the

[0046] In some embodiments, predicted pixel values

[0047]

number

[0048] but,

[0049]

number

[0050] and α and β are determined as follows:j,k (i j,k,t - (αμ j,k,t +β)) 2 is a constant determined to minimize μ j,k,t But the pixel in question is p j,k where j and k are iterated over a test set of pixels.

[0051] In some embodiments, μ x,y,t But the pixel in question is p j,k Pixel value i x,y,t is the estimated historical mean for

[0052] In some embodiments, Let the test set of pixels be image I t pixel p x,y contains between 1% and 25% of the total set.

[0053] In some embodiments, Let the test set of pixels be image I t pixel p x,y are geometrically uniformly distributed over the entire set of

[0054] In some embodiments, Estimated standard deviation σ x,y,t but,

[0055]

number

[0056] is determined by, ξ∈]0,1[.

[0057] In some embodiments, the method comprises: determining that at least one of α is further from 1 than a first threshold and β is further from 0 than a second threshold is true; 19. The method of claim 14, wherein the predicted pixel value is calculated based on the first threshold value of α and the second threshold value of β.

[0058]

number

[0059] determining Further includes:

[0060] In some embodiments, the method comprises: The pixel value i for which the first value is higher than the second value x,y,t With respect to, the following inequality exists:

[0061]

number

[0062] Pixel value i is true if and only if x,y,t storing the information indicating that i is part of the detected blob, x,y,t is the pixel value in question,

[0063]

number

[0064] is the predicted pixel value and B is an integer such that B>100.

[0065] In some embodiments, the method comprises: The method further includes using a Hoshen-Kopelman algorithm to group together adjacent pixels that are determined to be part of the same blob.

[0066] In some embodiments, The object is a golf ball.

[0067] The present invention provides a method for tracking a moving object, comprising: acquiring a series of digital images I from a digital camera, the digital images I representing optical input from a three-dimensional space within a field of view of the digital camera over time, each digital image I representing a corresponding pixel value i x,y Pixel p with x,y and In a computer, performing image segmentation for each image of a series of digital images I using a statistical model of the background of the optical input to detect blobs, wherein two or more pixel values ​​i of the image are x,y,t With respect to each of the Current Image I t pixel p x,y The current pixel value of i x,y,t , pixel p x,y The first S of the statistical model x,y,t value and second Q x,y,t value, as well as the confidence level value Z 2 determining the result of the inequality using the first S x,y,t value and second Q x,y,t The value is the current image I t Historical pixel values ​​i from an image from a series of previous digital images I x,y is calculated based on the current pixel value i x,y,t , the first S x,y,t value and second Q x,y,t value, as well as the confidence level value Z 2 each of which is stored in a memory of the computer as integer type data, and determining uses integer arithmetic in the computer; and In response to the result of the inequality, the current image I t The current pixel value of image pixel i x,y,t storing information in computer memory indicating that the detected blob is part of the detected blob; and using the stored information to correlate the detected blobs across the series of digital images I to determine a path of a moving object through three-dimensional space within the field of view of the digital camera; The method may be embodied as a method, including:

[0068] The present invention further provides a system for tracking a moving object, comprising a digital camera, a digital image analyzer, and a moving object tracker; A digital camera captures a series of digital images I at successive times t. t To generate a pixel p, the digital camera is positioned to represent an optical input from three-dimensional space within the field of view of the digital camera. x,y said digital image I having a corresponding set of t and the digital image is arranged to generate corresponding pixel values ​​i x,y,t a digital camera for capturing said series of digital images (I t ) is positioned so as not to move relative to the three-dimensional space during generation of the The digital image analyzer determines the pixel values ​​i x,y,t , wherein the first value is a pixel value i in question, and the second value is a pixel value i in question. x,y,t and predicted pixel values

[0069]

number

[0070] The second value is calculated as or based on the square of the difference between, first, the square of the number Z and, second, the pixel p in question. x,y Historical pixel values ​​of i x,y,{t-n,t-1} The estimated variance or standard deviation σ x,y,t The predicted pixel value calculated as or based on the product of

[0071]

number

[0072] But the pixel in question is p x,y Historical pixel values ​​of i x,y,{t-n,t-1} Z is calculated based on Z 2 is 10≦Z 2 is selected to be an integer such that ≦20; A digital image analyzer detects pixel values ​​i where the first value is higher than the second value. x,y,t With respect to pixel value i x,y,t is configured to store information in computer memory indicating that the detected blob is part of the detected blob; A moving object tracker tracks the sequence of digital images I based on information stored in computer memory to determine a path of a moving object through the three-dimensional space. t The present invention may also be embodied as a system configured to correlate detected blobs across a network.

[0073] Furthermore, the present invention, when implemented, A series of digital images from a digital camera t wherein the digital camera receives the digital images I at successive times t. t A digital camera is positioned to represent optical input from three-dimensional space to generate pixels p x,y said digital image I having a corresponding set of t and the digital image is arranged to generate corresponding pixel values ​​i x,y,t a digital camera for capturing said series of digital images (I t ) positioned so as not to move relative to the three-dimensional space during generation of the signal; The pixel value i x,y,t determining an inequality comparing a first value to a second value for two or more of the pixel values ​​i x,y,t and predicted pixel values

[0074]

number

[0075] The second value is calculated as or based on the square of the difference between, first, the square of the number Z and, second, the pixel p in question. x,y Historical pixel values ​​of i x,y,{t-n,t-1} The estimated variance or standard deviation σ x,y,t The predicted pixel value calculated as or based on the product of

[0076]

number

[0077] But the pixel in question is p x,y Historical pixel values ​​of i x,y,{t-n,t-1} Z is calculated based on Z 2 is 10≦Z 2 is selected to be an integer such that ≦20; and The pixel value i for which the first value is higher than the second value x,y,t With respect to pixel value i x,y,t storing information in computer memory indicating that is part of the detected blob; and extracting the sequence of digital images I based on information stored in a computer memory to determine a path of a moving object through the three-dimensional space. t Correlating blobs detected across The method may also be embodied as a computer software product configured to perform the steps of:

[0078] The computer software product may be implemented by a non-transitory computer-readable medium encoding instructions that cause one or more hardware processors located in at least one of the computer hardware devices in the system to perform digital image processing and object tracking.

[0079] In the following, the present invention will be described in detail with reference to exemplary embodiments thereof and the accompanying drawings. [Brief explanation of the drawings]

[0080] [Figure 1] 4 is a schematic diagram of a system 100 configured to perform a method of the type shown in FIG. 3. [Figure 2] FIG. 1 is a simplified diagram of a data processing device. [Figure 3] 1 is an overall flow chart for logically tracking a moving target object. [Figure 4] 2 is a flowchart of a method performed by the system 100 shown in FIG. 1. [Figure 5] FIG. 1 is a schematic diagram illustrating a noise model of the type described herein. [Figure 6] FIG. 2 is a diagram of an image frame illustrating a noise model. [Figure 7] FIG. 1 illustrates an example of clustering pixels into blobs. [Figure 8] FIG. 10 illustrates pixel intensity during a sudden exposure change event. DETAILED DESCRIPTION OF THE INVENTION

[0081] All figures share the same reference numbers for the same and corresponding parts.

[0082] 1, a method relates to tracking a moving target object 120. Generally, a system 100 may include one or several digital cameras 110, each positioned to represent optical input from a three-dimensional space 111 within the field of view of the digital camera 110, to generate digital images of such moving target object 120, the object moving through the space 111 and thus represented by the digital camera 110 in successive digital images. Such representation by the digital camera 110 will be referred to herein as a "representation" for simplicity.

[0083] The digital camera 110 captures a series of digital images (I t ) is generated. For example, the digital camera 110 may be fixed relative to the space 111, or if the digital camera 110 is movable, the digital camera 110 may capture a series of digital images (I t ) is kept stationary during the generation of the pixel p. Thus, the same part of the space 111 is depicted by the digital camera 110 each time, and the digital camera 110 captures the pixel p x,y a digital image I with a corresponding set of t and generating said generated digital image I t is the corresponding pixel value i x,y,t "x" and "y" represent coordinates in the image coordinate system, while "t" represents time.

[0084] Two or more different images t Pixel value i x,y,t correspond to each other, the individual pixels p x,y But the image in question is t This means that all of the above measures light entering the camera 110 from the same or substantially the same light cone. t There may be slight movements between pixels due to wind, thermal expansion, etc., but even if such slight noise-inducing movements exist, pixel p x,y It is understood that there is a substantial correspondence between any two consecutive images I t between any one same pixel p of the camera 110 x,y There may be at least 50% overlap between the light cones of the camera 110. In some cases, the camera 110 may be movable, such as pivotable. In such cases, the pixel p of the captured image x,y An image transformation may be applied to the captured image to make the pixels of the image correspond to pixels of a previous or future captured image.

[0085] If the system 100 includes two or more digital cameras 110, several such digital cameras 110 may be positioned to depict the same space 111 and, consequently, track the same moving target object 120 through said space 111. In such a case, several digital cameras 110 may be used to construct a stereoscopic view of the respective tracked paths of each target object 120.

[0086] As mentioned, the digital camera 110 captures a series of successive images I at different times. t Such images are sometimes referred to as image "frames." In some embodiments, digital camera 110 is a digital video camera arranged to generate digital moving images that include or are made up of a series of such digital image frames.

[0087] 1, system 100 includes a digital image analyzer 130 configured to analyze digital images received from digital camera 110 directly or through an intermediate system, either in the same or processed (reformatted, compressed, filtered, etc.) form. The analysis performed by digital image analyzer 130 may be performed in the digital domain. Digital image analyzer 130 may also be referred to as a "blob detector."

[0088] The system 100 further includes an object tracker 140 configured to track the moving target object 120 across multiple of the digital images based on information provided by the digital image analyzer 130. The analysis performed by the object tracker 140 may also be performed in the digital domain.

[0089] In an exemplary embodiment, system 100 is configured to track target object 120 in the form of a flying sports object, such as a flying ball, e.g., a flying baseball or golf ball. In some embodiments, system 100 is used at a golf driving range, such as a driving range, having multiple bays for hitting golf balls to be tracked using system 100. In other cases, system 100 may be installed in an individual golf range bay or on a golf tee and configured to track golf balls being hit from the bay or tee. System 100 may also be a portable system 100 configured to be positioned where it can track the moving target object 120. It will be understood that the monitored "space" described above, in each of these and other cases, will be the space through which a sports ball is expected to travel.

[0090] Various types of computers may be used in system 100. Digital image analyzer 130 and object tracker 140 constitute examples of such computers. In some cases, digital image analyzer 130 and object tracker 140 may be provided as software functions running on one and the same computer. One or more digital cameras 110 may also be configured to perform digital image processing, thus also constituting an example of such a computer. In some embodiments, digital image analyzer 130 and / or object tracker 140 are implemented as software functions configured to run on the hardware of one or more digital cameras 110. In other embodiments, digital image analyzer 130 and / or object tracker 140 are implemented on a standalone or combined hardware platform, such as on a computer server.

[0091] The one or more digital cameras 110, digital image analyzer 130, and object tracker 140 are configured to communicate digitally either through a communication path internal to the computer, such as through a computer bus, or through wired and / or wireless communication paths external to the computer, such as through an internet network 10 (e.g., the internet). In implementations requiring significant communication bandwidth, the camera 110 and digital image analyzer 130 can communicate through a direct wired digital communication route that does not go through the network 10. Alternatively, the digital image analyzer 130 and object tracker 140 may communicate with each other through the network 10 (e.g., a conventional internet connection).

[0092] The essential elements of a computer are generally a processor for executing instructions and one or more memory devices for storing instructions and data. As used herein, a "computer" may include a server computer, a client computer, a personal computer, an embedded programmable circuit, or a special-purpose logic circuit. Such a computer may be connected to one or more other computers through a network such as the Internet 10 or via any suitable peer-to-peer connection for digital communication, such as a Bluetooth® connection.

[0093] Each computer may include various software modules, which may be distributed between the application layer and the operating system. These may include executable and / or interpretable software programs or libraries, including, for example, various programs operating as the digital image analyzer 130 program and / or the object tracker 140 program. Other examples include digital image preprocessing and / or compression programs. The number of software modules used may vary from implementation to implementation and from one such computer to another. Each of the programs may be implemented in embedded firmware and / or as software modules distributed across one or more data processing devices connected by one or more computer networks or other suitable communication networks.

[0094] 2 shows an example of such a computer, a data processing device 300, which may include hardware or firmware devices including one or more hardware processors 312, one or more additional devices 314, a non-transitory computer-readable medium 316, a communication interface 318, and one or more user interface devices 320. The processor 312 can process instructions for execution within the data processing device 300, such as instructions stored on the non-transitory computer-readable medium 316, which may include a storage device such as one of the additional devices 314. In some implementations, the processor 312 is a single or multi-core processor, or two or more central processing units (CPUs). The data processing device 300 uses its communication interface 318 to communicate with one or more other computers 390, for example, via a network 380. Thus, in various implementations, the described processes may be executed in parallel, simultaneously, or serially on single or multi-core computing machines, and / or on a computer cluster / cloud, etc.

[0095] The data processing device 300 includes various software modules, which may be distributed between the application layer and the operating system, which may include executable and / or interpretable software programs or libraries including a program 330 that constitutes the digital image analyzer 130 described herein and is configured to perform the steps of the methods performed by such digital image analyzer 130. The program 330 may also constitute the object tracker 140 described herein and is configured to perform the steps of the methods performed by such object tracker 140.

[0096] Examples of user interface devices 320 include a display, a touchscreen display, a speaker, a microphone, a haptic feedback device, a keyboard, and a mouse. Furthermore, the user interface device need not be a local device 320, but can be remote from the data processing apparatus 300, e.g., a user interface device 390 accessible via one or more communications networks 380. The user interface device 320 can also be in the form of a standalone device having a screen, such as a conventional smartphone that is connected to the system 100 by a configuration or setup step. The data processing apparatus 300 can store instructions for performing the operations described herein on a non-transitory computer-readable medium 316, which can include, for example, one or more additional devices 314, e.g., one or more of a floppy disk device, a hard disk device, an optical disk device, a tape device, and a solid-state memory device (e.g., a RAM drive, flash memory, or EEPROM). Additionally, instructions for performing the operations described herein can be downloaded from one or more computers 390 (e.g., from the cloud) over a network 380 to a non-transitory computer-readable medium 316; in some implementations, the RAM drive is a volatile memory device into which instructions are downloaded each time the computer is powered on.

[0097] It is understood that the computer hardware described can be physical hardware, virtual hardware, or any combination thereof.

[0098] As mentioned, the system 100 is configured to perform a method according to one or more embodiments for optically tracking a moving target object 120 .

[0099] The present invention may also be embodied as a computer software product configured to perform the methods when executed on computer hardware of the type described herein, and thus may be deployed as part of system 100 to provide the functionality required to perform the methods.

[0100] Thus, both the system 100 and the computer software product are configured to track a moving target object 120 moving through the space 111 relative to one or several digital cameras 110 by including or embodying the above-mentioned digital image analyzer 130 and object tracker 140, which are further configured to perform the steps of the corresponding methods described herein.

[0101] In general, anything said in connection with the presently described methods is equally applicable to the system 100 and computer software products described herein, and vice versa.

[0102] FIG. 3 shows an overall flow chart for tracking a moving target object 120 based on digital image information received from one or several digital cameras 110 .

[0103] In computer vision, "image segmentation" is the process of separating an image into distinct regions that represent target objects therein. Generally, it is desirable to distinguish potential moving target objects from the background, which is generally changing, can be noisy, and is often extremely difficult to predict. For example, in the golf ball example, when such a ball is far away from the digital camera 110 that depicts the ball 120, the ball 120 will be captured by one single pixel p in the digital image frame generated by the digital camera 110. x,y It can even be as small as

[0104] For these reasons, it is generally not possible to separate a foreground object 120 from the background based solely on the detected shape relative to the expected shape of the target object 120. Instead, a statistical model of the background (hereafter referred to as the "noise model") is set up, and based on this model, a probability measure is used to identify pixels p that deviate from the expected value by more than a threshold. x,y It is proposed to identify neighboring pixels p in the detected digital image that deviate from the values ​​expected according to the model. x,y is the pixel p x,y are grouped together into "blobs" ("blob aggregation").

[0105] Such methods may result in a very large number of false positives, such as approximately 99.9% false positives. However, subsequent motion tracking analysis may detect the motion of consecutive digital image frames I t It is possible to cull the majority of all false positives, such as keeping only those blobs that appear to obey Newton's laws of motion.

[0106] As shown in Figure 3, the noise model step is used to suppress noise in an image frame, with the goal of reducing the number of detected blobs in the subsequent blob aggregation step. t Every pixel p x,y Multiple pixels p x,y , and therefore risks becoming a major bottleneck. These calculations, which aim to identify noise that does not fit the detected statistical pattern in order to identify outliers, can be handled by a high-performance GPU (Graphics Processing Unit), but performance may still prove to be an issue. The approach described herein is to analyze the pixel p x,yThis results in a dramatic reduction in the computational power required per image. This reduction can be exploited by using simpler hardware, lower power consumption, or a larger incoming image bit rate.

[0107] Turning now to FIG. 4, a method according to one or more embodiments is shown.

[0108] In a first step S1 the method starts.

[0109] In the next step S2, Z 2 A number Z is chosen such that is an integer. 2 is 10≦Z 2 Z may be selected to be an integer such that Z is ≦20. 2 It is noted that Z may be a non-integer value as long as Z is an integer value. This step S2 may be performed in advance, such as during the design process of system 100 or a calibration step of system 100.

[0110] In a subsequent step S3, the space 111 is divided into a series of digital images i at successive times t. t The space 111 is depicted using a digital camera 110 to generate a sequence of N digital images i at successive times t. t However, the procedure may be depicted using the digital camera 110 to generate digital images i at successive times t as long as the procedure is in progress. t It will be appreciated that the sequence of digital images i at successive times t can also be a continuous or semi-continuous procedure that continues to generate a sequence of digital images i at successive times t. Thus, in this case, the number of digital images N increases by one for each captured frame. In either case, t may be thought of as a stream of captured digital images, much like a digital video stream.

[0111] In the following step S4, the pixel value i x,y,tFor two or more (eg, some) of, an inequality is determined that involves comparing a first value to a second value.

[0112] The first value is the pixel value i in question. x,y,t and its pixel p x,y The calculated predicted pixel value of

[0113]

number

[0114] The second value is calculated based on the square of the difference between, on the one hand, the square of the selected number Z, where this square is an integer value, and, on the other hand, the pixel p in question. x,y Historical pixel values ​​of i x,y,{t-n,t-1} The estimated variance or standard deviation σ x,y,t Specifically, the second value is calculated based on the product of the estimated variance or estimated standard deviation σ x,y,t can be calculated based on the square of

[0115] predicted pixel values

[0116]

number

[0117] Also, the pixel in question p x,y Historical pixel values ​​of i x,y,{t-n,t-1} In other words, the image frame I captured by the camera 110 at a point in time prior to time t is t-Δt The predicted pixel value is calculated using information from

[0118]

number

[0119] is the estimated variance or standard deviation σx,y,t Historical pixel values ​​i, which may be the same or different from x,y,{t-n,t-1} can be calculated based on the set of

[0120] In the notation used herein, "n" refers to the number of historical pixel values ​​i considered by the noise model, counting backward from the image frame currently being considered. x,y,t Therefore, this notation represents the number of consecutive pixel values ​​i up to the currently considered image frame. x,y,t Assume that is used to calculate both the first and second values, but pixel value i x,y,t It will be appreciated that any suitable continuous or discontinuous, same or different intervals of may be used to calculate the first and second values, respectively.

[0121] Generally, the equations and formulas disclosed and discussed herein are provided as illustrative examples, and it will be understood that in practical embodiments, the equations and formulas may be tailored to particular needs, which may include, for example, the introduction of various constant and scaling factors, additional intermediate calculation steps such as filtering steps, etc.

[0122] In some embodiments, the inequality is:

[0123]

number

[0124] may also be written as

[0125]

number

[0126] is the predicted pixel value, and σ x,y,t is the pixel p in question x,yThe historical pixel value i x,y,{t-n,t-1} is the estimated standard deviation for

[0127] In general, the noise model currently described considers the noise x,y For i, we estimate the moving average and standard deviation based on the last n image frames, and then use these metrics to find the pixel value i at the same image location in the new frame. x,y,t deviates from an expected value by more than an acceptable limit.

[0128] This model is suitable for the image i under consideration, as long as the background contains only features that are considered to be stationary in a first approximation. t The Z-score can be designed to assume that every pixel in the background of t has inherent Gaussian noise. A normal distribution can be used to establish a suitable confidence interval. For example, if a Z-score of 3.464 is used, it can be seen that 99.95% of all samples that are not significantly different from the background fall within the corresponding confidence interval. Thus, at time t, the signal value i x,y Pixel p with x,y teeth,

[0129]

number

[0130] A pixel is considered significantly different from the background if k = 1 / k, where k iterates over the previous n frames. The limit is the (uncorrected) standard deviation

[0131]

number

[0132] The corrected (unbiased) standard deviation is the mathematically more correct choice, i.e., a more accurate estimate of σ results from dividing by n-1 instead of by n. However, this is not important for our purposes, since the limits used are multiples of the standard deviation that may be chosen freely. If the number n of previous image frames taken into account for the estimation of the standard deviation in the second value (used in the evaluation of the inequality) is chosen to be a power of two (e.g., 16, 32, 64, ...), then very low-cost and computationally efficient multiplications and divisions can be obtained by using shift operations.

[0133] Image Frame I i When processing, pixel value i from frame k∈[t - n, t - 1] x,y is used. A variation of the formula for calculating the standard deviation, which allows the standard deviation to be calculated in one go, is: where the formula for the estimate of the mean is also given.

[0134]

number

[0135]

number

[0136]

number

[0137] and S x,y,t =Σ k i x,y,k Then, equations (2) and (3) become

[0138]

number

[0139]

number

[0140] Looking back at (1), we see that both the left and right sides of this equation are non-negative, so it is OK to square both sides.

[0141]

number

[0142] Combining (4) and (6)

[0143]

number

[0144] This results in

[0145]

number

[0146] is equivalent to

[0147] It is noted that n≦N. Therefore, the inequality considered above can be expressed as (8),

[0148]

number

[0149] and

[0150]

number

[0151] is.

[0152] i x,y,t ,n,S x,y,t, and Q x,y,t Since all x, y, and z can be chosen to produce an appropriate or desired number of false positives, the entire computation can be done using only integers. This means that the computation can be performed without any loss of precision due to floating-point truncation errors. Also, integer operations are usually faster than their floating-point equivalents.

[0153] The table below shows various results for different chosen values ​​of Z.

[0154] [Table 1]

[0155] Equation (8) is the pixel value i x,y,t The sum S and the sum of squares Q of the most recent n observations of

[0156]

number

[0157]

number

[0158] It depends on the knowledge of each frame I t Each pixel value i x,y,t Although it is possible to compute statistics directly using (9) and (10) for a new frame I t is added to the noise model, and frame I from the nth frame t-n Recursive definitions where are removed, i.e.,

[0159]

number

[0160]

number

[0161] It is much more computationally efficient to use Q x,y,t Updating Q requires two multiplications to produce the square. x,y,t Since ∑ i = 1 ⁢ ... involves a difference of squares, it can be reduced to one single multiplication if rewritten as: Q x,y,t = Q x,y,t-1 + (i x,y,t + i x,y,t-n )(i x,y,t -i x,y,t-n ) (13) Z - = i x,y,t -i x,y,t-n and z + = i x,y,t + i x,y,t-n Then, S x,y,t = S x,y,t-1 + z - , and (14) Q x,y,t = Q x,y,t-1 + z + z - (15) Then, (14) and (15) are the total computations required to update the noise model. A straightforward implementation requires only three (int) additions, one (int) subtraction, and one (int) multiplication per pixel, making it very computationally efficient. Furthermore, these computations can be accelerated by using SIMD instruction sets such as AVX2 (on x86_64) or NEON (on aarch64), or can be performed on a GPU or even implemented on an FPGA.

[0162] Sequential image frames I t The calculations performed to update the noise model between "new frames" are conceptually shown in Figure 5, which illustrates how tis added to the noise model, and the frame I from n frames before is the last frame in the currently considered "queue of frames in the model" t-n As mentioned above, this indicates whether individual pixel values ​​i x,y,t and i x,y,t-n Considering z + and z - This can be efficiently done by computing the value of

[0163] From the above, 0≦S≦2 b n and 0 ≤ S ≤ 2 2b n, and b is the input pixel value i x,y,t is the bit depth of the data. In some embodiments, n is 300 or less, or 256 or less, or 128 or less, and 64 = 2 6 (The average is calculated over more than 64 consecutive image frames I t (n is not performed over a period of time). In some embodiments, n may be as few as 32, or as few as 16, or even as few as 10. In some embodiments, the n frames considered at each time are the n most recent frames captured and provided by the camera 110. In this case, the n frames may together cover a period of between 0.1 s and 10 s, such as between 0.5 s and 2 s of the captured video. In other words, the number of frames n considered may be relatively close to the frame rate used by the digital camera 110. The noise model then considers at least as many frames I as the length of the window size n. t In addition to keeping the actual image frame in memory with respect to pixel p x,y Furthermore, if the calculation (described in equation (19) below) is used, an additional single precision float may be required per pixel to store the estimated variance.

[0164] In some embodiments, pixel value i x,y,thas a bit depth ranging from 8 bits to 48 bits for one or several channels, such as a single channel (e.g., a gray channel) at 8 or 16 bits depth or three channels (such as RGB) at 16 or 24 bits depth.

[0165] The camera 110 captures pixel values ​​i across several color channels. x,y,t If you provide information, pixel value i x,y,t is the pixel value i by the digital image analyzer 130 x,y,t The image may be converted to a single channel (such as a grayscale channel) before processing. Alternatively, of several available channels, only one such channel may be used for analysis. Still alternatively, several channels may be analyzed separately in parallel, such that pixels detected to be blobs in at least one such analyzed channel are determined to be blobs at any one time.

[0166] The transformed pixel value i x,y,t can have a bit depth of at least 8 bits, and in some embodiments up to 24 bits, such as up to 16 bits. A bit depth of 12 bits has been found to provide a good balance between speed, memory requirements, and output quality. If the input data has a bit depth higher than necessary, the data from the camera 110 may be converted (downsampled) before processing by the digital image analyzer 130.

[0167] More generally, the number of bits required can be found as D + log2(n) for S and 2D + log2(n) for Q, where D is the bit depth of one single considered channel.

[0168] The following table shows the pixel values ​​i used when n = 64. x,y,t 1 shows the required storage space for S and Q depending on the bit depth of

[0169] [Table 2]

[0170] Generally, the method involves updating the noise model and performing blob detection on individual pixels p x,y This noise map may include storing in computer memory a collection of updated noise model information (S and Q) for each pixel p in the image. x,y can be updated and stored.

[0171] Using the calculations explained above, S x,y,t and Q x,y,t may be combined and stored in the computer memory as a single data type (such as a single structure, record, or tuple), the data type being a pixel p x,y Each contains 12 bytes or less, or 10 bytes or less, or even 8 bytes or less. (Image I t For every pixel p x,y for each analyzed pixel value i x,y,t Regarding S x,y,t and Q x,y,t This combined storage of updated values ​​of I, II, III, and IV as a single data type constitutes an example of the "noise model" described herein. Thus, the noise model is a noise model that stores the noise of successive image frames I, II, III, and IV generated and provided by (each) digital camera 110. t For each individual image frame in the set I t With respect to each analyzed digital image frame I t will be updated regarding.

[0172] In the same step S4, the pixel value i for which the first value is found to be higher than the second value is x,y,t information is stored in the computer memory with respect to pixel value i x,y,t is part of the detected blob.

[0173] This memory is used to store the generated pixmap, in other words, each pixel p x,y This can be done in a data structure that has such indication information for each pixel p x,y is the image frame I t Each pixel p belongs to or does not belong to the blob x,y The information about can be stored as a single binary bit and therefore can be stored very computationally efficiently.

[0174] One way to practically implement such a pixmap is to use a "noise map" of the usual type described below, where the pixmap is a noise map where each pixel p x,y With respect to the pixel p x,y The predicted pixel value of i x,y,t Also includes a value indicating

[0175] Therefore, for each frame, the noise model established above is applied to every pixel location p x,y Regarding the new frame I t That particular pixel value i x,y,t was outside the allowed limits (i.e., whether (6) or (8) was true). Furthermore, the noise map may be used to generate such a noise map for each pixel p at time t, such as based on the calculations performed in determining the noise model. x,y The predicted signal values ​​can be useful in downstream calculations, such as in the subsequent blob aggregation step, so it is computationally efficient to establish and store this information already at this point.

[0176] Figure 6 shows the latest available image frame I t The noise model after updating based on the information of frame I t is S at that time t x,y and Q x,y and how it relates to the value of

[0177] Each new image frame I arrives at the digital analyzer 130 t Even if it is possible to first generate the noise map of I and only afterwards update the noise model in the digital analyzer 130, both of them can be done in one go without unloading or overwriting information in memory during the calculations. t is loaded into CPU memory, and z + , z - , S x,y,t , Q x,y,t , σ x,y,t , and / or μ x,y,t However, before the loaded data is unloaded or overwritten in CPU memory, it is necessary to x,y The advantage realized then is that memory accesses avoid becoming a bottleneck: if the CPU incurs the penalty of loading data, all necessary calculations are performed before unloading or overwriting the data in CPU memory.

[0178] In the following example, the noise map is stored for pixel p x,y This information can be stored in a single 2-byte data type (such as a uint16).

[0179] Each noise map entry corresponds to a pixel p x,y The information indicating whether p is a blob pixel or not is given by the pixel in question p in the noise map. x,y In some embodiments, the most significant bit of each pixel p, such as the most significant bit of the exemplary two-byte structure, may be stored in the form of a single bit out of the total number of stored bits for each pixel p x,y The most significant bit of the data type used to store the noise map data of x,y,t indicates whether i is outside the blob generation limits. Then the lower 15 bits are the predicted (mean) pixel value i scaled to 15-bit precision. x,yThe signal can be encoded and stored in a fixed-point representation. x,y The signal is the predicted pixel value discussed above.

[0180]

number

[0181] In other words, pixel p x,y The predicted pixel value of i x,y,t The noise map value indicating the predicted pixel value

[0182]

number

[0183] to a 15-bit grayscale bit depth (if necessary).

[0184] In one example, for performance reasons, the encoding is as follows: First, the expected signal is scaled to 15 bits (0..32767). If n = 32 and the input pixel depth is 12 bits, this translates to S t is the value of each pixel i x,y,t This means that we use 17 bits for the pixel value i. A simple shift operation divides this number by 4, which puts the number into the 15-bit range. Second, x,y,t If p is within the bounds given in equation (6) (or its reformulated form (8)), then all bits are negated. Thus, a consumer of the noise map can generate a pixel p x,y You can iterate over the list and ignore all entries with the most significant bit set to 1.

[0185] Note that the pixmap for each pixel contains at least, or only, information about 1) whether the pixel is part of a blob and 2) the predicted pixel value for that pixel. In this case, the prediction is simply the arithmetic mean of the previous n frames, but we will later describe methods for predicting values ​​to be used when recent frames have undergone significant changes in capture parameters such as shutter time or gain, which may be alternatives to the methods described so far.

[0186] In some embodiments, the stored noise model is used to analyze the noise of an image frame I received from the camera 110 by the digital image analyzer 130. t In other words, the stored noise model incorporates all available information from Q x,y and S x,y To calculate the value of t n consecutive or non-consecutive image frames I up to t On the other hand, for each pixel p of the noise map x,y The estimated projection (predicted pixel value) stored for

[0187]

number

[0188] ) data is the second most recently received image frame I t Only using pixel values ​​i to be evaluated for blob classification. x,y,t The most recently received image frame I, containing t In practice, this can be updated without using the most recently received pixel value i x,y,t Before being updated using S x,y,t This may mean that the previous value of can be used to calculate the (transformed) predicted data, and the (transformed) predicted data is then stored in the pixmap.

[0189] In the above example, the predicted pixel value

[0190]

number

[0191] is the image frame I t Pixel p in the sequence x,y The historical pixel values ​​of the sampled set i x,y,t The estimated future mean pixel value μ is also determined based on x,y,t is determined as (or at least based on)

[0192] In an embodiment described in more detail below, the predicted pixel value

[0193]

number

[0194] teeth,

[0195]

number

[0196] and α and β are determined as follows: Σ j,k (i j,k,t - (αμ j,k,t + β)) 2 (16) is a constant determined to minimize j,k,t is the pixel p in question j,k are the estimated future mean pixel values ​​of image frame I, j and k ... t pixel p x,yThe determination of α and β may be done in any essentially conventional manner well known to those skilled in the art. As with the noise models described above, in some embodiments, μ x,y,t is the pixel p in question j,k Pixel value i x,y,t It is possible that the estimated historical average for .

[0197] The pure variance-based noise model described above has been found to produce good results in a wide range of environments. However, if the lighting conditions in an image change too rapidly, the noise map will initially be flooded with outliers. For image frame I following such changed lighting conditions, t In fact, the standard deviation estimates become inflated, which leads to some degree of blindness until the noise model stabilizes again.

[0198] The suitability of different variations of the presently described method may also vary depending on the hardware of the camera 110 used. For example, exposure and gain may be coarser or finer for different types of cameras, and aperture changes may be performed faster or slower.

[0199] It is then proposed to estimate a linear mapping between the mean intensity values ​​of the noise model and the pixel intensities of the new frame, i.e., find the values ​​of the variables α and β that minimize (16).

[0200] In (16), j is the image frame I t Pixel p at (geometrically) uniformly distributed pixel locations x,y A set of pixels p x,y may represent a sample or test set of

[0201] As a reminder, when defining the coefficients α and β, the same image frame I t Pixels p at different positions in x,yare considered, and such a pixel p x,y are compared with their corresponding positions in the noise model data.

[0202] In some embodiments, pixel p x,y The test set of images I t pixel p x,y In some embodiments, the pixel p x,y The test set of images I t pixel p x,y may include up to 80%, such as up to 50%, such as up to 25%, such as up to 10%, of the entire set of pixels p x,y The test set of images I t pixel p x,y For example, the set may be a set of images I t Spread throughout or image I t It is possible to form a uniform, sparse pattern that extends over at least 50% of the image I. t Whole or Image I t It is possible to form a sparse but evenly distributed set of vertical and / or horizontal full or broken lines distributed across at least 50% of the image. In some embodiments, overexposed pixels are not included in the test set. This can be determined by comparing pixel values ​​to a known threshold, often provided by the sensor manufacturer. If the threshold is not known, it can be easily determined experimentally.

[0203] Next, the equation for checking the bounds (i.e., (6) above or an equivalent formulation of this equation) is updated by:

[0204]

number

[0205]

number

[0206] Also, dispersion

[0207]

number

[0208] Since varies over time, the estimate of the variance also needs to be updated. Using the value from (4) is

[0209]

number

[0210] Unfortunately, this is not feasible, as its value would be inflated by changes in exposure, which is already compensated for by using . Instead, the variance estimate is updated by weighting it in the square of the current deviation.

[0211]

number

[0212] ξ∈]0,1[ is a coefficient that determines how much weight should be given to this deviation compared to the existing value. The larger ξ, the faster the noise model will adapt to fluctuations.

[0213] When applying a scaling factor α to the input, it is noted that the variances can be appropriately scaled and these are combined.

[0214]

number

[0215] This variant of the noise model is typically x,y Using one single-precision float (32 bits) for each

[0216]

number

[0217] In comparison, a pure variance noise model requires that each pixel p x,y The estimated variance of

[0218]

number

[0219] (as mentioned above)

[0220]

number

[0221] When this linear mapping model is used,

[0222]

number

[0223] is updated using (19). Since the definition is recursive, the variance of this pixel in the previous frame is either calculated from S and Q, or from the previous iteration's calculation of (19) for this pixel. The sum S x,y,t and the sum of squares Q x,y,t also uses image frame I to make the values ​​available as soon as the step effect ends. t It needs to be updated every time.

[0224] The predicted pixel value is

[0225]

number

[0226] To illustrate the case where the image is determined as shown in Figure 8, an example is now given as shown in Figure 8. In the chart shown in Figure 8, the Y axis represents the number of successive digital image frames I t One particular pixel p in the sequence x,y Pixel intensity of i x,y,t The x-axis shows the frame number. The window size n = 32, which means that the model is still initialized during the first 32 frames. After 32 frames have been processed, the model achieves the expected average μ x,y,t and dispersion

[0227]

number

[0228] contains enough information to make a prediction of . The line AVG shows the rolling average of the last 32 frames, which is the predictor determined according to (5).

[0229] As can be seen in the graph, the true signal value fluctuates around 2021 from the start until frame #60, where there is a sudden change in exposure time. The exposure time used can be provided as part of the frame's metadata. If the exposure time of a new frame differs significantly from the exposure time of the most recent frame, the levels will shift and while this occurs the model will be adversely affected, so the methods described in relation to (17)-(19) should be used.

[0230] As can also be seen on the graph, μ x,y,t takes 32 frames for μ to fully stabilize at the new level. Until that point is reached, μ x,y,tis not a particularly good predictor, since it lags behind. To compensate for this, the rolling average is subjected to a linear transformation according to (18). The result is shown in the graph as "Adj AVG". It can be clearly seen that this corresponds much better to the pixel values.

[0231] Similarly, as can be seen in the table below (corresponding to the graph in FIG. 8), the variance

[0232]

number

[0233] is around 250 before the exposure change, but swells to 4200 while the model is adapting. This is why the variance update method according to (19) is used. When processing frame 60, the process first uses a linear mapping to find the mean value μ x,y,t of

[0234]

number

[0235] The process is as follows:

[0236]

number

[0237] New pixel value i from x,y,t Calculate the deviation of and determine if it is out of bounds according to (20). If this is the first frame in which a change in exposure is noticed, then the variance of the previous frame

[0238]

number

[0239] is used, which is initially based on S and Q (determined as above), but where S and Q are still the pixel values ​​i from before the exposure change. x,y,t Finally, the next frame is used.

[0240]

number

[0241] is calculated according to (19), where # = frame number, PV = pixel value, AVG = AVG, and AAVG = Adj AVG.

[0242] [Table 3A]

[0243] [Table 3B]

[0244] S and Q can continue to be updated as above and used to allow the model to stabilize at new levels. Once we reach a point where α≈1 and β≈0, the mean and variance are again considered stable and we can return to the normal method of calculating variance.

[0245] Pixel value i assigned to the blob x,y,t Once the information about the pixel values ​​i assigned to the blob has been updated (and the noise map has also been updated), in the following step S6, x,y,t A blob is generated based on the

[0246] Blob generation is performed by dividing the individual pixels p x,yIt is important that the noise map generation is efficient, but the pixel values ​​i of interest in the noise map generation are x,y,t In blob generation, pixel p x,y More calculations per unit time can be provided.

[0247] The most recent pixel value i x,y,t Setting bounds based on the mean and sample standard deviation of image I works well in most cases, but t One notable difficulty arises when some of the pixel values ​​i are overexposed. In this case, the signal values ​​tend to saturate at some value near the upper end of the range, so that the affected pixel values ​​i x,y,t As θ stops varying over time, the standard deviation is also zero, which in turn means that even small changes will result in the creation of blobs.

[0248] To address this issue, an additional required minimum deviation can be added in step S5, which is used as an anti-saturation filter in the blob generation step,

[0249]

number

[0250] where μ x,y,t is the pixel value i x,y,t If the deviation is smaller than this, the pixel value i x,y,t is discarded as a non-blob pixel even though it exceeds the initial limit set up by the noise model.

[0251] Since square roots are expensive to compute,

[0252]

number

[0253] It is better to use q > 0 but << 1, and q 2 This suggests that is even smaller.

[0254]

number

[0255] where B is a positive number that controls the limit of filtering. Any number for B that gives a suitable filtering effect can be chosen, so it may be decided to choose an integer value. In some embodiments, B is at least 10, such as at least 50, such as at least 100. In some embodiments, B is at most 10,000, such as at most 1,000.

[0256] Then the condition is

[0257]

number

[0258] can be rewritten as:

[0259] Since the noise map and the noise model were updated in the same step, the noise model on which this currently considered noise map was based is already lost when arriving at the blob generation step. The noise model data is overwritten in the computer memory with each iteration of the method. However, μ x,y,t Since is stored in the noise map itself (with 15-bit precision in the example above), this value can be used instead when calculating (24). If the other terms are appropriately scaled (using fixed-point arithmetic), then (24) can also be calculated using only integer arithmetic.

[0260] After such a possible pair saturation filtering step, pixel values ​​i exceeding the limits of the noise model (as explained above in equations (1)-(24)) x,y,t are grouped together into multi-pixel blobs. This can be done using the Hoshen-Kopelman algorithm, which is a raster-scan method for forming such pixel groups that runs in linear time, and is well known in itself. During a first attempt, it x,y,t Check the pixel value i x,y,t exceeds the limit and pixel value i x,y,t is the neighboring pixel value i that belongs to the blob x±1,y±1,t If pixel value i x,y,t is added to the same blob. x,y,t Pixel values ​​i are classified into multiple adjacent blobs x±1,y±1,t , then they are stitched together into one single blob, with pixel value i x,y,t is added to the group. Finally, if there are no neighboring blobs, pixel value i x,y,t is registered as a new blob. For each blob, the following metrics can be aggregated: This provides different options for estimating the blob's center. One possibility is to use the absolute modulus of the deviation of the noise model:

[0261]

number

[0262] And another option is to weight the coordinates by the square of their deviation.

[0263]

number

[0264] [Table 4]

[0265] The experimental data so far shows that when the blob is small (blob size in pixels is 16 or less),

[0266]

number

[0267] but for larger blobs (number of pixels in blob > 32) use square-weighted

[0268]

number

[0269] We show that using the alternatives of , and interpolating between them for moderate sized blobs achieves good stereo matching.

[0270] Figure 7 shows the individual pixel values ​​i that were found to meet the criteria to be considered part of a blob at time t. x,y,t 1 shows an example clustering of four different detected blobs 1-4 based on

[0271] In a subsequent method step S7 performed by the target object tracker 140, the detected blobs are compared with the time-ordered sequence of digital images I in order to determine the path of the moving object through the space. t Such correlations can be used as a filtering mechanism, e.g., linear interpolation and / or implicit Newton's laws of motion, to remove blobs that do not move plausibly given a plausible model of the type of object being tracked.

[0272] If several cameras 110 are used, or if one or several cameras 110 are used together with sensors of another type of target object 120, the tracking information available from such available cameras 110 and any other sensors may be combined to determine the trajectories of one or several three-dimensional target objects 120 through space 111. This can be done, for example, using stereoscopic techniques that are well known per se.

[0273] In a subsequent step S8, one or several determined 2D and / or 3D trajectories of the target object 120 can be output to an external system and / or displayed graphically on a display of a trajectory monitoring device. For example, such displayed information can be used by a golfer using the system 100 to gain knowledge of the characteristics of a newly struck golf shot.

[0274] In a specific example, a user (such as a golfer) may be presented on a computer display screen with a visual 2D or 3D representation of the trajectory of a just-hit golf ball, detected using the above-described methods and systems, relative to a graphical representation, such as a virtual driving range. This provides the golfer with feedback that can be used to make decisions regarding various parts of a golf swing. The trajectory may also be part of a virtual experience in which the golfer plays, for example, a virtual golf hole, with the detected and displayed trajectory being represented as a golf shot in the virtual experience.

[0275] It is particularly noted that the amount of data required to process to achieve such a trajectory is substantial. For example, using a 10 Mpixel camera with an update rate of 100 images per second, 1 billion pixel values ​​per second need to be processed and evaluated for blob status. This analysis may be performed in a depicted space 111 that may include trees and other fine-grained objects that display rapidly changing light conditions, rapidly changing overall light conditions due to overcast conditions, etc. Using the systems and techniques described herein, it is possible to process data essentially in real time, such that, for example, a trajectory can be determined and output while the object is still in the air.

[0276] In the following step S9, the method ends.

[0277] As mentioned above, the present invention also relates to the system 100 itself, which includes the digital camera 110, the digital image analyzer 130, and the moving object tracker 140.

[0278] The digital camera 110 then captures a series of digital images I as described above. t The digital image analyzer 130 is positioned to render the space 111 to generate the pixel values ​​i as described above. x,y,t and determining the inequality with respect to one or several pixel values ​​i x,y,t The moving object tracker 140 is configured to store in computer memory information indicating that the sequence of digital images I is part of the detected blob. t The method is configured to correlate detected blobs across the network.

[0279] As also mentioned, the present invention also relates to a computer software product itself, which when executed on suitable hardware as described above, is configured to embody the digital image analyzer 130 and the moving object tracker 140. Thus, the computer software product processes a series of digital images I from the digital camera 110.t and performs the steps of the above-described method performed by the digital image analyzer 130 and the moving object tracker 140. For example, t may be provided as a continuous or semi-continuous stream of frames from digital camera 110 (and for each frame or set of frames received, the set of n most recent considered frames may be analyzed), or the entire set of N images may be received as one large batch and then analyzed. The computer software product may be executable on a computer belonging to system 100 and thus may form part of system 100.

[0280] Although several embodiments have been described above, it will be apparent to those skilled in the art that many modifications can be made to the disclosed embodiments without departing from the basic concept of the present invention.

[0281] For example, many additional data processing, filtering, conversion, etc. steps may be taken in addition to those described herein.

[0282] The generated blob data can be used in a variety of ways other than object tracking.

[0283] In general, anything that is said in relation to a method is equally applicable to a system and computer software product, and vice versa.

[0284] The present invention is therefore not limited to the described embodiments, but can be modified within the scope of the appended claims. [Explanation of symbols]

[0285] 10 Network 100 systems 110 Digital Camera 111 3D space 120 moving targets, balls, foreground objects 130 Digital Image Analyzer 140 Object Tracker 300 Data processing device 312 Hardware Processor 314 Additional Devices 316 Non-transitory computer-readable medium 318 Communication Interface 320 User Interface Devices 330 Programs 380 Networks, Communication Networks 390 Other Computers and User Interface Devices

Claims

1. 1. A method for tracking a moving object, comprising: A series of digital images (I) at successive times (t) t ), wherein said series of digital images (I t ) represents optical input from a three-dimensional space within the field of view of a digital camera, and the digital camera captures the series of digital images (I) having corresponding sets of pixels. t ), the sequence of digital images including corresponding pixel values, and the digital camera is arranged to capture the sequence of digital images (I t ) not moving relative to the three-dimensional space during generation; For two or more of the pixel values, determining an inequality that compares a first value to a second value, wherein the first value is a pixel of interest (p x,y ) pixel value (i x,y,t ) and predicted pixel value [Equation 1] and the second value is calculated as or based on the square of the difference between, firstly, the square of the number Z and, secondly, the pixel in question (p x,y ) estimated variance or standard deviation (σ x,y,t ) and the predicted pixel value is calculated as or based on the product of [Equation 2] The pixel in question (p x,y ) based on the historical pixel values, and the inequality is [Equation 3] and for each pixel value of said pixel values ​​and for the number n, [Equation 4] and [Equation 5] and for each pixel of said pixels, S x,y,t and Q x,y,t is stored in a computer memory; For pixel values ​​where the first value is higher than the second value, the pixel value (i x,y,t storing information in said computer memory indicating that the detected blob is part of the detected blob; and extracting the sequence of digital images (I) based on the information stored in the computer memory to determine a path of a moving object through the three-dimensional space. t ) and correlating the detected blobs across A method comprising:

2. If the inequality is [Equation 6] and [Equation 7] is the predicted pixel value, and σ x,y,t The pixel in question (p x,y ) the historical pixel values ​​(i x,y,{t-n,t-1} 2. The method of claim 1, wherein the estimated standard deviation is

3. Z is Z 2 is 10≦Z 2 10. The method of claim 1, wherein the number is selected to be an integer such that ≦20.

4. S x,y,t , Q x,y,t , or S x,y,t and Q x,y,t and are calculated recursively, whereby the pixel values ​​(i x,y,t ) is calculated for the pixel (p x,y ) but S at the previous time (t-1) x,y,t , Q x,y,t , or S x,y,t and Q x,y,t The method of claim 1 , wherein both and are calculated using previously stored calculated values.

5. S x,y,t But, S x,y,t = S x,y,t-1 + i x,y,t -i x,y,t-n It is calculated as follows, and Q x,y,t but, [Equation 8] The method of claim 4, wherein the calculated value is:

6. S x,y,t and Q x,y,t , pixels (p x,y 6. The method of claim 1, further comprising combining and storing in the computer memory a single data type containing 12 bytes or less for each of the plurality of data types.

7. For each pixel, the pixel value (i x,y,t 6. The method of claim 1, further comprising storing in the computer memory for a particular digital image a pixmap including the information indicating that the detected blob is part of the detected blob.

8. The pixel value (i x,y,t ) is part of the detected blob. x,y 8. The method of claim 7, wherein the first bit is indicated in a single bit for

9. The method of claim 7 , wherein the pixmap includes, for each pixel, a value indicating an expected pixel value for the pixel.

10. 10. The method of claim 9, wherein the value indicating the predicted pixel value of the pixel is achieved by storing the predicted pixel value as a fixed-point decimal number using a total of 15 bits for integer and fractional parts.

11. the predicted pixel value [Equation 9] , the estimated variance or standard deviation (σ x,y,t ), or both the predicted pixel value and the estimated variance or standard deviation are x,y ) n historical pixel values ​​(i x,y,{t-n,t-1} 6. The method of claim 1, wherein n is calculated based on a set of n=1, n=2, n=3, n=4, n=5, n=6, n=7, n=8, n=9, n=10, n=11, n=12, n=13, n=14, n=15, n=16, n=17, n=18, n=19, n=20, n=21, n=22, n=23, n=

12. The estimated variance or standard deviation (σ x,y,t 6. The method according to claim 1, wherein the number n of previous images taken into account for the estimation of σ ⁢ ...

13. 6. The method according to claim 1, wherein the pixel values ​​have a depth across one or several channels between 8 and 48 bits.

14. the predicted pixel value [Equation 10] is the set of digital images (I t ) pixel (p x,y ) and the estimated estimated future mean pixel value (μ x,y,t 6. The method of claim 1, wherein the value of the saturation level is determined based on the saturation level.

15. the predicted pixel value [0011] but, [0012] and α and β are determined as follows: j,k (i j,k,t - (αμ j,k,t +β)) 2 is a constant determined to minimize μ j,k,t The pixel in question (p j,k 15. The method of claim 14, wherein j and k are the estimated estimated future mean pixel values ​​of j and k, respectively, and j and k are iterated over a test set of pixels.

16. μ x,y,t The pixel in question (p x,y 16. The method of claim 15, wherein the estimated historical average of pixel values ​​of

17. The method of claim 15 , wherein the test set of pixels comprises between 1% and 25% of the entire set of pixels of a given image.

18. The method of claim 17 , wherein the test set of pixels is geometrically uniformly distributed across the entire set of pixels of the given image. 【Request 19】 【Number 13】 The estimated standard deviation (σ x,y,t 16. The method of claim 15, comprising determining ξ∈]0,1[.

20. determining that at least one of α is further from 1 than a first threshold and β is further from 0 than a second threshold is true; the predicted pixel value is calculated until it is determined that α is no longer farther from 1 than the first threshold and β is no longer farther from 0 than the second threshold. [0014] determining 16. The method of claim 15, comprising:

21. For the pixel values ​​where the first value is higher than the second value, the following inequality is satisfied: [Equation 15] The pixel value (i x,y,t ) is part of the detected blob, x,y,t is the pixel value in question, [0016] 2. The method of claim 1, comprising the step of: where B is the predicted pixel value and B is an integer such that B>100.

22. 22. The method of any one of claims 1 to 5 or 21, comprising using the Hoshen-Kopelman algorithm to group together individual adjacent pixels that are determined to be part of the same blob.

23. 22. The method of any one of claims 1 to 5 or 21, wherein the moving object is a golf ball.

24. 1. A system for tracking a moving object, comprising: A series of digital images (I) at successive times (t) t a digital camera (110) positioned to represent an optical input from a three-dimensional space within a field of view of the digital camera to generate a series of digital images (I) having corresponding sets of pixels; t ), the sequence of digital images including corresponding pixel values, and the digital camera is arranged to capture the sequence of digital images (I t a digital camera positioned so as not to move relative to the three-dimensional space during generation of the image; a computer having associated computer memory configured to execute the digital image analyzer; The digital image analyzer is configured to determine, for two or more of the pixel values, an inequality that compares a first value to a second value, the first value being a value that is greater than or equal to the pixel of interest (p x,y ) pixel value (i x,y,t ) and predicted pixel value [Equation 17] and the second value is calculated as or based on the square of the difference between, firstly, the square of the number Z and, secondly, the pixel in question (p x,y ) estimated variance or standard deviation (σ x,y,t ) and the predicted pixel value is calculated as or based on the product of [Equation 18] The pixel in question (p x,y ) based on the historical pixel values, and the inequality is [Equation 19] and for each pixel value of said pixel values ​​and for the number n, [Equation 20] and [0000] and for each pixel of said pixels, S x,y,t and Q x,y,t is stored in the computer memory, The digital image analyzer determines, for pixel values ​​where the first value is greater than the second value, whether the pixel value (i x,y,t a computer configured to store information in the computer memory indicating that the detected blob is part of the detected blob; and extracting the sequence of digital images (I) based on the information stored in the computer memory to determine a path of a moving object through the three-dimensional space. t a moving object tracker configured to correlate detected blobs across Including, the system.

25. 25. The system of claim 24, wherein the digital image analyzer is configured to perform the operations of any one of claims 2 to 23.

26. 24. A non-transitory computer-readable medium encoding a computer program product configured to perform the operations of any one of claims 1 to 23.

Citation Information

Patent Citations

  • Motion Based Pre-Processing of Two-Dimensional Image Data Prior to Three-Dimensional Object Tracking With Virtual Time Synchronization

    US20220051420A1