System and method for detecting motion in a 3D data reconstruction process

By generating temporal pixel images and calculating derived values, motion is detected using local correspondence density and temporal modulation. This solves the problem of insufficient motion detection capability in 3D data reconstruction of stereo vision systems, and improves detection speed and sorting efficiency.

CN112233139BActive Publication Date: 2026-04-14COGNEX CORP
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-06-28
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing stereo vision systems have limited motion detection capabilities during 3D data reconstruction, resulting in invalid picking points, reduced throughput in the sorting process, and increased error rates.

Method used

By generating temporal pixel images, calculating derived values ​​and correspondence data, and utilizing local correspondence density and temporal modulation to analyze whether motion exists in the scene, motion in the scene can be quickly detected.

Benefits of technology

It improves the speed and accuracy of motion detection, reduces invalid picking points, and enhances the efficiency and accuracy of the sorting process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112233139B_ABST
    Figure CN112233139B_ABST
Patent Text Reader

Abstract

The technology described herein relates to aspects of systems, methods, and computer-readable media for detecting motion in a scene. A first temporal pixel image is generated based on a first set of images of a scene over time, and a second temporal pixel image is generated based on a second set of images. One or more derived values are determined based on temporal pixel values in the first temporal pixel image, the second temporal pixel image, or both. Correspondence data is determined based on the first temporal pixel image and the second temporal pixel image to represent a set of correspondences between image points of the first set of images and image points of the second set of images. An indication is determined based on the one or more derived values and the correspondence data, the indication to indicate whether motion occurred in the scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The techniques described in this article typically involve detecting motion during three-dimensional (3D) reconstruction, and more particularly involve using two-dimensional images captured in the scene to detect motion of the scene during 3D reconstruction. Background Technology

[0002] Advanced machine vision systems and their underlying software are increasingly being used in various manufacturing and quality control processes. Machine vision enables faster, more accurate, and repeatable production of both mass-produced and customized products. A typical machine vision system includes: one or more cameras for focusing on an area of ​​interest; a frame capture / image processing element for capturing and transmitting images; a computer or onboard processing equipment; and a user interface for running the machine vision software application and manipulating the captured images. Appropriate lighting is also provided over the area of ​​interest.

[0003] One form of 3D vision system is based on stereo cameras, which employ at least two cameras positioned side-by-side with a baseline of one to several inches between them. Stereo vision-based systems typically rely on epipolar geometry constraints and image correction techniques. They use correlation-based methods, or in combination with relaxation techniques, to find correspondences from corrected images acquired by two or more cameras. However, conventional stereo vision systems have limitations in motion detection when creating 3D data reconstructions of objects. Summary of the Invention

[0004] In some aspects of this application, systems, methods, and computer-readable media for detecting motion during 3D reconstruction of a scene are provided.

[0005] In some aspects of this application, a system for detecting motion in a scene is disclosed. The system includes a processor communicating with a memory configured to execute instructions stored in the memory to: access a first set of images and a second set of images of the scene over time; generate a first time-pixel image comprising a first set of time pixels, wherein each time pixel in the first set of time pixels includes a set of pixel values ​​located at a relevant position in each of the first set of images; generate a second time-pixel image comprising a second set of time pixels, wherein each time pixel in the second set of time pixels includes a set of pixel values ​​located at a relevant position in each of the second set of images; determine one or more derived values ​​based on the time pixel values ​​of the first time-pixel image, the second time-pixel image, or both; determine correspondence data based on the first time-pixel image and the second time-pixel image, the correspondence data indicating a set of correspondences between image points in the first set of images and image points in the second set of images; and determine an indication based on the one or more derived values ​​and the correspondence data, the indication indicating whether motion has occurred in the scene.

[0006] In some examples, determining one or more derived values ​​includes: determining a first set of derived values ​​based on time pixel values ​​in a first time pixel image, and determining a second set of derived values ​​based on time pixel values ​​in a second time pixel image.

[0007] In some examples, determining one or more derived values ​​includes: for each time pixel of a first group of time pixels in a first time pixel image, determining first average value data to indicate the average value of the time pixel values, and for each time pixel of the first group of time pixels, determining first deviation value data to indicate the deviation value of the time pixel values.

[0008] In some examples, determining one or more derived values ​​further includes: for each time pixel of a second group of time pixels in the second time pixel image, determining second average data to indicate the average value of the time pixel values, and for each time pixel of the second group of time pixels, determining second deviation data to indicate a deviation value of the time pixel values. Calculating the first average data may include: calculating for each time pixel in the first group of time pixels: a time average value of the time pixel brightness value; and a root mean square deviation value of the time pixel brightness value.

[0009] In some examples, determining an indication includes: identifying multiple regions of a first time-pixel image, a second time-pixel image, or both; determining an average of one or more derived values ​​associated with each of the multiple regions; determining a correspondence indication based on a correspondence associated with the region; and determining a region indication based on the average and the correspondence indication, the region indication indicating whether motion has occurred in the region. Determining the region indication includes: determining that the average satisfies a first metric; determining that the correspondence indication satisfies a second metric; and generating a region indication to indicate the likelihood of motion in the region. An indication for indicating the likelihood of motion in a scene can be determined based on a set of region indications associated with each of the multiple regions.

[0010] In some examples, each of the first and second sets of images captures a relevant portion of the light pattern projected onto the scene, with each image in the first set having a first viewpoint of the scene and each image in the second set having a second viewpoint of the scene.

[0011] In some examples, each image in the first set of images is captured by a camera, and each image in the second set of images contains a portion of a pattern sequence projected onto the scene by a projector.

[0012] In some embodiments of this application, a computerized method for detecting motion in a scene is disclosed. The method includes: accessing a first set of images and a second set of images of the scene over time; generating a first time-pixel image based on the first set of images, comprising a first set of time pixels, wherein each time pixel in the first set of time pixels includes a set of pixel values ​​located at a relevant position in each of the first set of images; generating a second time-pixel image based on the second set of images, comprising a second set of time pixels, wherein each time pixel in the second set of time pixels includes a set of pixel values ​​located at a relevant position in each of the second set of images; determining one or more derived values ​​based on the time pixel values ​​of the first time-pixel image, the second time-pixel image, or both; determining correspondence data based on the first time-pixel image and the second time-pixel image, the correspondence data indicating a set of correspondences between image points in the first set of images and image points in the second set of images; and determining an indication based on the one or more derived values ​​and the correspondence data, the indication indicating whether motion has occurred in the scene.

[0013] In some examples, determining one or more derived values ​​includes: determining a first set of derived values ​​based on time pixel values ​​in a first time pixel image, and determining a second set of derived values ​​based on time pixel values ​​in a second time pixel image.

[0014] In some examples, determining one or more derived values ​​includes: for each time pixel of a first group of time pixels in a first time pixel image, determining first average data to indicate the average value of the time pixel values, and for each time pixel of the first group of time pixels, determining first deviation data to indicate a deviation value of the time pixel values. Determining one or more derived values ​​also includes: for each time pixel of a second group of time pixels in a second time pixel image, determining second average data to indicate the average value of the time pixel values, and for each time pixel of the second group of time pixels, determining second deviation data to indicate a deviation value of the time pixel values. Calculating the first average data may include: calculating for each time pixel in the first group of time pixels: a time average value of the time pixel brightness value; and a root mean square deviation value of the time pixel brightness value.

[0015] In some examples, determining an indication includes: identifying multiple regions of a first time-pixel image, a second time-pixel image, or both; determining an average of one or more derived values ​​associated with each of the multiple regions; determining a correspondence indication based on a correspondence associated with the region; and determining a region indication based on the average and the correspondence indication, the region indication indicating whether motion has occurred in the region. Determining the region indication includes: determining that the average satisfies a first metric; determining that the correspondence indication satisfies a second metric; and generating a region indication to indicate the likelihood of motion in the region. An indication for indicating the likelihood of motion in a scene can be determined based on a set of region indications associated with each of the multiple regions.

[0016] In some examples, each of the first and second sets of images captures a relevant portion of the light pattern projected onto the scene, with each image in the first set having a first viewpoint of the scene and each image in the second set having a second viewpoint of the scene.

[0017] In some aspects of this application, there is a relation to at least one non-transitory computer-readable storage medium. This at least one non-transitory computer-readable storage medium stores processor-executable instructions that, when executed by at least one computer hardware processor, cause the at least one computer hardware processor to perform the following actions: accessing a first set of images and a second set of images of a scene over time; generating a first time-pixel image including a first set of time pixels based on the first set of images, wherein each time pixel in the first set of time pixels includes a set of pixel values ​​located at a relevant position in each of the first set of images; generating a second time-pixel image including a second set of time pixels based on the second set of images, wherein each time pixel in the second set of time pixels includes a set of pixel values ​​located at a relevant position in each of the second set of images; determining one or more derived values ​​based on the time pixel values ​​of the first time-pixel image, the second time-pixel image, or both; determining correspondence data based on the first time-pixel image and the second time-pixel image, the correspondence data indicating a set of correspondences between image points in the first set of images and image points in the second set of images; and determining an indication based on the one or more derived values ​​and the correspondence data, the indication indicating whether motion has occurred in the scene.

[0018] Therefore, this application has already outlined the features of the disclosed subject matter quite broadly in order to better understand its subsequent detailed description and its contribution to the art. Of course, additional features of the disclosed subject matter, which will form the subject matter of the appended claims, will be described below. It should be understood that the wording and terminology used herein are for descriptive purposes and should not be considered restrictive. Attached Figure Description

[0019] The same or similar components shown in the various figures are indicated by the same reference numerals. For clarity, not every component is labeled in each figure. The figures are not necessarily drawn to scale, but are intended to illustrate various aspects of the technology and apparatus described herein.

[0020] Figure 1 An exemplary configuration with a projector and two cameras is shown according to some embodiments, wherein the two cameras are arranged to capture images of a scene including one or more objects in a consistent manner to produce stereoscopic image correspondences;

[0021] Figure 2 A pair of exemplary stereoscopic images corresponding to one of a series of projected light patterns are shown according to some embodiments;

[0022] Figure 3 A pair of exemplary stereoscopic images of a scene according to some embodiments are shown;

[0023] Figure 4 A pair of exemplary stereoscopic time-series images corresponding to a series of light patterns projected onto a scene, according to some embodiments, are shown;

[0024] Figure 5 An exemplary computerized method for determining whether motion is occurring in a scene, according to some embodiments, is illustrated;

[0025] Figure 6 An exemplary computerized method for determining the average value and deviation value of first and second time pixel images according to some embodiments is shown;

[0026] Figure 7 An exemplary computerized method for analyzing multiple regions to determine whether motion has occurred in the region is illustrated according to some embodiments;

[0027] Figure 8 A continuous high-quality image sequence is shown, displaying the first and second halves of a camera's original input image to illustrate example motion detected according to motion detection techniques of some embodiments;

[0028] Figure 9 An exemplary correspondence density mapping of different regions detected by motion detection techniques according to some embodiments is shown;

[0029] Figure 10 Exemplary average time values ​​of different regions detected by motion detection techniques according to some embodiments are shown;

[0030] Figure 11 Exemplary average deviation values ​​for different regions detected by motion detection techniques according to some embodiments are shown; and

[0031] Figure 12 An exemplary schematic diagram is shown, illustrating motion in different regions detected by motion detection techniques according to some embodiments. Detailed Implementation

[0032] The techniques described herein typically involve detecting motion within a three-dimensional (3D) scene during the reconstruction of a 2D image. The inventors have discovered and recognized that imaging applications, such as sorting applications (e.g., for picking, packaging, etc.) that utilize 3D data to separate multiple objects, can be affected by motion. For example, in an item picking application, 3D data can be used to locate an object and determine a picking point to approach that location to attempt to pick it up. Between the time interval from when the image of the scene used for 3D reconstruction begins to be captured to when an object is attempted to be picked up, if the object's position in the scene changes (e.g., if the chute has been filled by a new object during measurement), the picking point may be invalid. For example, if the particular object has not yet moved or has not moved far enough, the picking point may still be valid and will result in a successful pick. However, if the particular object has moved too far or has been covered by another object, the picking point may be invalid. If the picking point is invalid, the attempted picking action may not succeed (e.g., no picking), resulting in the wrong item being picked or potentially duplicate picking. Invalid picking points significantly reduce the throughput of the picking process. For example, it may require checking for picking errors, and another picking action must be performed only after the picking error has been resolved. Therefore, invalid picking points reduce throughput and increase the picking error rate.

[0033] The technique described herein can be used to identify object movement before providing a picking point to a customer. In some embodiments, the technique uses data acquired for 3D reconstruction to determine whether movement occurred in the scene when the image was acquired. If movement is detected in the scene, the technique allows skipping the determination of the picking point in the scene and instead recapturing data to obtain sufficient motion-free data. By utilizing information used in and / or computed as part of the 3D reconstruction process, this technique is much faster than other methods used to detect motion in images. For example, this technique can be executed in less than 2 ms, while optical-flow approaches may take 400 ms or longer. Optical-flow approaches are time-consuming, for example, because they require computation to track patterns in the scene (e.g., textures on boxes, such as letters or barcodes). For example, such techniques typically require segmenting multiple objects on the image over time and tracking those objects. This application avoids performing the computationally intensive processing described above and instead utilizes a subset of data generated during the 3D reconstruction process.

[0034] In some embodiments, these techniques include: detecting motion using structured light 3D sensing techniques that project structured light patterns onto a scene. When the scene is illuminated by the structured light pattern, the technique can obtain a sequence of stereoscopic images of the scene over time. The technique can utilize stereoscopic image correspondences of the temporal image sequence by leveraging local correspondence density, which reflects the number of correspondences found between stereoscopic image sequences in a particular region. In some embodiments, metrics, such as temporal averages and / or temporal deviations, can be calculated for each temporal image pixel. Spatial averages can be created for the temporal image sequence in the region for the time averages and / or temporal deviations. For example, a correspondence density value can be determined for each region by dividing the number of correspondences found in the region by the maximum number of possible correspondences in that region. For example, a quality criterion can be calculated for each region using previously calculated values. The calculated quality criterion can be used to determine the motion state (e.g., motion, some motion, no motion) for each region. For example, motion in a region can be determined if the correspondence density is less than a threshold, the average temporal deviation value is greater than a threshold, and the average temporal deviation value divided by the time average is still greater than a threshold.

[0035] In the following description, numerous specific details are set forth in relation to the systems and methods of the disclosed subject matter, as well as the environments in which such systems and methods operate, in order to provide a thorough understanding of the disclosed subject matter. Furthermore, it will be understood that the examples provided below are exemplary, and other systems and methods may be conceived within the scope of the disclosed subject matter.

[0036] Figure 1 An illustrative embodiment 100 of a machine vision system is shown, comprising a projector 104 and two cameras 106A, 106B (collectively referred to as cameras 106), arranged to capture images of an object or scene 102 in a consistent manner to produce stereoscopic image correspondences. In some embodiments, the projector 104 is employed to temporally encode a sequence of images of the object captured by the plurality of cameras 106. For example, the projector 104 may project a rotating random pattern onto the object, and each camera may capture a sequence of images including 12-16 images (or other number of images) of the object. Each image comprises a set of pixels constituting the image. In some embodiments, the light pattern may be shifted in the horizontal and / or vertical directions such that the pattern rotates on the object or scene (e.g., the pattern itself does not rotate clockwise or counterclockwise).

[0037] Each camera 106 may include a charge-coupled device (CCD) image sensor, a complementary metal-oxide-semiconductor (CMOS) image sensor, or another suitable image sensor. In some embodiments, each camera 106 may have a rolling shutter, a global shutter, or another suitable shutter type. In some embodiments, each camera 106 may have a GigE Vision interface, a Universal Serial Bus (USB) interface, a coaxial interface, a FireWire interface, or another suitable interface. In some embodiments, each camera 106 may have one or more intelligent functions. In some embodiments, each camera 106 may have a C-mount lens, an F-mount lens, an S-mount lens, or another suitable lens type. In some embodiments, each camera 106 may have a spectral filter suitable for a projector (e.g., projector 104) to block ambient light outside the projector's spectral range.

[0038] Figure 2 An exemplary pair of stereoscopic images 200 and 250 corresponding to one of a series of projected light patterns is shown. For example, projector 104 projects a light pattern onto an object, and camera 106 can capture stereoscopic images 200 and 250. In some embodiments, in order to reconstruct 3D data from a sequence of stereoscopic images from two cameras, it is necessary to find corresponding pixel pairs, such as pixels 202 and 252, between the images from each camera.

[0039] Figure 3 An exemplary pair of stereoscopic images 300 and 350 (and associated pixels) and corresponding pixels 302 and 352 are shown, representing the same portion of a pattern projected onto the two images 300 and 350. As described above, a projector 104 can project a light pattern onto a scene, and a camera 106 can capture stereoscopic images 300 and 350. In some embodiments, to reconstruct 3D data from a sequence of stereoscopic images from two cameras, it is necessary to find corresponding pixel pairs, such as pixels 302 and 352, between the images from each camera. In some embodiments, a sequence of stereoscopic images captured over time is used to identify the correspondence. Figure 3 and Figure 4The diagram shows a pair of stereo images. As the projector 104 continuously projects different light patterns onto the scene over time, cameras 106A and 106B can capture stereo temporal image sequences 400 and 450 with corresponding temporal pixels 402 and 452. These temporal pixels 402 and 452 capture the intensity fingerprint of the pixel as it changes over time. Each of cameras 106A and 106B can capture a sequence of 1, 2, 3, 4, ..., N images over time. Temporal pixels 402 and 452 are based on pixels (i, j) and (i', j') on the stereo temporal image sequences 400 and 450, respectively. Over time, each temporal pixel includes an ordered list of grayscale values: G_i_j_t, where t represents discrete-time instances 1, 2, 3, 4, ..., N.

[0040] In some embodiments, although in Figure 1 Two cameras 106 are shown, but the system can also be configured with only one camera. In the case of a single camera, for example, a projector can be configured to project a known light structure (e.g., such that the projector is used as a reverse camera). One temporal image sequence can be captured by this camera, while another temporal image sequence (e.g., used to establish correspondences, as described herein) can be determined based on the projected pattern structure or a sequence of pattern structures. For example, a correspondence can be established between an image sequence acquired by the single camera and a statically stored pattern sequence projected by the projector. Therefore, although some examples discussed herein involve the use of multiple cameras, this is for illustrative purposes only, and it should be understood that the techniques described herein can be used with systems having only one camera or more than two cameras.

[0041] In some embodiments, a normalized cross-correlation algorithm using temporal images or only a subset of temporal images can be applied to the two image sequences to determine pixel pairs corresponding to each image (e.g., having similar temporal gray values). For example, for each pixel of the first camera, a normalized cross-correlation of all feasible candidates along the epipolar line of the second camera with a threshold is performed to compensate for bias values ​​due to camera calibration, such as + / - one pixel or another suitable value.

[0042] In some aspects, the described system and method perform correspondence assignment between image points in a subset or all image points of paired images from a stereo image sequence in multiple steps. As a first exemplary step, an initial correspondence search is performed to coarsely estimate potential correspondences between image points derived from a subset or all paired images of the stereo image sequence. The initial correspondence search can be performed using temporal pixel values, thus allowing for pixel-level precision. As a second exemplary step, based on the potential correspondences derived from the first step, a correspondence refinement step is performed to more precisely locate the correspondences between image points in a subset or all paired images of the stereo image sequence. Correspondence refinement can be achieved by interpolating grayscale values ​​within the subset or all paired images of the stereo image sequence, these grayscale values ​​being located near the initial image points derived in the initial correspondence search. Correspondence refinement can be performed using subpixel values, thus providing greater accuracy than the pixel-level analysis in the first step. In one or two steps, the normalized cross-correlation algorithm described above can be applied to derive the latent and / or exact correspondences between image points in the two analyzed images. Triangulation can be performed on all found and established stereo correspondences (e.g., exceeding a certain metric, such as a similarity threshold) to compute 3D points for each correspondence, where the entire set of points can be referred to as 3D data. Further details can be found in PCT Publication (WO2017220598A1), the entire contents of which are incorporated herein by reference.

[0043] In some embodiments, as described herein, two cameras are used to capture stereo image sequences of an object, wherein after image acquisition, each image sequence will include 12-16 images of the object. To perform correspondence assignment on the stereo image sequences from the two cameras, the two steps described above can be performed. For the first step, an initial correspondence search is performed to associate each image point in the first image sequence with a corresponding image point in the second image sequence to find the image point with the highest correlation. In the example where each image sequence includes 16 images, this association is performed by using 16 temporal grayscale values ​​of each image point as a correlation "window" and associating suitable image point pairs from camera 1 and camera 2. At the end of the first step, the resulting coarse estimate provides potential candidates for potential correspondences, and these candidates will be pixel-level accurate because the search is performed using pixel values. For the second step, correspondence refinement can be performed to derive more accurate correspondences from the potential correspondences with sub-pixel precision. In the example where each image sequence comprises 16 images, based on the grayscale value sequence of each pixel in the images of the first image sequence, the correspondence refinement process interpolates grayscale values ​​into a subset or all pairs of images in the second image sequence. These grayscale values ​​are located near the initial image points obtained in the first step. In this example, performing correspondence refinement involves interpolating 16 interpolated values ​​at given subpixel positions into the images of the second image sequence. This association can be performed within a time window of the image points of camera 1 and an interpolation time window at the subpixel positions of camera 2.

[0044] The techniques described herein are used to analyze changes over time in captured image sequences within a region, where these image sequences are used to generate 3D data of a scene. The inventors have discovered and recognized that various metrics can be used to determine and / or estimate the presence of motion in a scene, such as temporal modulation (e.g., brightness varying over time due to a rotating pattern being projected onto the scene) and / or correspondence density (e.g., the ratio of the number of correspondences found between stereo image sequences during a correspondence search to the total number of possible correspondences). For example, temporal modulation can indicate the contrast of a pattern at a particular location over time (e.g., by representing maximum and minimum brightness over time), and a low temporal modulation can reflect motion. As another example, correspondence search can serve as an indicator of motion because it may fail to find corresponding pixels where motion exists. In some embodiments, these techniques compare the temporal modulation of image sequences captured in spatial regions (e.g., regions smaller than the total captured image size) with the correspondence density achieved in each of these regions to determine the presence of motion in the scene. For example, if static objects are present and the temporal sequence is well modulated, the image sequence may exhibit a high correspondence density, thus not indicating motion in the scene. However, if the object is moving, the time series may still be well modulated, but the correspondence density may be very low.

[0045] In some embodiments, the techniques described herein use temporal pixel images to determine whether motion exists in a scene. These techniques include: determining one or more derived values ​​of temporal pixels, correspondence data, or both (e.g., temporal average, temporal RMSD, etc.), and determining whether motion exists based on one or more temporal metrics and the correspondence data. Figure 5 An exemplary computerized method 500 for determining whether motion has occurred in a scene, according to some embodiments, is illustrated. In step 502, a computing device receives and / or accesses (e.g., from memory and / or local or remote memory or shared memory) first and second temporal image sequences of the scene. In step 504, the computing device generates a first temporal pixel image and a second temporal pixel image, respectively, for the first and second sequences. In step 506, the computing device determines derived values ​​for the first and second temporal pixel images. In step 508, the computing device determines correspondence data indicating a set of correspondences between image points in the first and second temporal pixel images. In step 510, the computing device determines whether motion exists in the scene based on the derived values ​​and the correspondence data.

[0046] In step 502, the computing device receives a time series of images of the scene as they change over time. Each image in the image time series can capture a relevant portion of the light pattern projected onto the scene, for example... Figures 1-2 The described rotating pattern. Each image in the first time-series image can be a first-view view of the scene (e.g., as shown by camera 106A), and each image in the second time-series image can be a second-view view of the scene (e.g., as shown by camera 106B).

[0047] In step 504, the computing device generates a time-pixel image. For example... Figure 4 As shown, each temporal pixel image includes a set of temporal pixels at a relevant location (e.g., pixel) in the original image sequence. For example, a first temporal pixel image includes a set of temporal pixels, each temporal pixel including a set of pixel values ​​located at a relevant pixel in each image of the first temporal image sequence. Similarly, for example, a second temporal pixel image includes a set of temporal pixels, each temporal pixel including a set of pixel values ​​located at a relevant pixel in each image of the second temporal image sequence.

[0048] In some embodiments, the computing device can process the temporal pixel images to generate various types of data. For example, the stereo image sequence acquired or captured during step 502 can be processed by normalizing this data and performing temporal correlation to construct 3D data. The computing device can generate one or more of a correspondence mapping map, a mean mapping map, and a deviation value mapping map (e.g., an RMSD mapping map). The correspondence mapping map can indicate the correspondence between temporal pixels in one temporal pixel image and temporal pixels in another temporal pixel image. For temporal pixel images, the mean mapping map can indicate the brightness variation of each temporal pixel. For temporal pixel images, the RMSD mapping map can indicate the RMSD value of the brightness of each temporal pixel.

[0049] In step 506, the computing device determines one or more derived values ​​based on the time pixel values ​​of a first time pixel image, a second time pixel image, or both. These derived values ​​can be determined, for example, using data determined in step 504 (such as a correspondence mapping map, a mean mapping map, and / or a RMSD mapping map). The derived values ​​may include, for example, a time average value indicating the average intensity value of a time pixel as it changes over time; a time deviation value (e.g., a root mean square deviation (RMSD) value indicating the deviation of the time pixel value); or both. For example, the time deviation value may indicate the pattern contrast as it changes over time (e.g., indicating the maximum and minimum brightness of a time pixel as it changes over time). Derived values ​​can be determined for each time pixel and / or set of time pixels in the time pixel image.

[0050] Figure 6An exemplary computer method 600 for determining average and deviation value data of first and second time-pixel images according to some embodiments is shown. In step 602, a computing device determines a first time-pixel average value of the first time-pixel image. In step 604, optionally, the computing device determines a second time-pixel average value of the second time-pixel image. In step 606, the computing device determines first time-pixel deviation value data of the first time-pixel image. In step 608, optionally, the computing device determines second time-pixel deviation value data of the second time-pixel image.

[0051] Reference Figure 6 The computing device can be configured to determine derived values ​​based solely on the first time-series pixel image (e.g., steps 602 and 606), and can be undetermined regarding the derived values ​​of the second time-series pixel image (e.g., steps 604 and 608). For example, a mean image map and / or a deviation image map can be generated separately for each camera. In some embodiments, the techniques described herein can be configured to use a mean image map and / or a deviation image map associated only with one camera. In such an example, the time-series image sequence acquired from the second camera can be used only to determine the corresponding pixels, and thus the correspondence density as described herein. In some embodiments, the techniques described herein can use a combination of derived values ​​generated for both cameras (e.g., a combination of mean image maps and / or deviation image maps from each camera).

[0052] In steps 508 and 510, Figure 7 An exemplary computerized method 700, according to some embodiments, for analyzing multiple regions to determine whether motion has occurred in those regions. In step 702, a computing device identifies multiple regions. In step 704, the computing device selects one region from the multiple regions. In step 706, the computing device determines average data based on time-averaged data associated with the selected region. In step 708, the computing device determines a correspondence indication based on a correspondence associated with the selected region. In step 710, the computing device determines whether motion has occurred in that region based on the average data and the correspondence indication. In step 712, the computing device determines whether there are more regions to analyze. If yes, the method returns to step 704. If no, the method continues to step 714, where the computing device determines whether possible motion exists in the scene.

[0053] In step 702, the computing device determines multiple regions. For example, the computing device may divide the original size of the images in an image sequence into a set of square regions, rectangular regions, and / or other regions useful for motion analysis. In some embodiments, the techniques described herein divide the original image into a set of N×N square images (e.g., where N is 30, 40, 50, 80 pixels, etc.). The techniques described herein can analyze the temporal information of the computation associated with each region. For example, (because such data can be specified based on time-pixel images) the techniques described herein can analyze the derived values ​​and / or correspondence data associated with each region. As shown in steps 704 and 712, the computing device may repeatedly analyze each region as described herein.

[0054] In step 706, the computing device determines spatial average data for the entire region based on temporal average data associated with the selected region. For example, the computing device may determine the spatial average data for that region. The techniques described herein may include generating one or more information maps for each region in a temporal pixel image. In some embodiments, the computing device may (e.g., using the mean of a mean map) determine the spatial average of the temporal average for each temporal pixel in that region, which the computing device may store as a representative average for that region. The spatial average of the temporal average can indicate, for example, how much the brightness of the entire region has changed over time. In some embodiments, the computing device may (e.g., using the RMSD value of an RMSD map) determine the spatial average of the temporal deviation value for each temporal pixel in that region, which the computing device may store as a representative average deviation value for that region. Since the average deviation value can provide an indication of brightness changes, it can, for example, be used as a confidence improvement for determination. For example, the average spatial deviation value can prevent the system from incorrectly determining the presence of motion in one or more specific regions due to low-light and / or dark areas in the scene.

[0055] In step 708, the computing device determines a correspondence indication based on correspondences associated with the selected region. In some embodiments, the techniques described herein can determine the correspondence indication based on the number of correspondences in the region. In some embodiments, the techniques described herein can determine the correspondence density by dividing the number of correspondences found in the region by the total number of possible correspondences in the region. For example, the number of correspondences can be divided by the number of temporal pixels in the region (e.g., N×N). In some embodiments, correspondences can be weighted before calculating the correspondence density. For example, to consider only higher-quality correspondences, the techniques described herein can determine a relevance score for each of the multiple correspondences found in the region and include only those correspondences with a relevance score exceeding a relevance threshold (e.g., 0.85, 0.90, or 0.95). This relevance score can be determined using normalized cross-correlation of a time window used herein, the value of which is taken from the interval [-1.0, 1.0].

[0056] In step 710, the computing device determines whether motion exists in the region based on the average data determined in step 706 and the correspondence indication determined in step 708. In some embodiments, the computing device compares the average data and / or the correspondence indication with a metric. For example, the computing device may determine whether the average satisfies a first metric, whether the correspondence indication satisfies a second metric, and generate a motion indication for the region based on the comparison result. In some embodiments, the techniques described herein may determine whether the region is marked as potentially containing motion by determining whether the corresponding density is less than a corresponding density threshold (e.g., 0.20, 0.25, 0.30, etc.), whether the average deviation value (e.g., the average RMSD value) is greater than an average deviation value threshold (e.g., 2.5, 3, 3.5, etc.), and / or whether the average deviation value divided by the average mean is greater than a relative average deviation value threshold (e.g., 0.190, 0.195, 0.196, 0.197, 0.200, etc.).

[0057] In step 714, the computing device analyzes the region indications of each region to determine whether motion exists in the scene. In some embodiments, the techniques described herein can utilize the region indications of motion to analyze the number of adjacent regions. For example, the computing device can sum and / or determine the size of the connected regions used to indicate motion and use the result as an indicator of motion throughout the scene. For example, the computing device can determine that if the number of multiple regions in a cluster of connected regions is greater than a threshold (e.g., 10 regions), then the scene may include motion. As described herein, if the techniques described herein identify a sufficient amount of potential motion in the scene, this information can be used downstream for purposes such as avoiding invalid pickup locations due to the motion of objects.

[0058] In some embodiments, the techniques described herein can use a mask to ignore one or more regions of a scene (e.g., a temporal pixel image). For example, a mask can be used to ignore regions within the camera's field of view that are irrelevant to the application (e.g., regions where movement may occur without causing errors). A mask can specify one or more regions of a captured image to be ignored from testing using the techniques described herein. For example, these techniques can be configured to ignore one or more motion regions caused by the movement of a robot, or to ignore regions in a conveyor belt with moving parts that are not of interest to the application (e.g., regions in the background).

[0059] Figures 8-11 Exemplary images illustrating various aspects of the techniques described herein are shown. Figure 8 Sequence maximum images 802 and 804 are shown, representing the pixel-wise maximum values ​​of the first and second halves of sixteen raw input images from a camera, respectively, to illustrate exemplary motion according to some embodiments of the technique. While capturing the raw images, a random pattern is projected onto and moved across the scene, as described herein. Regions 806A / 806B, 808A / 808B, and 810A / 810B illustrate example motion of the scene captured between maximum images 802 and 804. In this example, the envelope 812 shown in maximum image 802 is shifted downwards in the scene of maximum image 804.

[0060] Figure 9 An exemplary correspondence density mapping 900 for different regions is shown as a technical example according to some embodiments. The region in this example is 40×40 pixels. In this example, for each region, the number of valid correspondences with a correlation score higher than a threshold is divided by the region size (1600). In this example, the brighter the region, the higher the correspondence density. A black region indicates no correspondence. A completely white region indicates that all correspondences were found.

[0061] Figure 10 An exemplary average time mean map 1000 for different regions of a 40×40 pixel area is shown as a technical example according to some embodiments. This average map 1000 shows the spatial average of the time mean of pixels in different regions.

[0062] Figure 11 An exemplary average RMSD value mapping diagram 1100 for different regions of a 40×40 pixel area is shown as a technical example according to some embodiments. This average RMSD value mapping diagram 1100 shows the spatial average of the temporal RMSD values ​​of temporal pixels in different regions.

[0063] Figure 12 An exemplary motion map 1200 is shown, illustrating motion regions in different 40×40 pixel areas according to a technical example of some embodiments. The motion map 1200 highlights regions in the image where potential motion is detected, such as regions 1202 and 1204. As described herein, regions with motion can be regions with good image contrast / good pattern contrast, but few corresponding points, resulting in a low correspondence density. By graphically representing motion regions using the motion map 1200, the size of the connected motion regions can be visualized. The techniques described herein can sum the number of connected motion candidate regions, and if this number is greater than a threshold, it indicates that the image has motion.

[0064] While the techniques disclosed herein have been discussed in conjunction with stereo methods (e.g., temporal stereo methods, such as sequence acquisition), the techniques are not limited thereto. For example, these techniques can be used with single-image approaches (e.g., active and passive techniques).

[0065] The techniques operating according to the principles described herein can be implemented in any suitable manner. The processing and decision boxes in the flowcharts above represent steps and actions that can be included in algorithms that perform these various processes. Algorithms derived from these processes can be implemented as software integrated with one or more single-purpose or multi-purpose processors and used to direct its operation, or as functionally equivalent circuits (e.g., digital signal processing (DSP) circuits or application-specific integrated circuits (ASICs)), or can be implemented in any other suitable manner. It should be understood that the flowcharts included herein do not depict the syntax or operation of any particular circuit or any particular programming language or type of programming language. Rather, the flowcharts herein illustrate functional information for those skilled in the art to use in fabricating circuits or implementing computer software algorithms to perform processing procedures for a particular device used to perform the types of techniques described herein. It should also be understood that, unless otherwise indicated herein, the specific order of steps and / or actions described in each flowchart is merely an example of the implemented algorithms, and these algorithms can be modified in the implementation and implementation of the principles described herein.

[0066] Therefore, in some embodiments, the techniques described herein can be implemented as computer-executable instructions that are implemented as software, including application software, system software, firmware, middleware, embedded code, or any other suitable type of computer code. Such computer-executable instructions can be written using a variety of suitable programming languages ​​and / or programming or scripting tools, and can also be compiled into executable machine language code or intermediate code that executes on a framework or virtual machine.

[0067] When the techniques described herein are embodied as computer-executable instructions, these instructions can be implemented in any suitable manner, including as multiple functional facilities, each providing one or more operations to perform algorithms according to these techniques. However, a "functional facility" is instantiated, a structural component of a computer system, whose functionality enables one or more computers to perform specific operations when integrated with and executed by those computers. A functional facility can be part or all of a software element. For example, a functional facility can be implemented as a process, a functional component of a discrete process, or any other suitable processing unit. If the techniques described herein are implemented as multiple functional facilities, each functional facility can be implemented in its own way, and each functional facility need not be implemented in the same way. Furthermore, these functional facilities can be executed in parallel and / or serially as appropriate, and can exchange information with each other using shared memory on one or more computers on which they are executing, using message passing protocols or other suitable methods.

[0068] Typically, functional facilities include routines, programs, objects, components, data structures, etc., for performing specific tasks or implementing specific abstract data types. The functionality of functional facilities can generally be combined or distributed as needed within the system in which they operate. In some implementations, one or more functional facilities implementing the techniques described herein can together form a complete software package. In alternative embodiments, these functional facilities may be adapted to interact with other unrelated functional facilities and / or processes to enable software program applications.

[0069] This document has described some exemplary functional facilities for performing one or more tasks. However, it should be understood that the described division of functional facilities and tasks is only for illustrating the types of functional facilities that can implement the exemplary techniques described herein, and embodiments are not limited to implementation in any particular number, division, or type of functional facilities. In some implementations, all functions may be implemented in a single functional facility. It should also be understood that in some implementations, some of the functional facilities described herein may be implemented together with or separately from other functional facilities (i.e., as a single unit or separate units), or some of these functional facilities may not be implemented.

[0070] In some embodiments, computer-executable instructions for implementing the techniques described herein (when implemented as one or more functional facilities or in any other way) may be encoded on one or more computer-readable media to provide functionality to the media. Computer-readable media include magnetic media (e.g., hard disk drives), optical media (e.g., optical discs (CDs) or digital versatile disks (DVDs)), permanent or non-permanent solid-state storage (e.g., flash memory, magnetic RAM, etc.), or any other suitable storage media. Such computer-readable media may be implemented in any suitable manner. As used herein, a “computer-readable medium” (also referred to as a “computer-readable storage medium”) means a tangible storage medium. Tangible storage media are non-transitory and have at least one physical structural component. In the “computer-readable medium” as used herein, at least one physical structural component has at least one physical characteristic that can be altered in some way during processes including: the process of creating a medium with embedded information, the process of recording information thereon, or any other process of encoding the medium with information. For example, the magnetization state of a portion of the physical structure of the computer-readable medium may be altered during the recording process.

[0071] Furthermore, some of the aforementioned technologies include actions that store information (e.g., data and / or instructions) in certain ways for use by these technologies. In some implementations of these technologies, they are implemented, for example, as computer-executable instructions—the information can be encoded on a computer-readable storage medium. Where a particular structure is described herein as an advantageous format for storing the information, these structures can be used to give the information a physical organizational form when encoded on the storage medium. These advantageous structures can then provide functionality to the storage medium by influencing the operation of one or more processors interacting with the information, for example, by improving the efficiency of one or more processors in performing computer operations.

[0072] In some, but not all, implementations, the techniques described herein may be embodied as computer-executable instructions that can be executed by one or more suitable computing devices operating on any suitable computer system, or one or more computing devices (or one or more processors of one or more computing devices) may be programmed to execute these computer-executable instructions. When the instructions are stored in a computing device or processor in a manner accessible to the computing device or processor (e.g., stored in data memory (e.g., an on-chip cache or instruction register, a computer-readable storage medium accessible via a bus, a computer-readable storage medium accessible via one or more networks and accessible via the device / processor, etc.), the computing device or processor may be programmed to execute these instructions. Functional facilities including these computer-executable instructions may be integrated with and direct the operation of devices or systems including: a single multi-functional programmable digital computing device; a collaborative system of two or more multi-functional computing devices sharing processing power and jointly executing the techniques described herein; a single computing device dedicated to executing the techniques described herein, or a collaborative system of multiple computing devices (located in the same location or distributed in different locations); one or more field-programmable gate arrays (FPGAs) for executing the techniques described herein; or any other suitable system.

[0073] A computing device includes: at least one processor, a network adapter, and a computer-readable storage medium. The computing device may be, for example, a desktop or laptop computer, a personal digital assistant (PDA), a smartphone, a server, or any other suitable computing device. The network adapter may be any suitable hardware and / or software enabling the computing device to communicate wired and / or wirelessly with any other suitable computing device via any suitable computing network. The computing network may include wireless access points, switches, routers, gateways, and / or other networking devices, as well as any suitable wired and / or wireless communication media or media for exchanging data between two or more computers, including the Internet. The computer-readable medium may be adapted to store data to be processed by the processor and / or instructions to be executed. The processor processes the data and executes the instructions. The data and instructions may be stored on the computer-readable storage medium.

[0074] Additionally, a computing device may have one or more components and peripherals (including input and output devices). These devices can be used, in particular, to present a user interface. Examples of output devices that can be used to provide a user interface include printers or displays for outputting visual presentations, and speakers or other sound-generating devices for presenting auditory outputs. Examples of input devices that can be used to provide a user interface include keyboards and pointing devices (such as mice, touchpads, and digitizing tablets). As another example, a computing device may receive input information via speech recognition or other audible formats.

[0075] This document has described embodiments of implementing the technology in circuits and / or computer-executable instructions. It should be understood that some embodiments may be presented as methods, of which at least one example has been provided. Actions performed as part of a method may be ordered in any suitable manner. Thus, embodiments may be constructed in which actions are performed in an order different from the order shown, and some actions may be performed simultaneously even though they are shown as sequential actions in the illustrative embodiments herein.

[0076] Various aspects are described in this disclosure, including but not limited to the following:

[0077] (1) A system for detecting motion in a scene, the system including a processor in communication with a memory, the processor being configured to execute instructions stored in the memory, the instructions causing the processor to perform the following operations:

[0078] Access the first and second sets of images of the scene over time;

[0079] Based on the first set of images, a first time pixel image is generated, which includes the first set of time pixels, wherein each time pixel in the first set of time pixels includes a set of pixel values ​​located at a relevant position in each of the first set of images;

[0080] Based on the second set of images, a second time pixel image is generated, which includes the second set of time pixels, wherein each time pixel in the second set of time pixels includes a set of pixel values ​​located at a relevant position in each of the images in the second set of images;

[0081] One or more derived values ​​are determined based on the time pixel values ​​of a first time pixel image, a second time pixel image, or both.

[0082] Based on the first and second time-pixel images, correspondence data is determined, indicating the set of correspondences between image points in the first set of images and image points in the second set of images; and

[0083] Based on one or more derived values ​​and the corresponding relationship data, an indicator is determined to indicate whether motion has occurred in the scene.

[0084] (2) According to the system described in (1), determining one or more derived values ​​includes:

[0085] The first set of derived values ​​is determined based on the time pixel values ​​in the first time pixel image, and

[0086] The second set of derived values ​​is determined based on the time pixel values ​​in the second time pixel image.

[0087] (3) The system according to any one of (1) to (2), wherein determining one or more derived values ​​comprises:

[0088] For each time pixel in the first group of time pixels of the first time pixel image, determine first average value data to indicate the average value of the time pixel, and

[0089] For each time pixel in the first group of time pixels, determine the first deviation value data to indicate the deviation value of the time pixel value.

[0090] (4) The system according to any one of (1) to (3), wherein determining one or more derived values ​​further comprises:

[0091] For each time pixel in the second group of time pixels of the second time pixel image, determine second average data to indicate the average value of the time pixel, and

[0092] For each time pixel in the second group of time pixels, determine second deviation value data to indicate the deviation value of the time pixel value.

[0093] (5) The system according to any one of (1) to (4), wherein calculating the first average data includes:

[0094] Calculate for each time pixel in the first group of time pixels:

[0095] The time average of the pixel brightness value over time; and

[0096] The root mean square deviation of the pixel brightness value over time.

[0097] (6) The system according to any one of (1) to (5), wherein determining an instruction comprises:

[0098] Determine multiple regions of a first-time pixel image, a second-time pixel image, or both, and

[0099] For each of these multiple regions, the following is determined:

[0100] The average of one or more derived values ​​associated with this region;

[0101] A one-to-one correspondence indication determined based on the correspondence associated with the region; and

[0102] A region indicator is determined based on the average value and the correspondence indicator, and this region indicator is used to indicate whether movement has occurred in the region.

[0103] (7) The system according to any one of (1) to (6), wherein determining a region indication comprises:

[0104] This average value is determined to satisfy the first metric;

[0105] Determining that the correspondence indicates that the second metric is satisfied; and

[0106] Generate a region indicator to indicate the likelihood of movement within that region.

[0107] (8) The system according to any one of (1) to (7) further comprises:

[0108] Indicators for indicating the possibilities of motion in a scene are determined based on a set of area indicators associated with each of the multiple areas.

[0109] (9) The system according to any one of (1) to (8), wherein:

[0110] Each of the first and second sets of images captures a relevant portion of the light pattern projected onto the scene;

[0111] Each image in the first set of images has a first-person perspective of the scene; and

[0112] Each image in the second set of images presents a second perspective of the scene.

[0113] (10) The system according to any one of (1) to (9), wherein:

[0114] Each image in the first set was captured by a camera, and

[0115] Each image in the second set contains a sequence of patterns projected onto the scene by a projector.

[0116] (11) A computerized method for detecting motion in a scene, the method comprising:

[0117] Access the first and second sets of images of the scene over time;

[0118] Based on the first set of images, a first time pixel image is generated, which includes the first set of time pixels, wherein each time pixel in the first set of time pixels includes a set of pixel values ​​located at a relevant position in each of the first set of images;

[0119] Based on the second set of images, a second time pixel image is generated, which includes the second set of time pixels, wherein each time pixel in the second set of time pixels includes a set of pixel values ​​located at a relevant position in each of the images in the second set of images;

[0120] One or more derived values ​​are determined based on the time pixel values ​​of a first time pixel image, a second time pixel image, or both.

[0121] Based on the first and second time-pixel images, correspondence data is determined, indicating the set of correspondences between image points in the first set of images and image points in the second set of images; and

[0122] Based on one or more derived values ​​and the corresponding relationship data, an indicator is determined to indicate whether motion has occurred in the scene.

[0123] (12) According to the method of (11), determining one or more derived values ​​includes:

[0124] The first set of derived values ​​is determined based on the time pixel values ​​in the first time pixel image, and

[0125] The second set of derived values ​​is determined based on the time pixel values ​​in the second time pixel image.

[0126] (13) The method according to any one of (11) to (12), wherein determining one or more derived values ​​comprises:

[0127] For each time pixel in the first group of time pixels of the first time pixel image, determine first average value data to indicate the average value of the time pixel, and

[0128] For each time pixel in the first group of time pixels, determine the first deviation value data to indicate the deviation value of the time pixel value.

[0129] (14) The method according to any one of (11) to (13), wherein determining one or more derived values ​​further comprises:

[0130] For each time pixel in the second group of time pixels of the second time pixel image, determine second average data to indicate the average value of the time pixel, and

[0131] For each time pixel in the second group of time pixels, determine second deviation value data to indicate the deviation value of the time pixel value.

[0132] (15) The method according to any one of (11) to (14), wherein calculating the first average data comprises:

[0133] Calculate for each time pixel in the first group of time pixels:

[0134] The time average of the pixel brightness value over time; and

[0135] The root mean square deviation of the pixel brightness value over time.

[0136] (16) The method according to any one of (11) to (15), wherein determining an instruction comprises:

[0137] Determine multiple regions of a first-time pixel image, a second-time pixel image, or both, and

[0138] For each of these multiple regions, the following is determined:

[0139] The average of one or more derived values ​​associated with this region;

[0140] A one-to-one correspondence indication determined based on the correspondence associated with the region; and

[0141] A region indicator is determined based on the average value and the correspondence indicator, and this region indicator is used to indicate whether movement has occurred in the region.

[0142] (17) The method according to any one of (11) to (16), wherein determining a region indication comprises:

[0143] This average value is determined to satisfy the first metric;

[0144] Determining that the correspondence indicates that the second metric is satisfied; and

[0145] Generate a region indicator to indicate the likelihood of movement within that region.

[0146] (18) The method according to any one of (11) to (17) further comprises:

[0147] Indicators for indicating the possibilities of motion in a scene are determined based on a set of area indicators associated with each of the multiple areas.

[0148] (19) The method according to any one of (11) to (18), wherein:

[0149] Each of the first and second sets of images captures a relevant portion of the light pattern projected onto the scene;

[0150] Each image in the first set of images has a first-person perspective of the scene; and

[0151] Each image in the second set of images presents a second perspective of the scene.

[0152] (20) At least one non-transitory computer-readable storage medium storing processor-executable instructions that, when executed by at least one computer hardware processor, will cause the at least one computer hardware processor to perform the following actions:

[0153] Access the first and second sets of images of the scene over time;

[0154] Based on the first set of images, a first time pixel image is generated, which includes the first set of time pixels, wherein each time pixel in the first set of time pixels includes a set of pixel values ​​located at a relevant position in each of the first set of images;

[0155] Based on the second set of images, a second time pixel image is generated, which includes the second set of time pixels, wherein each time pixel in the second set of time pixels includes a set of pixel values ​​located at a relevant position in each of the images in the second set of images;

[0156] One or more derived values ​​are determined based on the time pixel values ​​of a first time pixel image, a second time pixel image, or both.

[0157] Based on the first and second time-pixel images, correspondence data is determined, indicating the set of correspondences between image points in the first set of images and image points in the second set of images; and

[0158] Based on one or more derived values ​​and the corresponding relationship data, an indicator is determined to indicate whether motion has occurred in the scene.

[0159] (21) The non-transitory computer-readable storage medium according to (20) is further configured to perform any one or more steps of (1) to (19).

[0160] The various aspects of the above embodiments can be used individually, in combination, or in various arrangements not specifically discussed in the foregoing embodiments. Therefore, their application is not limited to the details and arrangements of the components set forth in the foregoing description or shown in the accompanying drawings. For example, an aspect described in one embodiment can be combined in any way with aspects described in other embodiments.

Claims

1. A system for detecting motion in a scene, the system comprising a processor in communication with a memory, the processor being configured to execute instructions stored in the memory, the instructions causing the processor to perform the following operations: Access the first and second sets of images of the scene over time, where: Each image in the first set of images captures a relevant portion of a pattern sequence projected onto the scene over time, and each image in the first set of images has a first viewpoint of the scene; and Each image in the second set of images captures a relevant portion of a pattern sequence projected onto the scene over time, and each image in the second set of images has a second perspective of the scene; Based on the first set of images, a first time pixel image is generated, including a first set of time pixels, wherein each time pixel in the first set of time pixels includes a set of pixel values ​​located at a relevant position in each of the first set of images, and each of the first set of images captures a relevant portion of a pattern sequence projected onto the scene over time; Based on the second set of images, a second time pixel image is generated, which includes a second set of time pixels, wherein each time pixel in the second set of time pixels includes a set of pixel values ​​located at a relevant position in each of the second set of images, and each of the second set of images captures a relevant portion of a pattern sequence projected onto the scene over time; One or more derived values ​​are determined based on the time pixel values ​​of the first time pixel image, the second time pixel image, or both. Based on the first time-pixel image and the second time-pixel image, correspondence data is determined, wherein the correspondence indicates the set of correspondences between image points in the first group of images and image points in the second group of images; and Based on the one or more derived values ​​and the corresponding relationship data, an indication is determined, which is used to indicate whether motion has occurred in the scene; The determination of correspondence data includes determining the correspondence density based on the first time pixel image and the second time pixel image. The correspondence density indicates the relationship between the number of identified correspondences between image points of the first group of images and image points of the second group of images and the number of possible correspondences.

2. The system according to claim 1, further comprising: Based on the determination that motion occurs in the scene, the position of the object in the scene is determined at a future time after the time related to the time of capturing the first set of images and the second set of images.

3. The system according to claim 1, wherein, Determining one or more derived values ​​includes: The first set of derived values ​​is determined based on the time pixel values ​​in the first time pixel image, and The second set of derived values ​​is determined based on the time pixel values ​​in the second time pixel image.

4. The system according to claim 1, wherein, Determining one or more derived values ​​includes: For each time pixel in the first group of time pixels of the first time pixel image, a first average value data is determined to indicate the average value of the time pixel. For each time pixel in the first group of time pixels, a first deviation value data is determined to indicate the deviation value of the time pixel value.

5. The system according to claim 4, wherein, Determining one or more derived values ​​further includes: For each time pixel in the second group of time pixels of the second time pixel image, determine second average data to indicate the average value of the time pixel, and For each time pixel in the second group of time pixels, a second deviation value data is determined to indicate the deviation value of the time pixel value.

6. The system according to claim 4, wherein, The data used to calculate the first average include: Calculate for each time pixel in the first group of time pixels: The time average of the pixel brightness value over time; and The root mean square deviation of the pixel brightness value over time.

7. The system of claim 1, wherein determining an instruction comprises: Determine multiple regions of the first time pixel image, the second time pixel image, or both, and For each of the plurality of regions, the following is determined: The average of one or more derived values ​​associated with the region; A correspondence indication determined based on the correspondence associated with the region; and A region indicator is determined based on the average value and the correspondence indicator, and the region indicator is used to indicate whether movement has occurred in the region.

8. The system according to claim 7, wherein, Determining an area includes: The average value is determined to satisfy the first metric; Determining that the correspondence indicates that the second metric is satisfied; and Generate a region indicator to indicate the possibility of movement within the region.

9. The system according to claim 8, further comprising: Indicators for indicating the possibility of motion in the scene are determined based on a set of region indicators associated with each of the plurality of regions.

10. A computerized method for detecting motion in a scene, the method comprising: Access the first and second sets of images of the scene over time, where: Each image in the first set of images captures a relevant portion of a pattern sequence projected onto the scene over time, and each image in the first set of images has a first-person view of the scene; and Each image in the second set of images captures a relevant portion of a pattern sequence projected onto the scene over time, and each image in the second set of images has a second perspective of the scene; Based on the first set of images, a first time pixel image is generated, including a first set of time pixels, wherein each time pixel in the first set of time pixels includes a set of pixel values ​​located at a relevant position in each of the first set of images, and each of the first set of images captures a relevant portion of a pattern sequence projected onto the scene over time; Based on the second set of images, a second time pixel image is generated, which includes a second set of time pixels, wherein each time pixel in the second set of time pixels includes a set of pixel values ​​located at a relevant position in each of the second set of images, and each of the second set of images captures a relevant portion of a pattern sequence projected onto the scene over time; One or more derived values ​​are determined based on the time pixel values ​​of the first time pixel image, the second time pixel image, or both. Based on the first time-pixel image and the second time-pixel image, correspondence data is determined, wherein the correspondence data indicates the set of correspondences between image points in the first group of images and image points in the second group of images; and Based on the one or more derived values ​​and the corresponding relationship data, an indication is determined, which is used to indicate whether motion has occurred in the scene; The determination of correspondence data includes determining the correspondence density based on the first time pixel image and the second time pixel image. The correspondence density indicates the relationship between the number of identified correspondences between image points of the first group of images and image points of the second group of images and the number of possible correspondences.

11. The method of claim 10, further comprising: Based on the determination that motion occurs in the scene, the position of the object in the scene is determined at a future time after the time related to the time of capturing the first set of images and the second set of images.

12. The method of claim 10, wherein, Determining one or more derived values ​​includes: The first set of derived values ​​is determined based on the time pixel values ​​in the first time pixel image, and The second set of derived values ​​is determined based on the time pixel values ​​in the second time pixel image.

13. The method of claim 10, wherein determining one or more derived values ​​comprises: For each time pixel in the first group of time pixels of the first time pixel image, a first average value data is determined to indicate the average value of the time pixel. For each time pixel in the first group of time pixels, a first deviation value data is determined to indicate the deviation value of the time pixel value.

14. The method of claim 13, wherein, Determining one or more derived values ​​further includes: For each time pixel in the second group of time pixels of the second time pixel image, determine second average data to indicate the average value of the time pixel, and For each time pixel in the second group of time pixels, a second deviation value data is determined to indicate the deviation value of the time pixel value.

15. The method according to claim 13, wherein, The data used to calculate the first average include: Calculate for each time pixel in the first group of time pixels: The time average of the pixel brightness value over time; and The root mean square deviation of the pixel brightness value over time.

16. The method of claim 10, wherein determining an instruction comprises: Determine multiple regions of the first time-pixel image, the second time-pixel image, or both, and For each of the plurality of regions, the following is determined: The average of one or more derived values ​​associated with the region; A correspondence indication determined based on the correspondence associated with the region; and A region indicator is determined based on the average value and the correspondence indicator, and the region indicator is used to indicate whether movement has occurred in the region.

17. The method according to claim 16, wherein, Determining an area includes: The average value is determined to satisfy the first metric; Determining that the correspondence indicates that the second metric is satisfied; and Generate a region indicator to indicate the possibility of movement within the region.

18. The method of claim 17, further comprising: Indicators for indicating the possibility of motion in the scene are determined based on a set of region indicators associated with each of the plurality of regions.

19. At least one non-transitory computer-readable storage medium storing processor-executable instructions that, when executed by at least one computer hardware processor, cause the at least one computer hardware processor to perform the following actions: Access the first and second sets of images of the scene over time, where: Each image in the first set of images captures a relevant portion of a pattern sequence projected onto the scene over time, and each image in the first set of images has a first viewpoint of the scene; and Each image in the second set of images captures a relevant portion of a pattern sequence projected onto the scene over time, and each image in the second set of images has a second perspective of the scene; Based on the first set of images, a first time pixel image is generated, including a first set of time pixels, wherein each time pixel in the first set of time pixels includes a set of pixel values ​​located at a relevant position in each of the first set of images, and each of the first set of images captures a relevant portion of a pattern sequence projected onto the scene over time; Based on the second set of images, a second time pixel image is generated, which includes a second set of time pixels, wherein each time pixel in the second set of time pixels includes a set of pixel values ​​located at a relevant position in each of the second set of images, and each of the second set of images captures a relevant portion of a pattern sequence projected onto the scene over time; One or more derived values ​​are determined based on the time pixel values ​​of the first time pixel image, the second time pixel image, or both. Based on the first time-pixel image and the second time-pixel image, correspondence data is determined, wherein the correspondence data indicates the set of correspondences between image points in the first group of images and image points in the second group of images; and Based on the one or more derived values ​​and the corresponding relationship data, an indication is determined, which is used to indicate whether motion has occurred in the scene; The determination of correspondence data includes determining the correspondence density based on the first time pixel image and the second time pixel image. The correspondence density indicates the relationship between the number of identified correspondences between image points of the first group of images and image points of the second group of images and the number of possible correspondences.

20. At least one non-transitory computer-readable storage medium according to claim 19, wherein, Determining the one or more derived values ​​includes: For each time pixel in the first group of time pixels of the first time pixel image, determine a first average value data to indicate the average value of the time pixel value and a first deviation value data to indicate the deviation value of the time pixel value, and For each time pixel in the second group of time pixels of the second time pixel image, a second average value data for indicating the average value of the time pixel value and a second deviation value data for indicating the deviation value of the time pixel value are determined.

Citation Information

Patent Citations

  • Method for the three dimensional measurement of moving objects during a known movement

    WO2017220598A1

  • Method and device for determining correspondence, preferably for the three-dimensional reconstruction of a scene

    CN101443817A

  • Method and device for three-dimensional reconstruction of a scene

    US20090129666A1

  • Method and apparatus for registering a scene

    WO2013072807A1