Depth sensing computer vision system

By using a pair of 3D time-of-flight sensors to generate depth maps and process the output in an industrial environment, the problem of existing sensors' difficulty in detection in complex environments is solved, achieving efficient and reliable hazard detection that meets safety standards.

CN114721007BActive Publication Date: 2026-01-13SYMBOTIC LLC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202210358208.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-08-30
Filing Date
2019-08-28
Publication Date
2026-01-13
Estimated Expiration
2039-08-28

AI Technical Summary

Technical Problem

Existing 2D sensors are ineffective at detecting potential hazards in complex work cells in industrial environments, and existing 3D time-of-flight cameras have not passed rigorous safety ratings and cannot be used in applications with high safety requirements.

Method used

A pair of 3D time-of-flight sensors are used to generate a depth map through optical path overlay. Combined with a processor and calibration unit, the sensor output is processed to generate a reliable control signal that meets functional safety standards.

Benefits of technology

It enables efficient and reliable hazard detection of complex work units in industrial environments, meets stringent safety standards, and improves the robustness and reliability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114721007B_ABST
    Figure CN114721007B_ABST
Patent Text Reader

Abstract

In various embodiments, systems and methods for acquiring depth images utilize an architecture suitable for safety critical applications and can include multiple sensors (such as time-of-flight sensors) operating along different optical paths and comparison modules for ensuring correct sensor operation. Error metrics can be associated with pixel-level depth values to allow for safety control based on incompletely known depth.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This application claims priority and benefit to U.S. Provisional Patent Application No. 62 / 724,941, filed August 30, 2018, the entire contents of which are incorporated herein by reference. Technical Field

[0003] The field of this invention generally relates to monitoring industrial environments in which people and machinery interact or are in proximity, and more particularly to systems and methods for detecting unsafe conditions in monitored workspaces. Background Technology

[0004] Industrial machinery is generally dangerous to humans. Some machinery is dangerous unless completely shut down, while others may have multiple operating states, some dangerous and some not. In some cases, the level of danger may depend on a person's position or distance relative to the machinery. Therefore, many "protection" methods have been developed to separate people from machines and prevent injury. One very simple and common type of protection is a enclosure surrounding the machinery, constructed such that opening the enclosure's door would cause the circuitry to put the machinery into a safe state. If the door is placed far enough away from the machinery that a person cannot approach it before it is shut down, this ensures that a person will never approach the machinery while it is running. Of course, this prevents all interaction between people and machines and severely limits the use of workspace.

[0005] The problem becomes more acute when not only humans but also mechanical devices (e.g., robots) can move within the workspace. Both can change position and configuration rapidly and irregularly. Typical industrial robots are stationary but still possess powerful robotic arms that can cause injury within a large “envelope” of their possible motion trajectories. Furthermore, robots are often mounted on tracks or other types of external axes, and additional mechanical devices are typically contained within the robot’s end effector, both of which increase the robot’s effective total envelope.

[0006] Sensors such as light curtains can replace enclosures or other physical barriers, providing alternative methods to prevent contact between people and machinery. Sensors such as two-dimensional (2D) light detection and ranging (LIDAR) sensors can offer more sophisticated functionality, such as allowing industrial machinery or robots to slow down or issue warnings when intrusion is detected in the external area, and to stop only when intrusion is detected in the internal area. Furthermore, systems using 2D LIDAR can define multiple areas of various shapes.

[0007] Because they endanger personal safety, protective equipment must typically comply with stringent industry standards regarding functional safety, such as ISO 13849, IEC 61508, and IEC 62061. These standards specify the maximum failure rate of hardware components and define strict hardware and software development practices that must be followed to ensure the safe use of systems in industrial environments.

[0008] Such systems must ensure that they can detect dangerous situations and system failures with a very high probability, and that they respond to such events by switching controlled devices to a safe state. For example, a system for detecting intrusions in an area might be biased towards displaying an intrusion, risking a false positive to avoid the dangerous consequences of a false negative.

[0009] A new class of sensors shows great promise in machine protection, providing three-dimensional (3D) depth information. Examples of such sensors include 3D time-of-flight cameras, 3D LiDAR, and stereo vision cameras. These sensors have the ability to detect and locate areas around industrial machinery in 3D, offering several advantages over 2D systems. Particularly for complex work cells, it is difficult to determine the combination of 2D planes that effectively cover the entire space for monitoring; properly configured 3D sensors can alleviate this problem.

[0010] For example, when an intrusion is detected far beyond the robot's arm length ("Protective Separation Distance," or PSD), a 2D LiDAR system protecting the area occupied by an industrial robot will have to preemptively stop the robot because if the intrusion represents a person's leg, that person's arm might be closer and undetectable by the 2D LiDAR system. For sensors that cannot detect arms or hands, there is an additional term for PSD called intrusion distance, typically set at 850mm. In contrast, a 3D system allows the robot to continue operating until the person actually extends his or her arm towards the robot. This provides a tighter interlock between machine and human actions, avoiding premature or unnecessary shutdowns, facilitating many new safety applications and work cell designs, and saving factory floor space (which is always invaluable).

[0011] Another application of 3D sensing involves tasks that require collaboration between humans and machines to perform optimally. Humans and machines each have different strengths and weaknesses. Generally, machines may be more powerful, faster, more accurate, and have higher repeatability. Humans possess mobility, dexterity, and judgment capabilities far exceeding, and even surpassing, those of the most advanced machines. An example of a collaborative application is installing a dashboard in a car—the dashboard is heavy and difficult for humans to operate, but connecting it requires various connectors and fasteners, demanding human dexterity. 3D sensing-based safety systems allow industrial engineers to design processes that optimize the allocation of subtasks between humans and machines, best utilizing their different capabilities while ensuring safety.

[0012] 2D and 3D sensing systems may share underlying technologies. For example, RGB cameras and stereo vision cameras utilize a combination of lenses and sensors (i.e., cameras) to capture images of a scene, which are then algorithmically analyzed. Camera-based sensing systems typically include several key components. A light source illuminates the object to be inspected or measured. This light source can be part of the camera, as in active sensing systems, or independent of the camera, such as a light illuminating the camera's field of view, or even ambient light. The lens focuses the reflected light from the object and provides a wide field of view. The image sensor (typically a CCD or CMOS array) converts the light into electrical signals. The camera module typically integrates the lens, image sensor, and necessary electronics to provide electrical input for further analysis.

[0013] Signals from the camera module are fed to an image acquisition system, such as a frame capture system, which stores and further processes the 2D or 3D image signals. The processor runs image analysis software for the identification, measurement, and localization of objects within the captured scene. Depending on the specific design of the system, the processor can use a central processing unit (CPU), graphics processing unit (GPU), field-programmable gate array (FPGA), or any number of other architectures, and can be deployed in a standalone computer or integrated into the camera module.

[0014] 2D camera-based methods are well-suited for detecting defects or making measurements using known image processing techniques such as edge detection or template matching. 2D sensing is useful in unstructured environments, and, with the help of advanced image processing algorithms, can compensate for varying lighting and shadow conditions. However, algorithms for deriving 3D information from 2D images may lack robustness and applicability for security-critical applications because their failure modes are difficult to characterize.

[0015] While typical images provide 2D information about objects or space, 3D cameras add another dimension and estimate the distances to objects and other elements in the scene. Therefore, 3D sensing can provide 3D outlines of objects or space, which can be used to create a 3D map of the surrounding environment and to locate objects relative to that map. Robust 3D vision overcomes many of the problems of 2D vision because depth measurements can be used to easily separate the foreground from the background. This is particularly useful for scene understanding, where the first step is to distinguish the object of interest (foreground) from the rest of the image (background).

[0016] A widely used 3D camera-based sensing method is stereoscopic vision (or stereovisiosis). Stereoscopic vision typically uses two spaced-apart cameras, physically arranged similarly to the human eye. Given a point object in space, camera separation results in a measurable difference in the object's position between the two camera images. Using simple pinhole camera geometry, the object's position in 3D can be calculated from the image in each camera. This approach is intuitive, but its practical implementation is often not so simple. For example, features of the target must first be identified so that the two images can be compared for triangulation; however, feature recognition involves relatively complex computations and can consume considerable processing power.

[0017] Furthermore, 3D stereo vision is highly dependent on the ambient lighting environment, and its effectiveness can be reduced by shadows, occlusion, low contrast, lighting variations, or unexpected movement of objects or sensors. Therefore, it is common to use two or more sensors to acquire the surrounding field of view of the target, thus handling occlusion or providing redundancy to compensate for errors caused by attenuation and uncontrolled environments. Another common alternative is to use structured light patterns to enhance the system's ability to detect features.

[0018] Another approach to 3D imaging utilizes lasers or other active light sources and detectors. A light-detector system is similar to a camera-based system because it also integrates a lens and image sensor, converting light signals into electrical signals, but it does not capture an image. Instead, the image sensor measures the position and / or intensity of a tightly focused beam (typically a laser beam) over time. This detected change in the beam's position and / or intensity is used to determine object alignment, flux, reflection angle, time of flight, or other parameters to create an image or map of the observed space or object. Light-detector combinations include active triangulation, structured light, LiDAR, and time-of-flight sensors.

[0019] Active triangulation mitigates the environmental limitations of stereoscopic 3D by actively illuminating the object under study using a narrow-focused light source. The wavelength of the active illumination can be controlled, and the sensor can be designed to ignore other wavelengths of light, thus reducing interference from ambient light. Furthermore, the position of the light source can be varied, allowing the object to be scanned across multiple points and from multiple angles to provide a complete 3D image of the object.

[0020] 3D structured lighting is another approach based on triangulation and active light sources. In this method, a pre-designed light pattern (such as parallel lines, grids, or spots) is projected onto the target. The observed reflection pattern will be distorted by the target's contours, and the contours and distances to the object can be recovered by analyzing the distortion. Typically, sequential projection of the coded or phase-shifted pattern is required to extract a single depth frame, resulting in a low frame rate, which in turn means the object must remain relatively stationary during the projection sequence to avoid blurring.

[0021] Compared to simple active triangulation, structured light adds "feature points" to the target. Since these feature points are predetermined (i.e., spatially encoded) and very easy to identify, the structured light approach makes feature recognition easier, and triangulation more efficient and reliable. This technique shifts the complexity from the receiver to the light source, requiring more sophisticated light sources, but with simpler sensors and lower computational intensity.

[0022] Scanning LIDAR measures distance to an object or space by illuminating it with a pulsed laser beam and using sensors to measure the reflected pulses. By scanning the laser beam in 2D and 3D, the difference in laser return time and wavelength can be used to create a 2D or 3D representation of the scanned object or space. LIDAR uses ultraviolet (UV), visible, or near-infrared light, typically reflected by backscattering, to form an image or map of the space or object being studied.

[0023] 3D Time-of-Flight (ToF) cameras work by illuminating a scene with a modulated light source and observing the reflected light. The phase shift between the illumination and reflection is measured and converted into distance. Unlike LiDAR, it doesn't scan the light source; instead, it illuminates the entire scene simultaneously, resulting in a higher frame rate. Typically, the illumination comes from a solid-state laser or LED operating in the near-infrared range (approximately 800-1500 nm), invisible to the human eye. An imaging sensor, responding to the same spectrum, receives the light and converts the photon energy into current, then into charge, and finally into a digital value. The light entering the sensor has components due to ambient light and from the modulated illumination source. Distance (depth) information is embedded only in the component reflected from the modulated illumination. Therefore, the high ambient component reduces the signal-to-noise ratio (SNR).

[0024] To detect the phase shift between illumination and reflection, the light source in a 3D time-of-flight camera is pulsed or modulated using a continuous wave source (typically a sine or square wave). Distance is measured for each pixel in a 2D addressable array, resulting in a depth map or a set of 3D points. Alternatively, the depth map can be rendered in 3D space as a collection of points or a point cloud. These 3D points can be mathematically connected to form a mesh, onto which textured surfaces can be mapped.

[0025] 3D time-of-flight cameras are already used in industrial environments, but deployments to date have tended to involve applications where safety is not critical, such as boxing and palletizing. Because existing off-the-shelf 3D time-of-flight cameras lack safety ratings, they cannot be used in applications with stringent safety requirements, such as machine guarding or collaborative robot applications. Therefore, there is a need for architectures and technologies that enable 3D cameras, including time-of-flight cameras, to be available in applications requiring high safety and compliance with industry-recognized safety standards. Summary of the Invention

[0026] Embodiments of the present invention utilize one or more 3D cameras (e.g., time-of-flight cameras) in industrial safety applications. The 3D cameras generate depth maps or point clouds, which can be used by external hardware and software to classify objects in a work cell and generate control signals for mechanical equipment. In addition to meeting functional safety standards, embodiments of the present invention are also capable of processing the rich and complex data provided by 3D imaging to generate effective and reliable control outputs for industrial machinery.

[0027] Therefore, in a first aspect, the present invention relates to an image processing system. In various embodiments, the system includes first and second 3D sensors, each for generating an output array indicating pixel-wise values ​​of distance to an object within the sensor's field of view, the fields of view of the first and second 3D sensors overlapping along separated optical paths; at least one processor for combining multiple sequentially obtained output arrays from each 3D sensor into a single result (i.e., combined) output array; first and second depth calculation engines, executable by the processor, for processing consecutive result output arrays from the first and second 3D sensors, respectively, into pixel-wise arrays of depth values; and a comparison unit, executable by the processor, for (i) detecting pixel-wise differences in depth between the respective processed result output arrays from the first and second 3D sensors substantially simultaneously, and (ii) generating an alarm signal if the detected depth differences aggregated to exceed a noise metric. The depth calculation engines operate in a pipelined manner so that processing of a new combined output array begins before processing of a previous combined output array is completed.

[0028] In some embodiments, the 3D sensor is a time-of-flight (ToF) sensor. The first and second depth calculation engines and the comparison unit may be executed, for example, by a field-programmable gate array.

[0029] In various embodiments, the system further includes at least one temperature sensor, and the 3D sensors respond to the temperature sensor and modify their respective output arrays accordingly. Similarly, the system may further include at least one humidity sensor, in which case the 3D sensors will respond to the humidity sensor and modify their respective output arrays accordingly.

[0030] Multiple sequentially acquired output arrays can be combined into a single resulting output array using dark frames captured by a 3D sensor in unlit conditions. The pixel-wise output array may also include an optical intensity value for each value indicating the estimated distance to the object within the sensor's field of view, and the depth calculation engine may calculate an error metric for each depth value based at least in part on the associated optical intensity value. The error metric may be further based on sensor noise, dark frame data, and / or ambient light or temperature. In some embodiments, each depth calculation engine operates in a pipelined manner, whereby after each of the multiple computational processing steps, processing of the oldest combined output array is completed and processing of the newest combined output array begins.

[0031] In some embodiments, the system further includes a timer for storing the total cumulative running time of the system. The timer is configured to issue an alarm when a predetermined total cumulative running time is exceeded. The system may include a voltage monitor for monitoring all voltage rails of the system and for interrupting system power when a fault condition is detected.

[0032] In another aspect, the present invention relates to an image processing system, in various embodiments of which the image processing system includes a plurality of 3D sensors, each 3D sensor being configured to (i) illuminate a field of view of the sensor, and (ii) generate an output array of pixel-wise values ​​indicating distances to objects within the illuminated field of view; and a calibration unit being configured to (i) sequentially cause each 3D sensor to generate the output array while other 3D sensors illuminate their fields of view, and (ii) create an interference matrix from the generated output array. For each 3D sensor, the interference matrix indicates the degree of interference by other 3D sensors operating concurrently with it.

[0033] The system may further include a processor for operating the 3D sensors according to an interference matrix. The processor can suppress the simultaneous operation of one or more other 3D sensors during the operation of one of the 3D sensors. During the simultaneous operation of one or more other 3D sensors, the processor can correct values ​​obtained by one of the sensors.

[0034] In some embodiments, the system further includes an external synchronizer for enabling the 3D sensor to operate independently without interference. The system may further include a timer for storing the total cumulative runtime of the system. A calibration unit may respond to the total cumulative runtime and be configured to adjust a pixel-wise value indicating the distance based on it.

[0035] In various embodiments, the system further includes at least one temperature sensor, and the calibration unit responds to the temperature sensor and is configured to adjust a pixel-by-pixel value of the indicated distance based on it. Similarly, the system may further include at least one humidity sensor, in which case the calibration unit responds to the humidity sensor and is configured to adjust a pixel-by-pixel value of the indicated distance based on it. The system may include a voltage monitor for monitoring all voltage rails of the system and for interrupting system power in the event of a fault condition.

[0036] Another aspect of the invention relates to an image processing system, in various embodiments of which includes at least one 3D sensor for generating an output array of pixel-wise values, including light intensity values ​​and values ​​indicating estimated distances to objects within the sensor's field of view; a processor; and a depth calculation engine executable by the processor for processing a continuously combined output array from the at least one 3D sensor into a pixel-wise array of depth values. Each depth value has an associated error metric based at least in part on the associated intensity value. The error metric may further be based on sensor noise, dark frame data, ambient light, and / or temperature.

[0037] In some embodiments, the system further includes a controller for operating the machine within a safety envelope. The safety envelope has an error metric determined at least in part by pixels sensed by sensors and corresponds to the volume of a person near the machine. The system may include a voltage monitor for monitoring all voltage rails of the system and for interrupting system power in the event of a fault condition.

[0038] In another aspect, the present invention relates to a method for generating a digital representation of a 3D space and objects therein and detecting anomalies in that representation. In various embodiments, the method includes the steps of: arranging first and second 3D sensors in or near the space; causing each sensor to produce an array of pixel-wise values ​​indicating distances to objects in the 3D space and within the sensor's field of view, the fields of view of the first and second 3D sensors overlapping along separated optical paths; computationally combining multiple sequentially obtained output arrays from each 3D sensor into a single resulting output array; pipelinedly processing the successive resulting output arrays from the first and second 3D sensors, respectively, into a pixel-wise array of depth values; detecting pixel-wise differences in depth between the respective processed resulting output arrays from the first and second 3D sensors substantially simultaneously; and generating an alarm signal if the detected depth differences aggregate to exceed a noise metric.

[0039] The 3D sensor may be a time-of-flight (ToF) sensor. In some embodiments, the method further includes the step of providing at least one temperature sensor and modifying an output array in response to the output of the temperature sensor. Similarly, in some embodiments, the method further includes the step of providing at least one humidity sensor and modifying the output array in response to the output of the humidity sensor.

[0040] Using dark frames captured by a 3D sensor in unlit conditions, multiple sequentially acquired output arrays can be averaged or otherwise combined into a single resulting output array. The pixel-wise output array may also include an optical intensity value for each value, indicating an estimated distance to an object within the sensor's field of view, and an error metric may be based at least in part on the associated optical intensity value. Furthermore, the error metric may be further based on sensor noise, dark frame data, ambient light, and / or temperature.

[0041] In some embodiments, the method further includes the steps of: total cumulative runtime of the storage system, and issuing an alarm when a predetermined total cumulative runtime is exceeded. Execution may be pipelined, such that after each of the multiple computational processing steps, processing of the latest combined output array begins when processing of the oldest combined output array has finished.

[0042] In another aspect, the present invention relates to a method for calibrating a sensor array for 3D depth sensing. In various embodiments, the method includes the steps of: providing a plurality of 3D sensors, each sensor for (i) illuminating its field of view and (ii) generating an output array of pixel-wise values ​​indicating distances to objects within the illuminated field of view; sequentially causing each 3D sensor to generate its output array while other 3D sensors illuminate their fields of view; and creating an interference matrix from the resulting output array, the interference matrix representing the degree of interference with each 3D sensor relative to other 3D sensors operating concurrently.

[0043] The 3D sensors can be operated according to an interference matrix such that, during the operation of one of the 3D sensors, the simultaneous operation of one or more other 3D sensors is suppressed and / or the values ​​obtained by one of the sensors during the simultaneous operation of one or more other 3D sensors are corrected. The 3D sensors can be externally synchronized to allow them to operate independently without interference.

[0044] In various embodiments, the method further includes the step of: calculating the total cumulative runtime of the storage system and adjusting the per-pixel values ​​accordingly. The method may further include the step of: sensing temperature and / or humidity and adjusting the per-pixel values ​​accordingly.

[0045] Another aspect of the invention relates to a method for generating a digital representation of a 3D space and objects therein. In various embodiments, the method includes the steps of: providing at least one 3D sensor for generating an output array of pixel-wise values, wherein the values ​​include light intensity values ​​and values ​​indicating an estimated distance to an object within the sensor's field of view; and processing a successive combined output array from the 3D sensor into a pixel-wise array of depth values, each depth value having an associated error metric at least in part based on the associated intensity values.

[0046] Error metrics may be further based on sensor noise, dark frame data, ambient light, and / or temperature. In some embodiments, the method further includes the step of operating the machine within a safety envelope, the volume of which is determined at least in part by an error metric of pixels sensed by at least one sensor and corresponds to a person approaching the machine.

[0047] Generally, as used herein, the term "substantially" means ±10%, and in some embodiments, ±5%. Additionally, references to "an example," "example," "an embodiment," or "an embodiment" in the specification indicate that a particular feature, structure, or characteristic relating to the description of that example is included in at least one instance of the technology. Therefore, the phrases "in an example," "in a sample," "an embodiment," or "an embodiment" appearing in various places throughout this specification do not necessarily refer to the same example. Furthermore, specific features, structures, procedures, steps, or characteristics may be combined in any suitable manner in one or more examples of the technology. The headings provided herein are for convenience only and are not intended to limit or interpret the scope or meaning of the claimed technology. Attached Figure Description

[0048] In the accompanying drawings, the same reference numerals in different views generally represent the same parts. Furthermore, the drawings are not necessarily drawn to scale, but generally focus on illustrating the principles of the invention. In the following description, various embodiments of the invention are described with reference to the following drawings, wherein:

[0049] Figure 1 A camera architecture according to an embodiment of the present invention is illustrated schematically.

[0050] Figure 2 schematically shown Figure 1 The data stream of the embodiment shown in the figure. Detailed Implementation

[0051] The following discussion describes embodiments involving time-of-flight cameras; however, it should be understood that the invention can utilize any form of 3D sensor capable of recording a scene and typically assigning depth information to the recorded scene on a pixel-by-pixel basis. Functionally, the 3D camera generates depth maps or point clouds, which can be used by external hardware and software to classify objects in a work cell and generate control signals for mechanical equipment.

[0052] First refer to Figure 1 A representative system 100 is shown, which can be configured as a camera within a single housing or as multiple separate components. System 100 can be implemented as a camera within a single housing, and the system 100 includes a processing unit 110 and a pair of 3D sensors 115, one of which (sensor 115) M One sensor operates as the main sensor, while another (sensor 115) operates as the main sensor. SAs a sensor, camera 100 (or, in some embodiments, each of sensors 115) also includes a light source (e.g., a VCSEL laser source), a suitable lens, and a filter tuned to the light source. Reflected and backscattered light from the light source is captured by the lens and recorded by sensor 115. The light source may include a diffuser 120, although in low-power applications, a light-emitting diode (LED) may be used instead of the laser source and diffuser.

[0053] Processor 110 may be or include any suitable type of computing hardware, such as a microprocessor, but in various embodiments may be a microcontroller, peripheral integrated circuit element, CSIC (customer application integrated circuit), ASIC (application application integrated circuit), logic circuit, digital signal processor, programmable logic device, such as FPGA (field programmable gate array), PLD (programmable logic device), PLA (programmable logic array), RFID processor, graphics processing unit (GPU), smart chip, or any other device or device setup capable of implementing the steps of the process of the present invention.

[0054] In the illustrated embodiment, processor 110 operates the FPGA and can advantageously provide features supporting security-level operation, such as secure separation of the design flow to pinpoint the location and routing of security-critical parts of the design; clock checking; single-event upsets; CRC functionality for various data and communication paths across FPGA boundaries; and the use of security-level functions for individual submodules. Within the processor's integrated memory and / or in a separate main random access memory (RAM) 125 (typically dynamic RAM or DRAM) are instructions, conceptually shown as a set of modules, that control the operation of processor 110 and its interaction with other hardware components. These instructions can be encoded in any suitable programming language, including but not limited to high-level languages ​​such as C, C++, C#, Java, Python, Ruby, Scala, and Lua, and utilize, but not limited to, any suitable frameworks and libraries such as TensorFlow, Keras, PyTorch, or Theano. Additionally, the software can be implemented in assembly language and / or machine language for a microprocessor residing on the target device. An operating system (not shown) directs the execution of low-level basic system functions such as memory allocation, file management, and the operation of mass storage devices. At a higher level, a pair of conventional depth computing engines 1301, 1302 receive raw 3D sensor data and assign depth values ​​to each pixel of the recorded scene. Raw data refers to uncalibrated data from the sensor (e.g., 12 bits per pixel).

[0055] Two separate optical paths are created using two independent lenses and a 3D sensor module 115. Redundancy enables immediate detection should one of the camera modules 115 fail during operation. Furthermore, by not picking up exactly the same image from each lens and sensor combination, an additional level of processing can be performed by the image comparison module 135, projecting the response of a pixel from one optical path to the corresponding pixel in the other optical path. (This projection can be determined, for example, during the calibration phase.) Failure modes detectable through this comparison include false detections due to multiple reflections and sensor-to-sensor interference. The two independent images can also be used to reduce noise and / or improve resolution when the two sensors 115 agree on the camera's performance characteristics within an established noise metric. Redundant sensing for dual-channel imaging ensures that the level of reliability required for safety-critical operation in industrial environments is met.

[0056] If the comparison metric calculated by comparison module 135 is within an allowable range, the merged output is processed according to a network communication protocol for output. In the illustrated embodiment, the output is provided by a conventional low-latency Ethernet communication layer 140. This output can be utilized by a processor system for a security level of controlled mechanical equipment, as described, for example, in U.S. Provisional Application Serial No. 62 / 811,070, filed February 27, 2019, the entire disclosure of which is incorporated herein by reference.

[0057] System 100 may include one or more environmental sensors 145 to measure conditions such as temperature and humidity. In one embodiment, multiple on-board temperature sensors 145 are positioned at multiple locations on the camera housing, passing through sensor 115, for example, at the center of the illumination array and inside the camera housing (one near the master sensor and one near the slave sensor), for calibrating and correcting the 3D sensing module, as system-generated heat and variations or drift in ambient temperature affect camera operating parameters. For example, variations in camera temperature can affect the camera's baseline calibration, accuracy, and operating parameters. Calibration can be used to establish a sustainable operating temperature range; sensors detecting conditions outside these ranges may trigger shutdown to prevent dangerous malfunctions. Temperature correction parameters can be estimated during calibration and then applied in real time during operation. In one embodiment, system 100 identifies a stable background image and uses it to continuously verify the correctness of the calibration, and the temperature-corrected image remains stable over time.

[0058] A fundamental problem with using depth sensors in security-level systems is that the depth result from each pixel cannot be 100% certain. The actual distance to an object may differ from the reported depth. For well-lit objects, this difference is small and negligible. However, for poorly lit objects, the error between the reported and actual depth can become significant, manifesting as a mismatch between the object's actual and apparent positions, and this mismatch will be randomized based on each pixel. Pixel-level errors can be due to, for example, raw data saturation or clipping, unresolved ambiguous distances calculated from different modulation frequencies, large intensity mismatches between different modulation frequencies, prediction measurement errors above a certain threshold due to low SNR, or excessive ambient light levels. Security-level systems that require accurate distance knowledge cannot tolerate such errors. A typical approach used by time-of-flight cameras is to zero out the data for a given pixel if the received intensity is below a certain level. For pixels with moderate or low received light intensity, the system can either conservatively ignore the data and completely disregard the pixel, or accept the depth result reported by the camera—which may be off by a certain distance.

[0059] Therefore, the depth data provided in the output can include a predicted measurement error range for the depth result on a per-pixel basis, based on the original data processing and statistical models. For example, time-of-flight cameras typically output two values ​​per pixel: depth and light intensity. Intensity can be used as a rough measure of data confidence (i.e., the reciprocal of the error), so instead of outputting depth and intensity, the data provided in the output can be depth and an error range. The range error can also be predicted on a per-pixel basis based on variables such as sensor noise, dark frame data (described below), and environmental factors (e.g., ambient light and temperature).

[0060] Therefore, this method represents an improvement over the simple pass / fail criterion described above, ignoring all depth data from pixels with a signal-to-noise ratio (SNR) below a threshold. Using the simple pass / fail method, depth data appears to have zero measurement error; therefore, safety-critical processes that rely on this data integrity must set the SNR threshold high enough that the actual measurement error has no system-level safety impact. Despite the increased measurement error, pixels with low to medium SNR may still contain useful depth information and are either completely ignored (at high SNR thresholds) or used under the erroneous assumption of zero measurement error (at low SNR thresholds). Including a measurement error range on a per-pixel basis allows higher-level safety-critical processes to utilize the information provided by pixels with SNR levels ranging from low to medium, while appropriately limiting depth results from such pixels. This improves overall system performance and uptime compared to the simple pass / fail method, although it should be noted that the pass / fail criterion can still be used with this method for pixels with very low SNR.

[0061] According to embodiments of the invention, error detection can take different forms, but their common purpose is to prevent erroneous depth results from propagating to processes critical to higher levels of security on a pixel-by-pixel basis, without simply setting a threshold for a maximum permissible error (or an equivalent minimum required strength). For example, the depth of a pixel could be reported as 0 with a corresponding pixel error code. Alternatively, the depth calculation engine 130 could output a report of depth along with an expected range error, enabling downstream security level systems to determine whether the error is low enough to allow the use of the pixel.

[0062] For example, as described in U.S. Patent No. 10,099,372, the entire disclosure of which is incorporated herein by reference, a robot safety protocol may involve adjusting the robot's maximum speed (meaning the speed of the robot itself or any of its attachments) to be proportional to the minimum distance between any point on the robot and any point in the relevant set of sensed objects to be avoided. The robot is allowed to operate at its maximum speed when the nearest object is further away from a certain threshold distance, beyond which there is no concern about collision, and the robot comes to a complete stop if the object is within a certain minimum distance. Sufficient margin can be added to the specified distance to accommodate the movement of a relevant object or person toward the robot at a certain maximum practical speed. Thus, in one approach, an outer envelope or 3D region is generated around the robot by calculation. Outside this region, for example, all movements of a detected person are considered safe because, during the operating cycle, these movements do not bring the person close enough to the robot to pose a danger. Detection of any part of a person's body within a second 3D region defined in a first region does not prevent the robot from continuing to operate at full speed. However, if any part of the detected person exceeds the threshold of the second zone but remains outside the third inner danger zone within the second zone, the robot is signaled to operate at a slower speed. If any part of the detected person crosses the innermost danger zone—or is predicted to do so in the next cycle based on a human motion model—the robot stops operating.

[0063] In this scenario, the safe zone can be adjusted based on the estimated depth error (or the space assumed to be occupied by detected people can be expanded). The larger the detection error, the larger the envelope of the safe zone, or the space assumed to be occupied by detected people. In this way, the robot can continue operating based on the error estimate instead of shutting down, since too many pixels do not meet the pass / fail criteria.

[0064] Because any single image of a scene may contain low light and noise, in operation, after frame triggering, the two sensors 115 rapidly and sequentially acquire multiple images of the scene. These "subframes" are then averaged or otherwise combined to generate a single final frame for each sensor 115. The subframe parameters and timing relative to frame triggering can be programmed at the system level and can be used to reduce crosstalk between sensors. Programming may include subframe timing for time multiplexing and frequency modulation of the carrier.

[0065] like Figure 1As shown, an external synchronizer 150 can be provided for frame-level and, in some cases, subframe-triggered operation, to allow multiple cameras 100 to safely cover the same scene and to allow interlaced scanning of camera output. Frame-level and subframe-triggered operation can be timed multiplexed to avoid interference. One camera 100 can be designated as the master, controlling the overall timing of the camera to ensure that only one lighting scene is captured at a time. The master provides trigger signals to each camera to indicate when they should acquire the next frame or subframe.

[0066] Some embodiments utilize dark frames (i.e., images of the scene without illumination) for real-time correction of ambient noise and sensor offset. Typically, differential measurement techniques using multiple subframe measurements to eliminate noise sources are effective. However, by using dark subframes not only as measurements of the ambient level but also as measurements of inherent camera noise, the required number of subframes can be reduced, which increases the amount of signal available per subframe.

[0067] like Figure 2 As shown, when recording a set of subframes, a pipelined architecture can be used to facilitate efficient subframe aggregation and processing. Architecture 200 typically includes an FPGA 210 and a pair of master-slave time-of-flight sensors 215. M 215 S And multiple external DD R memory groups 2171, 2172 to support subframe aggregation from captured frame data. Since the subframes are generated by sensor 215 M 215 S The captures, which accumulate in DDR memory group 217 along data paths 2221 and 2222 respectively, reflect the difference between the rate of subframe capture and the rate of depth calculation processing.

[0068] Each data path 221 may have multiple DDR interfaces with error correction code (ECC) support to allow simultaneous read and write to memory, but the two data paths 221 are independent. Each depth computing pipeline 2301, 2302 operates in a pipelined manner, such that after each processing step, a new frame can begin when an earlier frame completes, with intermediate frames progressing incrementally through the processing path. Calibration-related data (e.g., temperature data) can be accumulated from environmental sensor 145 in DDR memory bank 217 and passed to the depth computing pipeline 230 along with concurrent sensor data, so that in each processing step, depth computing is performed based on the dominant environmental conditions at the time of frame acquisition.

[0069] As described above, the sensor comparison processing unit 235 compares the new images with depth information that appear from the depth calculation pipeline after each time step and outputs them as Ethernet data. Figure 2As shown, if needed, the Ethernet communication layer 240 can be implemented outside of the FPGA 210.

[0070] In a typical deployment, multiple 3D time-of-flight cameras are mounted and fixed around the workspace or object to be measured or imaged. An initial calibration step is performed by a calibration module 242 at each 3D time-of-flight camera (shown for convenience as part of system 200, but more typically implemented externally, e.g., as a separate component) to correct for the effects of structured noise, including temperature and camera-specific optical distortions. Other metadata, such as the expected background image for subframes, can also be captured, which can be used to monitor camera measurement stability in real time. Each camera 100 can trigger exposures frame-triggered or subframe-triggered by varying the illumination frequency and level (including the level of darkness captured by the camera in the absence of illumination). Multiple 3D time-of-flight cameras can be triggered at different frequencies and illumination levels via an external subframe external synchronizer 150 to minimize interference and reduce latency among all 3D time-of-flight cameras in the work unit. Latency between all cameras can be reduced and the acquisition frequency increased by a host unit with overall camera timing control (ensuring only one illumination scene at a time).

[0071] Data flows from each sensor 215 into the associated DDR 217 via the data receiving path in FPGA 210. Data is stored in the DDR 217 at the subframe level. Once the depth computing engine 230 recognizes that a complete subframe has accumulated in the associated DDR 217, it begins extracting data from it. These pixels flow through the depth computing engine 230 and are stored back in the associated DDR 217 as single-frequency depth values. These contain ambiguous depth results that need to be resolved later in the pipeline through comparison. Therefore, once the first three subframes required to compute the first single-frequency result are available in the DDR 217, the associated depth computing engine begins computing the ambiguous depth on a pixel-by-pixel basis using those three subframes. While this is happening, the next three subframes for the second single-frequency result are loaded into memory from the sensor 215, and they receive previously loaded data when the subframe queue is empty, thus avoiding wasted processing cycles during extraction. Once the first single-frequency result has been computed and fully loaded into memory, the depth computing engine begins computing the second single-frequency depth result in a similar manner. Meanwhile, the third set of subframes was loaded into memory.

[0072] However, instead of loading the second single-frequency depth result into memory during computation, it is processed pixel-by-pixel along with the first single-frequency depth result to produce a definite depth result. This result is then stored in memory as an intermediate value until it can be further compared with the second definite depth result obtained from the third and fourth single-frequency depth results. This process is repeated until all relevant subframes have been processed. Finally, all intermediate results are read from DDR, and the final depth and intensity values ​​are calculated.

[0073] Calibration can not only adjust for camera-specific performance differences but also characterize interference between cameras in multi-camera configurations. During initialization, one camera illuminates the scene at a time, while the other cameras determine how much signal they receive. This process helps create an interference matrix, which can be stored in DDR 217, determining which cameras can illuminate simultaneously. Alternatively, this method can also be used to create real-time corrections, similar to crosstalk correction techniques used for electronic signal transmission. In particular, the FPGAs 112 of multiple cameras can cooperate with each other (e.g., in an ad hoc network, or with one camera designated as a master and the others operating as slaves) to sequentially generate outputs for each camera while the other cameras illuminate their fields of view, and can share the resulting information to build and share the interference matrix from the generated outputs. Alternatively, these tasks can be performed by a supervisory controller operating all the cameras.

[0074] Camera parameters such as temperature, distortion, and other metadata are captured during calibration and stored in DDR 217; these are used during real-time recalibration and camera operation. The calibration data contains the optical characteristics of the sensor. As described above, the depth computing pipeline utilizes this data, along with stream frame data and data characterizing the sensor's fixed noise properties, when calculating depth and error. Camera-specific calibration data is collected during manufacturing and uploaded from non-volatile PROMs 2451 and 2452 to DDR3 storage upon camera startup. During runtime, the depth computing engine 230 accesses the calibration data from the DDR3 memory in real-time as needed. Specifically, real-time recalibration adjusts for drift in operating parameters such as temperature or illumination levels during operation in a conventional manner. Health and status monitoring information can also be sent after each frame of depth data and may include elements such as temperature, pipeline error codes, and FPGA processing latency margins required for real-time recalibration.

[0075] An operation timer 250 (shown again as an internal component for convenience, but can be implemented externally) may be included to keep track of the camera's running time, periodically sending this data to the user via communication layer 240. Calibration unit 242 may also receive this information to adjust operating parameters as the camera illumination system and other components age. Furthermore, once the VCSEL's aging limit is reached, timer 250 can generate an error status to warn the user that maintenance is required.

[0076] The aforementioned features address various possible failure modes of conventional 3D cameras or sensing systems, such as multiple exposures or common-mode failures, enabling operation in safety-level systems. The system may include additional features for safety-level operation. One such feature is monitoring each voltage rail by voltage monitor 160 (see...). Figure 1 This allows the camera to shut down immediately if a fault condition is detected. Another is a security level protocol used for data transmission between different elements of the 3D time-of-flight camera and the external environment (including external synchronizers). Broadly speaking, a security level protocol will include error checking to ensure that bad data does not propagate through the system. Security level protocols can be created around common protocols such as UDP, supporting high bandwidth but not inherently reliable. This can be achieved by adding security features such as packet enumeration, CRC error detection, and frame ID marking. These ensure that the current depth frame is the correct depth frame for further downstream processing after the frame data is output from the camera.

[0077] Some embodiments of the present invention have been described above. However, it is clearly stated that the present invention is not limited to these embodiments; rather, additions and modifications to the content explicitly described herein are also included within the scope of the present invention.

Claims

1. An image processing system, comprising: First and second 3D sensors, each for generating an output array of pixel-wise values ​​indicating the distance to an object within the sensor's field of view, the fields of view of the first and second 3D sensors overlapping along separate optical paths, the objects including robots and humans; The first and second depth calculation engines, which can be executed by at least one processor, are used to process the successive result output arrays from the first and second 3D sensors, respectively, into a pixel-wise array of depth values. The comparison unit, which can be executed by at least one processor, is used to detect pixel-by-pixel differences in depth between the result output arrays of the corresponding processing from the first and second 3D sensors substantially simultaneously. as well as The control processor is configured to (i) computationally generate a 3D safety envelope around the robot based on error metrics of pixels sensed by the first and second 3D sensors, (ii) control the robot’s operating speed based at least in part on a detected distance between the robot and the human in the 3D safety envelope, and (iii) adjust the detected distance based on pixel-wise differences in the detected depth.

2. The system according to claim 1, wherein, The detected distance is adjusted by expanding or contracting the 3D security envelope.

3. The system according to claim 1, wherein, The detected distance is adjusted by expanding or contracting the space occupied by a person.

4. The system according to claim 1, wherein, The deep computing engine operates in a pipelined manner so that it can begin processing a new output array before it has finished processing the previous output array.

5. The system according to claim 1, wherein, The first and second 3D sensors are time-of-flight (ToF) sensors.

6. The system of claim 1, further comprising at least one temperature sensor, wherein the processor responds to the at least one temperature sensor and further adjusts the detected distance based thereon.

7. The system of claim 1, further comprising at least one humidity sensor, wherein the processor responds to the at least one humidity sensor and further adjusts the detected distance based thereon.

8. The system according to claim 1, wherein, Larger pixel-wise differences in detected depth result in a larger downward adjustment of the detected distance.

9. A method for controlling a robot in a 3D workspace, the method comprising the following steps: Place the first and second 3D sensors in or near the workspace; Each of the sensors generates an output array that produces pixel-wise values ​​indicating the distance to an object in 3D space and within the sensor's field of view, with the fields of view of the first and second 3D sensors overlapping along separate optical paths, the objects including robots and humans; The sequential output arrays from the first and second 3D sensors are computationally processed into a pixel-by-pixel array of depth values. The detection outputs a pixel-by-pixel difference in depth between the corresponding processing results from the first and second 3D sensors, which are essentially processed simultaneously. Based on the error measure of pixels sensed by the first and second 3D sensors, a 3D safety envelope around the robot is computationally generated. The robot's operating speed is controlled at least in part based on the distance detected between the robot and the person within the 3D safety envelope; and The detected distance is adjusted based on the pixel-by-pixel difference in the detected depth.

10. The method according to claim 9, wherein, The detected distance is adjusted by expanding or contracting the 3D security envelope.

11. The method according to claim 9, wherein, The detected distance is adjusted by expanding or contracting the space occupied by a person.

12. The method according to claim 9, wherein, The first and second 3D sensors are time-of-flight (ToF) sensors.

13. The method of claim 9, further comprising adjusting the detected distance based on the sensed temperature.

14. The method of claim 9, further comprising adjusting the detected distance based on the sensed humidity.

15. The method according to claim 9, wherein, Larger pixel-wise differences in detected depth result in a larger downward adjustment of the detected distance.

Citation Information

Patent Citations

  • Detecting and classifying workspace regions for safety monitoring

    US10099372B2

  • Methods and systems for detecting and recognizing objects in a controlled wide area

    US20030235335A1

  • Robot safety system and a method

    US20110264266A1

  • 3D modeling with depth camera and surface normals

    US9137511B1