Device and method for image processing
By combining polarized images and ranging data, using neural networks and analytical formulas for depth estimation, the depth estimation problem and the infeasibility of active sensors in textureless scenes are solved, and a more accurate and robust depth estimation effect is achieved.
Patent Information
- Application Number
- CN202080103540.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-11-06
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2040-11-06
AI Technical Summary
Existing depth estimation methods are difficult to work effectively in textureless scenarios, and traditional active sensors are not feasible in some scenarios, with problems of noise and low resolution.
By combining polarized images and ranging data, depth estimation is performed using a trained neural network, and synthetic polarized images and synthetic ranging data are formed by analytical formulas to improve the accuracy and robustness of depth estimation.
This method can provide more accurate depth estimation in texture-free scenarios, reduce computational complexity, and implement real-time depth estimation on small devices.
Smart Images

Figure CN115997235B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to depth estimation in a scene. The scene can be represented by a visual image. The depth estimation process can be implemented by using data from the visual image and other data representing the depth of at least some positions in the scene. The scene may be within the field of view of one or more sensors. Background Art
[0002] A great deal of work has been done in the field of computer vision to develop systems that can estimate the depth of each position in a visual image. For example, by identifying features in the image, a properly trained neural network can estimate the distance from a reference point (e.g., the position of the camera that captured the image) to the position of an item depicted at a certain position in the image. Another area of research is sensor devices for directly estimating depth, e.g., by emitting a signal in a selected direction and estimating the time required for the signal to return. Each method relies on some known information to infer depth information. The known quantities are what distinguish the various sensor device methods. Examples of known information include the spatial distance between sensor pairs (stereo), known light patterns (encoded light or structured light), and the speed of light (LiDAR, time-of-flight measurements). In each case, known variables are used to estimate depth, such as the distance to a position in the image for which depth data was not previously known.
[0003] A common way to classify depth estimation methods is to distinguish between so-called passive methods and so-called active methods. Passive methods typically retrieve depth information from visible spectrum images and consider two-view images (spatial, "stereo"), multi-view images (also called "stereo" or "temporal", "current time") to perform image correspondence matching and triangulation / trilateration. A stereo depth camera has two sensors that are typically spaced relatively close together, and the system can compare the two images from these sensors, which image the same or at least overlapping fields of view. Since the distance between the sensors that capture the images is known, these comparisons can provide depth information.
[0004] For a single-viewpoint image, the distance to a line can be determined. For a second-viewpoint image, the correct distance to other positions can be inferred by comparing the image content. The maximum distance that a stereo device can reliably measure is directly related to the spacing between the two sensors. The wider the baseline, the more reliably the system can infer the distance. The distance error increases quadratically with distance. Decades of in-depth research have been carried out in the field of stereo technology, and this field remains an active research area, but there are still inherent problems that hinder practical applications. These problems include the need for precise image correction (i.e., computationally obtaining coplanar and horizontally aligned image planes) and the ill-conditioned nature of performing correspondence matching in textureless regions (regions of the image space) of the scene.
[0005] Another form of depth estimation is to use time-of-flight. This is called an active method. Light is projected onto the scene, and then depth information can be measured based on the echo signal. Techniques based on time-of-flight (ToF) can be regarded as the latest method of active depth sensing. A ToF camera determines depth information by measuring the phase difference between the emitted light and the reflected light or the time it takes for light to travel to and from the scene from the illumination source. ToF devices are faster than comparable laser distance scanners and are able to capture depth information of dynamic scenes in real time. Indirect ToF measurement: In the absence of spatial incidence between the illumination source and the receiving sensor, compared with consumer and high-end visible-spectrum passive cameras (which may have a resolution of millions of pixels), it is usually relatively noisy and has a lower image resolution (e.g., 200×200 pixels). Depending on the power and wavelength of the light, the time-of-flight sensor can measure depth at a long distance. LiDAR sensors utilize the knowledge of the speed of light and are essentially time-of-flight cameras that use lasers for depth calculation. Laser distance scanner devices are the earliest active methods and usually can achieve high precision. However, due to the layer-by-layer nature of laser scanning, they are very time-consuming and usually not suitable for dynamic scenes. Similar to other ToF cameras, these devices emit light beams and sweep the light beams across the scene to measure the time it takes for the light to return to the sensor on the camera. One disadvantage of time-of-flight cameras (low power) is that they are vulnerable to other cameras in the same space and may not work properly under outdoor conditions. Achieving strong performance under outdoor conditions requires higher energy but usually only provides sparse depth signals. If there may be a situation where the light recorded on the sensor may not be the light emitted from a specific relevant camera (e.g., from some other source, such as the sun or other cameras), this will have an adverse impact on the quality of the obtained depth estimation. The most important error source of direct ToF usually boils down to multi-path interference (MPI), that is, the situation where light is emitted from the correct (original) light source but is measured after multiple reflections within the scene, which seriously affects the distance measurement.
[0006] Another class of active sensors is based on the principle of structured light or coded light. These sensors rely on using a light emitter device to project a light pattern (usually from the invisible part of the spectrum, such as infrared) onto the scene. The projected pattern is a visual, current time pattern, or a combination thereof. Since the projected light constitutes a pattern known to the device, the nature of the pattern sensed by the camera sensor in the scene can provide depth information. By exploiting the difference between the expected image pattern and the actual image (what is seen through the camera), the distance to the camera sensor for each pixel (a dense "depth map") can be calculated. Structured light sensors can now be regarded as a rather mature technology, with commercial hardware and a range of consumer devices available on the market. This technology relies on accurately capturing the light projected onto the scene, so the device performs best at relatively short distances indoors (affected by the luminous power). Performance can also be affected if there is other noise in the environment due to other cameras or devices emitting light in the common part of the spectrum (such as infrared). The depth maps generated by these sensors may also contain holes, which are caused by occlusions due to the relative displacement between the light projection source and the (infrared) camera observing the light.
[0007] In addition to intensity, velocity, and color (wavelength), another source of light information that has received less extensive consideration in depth photometric recovery tasks is light polarization. Light polarization is affected by factors in the scene, such as surface shape, surface curvature, surface material, and the position of the object relative to the light source. Therefore, polarization can provide an additional information signal about surface geometry and scene depth. In particular, polarization imaging can be used for shape determination of specular reflections and transparent objects, where the intensity and wavelength of the reflection are less well-defined (for example, a transparent object will present the color of any object behind it). The basic assumption is that the scene is illuminated by unpolarized light, so it can be assumed that any detected polarization is caused by surface reflection. A related assumption is that the observed object has a smooth reflective surface. By measuring the degree of polarization of the light incident on the camera, the direction of the surface normal can be obtained, and then, by acquiring these surface normals at a sufficient number of points in the scene, the scene surface can be reconstructed. The nature of the signal provided by a polarization camera helps to provide reliable information in the physical inversion of the surface normal direction and can generally provide contrast enhancement and reflection removal. However, this mode is susceptible to absolute distance errors at each point on the surface.
[0008] Active distance sensors are typically used in applications where estimation accuracy and robustness are of great importance, such as robotics and other autonomous systems. However, many factors make it infeasible to rely solely on expensive active sensors in every scenario, namely: scene geometric constraints, size, power (active illumination), heat dissipation, and the expected lifetime / duration of passive and active components. Current learning-based methods have been combined with many input modalities, but reasonable performance can now be achieved by methods that rely only on passive sensor inputs. These models typically take RGB images (monocular or stereo) as input and utilize recent learning strategies. Recent work has utilized fully supervised convolutional neural networks (CNNs) to infer depth from passive stereo image pairs (see the article "Stereo Matching by Training a Convolutional Neural Network to Compare Image Patches" (17(1), pp. 2287-2318) by J. and Y. LeCun in the Journal of Machine Learning Research in 2016, and the article "StereoNet: Guided Hierarchical Refinement for Real-Time Edge-Aware Depth Prediction" (pp. 573-590) by S. Khamis, S. Fanello, C. Rhemann, A. Kowdle, J. Valentin, and S. Izadi in the Proceedings of the European Conference on Computer Vision (ECCV) in 2018) or even monocular images (see the article "Unsupervised Monocular Depth Estimation with Left-Right Consistency" (pp. 270-279) by C. Godard, O. Mac Aodha, and G. J. Brostow in the Proceedings of the IEEE International Conference on Computer Vision and Pattern Recognition in 2017, and the article "Digging into Self-Supervised Monocular Depth Estimation" (pp. 3828-3838) by C. Godard, O. Mac Aodha, M. Firman, and G. J. Brostow in the Proceedings of the IEEE International Conference on Computer Vision and Pattern Recognition in 2019), where prior knowledge and information about the structure of the world are encoded in the (learned) model weights.
[0009] Previous work has also incorporated visual patterns. The recent stereo camera devices (discussed earlier) can also have "active" components and project infrared light into the scene to further improve depth estimation accuracy. Compared to structured light or coded light cameras, stereo cameras can use any part of the spectrum to measure depth. Since these devices use any visual feature to measure depth, they are able to operate under most lighting conditions, including outdoors. Adding infrared emitters enables these devices to also operate under low light conditions because the camera can still perceive depth details (see the article "End-to-End Self-Supervised Learning for Active Stereo Systems" (pages 784 - 801) in the proceedings of the European Conference on Computer Vision (ECCV) 2018 by Y. Zhang, S. Khamis, C. Rhemann, J. Valentin, A. Kowdle, V. Tankovich, M. Schoenberg, S. Izadi, T. Funkhouser, and S. Fanello).
[0010] Figure 1 The pointwise distance error and surface orientation error are shown. Depth estimation is represented by the surface. The top row illustrates that the pointwise distance error statistics can distinguish the distance to the camera sensor (top right), but provide little information in distinguishing different surface normal directions (top left). The ToF sensor is sensitive to this depth estimation error. The bottom row illustrates that the surface orientation error statistics (bottom left) can distinguish normal direction differences, but it is difficult to resolve the ambiguity of the camera sensor distance differences (bottom right). The polarization sensor is more sensitive to this depth estimation error.
[0011] In summary, depth perception is one of the fundamental challenges in the field of computer vision. A large number of applications can be realized through accurate scene depth estimation. In these applications, robust, accurate, and real-time depth estimation devices (and technologies) will become useful enabling components.
[0012] Existing methods generally have the following problems. Laser scanners are too slow to be used in real time. Passive stereo cannot be used in textureless scenes. Time-of-flight sensors can provide real-time independent estimates for each pixel, but usually have low resolution, high noise, and poor calibration. Photometric stereo is prone to low-frequency distortion, and it may be difficult to obtain accurate absolute distances for polarization signals.
[0013] When multiple modalities are available for capturing common scenes, strategies involving sensor fusion can be utilized to improve depth estimation, that is, combining multiple complementary signal sources to enhance depth estimation accuracy. Summary of the Invention
[0014] According to one aspect, the present invention provides an image processing apparatus for estimating depth of field in a field of view. The apparatus includes one or more processors for: receiving a captured polarization image representing the polarization of light received at a first set of multiple positions in the field of view; processing the captured polarization image using a trained first neural network to form a first depth estimate for one or more positions in the field of view; receiving ranging data representing the environmental distance from a reference point to one or more positions in the field of view; processing the ranging data using a trained second neural network to form a second depth estimate for a second set of multiple positions in the field of view; forming a synthetic polarization image representing an estimated polarization of light received at a third set of multiple positions in the field of view by processing one or both of the first depth estimate and the second depth estimate using a first analytical formula; and forming synthetic ranging data representing an estimated environmental distance to one or more positions in the field of view by processing one or both of the first depth estimate and the second depth estimate using a second analytical formula.
[0015] Unlike other methods such as machine learning / neural networks, using analytical formulas to form the synthetic data can reduce the computational complexity of the corresponding tasks. This can make them easier to implement on small devices.
[0016] Once the synthetic polarization image and the synthetic ranging data are formed, they can be compared with the captured polarization image and the received ranging data, respectively. For regions of the image, this comparison can be used to select either the synthetic ranging data or the received ranging data to represent the depth estimate for the region of the image.
[0017] The first set, second set, and third set of multiple positions may be the same or different. It is convenient if they all include a common set of points or regions because this allows for easy comparison of the data.
[0018] The polarization image may represent the polarization of light received at multiple positions in the field of view under one or more predetermined polarizations. This can allow for the selection of a preferred depth estimate regardless of the actual polarization of the light captured from a given portion of the field of view.
[0019] The image processing apparatus may include active sensor means for forming the ranging data. The active sensor means may include a time-of-flight sensor. The sensors may be located at the same position. They may be connected together to image the same or overlapping scenes. This helps to achieve generality between the subjects of the captured data.
[0020] The device can be used to form a plurality of synthetic polarization images estimated by the first analytical formula for a plurality of assumed reflection characteristics. This can allow for the selection of a preferred depth estimate regardless of the actual polarization of the light captured from a given portion of the field of view.
[0021] The assumed reflection characteristics can include diffusivity and specularity. This can allow for the modeling of the behavior of different surfaces.
[0022] The device can be used to form a plurality of synthetic polarization images estimated by the first analytical formula for a plurality of assumed polarizations. This can enable the system to benefit from a camera that captures images at multiple polarizations.
[0023] The device can be used to form a plurality of synthetic polarization images estimated by the first analytical formula for a plurality of assumed colors. This can enable the system to adapt to the reflected light of different colors in the captured images.
[0024] The step of comparing the polarization image with the synthetic polarization image can include: reducing the plurality of synthetic polarization images to the synthetic polarization image by selecting, for each position in the field of view, the plurality of synthetic polarization images that preserve the estimated polarization information, the polarization information at that position having the minimum estimated error. This can allow for the selection of a preferred estimated depth.
[0025] The first analytical formula can be used to form distance estimates to a plurality of positions on the field of view based on the intensity of at least one polarization image at the corresponding positions. This helps to improve distance estimation.
[0026] The second analytical formula can be used to form a plurality of estimates of the distance to the corresponding position for each of a plurality of positions on the field of view based on the corresponding phase offset. This can allow for the selection of a preferred value among these estimates.
[0027] The image processing device can include a camera for capturing the captured polarization image. Then, the image captured by the camera can be enhanced based on the above calculations.
[0028] The captured polarization image can include a stereoscopic polarization image. This helps to form depth information.
[0029] The first neural network and the second neural network are the same. This can reduce computational complexity and memory requirements.
[0030] The first analysis formula can be used to calculate a polarization estimate based on trigonometric functions of an angle formed by a first sub-angle and a second sub-angle, where the first sub-angle is calculated based on the normal of a corresponding surface, and the second sub-angle represents a candidate polarization angle. In this way, such formulas can apply a model of reflection behavior.
[0031] The second analysis formula can be used to calculate a distance estimate based on trigonometric functions of an angle formed by phase values, where the phase values are calculated based on the depth of a corresponding surface. In this way, such formulas can apply a model of reflection behavior.
[0032] According to a second aspect, the present invention provides a computer-implemented method for estimating depth of field in a field of view, the method comprising: receiving a captured polarization image representing light polarization received at a first set of multiple positions in the field of view; processing the captured polarization image using a trained first neural network to form a first depth estimate for one or more positions in the field of view; receiving ranging data representing distances from a reference point to one or more environmental positions in the field of view; processing the ranging data using a trained second neural network to form a second depth estimate for a second set of multiple positions in the field of view; forming a synthetic polarization image representing an estimated value of light polarization received at a third set of multiple positions in the field of view by processing one or both of the first depth estimate and the second depth estimate using a first analysis formula; forming synthetic ranging data representing an estimated value of the distance to one or more environmental positions in the field of view by processing one or both of the first depth estimate and the second depth estimate using a second analysis formula. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] The present invention will now be described by way of example with reference to the accompanying drawings. In the drawings:
[0034] Figure 1 Depth estimation techniques are shown;
[0035] Figure 2 is a schematic diagram of a device for performing depth estimation;
[0036] Figure 3 shows a first schematic diagram of a network architecture for depth estimation according to a polarization, associated ToF input image pattern;
[0037] Figure 4 shows a second schematic diagram of a network architecture for depth estimation according to a polarization, associated ToF input image pattern. DETAILED DESCRIPTION
[0038] Figure 2The figure shows a device for implementing the present system. In this example, the device is a mobile phone, but it can be any suitable device and / or the described functions can be divided among multiple separate devices.
[0039] Figure 2 The device shown in the figure includes a housing 1 that houses other components. A camera 2 and an active depth sensor 3 are connected to the housing. The camera 2 and the sensor 3 are connected such that they image the same or at least overlapping fields of view. A processor 5 (there can be multiple processors) is communicatively coupled to the camera and the depth sensor to receive data from the camera and the depth sensor. A memory 6 is coupled to the processor. The memory stores code in a non-transitory form, and the code can be executed by the processor to perform the functions described herein. By making such code available for execution, the processor is adapted according to a general-purpose processor to perform these functions. Figure 2 The device shown in the figure is handheld and is powered by a battery or other local energy storage 7.
[0040] Figure 2 The camera 2 shown in the figure can include an image sensor and (optionally) some on-board processing capabilities. For example, the active depth sensor can be a ToF sensor. It can also include some on-board processing capabilities.
[0041] Figure 2 The device shown in the figure can have a transceiver 8 that is capable of communicating with other entities over a network. These entities can be physically remote from Figure 2 the camera device shown in the figure. The network can be a publicly accessible network, such as the Internet. The other entities can be cloud-based. These entities are logical entities. In fact, each of them can be provided by one or more physical devices (e.g., servers and data storage areas), and the functions of two or more of these entities can be provided by a single physical device. Each physical device that implements these entities includes a processor and a memory. Each device can also include a transceiver for sending data to Figure 2 the transceiver 8 of the device shown in the figure and receiving data from that transceiver. Each memory stores code in a non-transitory manner, and the code can be executed by the corresponding processor to implement the corresponding entity in an appropriate manner.
[0042] When performing processing that is beneficial to the Figure 2 device shown in the figure, the processing can be performed only at the device, or all or part of it can be offloaded to the other entities described above.
[0043] The systems described below can estimate scene depth from multiple visual information sources. A learning-based pipeline can combine multiple information sources to recover scene depth estimates. It has been found that this architecture can provide higher accuracy compared to cases where estimates from a single modality are used.
[0044] The depth information obtained from an image can essentially consist of information indicating two things: (1) surface orientation and (2) the point-by-point distance to the sensor that captured the image. The position of such a sensor when capturing the image can be regarded as a reference point for estimating the depth of the image. The methods described below fuse information from multiple modalities. It is necessary to leverage the advantages of each sensor and obtain consistent information from the components.
[0045] In this system, a learning-based strategy is used to perform depth estimation, which involves self-supervised consistency data from multiple modalities. The system combines information from multiple image modalities. In the example described below, these modalities are (1) direct measurement depth data from a time-of-flight (ToF) sensor and (2) polarization data from a visual image. The framework described below can learn (or provide a learned) depth estimation model. The learning of such a model may be based on the concept that the input signals represent a consistent external world and thus must be consistent between image modalities. Thus, for example, if the depth values of a part of the image are all the same, they represent a plane, and the normals extracted in that region need to be similar. It has been found that this results in a method for estimating depth that is trained using multiple image sources but only requires a single modality at inference.
[0046] An end-to-end pipeline for this purpose can be trained using signals of spatial, temporal, and physical model consistency without the need for ground truth annotation labels. The resulting model can utilize stereo, temporal signals, and data from different modalities (such as ToF and polarization images). For a specific example, it can be observed that ToF data is usually clearer at close range and can provide reliable information about the absolute depth distance between the scene surface (target) and the camera sensor (reference point). In this case, active sensors are more accurate because there is no corresponding matching requirement. In contrast, it has been noted that although polarization data is also clearer, it may not be accurate in this regard. However, using the methods described below, polarization data can provide benefits in correctly identifying the surface normal direction. Using this modality provides information about the polarization state of the diffuse light, which in turn allows for establishing correspondences on featureless surfaces, enabling stereo-based surface recovery in a typically challenging setup. The current learning-based strategy allows implicitly leveraging the different advantages of these modalities (short-range, long-range). Learning in multiple modalities also improves the quality of a single individual modality.
[0047] Learning-based methods can now obtain depth estimates from a single RGB image. However, a large number of existing methods treat depth prediction as a supervised regression problem, and thus require a large amount of corresponding ground truth data for model training. Obtaining high-quality depth data as ground truth labels in a series of environments can be considered very costly and usually infeasible. As an alternative to the challenging task of collecting ground truth depth data in a series of environments, self-supervised methods have recently been proposed, eliminating the requirement for per-pixel labels. Then, a self-supervised training signal is defined using (for acquisition) a stereo pair of cameras or a monocular video and an appropriate definition of an image reconstruction training loss. In this way, a self-supervised loss can be constructed using (1) two sensors or (2) a cyclic reconstruction method, where the comparison of the original input data with its reconstructed version can be considered.
[0048] Self-supervised strategies can be extended to multiple image modalities. For each of a set of considered image modalities (e.g., polarization, ToF), the present model utilizes an autoencoder-like architecture with a single decoder network head to perform depth and surface normal prediction tasks. By using an analytical (derivative) transformation from the predicted depth to the surface normal, consistency within the modality task heads can be enforced. By comparing and penalizing the lack of (1) depth prediction consistency and (2) surface normal prediction consistency between the polarization and ToF image network outputs, inter-modality output consistency can also be ensured. (Self-)supervision is achieved using a left-right stereo image pair and performing a left-right consistency check between the image reprojections. This enables the training of a model with geometric consistency without ground truth labels. The self-supervised spatial consistency adaptively utilizes a second sensor, while the self-supervised temporal consistency adaptively utilizes video data. In summary, the intrinsic properties of the combined modalities (polarization, ToF) improve distance and geometry estimation by establishing a physically consistent model for learning. This enables us to define the same model learning strategy for each (two) image modality. It can be noted that specular and diffuse masks are also identified as a by-product of the polarization process. Figure 3 An overview schematic of the current model architecture is provided. Figure 4 Describes how to combine multiple architectures.
[0049] Figure 3An overview of the model for the said processing is shown. The model input is a ToF-related image or a polarization image, and the output is a depth map and a normal map. The architecture consists of a traditional "U-net" (see the article "Unet: Convolutional Networks for Biomedical Image Segmentation" (pages 234 - 241), Springer, Cham.) by O. Ronneberger, P. Fischer, and T. Brox published in the International Conference on Medical Image Computing and Computer-Assisted Intervention in October 2015) and skip connections. The encoder component uses blocks in the "Resnet" (see the article "Image Recognition with Deep Residual Learning" (pages 770 - 778) by K. He, X. Zhang, S. Ren, and J. Sun in the Proceedings of the IEEE International Conference on Computer Vision and Pattern Recognition in 2016) style, and the decoder is a convolutional cascade with layer size adjustment. Each encoder has two decoders. Each decoder outputs a different target; one for the depth image and the other for the normal image. Finally, the depth is parsed and transformed (e.g., using the cross product of the image derivatives in the x and y directions at each position) to form a second normal image. Then, during training, the two normal images are used to enhance consistency, while the depth is the output (the final product).
[0050] Figure 4 Shows how two U-Net architectures can be combined. Using the parsing formula (e.g., see below), predictions for the input (ToF or polarization) can be formed based on the predicted depth. These predictions are then compared with the actual input to help guide the network to predict a more accurate depth. Thus, for example, the predicted ToF from the ToF predicted depth and the predicted ToF from the polarization predicted depth are used to guide the network that takes ToF as input and returns depth, and vice versa.
[0051] The parsing formula is used for (i) estimating synthetic depth data from measured polarization data (in the form of one or more images, which can be captured through a polarization filter at an appropriate angle), and (ii) estimating synthetic polarization data from measured depth data. Then, the synthetic polarization data can be compared with the measured polarization data, and the synthetic depth data can be compared with the measured depth data. Multiple regions of the relevant field of view can be identified. For each such region, a selection can be made based on a comparison of the depth data sources that are most consistent internally, and the depth indicated by that source can be considered as the depth of the corresponding region. Alternatively, another means of fusing data from multiple sources can be used, such as by simple averaging, by selecting the value that is most consistent with adjacent regions, or by using other information in the image (such as brightness).
[0052] The first analysis formula can be used to form synthetic polarization data based on measured or estimated depth data.
[0053] Most light sources emit unpolarized light. When the light is emitted onto an object, it becomes polarized light. The polarization camera captures the polarization intensity in different directions, for example: Capture the polarization intensity, for example:
[0054]
[0055] Where, Represents the polarization angle, i un Is the intensity of unpolarized light, ρ is the degree of linear polarization, and φ is the phase angle.
[0056] The polarization parameters ρ and φ come from the diffuse surface (d) or the specular reflection surface (s), as follows:
[0057]
[0058] Where, Is the viewing angle, η is the refractive index of the object, and
[0059]
[0060] Where, α is the azimuth of the normal Note that the π ambiguity stems from the fact that adding π to φ leaves equation (1) unchanged.
[0061] Finally, the azimuth angle α and the viewing angle θ are obtained as follows:
[0062]
[0063] Where, the view vector pointing from the considered point to the center of the camera is obtained by the following formula
[0064] And
[0065] Where, c x And c y Are the coordinates of the image center, and Z is the estimated depth map.
[0066] The indirect time-of-flight sensor measures the correlation between the transmitted known signal and the received measured signal. By adopting a four-bucket sampling strategy, the distance from the sensor to the object can be recovered. The four correlation measurements are modeled as follows:
[0067] Where, i ∈ {0, 1, 2, 3} (6)
[0068] Among them, A is the amplitude, I is the intensity, and φ is the phase difference between the transmitted signal and the received signal:
[0069]
[0070] Equation 6 shows four images for all i values.
[0071] When comparing the synthesized and measured depth and polarization data, the best image from each of the 24 images from the first formula and the four best images from the second formula can be selected to maintain consistency.
[0072] To train a suitable neural network to implement this system, training data can be formed in an input format, including: video rate I-ToF and stereo polarization (synchronous) imaging sources. To obtain these data, a proof-of-concept embodiment was constructed. It has a hardware camera device that tests the proposed training and inference concepts through real-world (indoor, outdoor) image data. The hardware device consists of a polarized stereo camera, a time-of-flight (ToF) sensor, and a structured light active sensor. It is capable of capturing 1280×720 depth images (from the active sensor), 2448×2048 color polarized raw images (×2), and 4×640×480 8-bit correlation images (ToF). A hardware trigger is used to synchronize the image capture mode, such that one main camera triggers other main cameras through its exposure active side wings, and this enables data to be captured at a speed of approximately 10 frames per second. Other manual controls are available for camera device exposure, gain, focus, and aperture settings.
[0073] The applicant hereby separately discloses each individual feature described herein and any combination of two or more such features. With the ordinary knowledge of those skilled in the art, such features or combinations can be implemented as a whole based on this specification, regardless of whether such features or combinations of features can solve any of the problems disclosed herein, and are not limited to the scope of the claims. This application shows that various aspects of the present invention can be constituted by any such individual feature or combination of features. In view of the foregoing description, various modifications within the scope of the present invention will be obvious to those skilled in the art.
Claims
1. An image processing apparatus for estimating depth of field in a field of view, characterized in that, the apparatus includes one or more processors (5) for: receiving a captured polarization image, the polarization image representing the polarization of light received at a first set of a plurality of positions in the field of view; processing the captured polarization image using a trained first neural network to form a first depth estimate for one or more positions in the field of view; receiving ranging data, the ranging data representing the environmental distance from a reference point to one or more positions in the field of view; processing the ranging data using a trained second neural network to form a second depth estimate for a second set of a plurality of positions in the field of view; forming a synthetic polarization image by processing one or both of the first depth estimate and the second depth estimate using a first analytical formula, the synthetic polarization image representing an estimated value of the polarization of light received at a third set of a plurality of positions in the field of view; forming synthetic ranging data by processing one or both of the first depth estimate and the second depth estimate using a second analytical formula, the synthetic ranging data representing an estimated value of the environmental distance to one or more positions in the field of view.
2. The image processing apparatus according to claim 1, characterized in that, the polarization image represents the polarization of light received at a plurality of positions in the field of view under one or more predetermined polarizations.
3. The image processing apparatus according to claim 1 or 2, characterized in that, the image processing apparatus includes an active sensor device (3) for forming the ranging data, the active sensor device including a time-of-flight sensor.
4. The image processing apparatus according to claim 1, characterized in that, the apparatus is configured to form a plurality of synthetic polarization images, the synthetic polarization images being estimated by the first analytical formula for a plurality of assumed reflection characteristics.
5. The image processing apparatus according to claim 4, characterized in that, the assumed reflection characteristics include diffusivity and specular reflectivity.
6. The image processing apparatus according to claim 1, characterized in that, the apparatus is configured to form a plurality of synthetic polarization images, the synthetic polarization images being estimated by the first analytical formula for a plurality of assumed polarizations.
7. The image processing apparatus according to claim 1, characterized in that, the apparatus is configured to form a plurality of synthetic polarization images, the synthetic polarization images being estimated by the first analytical formula for a plurality of assumed colors.
8. The image processing apparatus according to any one of claims 4 to 7, characterized in that, the step of comparing the polarization image with the synthetic polarization images includes: reducing the plurality of synthetic polarization images to the synthetic polarization image by selecting, for each position in the field of view, the synthetic polarization image that stores the estimated polarization information with the smallest estimated error for that position.
9. The image processing apparatus according to any one of claims 1, 2, 4 - 7, characterized in that, The first analysis formula is used to form distance estimates to multiple positions on the field of view based on the intensity of at least one polarized image at corresponding positions.
10. The image processing device according to any one of claims 1, 2, 4-7, wherein, the second analysis formula is used to form multiple estimates of the distance to the corresponding position for each of the multiple positions on the field of view based on the corresponding phase offset.
11. The image processing device according to any one of claims 1, 2, 4-7, wherein, the image processing device includes a camera (2) for capturing the captured polarized image.
12. The image processing device according to any one of claims 1, 2, 4-7, wherein, the captured polarized image includes a stereoscopic polarized image.
13. The image processing device according to any one of claims 1, 2, 4-7, wherein, the first neural network and the second neural network are the same.
14. The image processing device according to any one of claims 1, 2, 4-7, wherein, the first analysis formula is used to calculate a polarization estimate based on a trigonometric function of an angle formed by a first sub-angle and a second sub-angle, the first sub-angle is calculated based on the normal of the corresponding surface, and the second sub-angle represents a candidate polarization angle.
15. The image processing device according to any one of claims 1, 2, 4-7, wherein, the second analysis formula is used to calculate a distance estimate based on a trigonometric function of an angle formed by a phase value, and the phase value is calculated based on the depth of the corresponding surface.
16. A computer-implemented method for estimating depth of field on a field of view, wherein, the method includes: receiving a captured polarized image, the polarized image representing the polarization of light received at a first set of multiple positions on the field of view; processing the captured polarized image using a trained first neural network to form a first depth estimate for one or more positions on the field of view; receiving ranging data, the ranging data representing the environmental distance from a reference point to one or more positions on the field of view; processing the ranging data using a trained second neural network to form a second depth estimate for a second set of multiple positions on the field of view; forming a synthetic polarized image by processing one or both of the first depth estimate and the second depth estimate using a first analysis formula, the synthetic polarized image representing the estimated polarization of light received at a third set of multiple positions on the field of view; forming synthetic ranging data by processing one or both of the first depth estimate and the second depth estimate using a second analysis formula, the synthetic ranging data representing the estimated environmental distance to one or more positions on the field of view.
Citation Information
Patent Citations
Information acquiring device and information acquiring method
CN108028892A
Strong reflection workpiece vision measurement method based on polarization imaging
CN111127384A