Systems and methods of synthesizing data for training models under different imaging modalities
By synthesizing images in a 3D virtual scene and generating training data using polarization and thermal imaging systems, the problem of time-consuming and expensive manual annotation in existing technologies is solved, enabling efficient training and performance improvement of deep learning models under multiple imaging modalities.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- INTRINSIC INNOVATION LLC
- Filing Date
- 2021-01-04
- Publication Date
- 2026-05-19
AI Technical Summary
In training computer vision models, the current technology involves time-consuming and expensive manual collection and annotation of photos of different scenes. Furthermore, existing 3D rendering software cannot simulate invisible light behaviors such as polarization and thermal radiation, resulting in poor model performance in specific imaging modalities.
By using a synthetic data generator to place object models in a 3D virtual scene, adding lighting and imaging modality-specific materials, rendering synthetic images, capturing multi-angle images using systems such as polarization cameras and thermal imagers, generating a tensor of polarization feature space, and applying techniques such as style transfer to generate training data.
It generates realistic training data, which can effectively train deep learning models for computer vision tasks under various imaging modalities such as polarization and thermal imaging, thereby improving the performance of the models in specific environments.
Smart Images

Figure CN115428028B_ABST
Abstract
Description
[0001] Cross-reference to related applications
[0002] This application claims priority and benefit to U.S. Provisional Patent Application No. 62 / 968,038, filed January 30, 2020, with the United States Patent and Trademark Office, the entire disclosure of which is incorporated herein by reference. Technical Field
[0003] Various aspects of embodiments of this disclosure relate to machine learning techniques, particularly the synthesis or generation of data for training machine learning models. Background Technology
[0004] Large amounts of data are typically used to train statistical models (such as machine learning models). In the field of computer vision, training data often includes labeled images used to train deep learning models (such as convolutional neural networks) to perform computer vision tasks (such as image classification and instance segmentation). However, manually collecting and labeling photographs of various scenes is time-consuming and expensive. Some techniques used to augment these datasets include generating synthetic training data. For example, 3D computer graphics rendering engines (e.g., scanline rendering engines and ray tracing rendering engines) are able to generate realistic 2D images of virtual environments with 3D models of objects that can be used to train deep learning models. Summary of the Invention
[0005] Various aspects of embodiments of this disclosure relate to machine learning techniques, particularly the synthesis or generation of data for training machine learning models. Specifically, various aspects of embodiments of this disclosure relate to synthesizing images for training machine learning models to perform computer vision tasks on input images captured based on imaging modalities different from those of visible light in a scene.
[0006] According to one embodiment of this disclosure, a method for generating a synthetic image of a virtual scene includes: placing 3D models of objects in a three-dimensional (3-D) virtual scene using a synthetic data generator implemented by a processor and a memory; adding illumination to the 3-D virtual scene using the synthetic data generator, the illumination including one or more illumination sources; applying imaging modality-specific materials to the 3D models of the objects in the 3-D virtual scene according to a selected imaging modality using the synthetic data generator, each of the imaging modality-specific materials including an empirical model; setting a scene background according to the selected imaging modality using the synthetic data generator; and rendering a two-dimensional image of the 3-D virtual scene based on the selected imaging modality using the synthetic data generator to generate a synthetic image according to the selected imaging modality.
[0007] The empirical model can be generated based on sampled images captured using an imaging system on the surface of the material being used. The imaging system is configured to capture images using the selected imaging modality. The sampled images may include images of the surface of the material captured from multiple different poses relative to the normal direction of the surface of the material.
[0008] The selected imaging mode may be polarization, and the imaging system includes a polarization camera.
[0009] The selected imaging mode can be thermal, and the imaging system can include a thermal imager. The thermal imager can include a polarization filter.
[0010] Each of the sampled images can be stored in association with a corresponding angle of its pose relative to the normal direction of the surface of the material.
[0011] The sampled images may include: a plurality of first sampled images of the surface of the material illuminated by light having a first spectral profile; and a plurality of second sampled images of the surface of the material illuminated by light having a second spectral profile different from the first spectral profile.
[0012] The empirical model may include a surface light field function calculated by interpolation between two or more of the sampled images.
[0013] The empirical model may include a surface light field function calculated by a deep neural network trained on the sampled image.
[0014] The empirical model may include a surface light field function computed by a generative adversarial network trained on the sampled image.
[0015] The empirical model may include a surface light field function calculated using a mathematical model generated based on the sampled image.
[0016] This method may also include applying style transfer to the synthesized image.
[0017] According to one embodiment of this disclosure, a method for generating a tensor of a polarization feature space for a 3-D virtual scene includes: rendering an image of surface normals of a 3-D virtual scene comprising a 3-D model of a plurality of objects by a synthetic data generator implemented by a processor and a memory, the surface normals including azimuth and zenith components; determining the material of the objects for the surface of the 3-D model of the objects in the 3-D virtual scene by the synthetic data generator; and calculating the tensor of the polarization feature space by the synthetic data generator based on the azimuth and zenith components of the surface normals, the tensor of the polarization feature space including: linear polarization degree; and linear polarization angle of the object surface.
[0018] The method further includes: determining whether the surface of the 3-D model of the object is specularly dominant; in response to determining that the surface of the 3-D model of the object is specularly dominant, calculating a tensor of the polarization feature space based on a specular polarization equation; and in response to determining that the surface of the 3-D model of the object is specularly dominant, calculating a tensor of the polarization feature space based on a diffuse polarization equation.
[0019] This method also includes: calculating the tensor of the polarization feature space based on the diffuse polarization equation.
[0020] This method also includes applying style transfer to the tensor of the polarization feature space.
[0021] According to one embodiment of this disclosure, a method for synthesizing a training dataset is provided, the method being based on generating a plurality of synthetic images generated according to any one of the methods described above.
[0022] According to one embodiment of this disclosure, a method for training a machine learning model includes: generating a training dataset according to any one of the above methods; and calculating parameters of the machine learning model based on the training dataset.
[0023] According to one embodiment of this disclosure, a system for generating a synthetic image of a virtual scene includes: a processor; and a memory storing instructions, which, when executed by the processor, cause the processor to implement a synthetic data generator to: place 3D models of objects in a three-dimensional (3-D) virtual scene; add lighting to the 3-D virtual scene, the lighting including one or more lighting sources; apply imaging modality-specific materials to the 3D models of the objects in the 3-D virtual scene according to a selected imaging modality, each of the imaging modality-specific materials including an empirical model; set a scene background according to the selected imaging modality; and render a two-dimensional image of the 3-D virtual scene according to the selected imaging modality to generate a synthetic image according to the selected imaging modality.
[0024] The empirical model can be generated based on sampled images of the surface of the material captured using an imaging system configured to capture images using the selected imaging modality, and the sampled images can include images of the surface of the material captured from multiple different poses relative to the normal direction of the surface of the material.
[0025] The selected imaging mode may be polarization, and the imaging system may include a polarization camera.
[0026] The selected imaging mode can be thermal, and the imaging system can include a thermal imager. The thermal imager can include a polarization filter.
[0027] Each of the sampled images can be stored in association with a corresponding angle of its pose relative to the normal direction of the surface of the material.
[0028] The sampled images may include: a plurality of first sampled images captured by a material surface illuminated by light having a first spectral profile; and a plurality of second sampled images captured by a material surface illuminated by light having a second spectral profile different from the first spectral profile.
[0029] The empirical model may include a surface light field function calculated by interpolation between two or more of the sampled images.
[0030] The empirical model may include a surface light field function calculated by a deep neural network trained on the sampled image.
[0031] The empirical model may include a surface light field function computed by a generative adversarial network trained on the sampled image.
[0032] The empirical model may include a surface light field function calculated using a mathematical model generated based on the sampled image.
[0033] The memory may also store instructions that, when executed by the processor, cause the synthetic data generator to apply style transfer to the synthetic image.
[0034] According to one embodiment of this disclosure, a system for generating a tensor of a polarization feature space for a 3-D virtual scene includes: a processor; and a memory storing instructions, which, when executed by the processor, cause the processor to implement a synthetic data generator to: render an image of surface normals of a 3-D virtual scene comprising a 3-D model of a plurality of objects, the surface normals including azimuth and zenith components; determine the material of the objects for the surfaces of the 3-D models of the objects in the 3-D virtual scene; and calculate the tensor of the polarization feature space based on the azimuth and zenith components of the surface normals, the tensor of the polarization feature space including: a degree of linear polarization; and a linear polarization angle of the object surface.
[0035] The memory may also store instructions that, when executed by the processor, cause the synthetic data generator to: determine whether the surface of the 3D model of the object is specularly dominant; in response to determining that the surface of the 3D model of the object is specularly dominant, calculate the tensor of the polarization feature space based on the specular polarization equation; and in response to determining that the surface of the 3D model of the object is specularly dominant, calculate the tensor of the polarization feature space based on the diffuse polarization equation.
[0036] The memory may also store instructions that, when executed by the processor, cause the synthetic data generator to calculate the tensor of the polarization feature space based on the diffuse polarization equation.
[0037] The memory may also store instructions that, when executed by the processor, cause the synthetic data generator to apply style transfer to the tensor of the polarization feature space.
[0038] According to one embodiment of the present disclosure, a system for synthesizing a training dataset is provided, the system being configured to synthesize the training dataset using any one of the systems described above.
[0039] According to one embodiment of this disclosure, a system for training a machine learning model includes: a processor; and a memory storing instructions that, when executed by the processor, cause the processor to: receive a training dataset generated by any of the above-described systems; and calculate parameters of the machine learning model based on the training dataset. Attached Figure Description
[0040] The accompanying drawings, together with the specification, illustrate exemplary embodiments of the present invention and, together with the specification, serve to explain the principles of the invention.
[0041] Figure 1This is a block diagram describing a system for training a statistical model to perform computer vision tasks based on images of various modalities, according to embodiments of the present disclosure, wherein training is performed using generated data.
[0042] Figure 2 This is a schematic block diagram of a computer vision system according to an embodiment of the present invention, which is configured to use polarization imaging and can be trained using synthetic polarization image data generated therefrom.
[0043] Figure 3A It is an image or intensity image of a scene, in which a real transparent sphere is placed on top of a printout of a photograph depicting another scene containing two transparent spheres (“deception”) and some background clutter.
[0044] Figure 3B Depicting Figure 3A The intensity image of the superimposed segmentation mask, calculated by a convolutional neural network based on comparison mask regions (Mask R-CNN), is given by the entity with the transparent sphere identified as the entity, where the real transparent sphere is correctly identified as the entity and two deception objects are incorrectly identified as the entity.
[0045] Figure 3C It is the angle of the polarization image calculated from the captured polarization raw frame of the scene according to an embodiment of the present invention.
[0046] Figure 3D An embodiment of the present invention is depicted. Figure 3A An intensity image with an overlay segmentation mask calculated using polarization data, where the real transparent sphere is correctly identified as an entity and the two deceptions are correctly excluded as entities.
[0047] Figure 4 It is a high-level description of the interaction between light and transparent and non-transparent (e.g., diffuse and / or reflective) objects.
[0048] Figure 5 It is a diagram showing the energy of light transmitted and reflected through a surface with a refractive index of approximately 1.5 within the incident angle range.
[0049] Figure 6 This is a flowchart depicting a pipeline for generating a synthetic image according to an embodiment of the present disclosure.
[0050] Figure 7 This is a schematic diagram illustrating the use of a polarization camera system to sample real materials from multiple angles according to an embodiment of the present disclosure.
[0051] Figure 8 This is a flowchart depicting a method for capturing images of a material from different perspectives using a specific imaging modality to be modeled, according to an embodiment of the present disclosure.
[0052] Figure 9 This is a flowchart depicting a portion of a method for rendering a virtual object using a material-based empirical model, according to an embodiment of the present disclosure.
[0053] Figure 10 This is a flowchart depicting a method for calculating a tensor of a synthetic feature or polarization representation space of a virtual scene according to an embodiment of the present disclosure.
[0054] Figure 11 This is a flowchart depicting a method for generating a training dataset according to an embodiment of the present disclosure. Detailed Implementation
[0055] In the following detailed description, only certain exemplary embodiments of the invention are shown and described by way of example. As those skilled in the art will recognize, the invention can be embodied in different forms and should not be construed as limited to the embodiments set forth herein. Throughout the specification, similar reference numerals denote similar elements.
[0056] Various aspects of embodiments of this disclosure relate to systems and methods for synthesizing or generating data for training machine learning models to perform computer vision tasks on images captured based on modalities other than standard modalities, such as color or monochrome cameras configured to capture images based on the intensity of visible light. Examples of other modalities include images captured based on polarized light (e.g., images captured using a polarizing filter or polarization filter in the optical path of a camera for capturing circularly and / or linearly polarized light), non-visible or invisible light (e.g., light in the infrared or ultraviolet range), and combinations thereof (e.g., polarized infrared light). However, embodiments of this disclosure are not limited thereto and can be applied to other multispectral imaging techniques.
[0057] More specifically, aspects of embodiments of this disclosure relate to generating synthetic images using different imaging modalities for training machine learning models to perform computer vision tasks.
[0058] Typically, computer vision systems used to compute segmentation maps that classify objects depicted in a scene can include trained convolutional neural networks that take two-dimensional images (e.g., captured by a color camera) as input and output segmentation maps based on those images. Such convolutional neural networks can be pre-trained on existing datasets (see, for example, J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, ImageNet: A Large-Scale Hierarchical Image Database. IEEE Computer Vision and Pattern Recognition (CVPR), 2009). However, these existing datasets may contain images that do not represent the images expected to be encountered in the specific application of the computer vision system, and therefore these pre-trained models may perform poorly on the specific computer vision task the computer vision system is intended to perform. For example, a computer vision system for a manufacturing environment is more likely to encounter images of tools, partially assembled products, manufactured components, and the like, rather than images of people, animals, household items, and outdoor environments that might be found in a more "general target" dataset.
[0059] Therefore, "retraining" involves updating the parameters (e.g., connection weights) of a pre-trained model based on additional training data from a specific target domain associated with the task to be performed by the retrained model. Continuing the example above, labeled images of tools, partially assembled products, components, and the like from a specific manufacturing environment can be used as training data to retrain a pre-trained model (e.g., a pre-trained convolutional neural network) to improve its performance in detecting and classifying objects encountered in that manufacturing environment. However, manually collecting different images of typical scenes in that manufacturing environment and labeling these images based on their underlying ground truth values (e.g., identifying pixels corresponding to different categories of objects) is typically a time-consuming and expensive task.
[0060] As described above, 3D rendering computer graphics software can be used to generate training data for training machine learning models to perform computer vision tasks. For example, existing 3D models of tools, partially assembled products, and manufacturing components can be arranged in a virtual scene based on various ways such objects might be encountered in the real world (e.g., including lighting conditions and 3D models supporting surfaces and fixtures in the environment). For instance, partially assembled products can be placed on 3D models of conveyor belts, components can be located in parts bins, and tools can be placed on workbenches and / or in a scene depicting the process of positioning components within partially assembled products. Therefore, 3D computer graphics rendering systems are used to generate realistic images of typical arrangements of objects in a specific environment. These generated images can also be automatically labeled. In particular, segmentation maps can be automatically generated (e.g., by mapping object surfaces to their specific category labels) when specific 3D models used to depict each of different types of objects are already associated with category labels (e.g., screws of different sizes, pre-assembled components, products at various assembly stages, specific types of tools, etc.).
[0061] However, 3D rendering computer graphics software systems are typically tailored to generate images representing typical imaging modalities based on the intensity of visible light (e.g., the intensity of red, green, and blue light). Such 3D rendering software (such as Blender Foundation's...) The behavior of electromagnetic radiation that may be invisible or negligible when rendering realistic scenes is typically not considered. Examples of these additional behaviors include the polarization of light (e.g., when polarized light interacts with transparent and reflective objects in the scene, as detected by a camera with a polarization filter in its optical path), thermal or infrared radiation (e.g., as emitted by warm objects in the scene and detected by a camera system sensitive to infrared light), ultraviolet radiation (e.g., as detected by a camera system sensitive to ultraviolet light), and combinations thereof (e.g., polarized and thermal radiation, polarized and visible light, polarized and ultraviolet light, etc.).
[0062] Therefore, aspects of embodiments of this disclosure relate to systems and methods for modeling the behavior of various materials during polarization-based or other imaging modalities. Data (e.g., images) generated according to embodiments of this disclosure can then be used as training data for training deep learning models (such as deep convolutional neural networks) to compute predictions based on imaging modalities other than standard imaging modalities (e.g., the intensity of the visible portion of the visible light or electromagnetic spectrum).
[0063] As an illustrative example, embodiments of this disclosure are described in the context of generating synthetic images of objects captured by polarization filters (referred to herein as “polarized raw frames”), where these images can be used to train deep neural networks (such as convolutional neural networks) to perform tasks based on polarized raw frames. However, embodiments of this disclosure are not limited to generating synthetic polarized raw frames for training convolutional neural networks that take polarized raw frames (or features extracted from them) as input data.
[0064] Figure 1 This is a block diagram depicting a system for training a statistical model to perform computer vision tasks based on images of various modalities, wherein training is performed using data generated according to embodiments of this disclosure. Figure 1 As shown, training data 5 is provided to a model training system 7, which takes model 30 (e.g., a pre-trained model or a model structure with initial weights) and uses the training data 5 to generate a trained model (or a retrained model) 32. Model 30 and the trained model 32 can be statistical models (such as deep neural networks, which include convolutional neural networks). A synthetic data generator 40, according to embodiments of this disclosure, generates synthetic data 42, which may include the training data 5 used to generate the trained model 32. The model training system 7 may apply an iterative process to update the parameters of model 30 to generate the trained model 32 based on the provided training data 5 (e.g., including the synthetic data 42). Updating the parameters of model 30 may include, for example, applying gradient descent (and in neural networks, backpropagation) based on a loss function that measures the difference between the labels and the model's output in response to the training data. The model training system 7 and the synthetic data generator 40 may be implemented using one or more electronic circuits.
[0065] According to various embodiments of this disclosure, the model training system 7 and / or the synthetic data generator 40 are implemented using one or more electronic circuits configured to perform various operations as described in more detail below. Types of electronic circuits may include a central processing unit (CPU), a graphics processing unit (GPU), an artificial intelligence (AI) accelerator (e.g., a vector processor that may include a vector arithmetic logic unit configured to efficiently perform common neural network operations such as dot product and softmax), a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), a digital signal processor (DSP), etc. For example, in some cases, aspects of embodiments of this disclosure are implemented as program instructions stored in a non-volatile computer-readable memory that, when executed by electronic circuits (e.g., a CPU, GPU, AI accelerator, or a combination thereof), perform the operations described herein to compute a segmentation map 20 from an input polarized raw frame 18. The operations performed by the model training system 7 and the synthetic data generator 40 may be performed by a single electronic circuit (e.g., a single CPU, a single GPU, etc.) or may be distributed among multiple electronic circuits (e.g., multiple GPUs or a CPU combined with a GPU). Multiple electronic circuits can be local to each other (e.g., located on the same die, within the same package, or within the same embedded device or computer system) and / or can be remote to each other (e.g., via a network such as a local personal area network such as a local area network). (This includes communication via a local area network (such as a local wired and / or wireless network) and / or via a wide area network (such as the Internet), such as performing some operations locally and other operations on a server hosted by a cloud computing service). One or more electronic circuits for implementing the operation of the model training system 7 and the synthetic data generator 40 may herein be referred to as a computer or computer system, which may include a memory storing instructions that, when executed by the one or more electronic circuits, implement the system and methods described herein.
[0066] Figure 2 This is a schematic block diagram of a computer vision system according to an embodiment of the present invention, which is configured to use polarization imaging and can be trained based on generated synthetic polarization imaging data.
[0067] In context, Figure 2 This is a schematic diagram of a system in which a polarization camera images a scene and provides a polarization raw frame to a computer vision system, which includes a model trained to perform computer vision tasks based on the polarization raw frame or polarization features calculated based on the polarization raw frame.
[0068] The polarization camera 10 has a lens 12 with a field of view, wherein the lens 12 and the camera 10 are oriented such that the field of view surrounds the scene 1. The lens 12 is configured to guide light (e.g., focused light) from the scene 1 to a photosensitive medium (such as an image sensor 14 (e.g., a complementary metal-oxide-semiconductor (CMOS) image sensor or a charge-coupled device (CCD) image sensor)).
[0069] The polarization camera 10 also includes a polarizer, polarization filter, or polarization mask 16 placed in the optical path between the scene 1 and the image sensor 14. According to various embodiments of the present disclosure, the polarizer or polarization mask 16 is configured to enable the polarization camera 10 to capture an image of the scene 1, wherein the polarizer is set at various specified angles (e.g., rotated at 45° or at 60° or at a non-uniform spacing).
[0070] As an example, Figure 2 An embodiment is depicted in which the polarization mask 16 is a polarization mosaic aligned with the pixel grid of the image sensor 14 in a manner similar to the red-yellow-blue (RGB) color filters (e.g., Bayer filters) of a color camera. Similar to how color filters filter incident light based on wavelength, causing each pixel in the image sensor 14 to receive a specific portion of the spectrum (e.g., red, green, or blue) according to the pattern of the color filters in the mosaic, the polarization mask 16 using the polarization mosaic filters light based on linear polarization, causing different pixels to receive light at different angles of linear polarization (e.g., at 0°, 45°, 90°, and 135°, or at 0°, 60°, and 120°). Thus, such as Figure 2 The polarization camera 10 shown, using a polarization mask 16, is capable of capturing four different linearly polarized lights in parallel or simultaneously. An example of a polarization camera is from Wilsonville, Oregon. Systems Company Production S-polarization camera.
[0071] While the foregoing description relates to some possible implementations of a polarization camera using polarization mosaic, embodiments of this disclosure are not limited thereto and include other types of polarization cameras capable of capturing images under multiple different polarizations. For example, the polarization mask 16 may have fewer than or more than four different polarizations, or it may have polarizations at different angles (e.g., polarization angles of 0°, 60°, and 120°, or polarization angles of 0°, 30°, 60°, 90°, 120°, and 150°). As another example, the polarization mask 16 may be implemented using an electronically controlled polarization mask (such as an electro-optic modulator (e.g., possibly including a liquid crystal layer)) where the polarization angle of individual pixels of the mask can be controlled independently, such that different portions of the image sensor 14 receive light with different polarizations. As another example, the electro-optic modulator may be configured to transmit light with different linear polarizations when capturing different frames, for example, so that the camera captures an image using a polarization mask sequentially set to different linear polarizer angles (e.g., sequentially set to: 0 degrees; 45 degrees; 90 degrees; or 135 degrees). As another example, the polarization mask 16 may include a mechanically rotated polarization filter, such that the polarization camera 10 uses the polarization filter, which is mechanically rotated relative to the lens 12, to transmit light to the image sensor 14 at different polarization angles to capture different polarized raw frames.
[0072] A polarization camera can also refer to a multi-camera array with substantially parallel optical axes, such that each camera captures an image of the scene from substantially the same orientation. The optical path of each camera in the array includes a polarization filter, where the polarization filter has a different polarization angle. For example, a 2x2 array of four cameras could include a camera with a polarization filter set at 0°, a second camera with a polarization filter set at 45°, a third camera with a polarization filter set at 90°, and a fourth camera with a polarization filter set at 135°.
[0073] Therefore, the polarization camera captures multiple input images 18 (or polarization raw frames) of scene 1, where each polarization raw frame 18 corresponds to a different polarization angle φ. pol Images captured after a polarization filter or polarizer (e.g., 0°, 45°, 90°, or 135°). Each polarization raw frame is captured from essentially the same pose relative to scene 1 (e.g., images captured using polarization filters at 0°, 45°, 90°, or 135° are all captured by the same polarization camera located and oriented in the same position), rather than from different positions and orientations relative to the scene. Polarization camera 10 can be configured to detect light in various different portions of the electromagnetic spectrum, such as the human-visible portion of the electromagnetic spectrum, the red, green, and blue portions of the human-visible spectrum, and the invisible portions of the electromagnetic spectrum (such as infrared and ultraviolet light).
[0074] Figure 3A , Figure 3B , Figure 3C and Figure 3D Background information is provided for illustrating segmentation maps computed through comparison methods and semantic segmentation or entity segmentation according to embodiments of the present disclosure. More specifically, Figure 3A It is an image or intensity image of a scene, in which a real transparent sphere is placed on top of a printout of a photograph depicting another scene containing two transparent spheres (“deception”) and some background clutter. Figure 3B The diagram depicts the use of different line pattern identifiers computed by a convolutional neural network based on comparison mask regions (Mask R-CNN). Figure 3A The algorithm uses a segmentation mask overlaid on the intensity image of a transparent sphere, where the truly transparent sphere is correctly identified as an entity and two deception objects are incorrectly identified as entities. In other words, the Mask R-CNN algorithm has been fooled into labeling the two deception transparent spheres as entities that are actually transparent spheres in the scene.
[0075] Figure 3C This is a linear polarization angle (AOLP) image calculated from a captured polarized raw frame of a scene according to an embodiment of the present invention. Figure 3C As shown, transparent objects possess a highly distinctive texture in polarization space (such as the AOLP domain), where geometrically dependent markings exist at the edges, and a distinct, unique, or specific pattern is presented on the surface of the transparent object with a linear polarization angle. In other words, the intrinsic texture of a transparent object (e.g., completely different from the extrinsic texture adopted by the background surface visible through the transparent object) in... Figure 3C In the polarization angle image, compared to Figure 3A It is more visible in the intensity image.
[0076] Figure 3D An embodiment of the present invention is depicted. Figure 3A An intensity image with an overlay segmentation mask calculated using polarization data, wherein the overlay line pattern correctly identifies the real transparent sphere as a solid and correctly excludes two deceptions as solids (e.g., with...). Figure 3B compared to, Figure 3D (Not including the overlapping line patterns on the two deceptions). Although Figure 3A , Figure 3B , Figure 3C and Figure 3D Examples illustrating the detection of a real transparent object in the presence of a deceptively transparent object are provided, but embodiments of this disclosure are not limited thereto and can also be applied to other optically challenging objects, such as transparent, translucent, and non-matte or non-Lambertian objects, as well as non-reflective (e.g., matte black objects) and multipath-induced objects.
[0077] Polarization feature representation space
[0078] Some aspects of embodiments of this disclosure relate to systems and methods for extracting features from polarized raw frames, wherein these extracted features are processed by system 100 for robust detection of optically challenging properties in the surface of an object. In contrast, comparison techniques that rely solely on intensity images may fail to detect these optically challenging features or surfaces (e.g., [missing information - likely referring to a specific surface or feature).) Figure 3A Intensity image and Figure 3C (Comparison of AOLP images, as described above). The term "first tensor" in "first representation space" will be used herein to refer to features calculated (e.g., extracted) from the polarized raw frame 18 captured by the polarization camera, wherein these first representation spaces include at least polarization feature spaces (e.g., feature spaces such as AOLP and DOLP containing information about the polarization of light detected by the image sensor), and may also include non-polarization feature spaces (e.g., feature spaces that do not require information about the polarization of light arriving at the image sensor, such as images calculated solely based on intensity images captured without any polarization filters).
[0079] The interaction between light and transparent objects is rich and complex; however, the material of an object determines its transparency in the visible light spectrum. For many transparent household items, most visible light passes directly through, while a small fraction (~4% to ~8%, depending on reflectivity) is reflected. This is because light in the visible portion of the spectrum does not have enough energy to activate the atoms in a transparent object. Therefore, the texture (e.g., appearance) of the object behind (or visible through) a transparent object dominates the appearance of the transparent object. For example, when observing a clear glass and a flat-bottomed glass on a table, the appearance of the object on the other side of the flat-bottomed glass (e.g., the surface of the table) usually dominates what is seen through the glass. This property presents some difficulties when attempting to detect surface features of transparent objects (such as windows and smooth, transparent coatings) solely based on intensity images.
[0080] Figure 4 It is a high-level description of the interaction between light and transparent and non-transparent (e.g., diffuse and / or reflective) objects. Figure 4 As shown, polarization camera 10 captures a raw polarization frame of a scene that includes a transparent object 402 in front of an opaque background object 403. Light 410 striking the image sensor 14 of polarization camera 10 contains polarization information from both the transparent object 402 and the background object 403. A small portion of the reflected light 412 from the transparent object 402 is strongly polarized, thus having a significant impact on polarization measurements, unlike the light 413 reflected away from the background object 403 and passing through the transparent object 402.
[0081] Similarly, light rays striking the surface of an object can interact with the surface shape in various ways. For example, a surface with a smooth coating can behave in a manner similar to... Figure 4 The opaque object shown is largely similar to the transparent object in front of it, where the interaction between light and a transparent or translucent layer (or varnish layer) of a smooth coating causes the light reflected off the surface to be polarized based on the properties of the transparent or translucent layer (e.g., based on the layer thickness and surface normal), which are encoded in the light that strikes the image sensor. Similarly, as discussed in more detail below with respect to the theory of shape of polarization (SfP), variations in the shape of a surface (e.g., the direction of the surface normal) can lead to significant changes in the polarization of light reflected from the object's surface. For example, a smooth surface typically exhibits the same polarization properties overall, but scratches or dents in the surface alter the direction of the surface normal in those areas, and light striking a scratch or dent may be polarized, attenuated, or reflected in a manner different from that in other parts of the object's surface. Models of the interaction between light and matter typically consider three fundamental elements: geometry, illumination, and material. Geometry is based on the shape of the material. Illumination includes the direction and color of the light. Material can be parameterized by the refractive index of light or angular reflection / transmission. This angular reflection is known as the bidirectional reflection distribution function (BRDF), although other functional forms can more accurately represent certain scenarios. For example, the two-way subsurface scattering distribution function (BSSRDF) will be more accurate in the case of materials that exhibit subsurface scattering (e.g., marble or wax).
[0082] The light 410 of the image sensor 16 of the impact polarization camera 10 has three measurable components: the intensity of the light (intensity image / I), the percentage or proportion of linear polarization of the light (degree of linear polarization / DOLP / ρ), and the direction of that linear polarization (angle of linear polarization / AOLP / φ). These attributes encode information about the surface curvature and material of the imaged object, which can be used by the predictor 800 to detect transparent objects, as described in detail below. In some embodiments, the predictor 800 can detect other optically challenging objects based on similar polarization properties of light passing through a translucent object and / or light interacting with a multipath-induced object or a non-reflective object (e.g., a matte black object).
[0083] Therefore, some aspects of embodiments of this disclosure relate to synthesizing polarized raw frames that can be used to compute a first tensor in one or more first representation spaces, the first tensor including derived feature maps based on intensity I, DOLPρ, and AOLPφ. Some aspects of embodiments of this disclosure also relate to directly synthesizing tensors (such as DOLPρ and AOLPφ) of one or more representation spaces for training a deep learning system to perform computer vision tasks based on information about the polarization of light in a scene (and, in some embodiments, based on other imaging modalities, such as thermal imaging and a combination of thermal and polarization imaging).
[0084] Measuring the intensity I, DOLPρ, and AOLPφ at each pixel requires measurement at different angles φ after a polarization filter (or polarizer). pol Three or more raw polarized frames of the scene being filmed (e.g., because there are three unknowns to be determined: intensity I, DOLPρ, and AOLPφ). For example, as described above. The S-polarization camera captures the polarization angle φ. pol Four polarization original frames are generated by using polarization original frames at 0 degrees, 45 degrees, 90 degrees, or 135 degrees. In this paper, they are represented as I0 and I. 45 I 90 , and I 135 .
[0085] At each pixel The relationship between strength I, DOLPρ, and AOLPφ can be expressed as:
[0086]
[0087] Therefore, through four different polarization original frames (I0、I 45 I 90 , and I 135 The intensity I, DOLPρ, and AOLPφ can be solved using a system of four equations.
[0088] The polarization shape (SfP) theory (see, for example, Gary A. Atkinson and Edwin R. Hancock, Recovery of surface orientation from diffuse polarization. IEEE Transactions on Image Processing, 15(6):1653-1664, 2006.) states that when diffuse polarization dominates, the refractive index (n) of the surface normal of an object and the azimuth angle (θ) are related to the polarization shape (SfP). a ) and zenith angle (θ) zThe relationship between the φ and ρ components of the light rays from the object follows the following characteristics:
[0089]
[0090] φ=θ a (3)
[0091] And when specular reflection is dominant:
[0092]
[0093]
[0094] Note that in both cases, ρ changes with θ z The polarization increases exponentially with increasing refractive index, and if the refractive index is the same, specular reflection is more polarized than diffuse reflection.
[0095] Therefore, some aspects of embodiments of this disclosure involve applying SfP theory to generate synthetic original polarization frames 18 and / or AOLP and DOLP images based on the shape (e.g., orientation) of surfaces in a virtual environment.
[0096] Light rays from a transparent object have two components: the reflected component includes the reflected intensity I. r , reflection DOLPρ r and reflection AOLPφ r The refractive part includes refractive intensity I t , refracted DOLPρ t and refraction AOLPφ t The intensity of a single pixel in the resulting image can be written as:
[0097] I = I r +I t (6)
[0098] When it has a linear polarization angle φ pol When the polarization filter is placed in front of the camera, the value at a given pixel is:
[0099]
[0100] According to I r ρ r φ r I t ρ t and φ t Solve the above expressions for the pixel values in the DOLPρ image and the AOLPφ image:
[0101]
[0102]
[0103] Therefore, according to one embodiment of this disclosure, the above formulas (7), (8) and (9) provide a model for forming a first tensor 50 comprising an intensity image I, a DOLP image ρ, and an AOLP image φ, wherein the use of polarization images or tensors (including the DOLP image ρ and AOLP image φ based on formulas (8) and (9)) in the polarization representation space enables a trained computer vision system to reliably detect optically challenging surface properties of objects that are typically undetectable by a comparison system that uses only the intensity I image as input.
[0104] More specifically, the first tensor in the polarization representation space (such as polarization images DOLPρ and AOLPφ) can reveal surface properties of an object that may otherwise appear lacking in texture in the intensity domain I. Transparent objects can possess textures invisible in the intensity domain I because the intensity strictly depends on I. r / I t The ratio (see formula (6)). Different from I t An opaque object has a value of 0, while a transparent object transmits most of the incident light and reflects only a small portion of it. As another example, thin or small deviations in the shape of other smooth surfaces (or smooth portions of other rough surfaces) may be substantially invisible or have low contrast in the intensity I domain (e.g., a domain that does not consider the polarization of light), but may be clearly visible or have high contrast in polarization representation spaces (such as DOLPρ and AOLPφ).
[0105] Therefore, an exemplary method for acquiring surface topography is to use polarization cues in conjunction with geometric regularization. Fresnel equations relate DOLPρ and AOLPφ to surface normals. These equations can be used for anomaly detection by utilizing so-called surface polarization modes. A polarization mode is a tensor of size [M, N, K], where M and N are the horizontal and vertical pixel dimensions, respectively, and where K is the polarization data channel, the size of which can vary. For example, if circular polarization is ignored and only linear polarization is considered, K will equal 2, because linear polarization has both polarization angle and polarization degree (DOLPρ and AOLPφ). Similar to moiré patterns, in some embodiments of this disclosure, the feature extraction module 700 extracts polarization patterns in a polarization representation space (e.g., DOLPρ space and AOLPφ space), as shown above. Figure 1 A and Figure 1In example feature output 20 of B, the horizontal and vertical dimensions correspond to the lateral field of view of a narrow strip or patch of the surface of an object captured by polarization camera 10. However, this is an exemplary case: in many embodiments, the narrow strip or patch of the surface may be vertical (e.g., taller than the width), horizontal (wider than the height), or have a more conventional field of view (FoV) that tends to be close to a square (e.g., with an aspect ratio of 4:3 or 16:9).
[0106] While the foregoing discussion provides specific examples of linear polarization-based polarization representation spaces when using a polarization camera with one or more linear polarization filters to capture polarization-based raw frames corresponding to different angles of linear polarization and to compute tensors of polarization representation spaces (such as DOLP and AOLP), embodiments of this disclosure are not limited thereto. For example, in some embodiments of this disclosure, the polarization camera includes one or more circular polarization filters configured to pass only circularly polarized light, and wherein a first tensor of polarization mode or circular polarization representation space is further extracted from the polarization-based raw frames. In some embodiments, these additional tensors of the circular polarization representation space are used alone, and in other embodiments they are used in conjunction with tensors of linear polarization representation spaces (such as AOLP and DOLP). For example, a polarization pattern that includes tensors of polarization representation spaces may include tensors in the circular polarization space, AOLP, and DOLP, wherein the polarization pattern may have dimensions [M, N, K], where K is 3, to further include tensors of the circular polarization representation space.
[0107] Figure 5 It is a diagram showing the energy of light transmitted and reflected through a surface with a refractive index of approximately 1.5 within the incident angle range. For example... Figure 5 As shown, at low incident angles (e.g., at angles close to a plane perpendicular to the surface), the transmitted energy (such as...) Figure 5 (as shown by the solid line in the image) and reflected energy (such as...) Figure 5 The slope of the reflected energy (shown by the dashed line in the diagram) is relatively small. Therefore, when the angle of incidence is small (e.g., close to perpendicular to the surface, in other words, close to the surface normal), small differences in the surface angle may be difficult to detect in the polarization mode (low contrast). On the other hand, as the angle of incidence increases, the slope of the reflected energy increases gradually, while the slope of the projected energy decreases gradually (with a larger absolute value) as the angle of incidence decreases. Figure 5 In the example shown with a refractive index of 1.5, the slopes of the two lines begin to steepen substantially at an incident angle of approximately 60°, and their slopes become very steep at an incident angle of approximately 80°. For different materials, the specific shape of the curves may vary depending on the material's refractive index. Therefore, the incident angles corresponding to the steeper portions of the curves (e.g., angles close to parallel to the surface, such as approximately 80° in the case of a refractive index of 1.5) will vary. Figure 5(As shown) Capturing an image of the surface being detected can improve the contrast and detectability of changes in surface shape in the polarization original frame 18, and can also improve the detectability of such features in the tensor of the polarization representation space, because small changes in the angle of incidence (due to small changes in the surface normal) can lead to large changes in the captured polarization original frame.
[0108] The use of polarization cameras to detect the presence and shape of optically challenging objects and surfaces is described in more detail in, for example, PCT patent application No. US / 2020 / 048604, filed August 28, 2020, and PCT patent application No. US / 2020 / 051243, filed September 17, 2020, the entire disclosures of which are incorporated herein by reference. Such computer vision systems can be trained to perform computer vision tasks on polarization data based on training data generated according to embodiments of this disclosure. In some embodiments, these computer vision systems use machine learning models (such as deep neural networks (e.g., convolutional neural networks)) to perform computer vision tasks, wherein the deep learning model is configured to take features in the polarization raw frame and / or polarization representation space as input.
[0109] Simulating the polarization physics of different materials is a complex task, requiring knowledge of material properties, the spectral distribution and polarization parameters of the illumination used, and the angle at which the observer views the reflected light. Realistically simulating the physics of light polarization and its effects on object illumination is not only a complex task but also a computationally intensive one, involving complex forward models that often produce highly inaccurate (unrealistic) images. Therefore, various comparative 3D computer imaging systems typically fail to accurately model the physics of light polarization and its effects on object illumination, and consequently, if imaging is performed using a camera with polarization filters in its optical path (e.g., a polarizing camera), images of virtual environments cannot be synthesized or rendered in a way that realistically represents the corresponding real-world environment. Consequently, comparative techniques for generating synthetic data for training computer vision systems operating on standard imaging modalities (such as visible light images without polarization filters) are generally incapable of generating training data for training computer vision systems operating on other imaging modalities (e.g., polarizing cameras, thermal imagers, etc.).
[0110] As described above, various aspects of embodiments of this disclosure relate to generating or synthesizing data for training machine learning models to take data captured using imaging modalities other than those captured by a standard camera (e.g., a camera configured to capture the intensity of visible light without using filters such as polarization filters) as input, which will be referred to herein as multimodal images or all-optical images. The term "multimodal" refers to the all-optical theory of light, where each dimension of the all-optical domain (e.g., wavelength, polarization, angle, etc.) is an example of a mode of light. Therefore, multimodal or all-optical imaging includes, but is not limited to, using multiple imaging modalities simultaneously. For example, the term "multimodal" may be used herein to refer to a single imaging modality, where the single imaging modality is a modality different from the intensity of visible light without using filters such as polarization filters. Polarized raw frames and / or tensors of polarization representation spaces captured by one or more polarization cameras are an example of one type of input in multimodal or all-optical imaging modalities (e.g., using multimodal imaging or all-optical imaging).
[0111] Generally, various aspects of embodiments of this disclosure relate to four techniques that can be used individually or in combination as part of a pipeline for generating synthetic training data based on multimodal or all-optical imaging modalities, such as polarization imaging modalities. These techniques include: domain randomization, texture mapping, normal mapping, and style transfer, which are discussed in more detail below.
[0112] Figure 6 This is a flowchart describing a pipeline for generating a composite image according to one embodiment of the present disclosure. In some embodiments of the present disclosure, Figure 6 The operations are performed by the synthetic data generator 40, for example, by dedicated program instructions stored in the memory of the synthetic data generator 40. When executed by the processor of the synthetic data generator 40, these dedicated program instructions cause the synthetic data generator to perform the dedicated operations described herein for generating synthetic images based on physical simulations of optical phenomena. For convenience, various aspects of embodiments of this disclosure will be described in the context of applying polarization imaging to perform computer vision tasks on optically challenging manufactured components and tools, such as objects with transparent surfaces, shiny metallic surfaces, and / or dark matte surfaces, within a manufacturing context.
[0113] In operation 610, the composite data generator 40 places a 3-D model of the object in a virtual scene. The 3-D model of the object can be readily obtained from computer-aided design (CAD) models of components and partially or fully assembled manufactured products, within the context of generating a composite image of the scene in the manufacturing environment. These CAD models may have been previously generated during the product design phase and can be obtained from, for example, component suppliers (e.g., from suppliers providing components to manufacturers), publicly available information (e.g., data sheets), or from in-house product designers employed by the manufacturer. In some cases, CAD models can be manually generated based on the technical specifications of the components.
[0114] In some embodiments of this disclosure, 3D models of objects are placed in a virtual scene in a manner similar to the arrangement that those objects are expected to encounter for a specific computer vision task performed to train a machine learning model.
[0115] In the above example of computer vision in a manufacturing context, one task is to perform entity segmentation on a box of components, where the components may be homogeneous (e.g., all components in the box are identical, such as a spring box or a screw box) or heterogeneous (e.g., a mixture of different types of components, such as screws of different sizes or screws mixed with matching nuts). Objects may be randomly arranged within the box, where components can be oriented in many different directions, and where, in heterogeneous component boxes, different types of components are mixed together rather than separated into different parts of the box. A computer vision system can be trained to compute a segmentation map of the box to identify the location and orientation of individual components within the box (and, in the case of heterogeneous component boxes, to identify the type of object). This segmentation map can then be used by an actuator system (such as a robotic arm) to pick up components from the box and add the picked-up components to a partially assembled product.
[0116] Therefore, in some embodiments of this disclosure, the composite data generator 40 generates a scene of the component within the box by placing a 3-D model of the virtual box in the scene and casting 3-D models of the components into the virtual box, such as by using a physics simulation engine (e.g., a physics engine incorporated into a 3-D computer graphics rendering system). For example, 3D rendering software includes physical systems that simulate various real-world physical phenomena, such as the movement, collisions, and potential deformation of rigid bodies, cloth, soft bodies, and fluids under the influence of gravity or other forces. Therefore, rigid body simulation can be used to simulate rigid components (e.g., screws, bolts, relatively stiff springs) falling into a rigid virtual box, and soft body simulation can be used to simulate elastic or deformable components (e.g., ropes, wires, plastic sheets, etc.) falling into a rigid simulation box.
[0117] More specifically, various different scenarios representing different potential states of a box can be generated, such as by dropping various numbers of entities of 3-D models of components into the virtual box. For example, if a typical box has a maximum capacity of 1000 screws, various scenarios can be generated by dropping 1000 screws, 900 screws, 500 screws, 100 screws, and 10 screws into the virtual box to generate different scenarios representing different potential fullness states of the virtual box. Furthermore, multiple scenarios can be generated for any given number of screws (or the number of screws may be random between the generation of different scenarios), where the arrangement of components within the box is also random, such as by dropping components into the box from different random positions above the box at once.
[0118] Therefore, in operation 610, the synthetic data generator 40 generates a scene including the arrangement of representative objects.
[0119] In operation 630, the synthetic data generator 40 adds lighting to the virtual scene generated in operation 610. Specifically, the synthetic data generator 40 adds one or more light sources to the virtual scene, where the light sources illuminate part or all of the surface of objects in the bin. In some embodiments, the positions of the one or more light sources are randomized, and multiple scenes are generated using light sources at different locations (e.g., different angles and distances) relative to the parts bin to improve the robustness of training. In some embodiments of this disclosure, the virtual lighting includes virtual light sources representing light sources found in the environment in which the computer vision system is trained. Examples of potential representative light sources include different color temperatures corresponding to, for example, incandescent lamps, fluorescent lamps, light-emitting diode (LED) bulbs, natural light from simulated windows in the environment, and other forms of lighting technology, where the shape of the virtual light (e.g., the direction of the light emitted by the lamp) can range from direct light to diffused light. In some embodiments of this disclosure, the characteristics of the light (e.g., color temperature and shape) are also randomized to generate different scenes with different types of lighting.
[0120] In operation 650, the synthetic data generator 40 applies a modality-specific material to objects in a 3-D virtual scene. For example, in the case of generating synthetic polarization imaging data, a polarization-specific material is applied to objects in the virtual scene, while in the case of generating synthetic thermal imaging data, a thermal imaging-specific material can be applied to objects in the virtual scene. For ease of illustration, polarization-specific materials will be described in detail herein, but embodiments of this disclosure are not limited thereto, and materials specific to multimodal imaging and / or all-optical imaging modality generation and application can also be applied.
[0121] Some aspects of embodiments of this disclosure involve domain randomization, where the material appearance of objects in a scene is randomized beyond the typical appearance of the objects. For example, in some embodiments, a large number of materials with random colors (e.g., thousands of different materials of different colors randomly selected) are applied to different objects in a virtual scene. In a real-world environment, objects in a scene typically have well-defined colors (e.g., rubber washers typically appear as matte black and screws may be specific shades of glossy black, matte black, gold, or bright metallic). However, real-world objects may often have different appearances due to variations in lighting conditions (such as the color temperature of light, reflection, specular highlights, etc.). Therefore, applying randomization to the colors of the materials applied to objects when generating training data expands the domain of the training data to include unrealistic colors, thereby increasing the diversity of training data for training more robust machine learning models that can make more accurate predictions (e.g., more accurate entity segmentation maps) under a wider range of real-world conditions.
[0122] Some aspects of embodiments of this disclosure involve performing texture mapping to generate a model of a material (parametric material) that depends on one or more parameters based on an imaging modality. For example, as described above, the appearance of a given surface in a scene imaged by a polarizing camera system can be based on the material properties of the surface, the spectral profiles and polarization parameters of one or more illumination sources (light sources) in the scene, the angle of incidence of light onto the surface, and changes in the viewpoint angle of the observer (e.g., the polarizing camera system). Therefore, simulating the physics of polarization in different materials is a complex and computationally intensive task.
[0123] Therefore, some aspects of embodiments of this disclosure relate to simulating the physical phenomena of various imaging modalities based on empirical data, such as captured real-world images of real-world materials. More specifically, an imaging system (e.g., a polarization camera system) implementing a particular imaging modality of interest is used to collect sample images from objects made of that particular material. In some embodiments, the collected sample images are used to calculate empirical models of the material, such as its surface optical field function (e.g., a bidirectional reflectance density function or BRDF).
[0124] Figure 7 This is a schematic diagram illustrating the use of a polarization camera system to sample real material from multiple angles according to an embodiment of the present disclosure. Figure 8 This is a flowchart depicting a method 800 for capturing images of a material from different perspectives using a specific imaging modality to be modeled, according to an embodiment of the present disclosure. Figure 7As shown, the surface 702 of the physical object (e.g., a washer, screw, etc.) is made of a material of interest (e.g., black rubber, chrome-plated stainless steel, etc., respectively). In operation 810, this material is placed in a physical scene (e.g., on a laboratory workbench). In operation 830, a physical lighting source 704 (such as an LED light or fluorescent light) is placed in the scene and arranged to illuminate at least a portion of the surface 702. For example, as... Figure 7 As shown, light 706 emitted from physical lighting source 704 is incident on surface 702 at a specific point 708 at an angle α relative to the normal direction 714 of surface 702 at that specific point 708.
[0125] In operation 850, the imaging system is used to capture images of the surface 702 of the object from multiple poses relative to the normal direction of the surface. Figure 7 In the illustrated embodiment, a polarization camera system 710 is used as an imaging system to capture images of surface 702 (including portions illuminated by a physical illumination source 704, such as specific points 708). The polarization camera system 710 captures images of surface 702 from different poses 712, such as by moving the polarization camera system 710 from one pose to the next, and capturing polarized raw frames from each pose. Figure 7 In the illustrated embodiment, the polarization camera system 710 images the surface 702 at a 0° frontal parallel observer angle β (e.g., a frontal parallel view aligned with the surface normal 714 at point 708) in a first posture 712A, at an intermediate observer angle β (such as 45° relative to the surface normal 714) in a second posture 712B, and at a shallow observer angle β (e.g., slightly less than 90 degrees, such as 89°) relative to the surface normal 714 in a third posture 712C.
[0126] As described above, the polarization camera system 710 is typically configured to capture polarized raw frames using polarization filters of different angles (e.g., a polarization mosaic with four different polarization angles in the optical path of a single lens and sensor system, an array of four cameras, each with a linear polarization filter of different angles, polarization filters with different angles set for different frames captured from the same pose at different times, etc.).
[0127] In operation 870, the image captured by the imaging system is stored along with the relative pose of the camera with respect to the normal direction of the surface (e.g., observer angle β). For example, the observer angle β may be stored in metadata associated with the image and / or the image may be indexed in part based on the observer angle β. In some embodiments, the image may be indexed by parameters including: observer angle β (or the angle of the camera position relative to the surface normal), material type, and illumination type.
[0128] Therefore, in such Figure 7 The arrangement shown uses, for example Figure 8 The method involves a polarization camera system 710 capturing multiple images of a material under given illumination conditions (e.g., in the case of a known spectral profile of a physical illumination source 704) at different reflection angles (e.g., at different orientations 712). These images include four images with linear polarization angles of 0°, 45°, 90°, and 135°.
[0129] Due to the nature of the physical phenomenon of polarization, each of these viewpoints or poses 712 gives a different polarization signal. Therefore, by capturing images of surface 702 from different observer angles, a model of the material's BRDF can be estimated based on interpolation between images captured by a camera system at one or more poses 712 with the closest corresponding observer angle β using physical illumination sources 704 at one or more of the closest corresponding incident angles α.
[0130] Although for convenience Figure 7 The embodiments depict only three poses 712, but the embodiments of this disclosure are not limited to this, and the material can be sampled at higher rates (such as with a 5° interval or smaller between adjacent poses). For example, in some embodiments, the polarization camera system 712 is configured to operate as a video camera system, wherein polarized raw frames are captured at high rates (such as 30 frames per second, 60 frames per second, 120 frames per second, or 240 frames per second) to obtain high-density images captured at a large number of angles relative to the surface normal.
[0131] Similarly, in some embodiments, the orientation of the physical illumination source 704 relative to the surface 702 is modified such that light emitted from the physical illumination source 704 is incident on the surface 702 at different angles α, wherein multiple images of the surface are similarly captured by the polarization camera system 710 from different orientations 712.
[0132] The sampling rates for different angles (e.g., the incident angle α and the observer or polarizing camera system angle β) can be selected such that intermediate viewpoints can be interpolated (e.g., bilinearly interpolated) without significant loss of realism. In various embodiments of this disclosure, the spacing of the intervals may depend on the physical characteristics of the imaging modalities, some of which exhibit greater angular sensitivity than others, and therefore higher accuracy can be achieved with fewer poses (wider spacing) for modalities with lower angular sensitivity, while a larger number of poses (closer spacing) can be used for modalities with higher angular sensitivity. For example, in some embodiments, when capturing polarized raw frames for a polarizing imaging modal, the pose of the polarizing camera system 710 is set to an interval angle of approximately five degrees (5°), and images of surface 702 can also be captured by physical illumination source 704 at various locations (similarly spaced at an angle of approximately five degrees (5°)).
[0133] In some cases, the appearance of a material in the imaging mode of an empirical model also depends on the type of lighting source, such as incandescent lamps, fluorescent lamps, light-emitting diode (LED) bulbs, sunlight, and therefore parameters of one or more lighting sources used to illuminate a real-world scene are included as parameters of the empirical model. In some embodiments, different empirical models are trained for different lighting sources (e.g., one model of the material under natural lighting or sunlight and another model of the material under fluorescent lighting).
[0134] Reference Figure 6 In some embodiments, during operation 670, the synthetic data generator 40 sets a virtual background for the scene. In some embodiments, the virtual background is an image captured using the same imaging modality as the modality simulated by the synthetic data generator 40. For example, in some embodiments, when generating a synthetic polarized image, the virtual background is a real image captured using a polarization camera, and when generating a synthetic thermal image, the virtual background is a real image captured using a thermal imager. In some embodiments, the virtual background is an image of an environment similar to the environment in which the trained machine learning model is intended to operate (e.g., a manufacturing site or factory in the case of a computer vision system used to manufacture robots). In some embodiments, the virtual background is randomized, thereby increasing the diversity of the synthetic training dataset.
[0135] In operation 690, the synthetic data generator 40 renders a 3D scene based on a specified imaging modality (e.g., polarization, thermal, etc.) using one or more of an empirically derived modality-specific model of the material. Some aspects of embodiments of this disclosure relate to rendering images based on an empirical model of a material according to one embodiment of this disclosure. The empirical model of the material can be developed, as described above, based on samples collected from captured images of real-world objects made of the material of interest.
[0136] Typically, 3D computer graphics rendering engines generate a 2D rendering of a virtual scene by calculating the color of each pixel in the output image based on the colors of the surfaces of the virtual scene depicted by pixels. For example, in a ray tracing rendering engine, virtual rays are emitted from a virtual camera into the virtual scene (the opposite of the typical path of light in the real world), where the virtual rays interact with the surfaces of 3D models of objects in the virtual scene. These 3D models are typically represented using geometry such as a mesh of points defining flat surfaces (e.g., triangles), where these surfaces can be specified with materials describing how the virtual rays interact with the surfaces (e.g., reflection, refraction, scattering, dispersion, and other optical effects) and textures representing the colors of the surfaces (e.g., the texture can be a solid color or can be, for example, a bitmap image applied to the surface). The path of each virtual ray is traced (or “traced”) through the virtual scene until it reaches a light source (e.g., a virtual luminaire) in the virtual scene, and the cumulative modifications of the textures encountered along the optical path from the camera to the light source are combined with the characteristics of the light source (e.g., the color temperature of the light source) to calculate the color of the pixel. As those skilled in the art will understand, this general process can be modified, such as by performing anti-aliasing (or smoothing) by tracing multiple rays through different parts of each pixel and calculating the pixel's color based on a combination (e.g., average) of different colors calculated by tracing different rays interacting with the scene.
[0137] Figure 9 This is a flowchart depicting a method 900 for rendering a portion of a virtual object using a material-based empirical model according to an embodiment of this disclosure. Specifically, Figure 9 Embodiments involving calculating color when tracing a ray of light through a pixel in a virtual scene (such as when the ray interacts with a surface having a material modeled according to an embodiment of this disclosure). However, those skilled in the art would understand prior to the effective filing date of this application how the techniques described herein can be applied as part of a larger rendering process, where multiple colors are calculated and combined for a given pixel of an output image, or where a scanline rendering process is used instead of ray tracing.
[0138] More in detail, Figure 9The embodiments described herein depict a method for rendering the surface of an object in a virtual scene based on the view from a virtual camera in the virtual scene, wherein the surface has a material modeled according to embodiments of this disclosure. Given that objects are being rendered, and the compositing data generator 40 has access to the ground-real geometry of each object being rendered, the per-pixel normal, material type, and lighting type are known parameters for the graphical rendering of the material, appropriately modulating its use. During the rendering process, camera rays are traced from the optical center of the virtual camera to each 3-D point on the object visible from the camera. Each 3-D point on the object (e.g., having XYZ coordinates) is mapped to 2-D coordinates (e.g., having UV coordinates) on the object's surface. Each UV coordinate of the object's surface has its own surface light field function (e.g., a bidirectional reflection function or BRDF), which is represented as a model based on the image generation of the real material, as described above, for example regarding... Figure 7 and Figure 8 As stated above.
[0139] In operation 910, the compositing data generator 40 (e.g., running a 3D computer graphics rendering engine) determines the normal direction of a given surface (e.g., relative to a global coordinate system). In operation 930, the compositing data generator 40 determines the material of the object's surface and assigns it to the surface as part of the design of the virtual scene.
[0140] In operation 950, the synthetic data generator 40 determines the observer angle β of the surface, for example, based on the direction of light reaching the surface (e.g., the angle from the virtual camera to the surface if the surface is the first surface reached by light from the camera, otherwise the angle from another surface in the virtual scene to that surface). In some embodiments, in operation 950, the incident angle α is also determined based on the angle of light leaving the surface (e.g., in the direction toward a virtual light source in the scene, due to the reversal of the light direction during ray tracing). In some cases, the incident angle α depends on the properties of the material determined in operation 930, such as whether the material is transparent, reflective, refractive, diffuse (e.g., matte), or a combination thereof.
[0141] In operations 970 and 990, the synthetic data generator 40 configures a model of the material based on the observer angle β (and, if applicable, the incident angle α and other conditions such as the spectral profile or polarization parameters of the lighting sources in the scene), and calculates the color of the pixels in part based on the configured model of the material. The model of the material can be retrieved from a collection or database of models of different standard materials (e.g., model materials have been empirically generated based on the type of material expected to be depicted in a virtual scene generated by the synthetic data generator 40 for generating training data for a specific application or use case (such as materials used to manufacture components for a specific electronic device in the case of computer vision for supporting the manufacture of electronic devices in robotics)). The model is generated based on images of real materials captured as described above according to embodiments of this disclosure. For example, in operation 930, the synthetic data generator 40 may determine that the surface of an object in the virtual scene is made of black rubber; in this case, in operation 970, a model of the material generated from a captured image of a real surface made of black rubber is loaded and configured.
[0142] In some cases, virtual scenes include objects having materials made of materials not represented in a database or collection of material models. Therefore, some aspects of embodiments of this disclosure involve simulating the appearance of materials that do not have a perfect or similar match in a database of material models by interpolating between predictions made by different real-world models. In some embodiments, existing materials are represented in an embedding space based on a set of parameters characterizing the material. More formally, an interpretable material embedding M such that F(M) glass θ out φ out The light field at the polarized surface of the glass is given by (θ, x, y), where the observer angle β is given by (θ, x, y). out φ out ) represents, and at the position (x, y) on the surface (mapped to the (u, v) coordinate space on the 3-D surface), and a similar embedding F(M) can be performed for another material (such as rubber). rubber θ out φ out (x, y). This embedding of the material in the embedding space can then be parameterized in an interpretable manner using, for example, a beta variational autoencoder (VAE), and then interpolated to generate new material not directly based on empirically collected samples, but rather interpolated between multiple different models built separately based on their own empirically collected samples. The generation of additional material in this manner further extends the domain randomness of the synthetic training data generated according to embodiments of this disclosure and improves the robustness of deep learning models trained on such synthetic data.
[0143] The various embodiments of this disclosure relate to different ways in which the model of the material can be implemented.
[0144] In various embodiments of this disclosure, the model representing the surface light field function or BRDF of the material is used, for example, a deep learning-based BRDF function (e.g., based on a deep neural network (such as a convolutional neural network)), a mathematically modeled BRDF function (e.g., a set of one or more closed equations or one or more mathematically solvable open equations), or a data-driven BRDF function using linear interpolation.
[0145] In operation 970, the synthetic data generator 40 configures a model of the material identified in operation 950 based on current parameters (such as the incident angle α and the observer angle β). In the case of using a data-driven BRDF function with linear interpolation, in operation 970, the synthetic data generator 40 retrieves an image of the material identified in operation 950 whose parameters are closest to the current ray in the parameter space. In some embodiments, the material is indexed (e.g., stored in a database or other data structure) and accessed based on the material type, illumination type, incident angle of the light, and the angle of the camera relative to the surface normal of the material (e.g., the observer angle). However, embodiments of this disclosure are not limited to the parameters listed above, and other parameters may be used, depending on the characteristics of the imaging modality. For example, for some materials, the incident angle and / or illumination type may have no effect on the appearance of the material, and therefore these parameters may be omitted and do not need to be determined as part of method 900.
[0146] Therefore, in the case of a data-driven BRDF function with linear interpolation, in operation 970, the synthetic data generator 40 retrieves one or more images that are closest to the current ray associated with the current pixel being rendered, given parameters. For example, the observer angle might be 53° relative to the surface normal of an object, and a sample of real-world material might include images captured at observer angles spaced 5° apart, in this example, images captured at 50° and 55° relative to the surface normal of a real-world object made of the material of interest. Thus, images of real-world material captured at 50° and 55° will be retrieved (in the case of additional parameters, these parameters, such as incident angle and illumination type, will further identify the specific images to be retrieved).
[0147] Continuing with the example of a data-driven BRDF function with linear interpolation, in operation 990, the synthetic data generator 40 calculates the surface color of a pixel based on the closest image. In the case of only one matching image (e.g., if the observer angle in the virtual scene matches an observer angle in one of the sampled images), that sampled image is used directly to calculate the surface color. In the case of multiple matching images, the synthetic data generator 40 interpolates the colors of the multiple images. For example, in some embodiments, linear interpolation is used to interpolate between the multiple images. More specifically, if there are four different sample images with different observer angles in the azimuth relative to the surface normal and the polar angle relative to the incident angle of the illumination source, bilinear interpolation can be used to interpolate between the four images along the azimuth and polar directions. As another example, if the appearance of the material also depends on the incident angle, further interpolation can be performed based on images captured at different incident angles (while interpolating between images captured at different observer angles for each of the different incident angles). Thus, in operation 990, the surface color of the scene is calculated for the current pixel based on a combination of one or more images captured from real-world materials.
[0148] In some embodiments of this disclosure where the model is a deep learning network, the surface light field function of the material is implemented using a model that includes training a deep neural network to predict the value of a bidirectional reflection function directly from a set of parameters. More specifically, images of the real material captured from multiple different poses (as shown above, for example, regarding...) Figure 7 and Figure 8 The described data is used to generate training data relating to parameters (such as observer angle β, incident angle α, spectral characteristics of the illumination source, etc.) of the observed appearance of a portion of the material (e.g., the center of an image). Therefore, in some embodiments, a deep neural network is trained (e.g., by applying backpropagation) to estimate the BRDF function based on training data collected from images of the collected real material. In these instances, in operation 970, the model is configured such that if multiple deep neural networks are available, one deep neural network is selected from among those associated with the model (e.g., selected based on matching the parameters of the virtual scene with the parameters of the data used to train the deep neural network, such as the parameters of the illumination source), and the parameters are provided to the selected deep neural network (or a single deep neural network, if only one deep neural network is associated with the model) with inputs (such as observer angle β, incident angle α, etc.). In operation 990, the synthetic data generator 40 calculates the color of the surface of a virtual object from the configured model by forward propagation through the deep neural network to calculate the color at the output, wherein the calculated color is the color of the surface of the virtual scene as predicted by the deep neural network of the configured model.
[0149] In some embodiments of this disclosure where the model is a deep neural network, the surface light field function of the material is implemented using a model including one or more conditional generative adversarial networks (see, for example, "Generative adversarial nets." Advances in Neural Information Processing Systems. 2014.) Each conditional generative adversarial network can be trained to generate images of the material based on random input and one or more conditions, where the conditions include current parameters of the observed surface (e.g., observer angle β, incident angle α of each illumination source, polarization state of each illumination source, and material properties of the surface). According to some embodiments, a discriminator is trained adversarially to determine whether the input image is a real image captured under given conditions or a real image generated by the conditional generator based on that set of conditions, based on the input image and a set of conditions associated with the image. By alternately retraining the generator to generate images that can "fool" the discriminator and training the discriminator to distinguish between generated and real images, the generator is trained to generate real images of the material captured under various capture conditions (e.g., at different observer angles), thereby enabling the trained generator to represent the surface light field function of the material. In some embodiments of this disclosure, different generative adversarial networks (GANs) are trained for different conditions of the same material (e.g., for different types of lighting sources, different polarization states of the lighting sources, etc.). In these embodiments, in operation 970, the model is configured such that if there are multiple conditional generative adversarial networks (GANs) associated with the model, one conditional GAN is selected from the multiple conditional GANs associated with the model (e.g., selected based on matching the parameters of the virtual scene with the parameters of the data used to train the deep neural network, such as the parameters of the lighting source), and the parameters of the virtual scene are provided as conditions for the conditional GAN, such as the observer angle β, the incident angle α, etc. In operation 990, the synthetic data generator 40 calculates the color of the surface of the virtual object from the configured model by forward propagation through the conditional GAN to calculate the color at the output (e.g., a synthetic image of the surface of the object based on the current parameters), wherein the calculated color is the color of the surface of the virtual scene as generated by the conditional GAN of the configured model.
[0150] In some embodiments of this disclosure, the surface optical field function is modeled using a closed-form mathematically derived bidirectional reflectance distribution function (BRDF), which is derived from (such as according to...) Figure 8The described method is configured using empirically collected samples (e.g., images or photographs) of real material captured from different angles. For example, examples of techniques for configuring a BRDF based on images or photographs of material collected from different angles or poses are described in Ramamoorthi, Ravi, and Pat Hanrahan, "A Signal-Processing Framework for Inverse Rendering." Proceedings of the 28th Annual Conference on Computer Graphics and Interactive Techniques. 2001. and Ramamoorthi, Ravi, "A Signal-Processing Framework for Forward and Inverse Rendering." Stanford University, 2002, 52-79. Therefore, in some embodiments, the mathematically derived BRDF of closed form is configured using empirically collected samples of real material and is included as a component of a material model for modeling the multimodal and / or all-optical properties of the material, which is used for computer rendering of multimodal and / or all-optical images of virtual scenes.
[0151] In some embodiments, virtual objects made of multiple materials, whether bimodal or multimodal, will have a similar set of images for each material type used in the virtual object in question. The appearances of the different materials are then combined in the final rendering of the images (e.g., additively combined according to weights associated with each material in the virtual model). In some embodiments, this same method of combining multiple materials is also applied to multilayer materials, such as transparent layers on a shiny material. In some embodiments, multilayer materials are modeled by individually sampling the multilayer materials (e.g., capturing their images).
[0152] The final effect of the rendering process using an empirical model of materials according to embodiments of this disclosure is that the final render has a simulated polarization signal that closely approximates the real polarization signal in the real environment. The accuracy of the material depicted by the empirical model in the rendered virtual environment depends on the degree of matching between the conditions of the virtual environment and the conditions for capturing samples of real-world materials (e.g., the degree of matching between the spectral profile of the lighting source in the virtual scene and the real-world lighting source, the degree of matching between the observer angle in the virtual scene and the observer angle used when performing sampling, etc.).
[0153] As described above, while aspects of embodiments of the present disclosure are described herein in the context of simulating or mimicking the appearance of polarization, embodiments of the present disclosure are not limited thereto. The appearance of materials in multimodal imaging modes and / or all-optical imaging modes (such as thermal imaging, thermal imaging with polarization, etc.) can also be captured according to embodiments of the present disclosure. For example, the behavior of materials in thermal imaging modes (e.g., infrared imaging) can similarly be captured by using... Figure 7 Thermal imagers with similar arrangements as shown are Figure 8 The method described herein models the material by capturing images from multiple poses. Based on these captured images, the corresponding images can then be retrieved and, if necessary, compared with... Figure 9 The method shown is similar to interpolating images in a 3D rendering engine to simulate the appearance of materials under thermal imaging.
[0154] Therefore, some embodiments of this disclosure relate to systems and methods for generating synthetic image data of virtual scenes when they appear in various imaging modalities (such as polarization imaging and thermal imaging) by rendering images of virtual scenes using empirical models of materials in the virtual scene. In some embodiments, these empirical models may include images of real-world objects captured using one or more imaging modalities (such as polarization imaging and thermal imaging). This synthetic image data can then be used to train machine learning models to operate on image data captured by imaging systems using these imaging modalities.
[0155] Some aspects of embodiments of this disclosure relate to generating synthetic data about image features, typically derived from imaging data. As a specific example, some aspects of embodiments of this disclosure relate to generating synthetic features or tensors representing polarization space (e.g., linear degree of polarization or DOLPρ and linear angle of polarization or AOLPφ). As described above, the polarization shape (SfP) provides DOLPρ and AOLPφ with respect to the refractive index (n) of the surface normal of the object, the azimuth angle (θ)... a ) and zenith angle (θ) z The relationship between ).
[0156] Therefore, some aspects of the embodiments relate to the refractive index (n) and azimuth angle (θ) of the surface based on the virtual scene. a ) and zenith angle (θ) z (These are all known parameters of the virtual 3D scene) The linear polarization degree or DOLPρ and linear polarization angle or AOLPφ of the synthesized surfaces visible to the virtual camera in the virtual scene are generated.
[0157] Figure 10This is a flowchart depicting a method 1000 for calculating synthetic features or tensors of a polarization representation space of a virtual scene according to an embodiment of the present disclosure. In operation 1010, a synthetic data generator 40 renders a normal image (e.g., an image of the direction of surface normals of the virtual scene at that pixel, with each pixel corresponding to that pixel). The normal vector at each component includes an azimuth angle θ. a Components and zenith angle θ z Components. In operation 1030, the synthetic data generator 40 divides the normal vector at each point of the normal image into two components: the azimuth angle θ at that pixel. a and zenith angle θ z As described above, these components can be used to calculate estimates of DOLPρ and AOLPφ by using the shapes from polarization equations (2) and (3) for the diffuse case and equations (4) and (5) for the specular case. To simulate real polarization errors, in some embodiments of this disclosure, the synthetic data generator 40 applies a semi-global perturbation to the normal map before applying the polarization equations (e.g., polarization equations (2), (3), (4), and (5)). This perturbation changes the size of the normal while preserving the gradient of the normal. This simulates the error caused by the material properties of objects and their interaction with polarization. In operation 1050, for a given pixel, the synthetic data generator 40 determines the material of the object's surface (e.g., the material associated with the surface at each pixel in the normal map) based on the parameters of the object in the virtual scene, and uses this material in conjunction with the geometry of the scene according to 3-D rendering techniques to determine whether the given pixel is specularly dominant. If so, the synthetic data generator 40 calculates DOLPρ and AOLPφ based on the specular equations (4) and (5) in operation 1092. If not, the synthetic data generator 40 calculates DOLPρ and AOLPφ based on the diffuse equations (2) and (3) in operation 1094.
[0158] In some embodiments, it is assumed that all surfaces are diffuse, and therefore operations 1050 and 1070 can be omitted, and the synthesized DOLPρ and AOLPφ are calculated based on the shapes derived from the polarization equations (2) and (3) for diffuse.
[0159] In some embodiments, the synthesized DOLPP1 and AOLPφ data are rendered into color images by applying color maps (such as "viridis" color maps or "Jet color maps"). (See, for example, "Somewhere over the rainbow: An empirical assessment of quantitative color maps" by Liu, Yang, and Jeffrey Heer, Proceedings of the 2018 CHI Conference on Human Factors in Computing) Systems. 2018. These color-mapped versions of the synthesized tensors in polarization space can be more easily provided as input for retraining pre-trained machine learning models (such as convolutional neural networks). In some embodiments, when synthesizing DOLPρ and AOLPφ data, random color maps are applied to various synthetic data such that the synthesized training dataset includes color images representing DOLPρ and AOLPφ data in various different color maps, enabling the network to perform predictions at inference time regardless of the specific color map used to encode the real DOLPρ and AOLPφ data. In other embodiments of this disclosure, the same color map is applied to all synthesized DOLPρ and AOLPφ data (or a first color map is used for DOLPρ and a different second color map is used for AOLPφ), and at inference time, color maps are applied to the tensors extracted in the polarization representation space to match the synthesized training data (e.g., the same first color map is used to encode DOLPρ extracted from captured real polarization raw frames and the same second color map is used to encode AOLPφ extracted from captured real polarization raw frames).
[0160] Therefore, some aspects of embodiments of this disclosure relate to synthesizing features of a representation space specific to a particular imaging modality, such as by synthesizing the DOLPρ and AOLPφ of the polarization representation space of a polarization imaging modality.
[0161] Some aspects of embodiments of this disclosure relate to combinations of the above-described techniques for generating synthetic images for training machine learning models. Figure 11 This is a flowchart depicting a method for generating a training dataset according to an embodiment of the present disclosure. One or more virtual scenes representing a target domain can be generated as described above (e.g., for generating an image of a component box, by selecting one or more 3-D models of the components and casting the entities of the 3-D models into the container). For example, some aspects of embodiments of the present disclosure relate to forming a training dataset based on: (1) images generated solely by domain randomization in operation 1110, and (2) images generated solely by texture mapping in operation 1112 (e.g., according to...). Figure 9The image generated by the embodiment, and (3) the image generated in operation 1114 solely by normal mapping (e.g., according to Figure 10 (Image generated by the embodiment).
[0162] In addition, the training dataset may include images generated by models using materials generated through interpolation between different empirical generative models (as described above, parameterized embedding spaces).
[0163] In some embodiments of this disclosure, the generated images based on (1) domain randomization, (2) texture mapping, and (3) normal mapping are further processed by applying style transfer or other filters to the generated images in operations 1120, 1122, and 1124, respectively, before the images are added to the training dataset. Applying style transfer results in a more consistent appearance compared to images generated using the three techniques described above. In some embodiments, the style transfer process transforms the synthesized input image into one that looks more like the original polarized frame based on the imaging modality of interest (e.g., causing the image generated using (1) domain randomization and the feature map generated using (3) normal mapping to look more like the original polarized frame) or by making the synthesized input image look more like an artificial one (such as by applying an unrealistic painting style to the input image (e.g., causing the image generated using (1) domain randomization, the rendering using (2) texture mapping, and the feature map generated using (3) normal mapping to look like a painting drawn on a canvas with a brush)).
[0164] In some embodiments, a neural style transfer network is trained and used to perform style transfer on images selected for the training dataset in operation 1122, such as SytleGAN for complex global style transfer (see, for example, "Analyzing and improving the image quality of stylegan." Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2020.); a patch-based network for local style transfer (see, for example, "Fastpatch-based style transfer of arbitrary style." arXiv preprint arXiv:1612.04337 (2016.); and a network using domain adaptation (see, for example, "Domainstylization: A strong, simple baseline for synthetic to real image domain adaptation." arXiv preprint arXiv:1807.09384 (2018.)). Therefore, all images in the training dataset can have a similar style or appearance, regardless of how the images were obtained (e.g., transformed by style transfer operations) (e.g., whether captured by (1) domain randomization, (2) texture mapping, (3) normal mapping or other sources, such as real images of objects captured using an imaging system that implements the modality of interest (e.g., polarization imaging or thermal imaging).
[0165] In some embodiments of this disclosure, the images in the training dataset are sampled from synthetic datasets (1), (2), and (3) based on hard example mining (see, for example, "Hard example mining with auxiliary embeddings" by Smirnov, Evgeny, et al., Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, 2018). Sampling the synthetic dataset using hard example mining can improve the efficiency of the training process by reducing the size of the training set, removing a large number of redundant images that have little impact on the training process, and retaining the "hard examples" that have a greater impact on the resulting trained model.
[0166] As briefly mentioned above, when generating training data for supervised learning, the synthetic data generator 40 also automatically generates labels for the synthesized images (e.g., the expected output). For example, when generating training data for training a machine learning model to perform an image classification task, the labels generated for a given image may include the categories of objects depicted in the image. These category labels can be generated by identifying each unique category of objects visible in the virtual scene. As another example, when generating training data for training a machine learning model to perform an entity segmentation task, the generated labels may include a segmentation map in which each entity of each object, along with its category (e.g., objects of the same category have the same category identifier), is uniquely identified (e.g., with different entity identifiers). For example, the segmentation map can be generated by tracing light rays from the camera to the virtual scene, where each ray may intersect some first surface of the virtual scene. Each pixel of the segmentation map is labeled accordingly based on the entity identifier and category identifier of the object that the light rays emitted from the camera strike through the pixel on the surface.
[0167] As described above, and with reference Figure 1 The resulting training dataset 42, generated by the synthetic data generator 40, is then used by the model training system 7 as training data 6 to train model 30 (such as a pre-trained model or a model initialized with random parameters) to produce trained model 32. Continuing with the example above, where training data is generated based on polarization imaging modalities, training dataset 5 can be used to train model 30 to operate on polarization input features (such as polarization original frames (e.g., images generated via texture mapping) and tensors of the polarization representation space (e.g., images generated via normal mapping)).
[0168] Therefore, training data 5, including synthetic data 42, is used to train or retrain machine learning model 30 to perform computer vision tasks based on specific imaging modalities. For example, synthetic data based on polarization imaging modalities can be used to retrain convolutional neural networks that may have been pre-trained to perform entity segmentation based on standard color images to perform entity segmentation based on polarization input features.
[0169] In deployment, a trained model 32, trained based on training data generated according to embodiments of this disclosure, is then configured to take inputs similar to the training data (such as polarization original frames and / or tensors of polarization representation spaces) (where these input images are further modified by applying the same style transfer (if any) when generating the training data) to generate predicted outputs (such as segmentation maps).
[0170] While this document describes some embodiments of the present disclosure with respect to polarization imaging modalities, embodiments of the present disclosure are not limited thereto and include multimodal imaging modalities and / or all-optical imaging modalities, such as thermal imaging, thermal imaging with polarization (e.g., with polarization filters), and ultraviolet imaging. In these embodiments using different modalities, real-world image samples captured from real-world materials using imaging systems implementing these modalities are used to generate models of the materials as they will appear in these imaging modalities, and the surface light field functions of the materials relative to these modalities are modeled as described above (e.g., using deep neural networks, generative networks, linear interpolation, explicit mathematical models, etc.) and used to render images according to these modalities using a 3-D rendering engine. The rendered images in the modalities can then be used to train or retrain one or more machine learning models (such as convolutional neural networks) to perform computer vision tasks based on input images captured using these modalities.
[0171] Therefore, aspects of embodiments of this disclosure relate to systems and methods for generating simulated or synthetic data, which represent image data captured by imaging systems using various imaging modalities, such as polarization, thermal, ultraviolet, and combinations thereof. Simulated or synthetic data can be used as training datasets and / or to augment training datasets for training machine learning models to perform tasks, such as computer vision tasks, on data captured using imaging modalities corresponding to the imaging modalities of the simulated or synthetic data.
[0172] While the invention has been described in conjunction with certain exemplary embodiments, it should be understood that the invention is not limited to the disclosed embodiments, but is instead intended to cover various modifications and equivalent arrangements included within the spirit and scope of the appended claims and their equivalents.
[0173] In some embodiments of this disclosure, the order of operations performed may differ from the order depicted in the figures and described herein. For example, although Figure 6 An example of a method for generating synthetic images is described, but embodiments of this disclosure are not limited thereto. For example, Figure 6 Some of the operations shown can be performed in different orders or simultaneously. As a specific example, in various embodiments of this disclosure, operations such as placing 3D models of objects in virtual scene 610, adding lighting to virtual scene 630, applying modality-specific materials to objects in virtual scene 650, and setting scene background 670 can be performed in various orders before rendering the 3D scene based on a specified imaging modality in operation 690. As another example, although... Figure 8 An embodiment in which lighting is placed in a real-world scene after a real-world object is placed in the scene is described; however, the embodiments disclosed herein are not limited thereto, and lighting may be added to the scene before a real-world object is placed in the scene.
[0174] In some embodiments of this disclosure, certain operations may be omitted or not performed, and in some embodiments, additional operations not described herein may be performed before, after, or between the various operations described herein.
Claims
1. A method for generating synthetic images of a virtual scene, comprising: A 3D model of an object in a 3D virtual scene is obtained through a synthetic data generator implemented by a processor and memory. The synthetic data generator adds lighting to the 3D virtual scene, the lighting including one or more virtual lighting sources; To obtain a model that simulates the polarization characteristics of an object with a specific surface material; Determine the observer's angle of the objects in the virtual scene. The synthetic data generator generates linear polarization degree (DOLP) and linear polarization angle (AOLP) images from an observer's perspective of an object having the specific surface material, including: The observer's angle is provided as input to the empirical model to generate data representing the corresponding polarization signal of a specific surface material at each location of the object in the scene, and The DOLP image and the AOLP image are calculated from data representing the corresponding polarization signals at each location in the scene corresponding to the object; and The machine learning model is trained using the DOLP and AOLP images generated by the synthetic data generator.
2. The method of claim 1, wherein the empirical model is generated based on sampled images of the surface of the material captured using an imaging system configured to capture polarization signals, and The sampled images include images of the surface of the material captured from multiple different poses relative to the normal direction of the surface of the material.
3. The method of claim 2, wherein the imaging system includes a polarization camera.
4. The method of claim 2, wherein each of the sampled images is stored in association with a corresponding angle of its pose relative to the normal direction of the surface of the material.
5. The method of claim 2, wherein the sampled image comprises: Multiple first sampled images captured by light illuminating the surface of the material with a first spectral profile; Multiple second sampled images captured by light illuminating the surface of the material with a second spectral profile different from the first spectral profile.
6. The method of claim 2, wherein the empirical model comprises a surface light field function calculated by interpolation between two or more of the sampled images.
7. The method of claim 2, wherein the empirical model comprises a surface light field function computed by a deep neural network trained on the sampled image.
8. The method of claim 2, wherein the empirical model comprises a surface light field function computed by a generative adversarial network trained on the sampled image.
9. The method of claim 2, wherein the empirical model comprises a surface light field function calculated by a mathematical model generated based on the sampled image.
10. A system for generating synthetic images of virtual scenes, comprising: processor; as well as A memory storing instructions that, when executed by the processor, cause the processor to implement a synthetic data generator to perform operations including: The synthetic data generator obtains a 3D model of an object in a 3D virtual scene. The synthetic data generator adds lighting to the 3D virtual scene, the lighting including one or more virtual lighting sources; To obtain a model that simulates the polarization characteristics of an object with a specific surface material; Determine the observer's angle of the objects in the virtual scene. The synthetic data generator generates linear polarization degree (DOLP) and linear polarization angle (AOLP) images from an observer's perspective of an object having the specific surface material, including: The observer's angle is provided as input to the empirical model to generate data representing the corresponding polarization signal of a specific surface material at each location of the object in the scene, and The DOLP image and the AOLP image are calculated from data representing the corresponding polarization signals at each location in the scene corresponding to the object; and The machine learning model is trained using the DOLP and AOLP images generated by the synthetic data generator.
11. The system of claim 10, wherein the empirical model is generated based on sampled images of the surface of the material captured using an imaging system configured to capture polarization signals, and The sampled images include images of the surface of the material captured from multiple different poses relative to the normal direction of the surface of the material.
12. The system of claim 11, wherein the imaging system includes a polarization camera.
13. The system of claim 11, wherein each of the sampled images is stored in association with a corresponding angle of its pose relative to the normal direction of the surface of the material.
14. The system of claim 11, wherein the sampled image comprises: Multiple first sampled images captured by light illuminating the surface of the material with a first spectral profile; as well as Multiple second sampled images captured by light illuminating the surface of the material with a second spectral profile that is different from the first spectral profile.
15. The system of claim 11, wherein the empirical model comprises a surface light field function calculated by interpolation between two or more of the sampled images.
16. The system of claim 11, wherein the empirical model comprises a surface light field function computed by a deep neural network trained on the sampled image.
17. The system of claim 11, wherein the empirical model comprises a surface light field function computed by a generative adversarial network trained on the sampled image.
18. The system of claim 11, wherein the empirical model comprises a surface light field function calculated by a mathematical model generated based on the sampled image.