Method and device for generating third-person perspective images and method for training a neural network
By projecting input images onto virtual surfaces and a common coordinate frame using a machine learning model, the method generates accurate third-person perspective images, addressing distortion issues in automotive surround view systems and improving driver assistance.
Patent Information
- Application Number
- JP2025511455
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-08-23
- Filing Date
- 2023-08-16
- Publication Date
- 2025-08-15
AI Technical Summary
Existing automotive surround view systems suffer from distortion and object doubling due to incorrect projection of camera images onto a fixed bowl approximation, and neural network-based approaches struggle to generate accurate third-person perspective images, especially when cameras are mounted low on vehicles, and fail to generalize across different vehicle types.
A method involving projecting input images onto virtual surfaces around an object, then onto a common coordinate frame, using a machine learning model to generate a third-person perspective image without relying on depth estimation, and employing multiple projections to determine the most accurate viewpoint.
This approach provides accurate third-person perspective images, enhancing driver situational awareness and maneuvering capabilities by avoiding distortions and generalizing across various vehicle types.
Smart Images

Figure 2025526975000001_ABST
Abstract
Description
[Technical Field]
[0001] Various embodiments relate to methods and devices for generating third-person perspective images, and methods for training neural networks. [Background technology]
[0002] Automotive surround view (SV) systems play an important role in assisting drivers in driving functions such as parking and other maneuvers. SV systems also improve road safety by providing drivers with perspectives around the vehicle, thereby eliminating blind spots. SV systems can include multiple cameras and a processor that stitches camera data from multiple cameras to generate a 360° view around the vehicle. These SV systems typically project camera images onto a fixed bowl approximation and stitch the images together. Because these camera images are not projected onto a correct model of the environment, the generated view may show distortion and doubling of objects. One approach to reducing distortion is to deform the bowl approximation based on static object information in the environment for a more accurate representation of the environment. Implementing such a solution requires information about the distance between static objects and the vehicle. Most neural network-based approaches assume the existence of at least some estimate of the depth of the surrounding scene, which can provide distance information. However, such approaches have various drawbacks. For example, monocular and binocular depth estimation can be inaccurate when the cameras are mounted low on the vehicle. Also, for outdoor scenes with a wide range of different distance scales to consider, it can be difficult for a single neural network to accurately render a projection. Furthermore, providing a driver with a third-person perspective, also referred to herein as a third-person view image, that allows the driver to view their vehicle and its surroundings as if they were outside the vehicle, is also useful in driver assistance applications. However, the neural network-based approaches described above do not apply to generating third-person perspectives. These approaches also do not generalize well to multiple vehicle types.
[0003] In view of the above, there is a need for an improved method for generating third-person perspective images that can address at least some of the above problems. Summary of the Invention
[0004] According to various embodiments, a computer-implemented method for generating a third-person perspective image is provided. The method can include receiving a plurality of input images capturing a surrounding of an object. The method can further include projecting each input image of the plurality of input images onto a respective virtual surface of a set of virtual surfaces around the object to generate a set of surface images. The method can further include projecting each surface image of the set of surface images onto a common coordinate frame to generate a transformed dataset. The method can further include generating the third-person perspective image based on the transformed dataset with a machine learning model.
[0005] According to various embodiments, there is provided a device for generating a third-person perspective image, which may include a processor configured to perform the above-described method.
[0006] According to various embodiments, there is provided a use of the above-described device for remotely operating a vehicle.
[0007] According to various embodiments, a computer-implemented training method for training a machine learning model is provided. The training method may include projecting each training image of a plurality of training images capturing a surrounding of a training object onto a respective virtual surface of a set of virtual surfaces surrounding the training object to generate a set of surface images. The training method may further include projecting each surface image of the set of surface images onto a common coordinate frame to generate a training dataset. The training method may further include training the machine learning model using the training dataset as an input to the machine learning model and further using a set of ground truth third-person perspective images corresponding to the plurality of training images as a training signal.
[0008] According to various embodiments, a data structure generated by the training method described above is provided.
[0009] Further features for advantageous embodiments are provided in the dependent claims.
[0010] In the drawings, like reference characters generally refer to the same parts throughout the different views. The drawings are not necessarily to scale, emphasis instead generally being placed upon illustrating the principles of the invention. In the following description, various embodiments are described with reference to the following drawings: [Brief explanation of the drawings]
[0011] [Figure 1A] FIG. 1 illustrates a simplified functional block diagram of a device for generating a third-person perspective image, according to various embodiments. [Figure 1B] 1 illustrates a simplified hardware block diagram of a device according to various embodiments. [Figure 2A] 1A-1D show top views of an object equipped with an SV system according to various embodiments. [Figure 2B] 2B illustrates a 3D representation of a set of virtual surfaces around the object of FIG. 2A, according to various embodiments. [Figure 3] 1 illustrates a flow diagram of a method for generating a third-person perspective image 120, according to various embodiments. [Figure 4] FIG. 1 illustrates a block diagram of an example neural network, according to various embodiments. [Figure 5] FIG. 1 illustrates a block diagram of a machine learning model, according to various embodiments. [Figure 6] 1 illustrates a flow diagram of a method for training a machine learning model, according to various embodiments. [Figure 7A] 1 shows an example of an equipment setup for collecting ground truth third-person perspective images. [Figure 7B] 1 shows an example of an equipment setup for collecting ground truth third-person perspective images. DETAILED DESCRIPTION OF THE INVENTION
[0012] Embodiments described below in relation to devices are equally valid for the respective methods, and vice versa. Furthermore, it will be understood that the embodiments described below may be combined, e.g., parts of one embodiment may be combined with parts of another embodiment.
[0013] It will be understood that any characteristic described herein for a particular device may also hold for any device described herein. It will be understood that any characteristic described herein for a particular method may also hold for any method described herein. Furthermore, it will be understood that for any device or method described herein, not all components or steps described must necessarily be enclosed in the device or method, and that only some (but not all) components or steps may be enclosed.
[0014] The term "coupled" (or "connected") in this specification may be understood to mean electrically coupled, e.g., communicatively coupled to transmit and receive data wirelessly or by wire, or mechanically coupled, e.g., attached or fixed, or simply in contact without any fixation, and it will be understood that both direct coupling and indirect coupling (in other words, coupling without direct contact) may be provided.
[0015] The term "third-person perspective image" may be referred to interchangeably as "third-person image" or "third-person view image."
[0016] The term "ground truth image" can refer to a real image, i.e., an image captured directly by a camera, while a generated viewpoint can refer to a synthetic image generated by a system or machine learning model.
[0017] The term "coordinate frame" can refer to a set of three vectors that have unit length and are perpendicular to each other, and can serve as a reference for defining positions.
[0018] In this regard, the devices described herein may include memory, for example, for use in processes performed on the device. The memory used in the embodiments may be volatile memory, for example, DRAM (Dynamic Random Access Memory), or non-volatile memory, for example, PROM (Programmable Read Only Memory), EPROM (Erasable PROM), EEPROM (Electrically Erasable PROM), or flash memory, for example, floating gate memory, charge trap memory, MRAM (Magnetoresistive Random Access Memory), or PCRAM (Phase Change Random Access Memory).
[0019] In order that the invention may be readily understood and put into practice, various embodiments will now be described by way of example, and not by way of limitation, with reference to the drawings.
[0020] According to various embodiments, a visualization method for an advanced driver assistance system (ADAS) can be provided. The visualization method can include synthesizing a viewpoint based on input from a vehicle's surround view system. The synthesized viewpoint can allow a driver to view their vehicle in their environment from a virtual location. The synthesized viewpoint is sometimes referred to as "novel" because the vehicle itself does not have sensors positioned at the virtual location. The visualization method can synthesize a novel viewpoint by projecting images captured by the surround view system onto a virtual surface surrounding the vehicle and then reprojecting all of the projected images to a common viewpoint or coordinate frame. The reprojected image data, along with ground truth images from the common viewpoint, can be used to train a machine learning model. The synthesized viewpoint can include a third person view, also referred to herein as an outsider perspective view. The common viewpoint can be an outsider perspective view.
[0021] Advantageously, the visualization method can provide the driver with a view of their vehicle relative to its environment, which can assist the driver in maneuvering the vehicle. For example, the synthesized perspective can allow the driver to see what is behind or to the side of the vehicle so that the driver can maneuver in tight spaces without damaging their vehicle. The visualization method can include a method 300 for generating a third-person perspective image, as described below with respect to FIG.
[0022] In addition to being used in the context of ADAS, visualization methods can also be useful in other applications, e.g., remote operation of automobiles, robots, and other machines, by providing the machine's controller or driver with a different perspective that is not directly available via sensors installed on the machine.
[0023] The visualization methods may also be useful in generating simulations of autonomous vehicles, for example, different perspectives of an autonomous vehicle showing the position of the vehicle relative to their surroundings.
[0024] FIG. 1A shows a simplified functional block diagram of a device 100 for generating a third-person perspective image according to various embodiments. The device 100 may also be referred to herein as a visualization device. The device 100 may be capable of performing the visualization methods described above. The device 100 may include a first projection module 102. The first projection module 102 may be configured to receive a plurality of input images 110. The plurality of input images 110 may capture the surroundings of an object 202 (shown in FIG. 2). The first projection module 102 may be configured to project each input image 110 of the plurality of input images 110 onto a respective virtual surface of a set of virtual surfaces around the object 202 to generate a set of surface images 112. The set of virtual surfaces may define a three-dimensional (3D) space around the object 202. The 3D space may at least partially surround the object 202. The device 100 may further include a second projection module 104 configured to project each surface image 112 of the set of surface images 112 onto a common coordinate frame to generate a transformed dataset 114. The transformed dataset 114 may approximate an image captured by a virtual camera in the common coordinate frame. In one embodiment, the common coordinate frame may be the coordinate frame of a virtual camera capturing the third-person perspective image 120. In another embodiment, the common coordinate frame may represent an overhead view. The device 100 may further include a machine learning model 106 configured to generate an outsider perspective image 120, also referred to interchangeably herein as a third person view image, based on the transformed dataset 114. The third-person perspective image 120 may approximate a ground truth third-person view of the object 202 and its surroundings. The third-person perspective image 120 may approximate the perspective of a person standing outside the object 202 and looking toward the object 202.
[0025] The machine learning model 106 may be trained with training data including images projected and reprojected by the first projection module 104 and the second projection module, as well as the transformed dataset 114. In other words, the method may include structuring multiple reprojections of the input image as input data to the machine learning model 106. By doing so, the machine learning model 106 may be exposed to many different projections of the input image 110, and the machine learning model 106 may cross-correlate these different projections to determine the depth of objects captured in the input image 110.
[0026] In prior art solutions, a standard depth estimation neural network can detect objects in an image, determine the size of the object in the image, infer the actual size of the object by recognizing the object's type, and estimate the object's depth based on the object's size in the image and the inferred actual size. The prior art solution can then use the estimated depth information to perform a geometric transformation of the image to a new perspective, such as a third-person perspective image. However, errors in depth estimation can cause distortions in the rendered image.
[0027] However, the device 100 can perform multiple projections of the input image 110 to generate “interim” portions of the third-person perspective image 120. If the “virtual surface” matches the actual depth map of the scene, the final projection can be accurate. The machine learning model 106 may differ from standard depth estimation neural networks of the prior art in that it does not estimate depth, but rather estimates which of the interim projections may be most accurate for the desired viewpoint. In general, the task of estimating which interim projection matches the desired viewpoint may be easier for the machine learning model 106 and therefore less computationally intensive while achieving better accuracy compared to estimating the depths of various objects in the input image 110. The machine learning model 106 may stitch the interim projections that match the desired viewpoint to form the third-person perspective image 120. The machine learning model 106 can also make some adjustments, for example, to fill in missing pixels or add a synthetic image of an object to the third-person perspective image 120.
[0028] 1B illustrates a simplified hardware block diagram of a device 100 according to various embodiments. The device 100 may include at least one processor 130. The device 100 may further include a plurality of sensors 204. Each sensor 204 of the plurality of sensors 204 may be configured to capture a respective input image 110 of the plurality of input images 110. The at least one processor 130 may be configured to perform the functions of at least one of the first projection module 102, the second projection module 104, and the machine learning model 106. The device 100 may further include at least one memory 134. The at least one memory 134 may include a non-transitory computer-readable medium. The at least one processor 130, the plurality of sensors 204, and the at least one memory 134 may be coupled to each other, for example, mechanically or electrically, via coupling wires 140.
[0029] In one example, device 100 may be suitable for use in conjunction with at least one vehicle and / or at least one vehicle-related application. In another example, device 100 may be suitable for use in conjunction with a vehicle computer (e.g., a vehicle dashboard computer) that may be configured to render one or more alternative viewpoints, for example, based on data communicated from a surround-view camera, for purposes of providing driver assistance. In yet another example, device 100 may be suitable for use in conjunction with a vehicle remote control console that may be configured to render one or more (suitable) viewpoints, for example, to enable remote operation based on data from a surround-view camera (e.g., communicable over a network, such as a communications network). In yet another additional example, device 100 may be suitable for use in conjunction with an autonomous vehicle simulation system that may be configured to render one or more current virtual viewpoints of a simulated vehicle in a scene, for example, based on multiple camera images associated with the actual scene, according to an embodiment of the present disclosure. Other examples (e.g., that may not be related to vehicles and / or vehicle-related applications) such as computer games, virtual 3D (three-dimensional) touring (e.g., of one or more places of interest), and / or providing 3D virtual experiences (e.g., via a browser, via a desktop application, and / or via a mobile device application, etc.) may possibly be useful according to embodiment(s) of the present disclosure.
[0030] According to an embodiment that can be combined with any of the above-described embodiments or further embodiments described below, object 202 may be a vehicle. Accordingly, device 100 may be configured to generate a third-person perspective image that provides a view of the vehicle relative to its environment. The third-person perspective image may provide a driver or operator (in the case of a remotely operated vehicle) with enhanced situational awareness for maneuvering the vehicle.
[0031] According to an embodiment that may be combined with any of the above-described embodiments or further embodiments described below, device 100 may include object 202. For example, device 100 may be a driving system that includes a vehicle and a visualization module that generates a third-person perspective image to facilitate control of the vehicle.
[0032] According to an embodiment that can be combined with any of the above-described embodiments or any further embodiments described below, device 100 may include software that generates third-person perspective images. Device 100 may include a server that runs the software. The server may be external to object 202, for example, a cloud-based computing server that renders images offline. Rendering images offline may enable the use of more powerful processors that can be shared across multiple devices 100 or vehicles, without the need to install a dedicated, expensive processor in each device 100.
[0033] According to an embodiment that can be combined with any of the above-described embodiments or any further embodiments described below, a vehicle may include a dashboard. The dashboard may be configured to display a third-person perspective image to provide driver assistance for maneuvering in tight spaces. Optionally, the dashboard may include a touchscreen display. Displaying the third-person perspective image on the dashboard may allow a driver to view the vehicle's position without taking their eyes off the road. The touchscreen display may allow the driver to quickly select their preferred perspective.
[0034] According to an embodiment that may be combined with any of the above-described embodiments or any further embodiments described below, the processor 130 of the device 100 may be configured to perform a method 300 (shown in FIG. 3).
[0035] FIG. 2A shows a top view of an object 202 equipped with a surround view (SV) system according to various embodiments. The object 202 may be a vehicle. The object 202 may be coupled to or include the device 100. The SV system may include multiple sensors 204. The sensors 204 may include at least one of a camera, a LiDAR, and a radar. The multiple sensors 204 may be positioned at corresponding multiple different locations on the object 202 such that each sensor 204 faces a different direction than the other sensors 204. Each sensor 204 may have a respective field of view (FOV) 206. The multiple sensors 204 may include various sensors, for example, a mixture of cameras and radars. The multiple sensors 204 may also have various FOVs 206. For example, the sensors 204 on the left and right sides of the object 202 may have smaller FOVs 206 compared to the sensors 204 at the front and rear ends of the object 202. The FOVs 206 of the multiple sensors 204 may overlap such that the combined FOV 206 of the multiple sensors 204 covers an area surrounding the object 202. The multiple sensors 204 may also be referred to herein as surround view (SV) sensors because they collectively capture information about the surroundings of the object 202. The images captured by the sensors 204 may be referred to herein as SV images and may form the input image 110.
[0036] According to various embodiments, the object 200 may be a vehicle. In the context of automotive applications, those skilled in the art will understand that SV can refer to automotive technology that provides the vehicle driver with a 360-degree view of the area surrounding the vehicle. An SV system may typically include four to six fisheye cameras mounted on the vehicle. Multiple sensors 204 may be part of the SV system, such as the fisheye cameras of the SV system. The device 100 may be part of the SV system and may be configured to generate a third-person perspective image 120 based on image data captured by the sensors 204.
[0037] 2B shows a 3D representation of a set of virtual surfaces 240 around an object 202, according to various embodiments. The set of virtual surfaces 240 can at least partially define a 3D region 230 around the object 202. For example, the 3D region 230 can be a frusto-conical space around the object 202. A circumferential or peripheral surface of the 3D region 230 can include the set of virtual surfaces 240. The surface of the 3D region 230 can be divided into subregions, referred to herein as virtual surfaces 240. Each virtual surface 240 can be represented as a mesh between mesh lines 242. Each mesh can be represented by mesh parameters, including vertex positions.
[0038] According to various embodiments, device 100 may include a virtual surface generation neural network. A process for training the virtual surface generation neural network may include specifying each intermediate surface as a mesh, where the parameters of the mesh are vertex positions (i.e., the edge connections between vertices are fixed). The training process may further include performing projections and obtaining a loss function. The training process may include backpropagating the projections onto mesh coordinates. Information regarding the training process may be found in "Advances in Neural Rendering" by Tewari et al. The virtual surface generation neural network may be configured to generate a set of virtual surfaces. Advantageously, using the virtual surface generation neural network to generate a set of virtual surfaces, instead of manually specifying the virtual surfaces, may optimize surface selection, especially when there are many intermediate surfaces, thereby resulting in a larger number of provisional projections that can become part of the final output image.
[0039] FIG. 3 shows a flow diagram of a method 300 for generating a third-person perspective image 120 according to various embodiments. The method 300 may include processes 302, 304, 306, and 308. The process 302 may include receiving a plurality of input images 110 capturing the surroundings of the object 202. The plurality of input images 110 may be captured by a plurality of sensors 204 coupled to the object 200, for example, as shown in FIG. 2. The process 304 may include performing a first projection. The first projection may include projecting each input image 110 of the plurality of input images 110 onto a respective virtual surface of a set of virtual surfaces around the object 200. The object 200 may be represented in a virtual 3D space, for example, a computer model, and the set of virtual surfaces may be represented in the virtual 3D space as at least partially surrounding the object 200. The first projection may be performed by the first projection module 102. As a result of the first projection, a set of surface images 112 may be generated. The process 306 may include a second projection. The second projection may further include projecting each surface image 112 of the set of surface images 112 onto a common coordinate frame to generate a transformed dataset 114. The second projection may be performed by the second projection module 104. The process 308 may include generating an output third-person perspective image 120 by providing the transformed dataset 114 to the machine learning model 106.
[0040] Advantageously, the method 300 can generate the third-person perspective image 120 without relying on depth estimates. Depth estimates obtained from camera images can be inaccurate due to inaccuracies in the camera mounting position or camera calibration. The first and second projections in the method 300 can transform the input image 110 into various different new perspectives to generate interim images. The machine learning model 106 can generate the third-person perspective image 120 by determining which of these interim images is most accurate.
[0041] The method 300 may include converting the input image 110 into a form suitable for input to the machine learning model 106, for example, through a first projection process and a second projection process. The method 300 may include defining a set of surfaces in 3D space around the object 202. One example of the set of surfaces may include a set of inclined planes at different elevation angles, each of which may be constrained to a finite spatial region in a top view. In another example, the set of surfaces may include a set of curved surfaces that collectively form a "bowl" shape surrounding the object 202. If depth information around the object 202 is known, the set of surfaces may be adaptively calculated based on the depth information using piecewise smooth fitting techniques such as at least one of expectation maximization, meanshift, random sample consensus (RANSAC), spectral clustering, Dirichlet processes, and robustized nonlinear least squares. In the first projection process, the input image may be projected onto the set of surfaces using standard projection-based operations. In an alternative embodiment, instead of projecting raw image values, e.g., red-green-blue (RGB), onto a set of surfaces, intermediate feature vectors extracted by the machine learning model 106 can be projected onto a set of surfaces. In a second projection process, the data from each surface may be projected onto a common coordinate frame, which may then be used as input to the machine learning model 106.
[0042] According to various embodiments, at least one of the first projection module 102 and the second projection module 104 may include a viewpoint-rendering neural network. The viewpoint-rendering neural network can be configured to render new viewpoints based on available image data using viewpoint-rendering methods such as image-to-image transformation, neural point-based rendering, online hypernetwork-based neural radiance field (NeRF), online image-based NeRF, and geometry-free autoregressive modeling. These viewpoint-rendering methods are commonly known to those skilled in the art and are briefly described in the following paragraphs.
[0043] An image-to-image transformation method can include converting an image to a point cloud and projecting each pixel from equirectangular coordinates to Cartesian coordinates. Multiple closest views can be selected for a desired target view. For each selected view, the point cloud can be transformed with a rigid transformation and projected onto the equirectangular image. Another example of an image-to-image transformation method can include constructing a 3D scene mesh by wrapping a lattice grid, also referred to herein as a mesh sheet, onto the scene geometry, and then generating new views by moving a virtual camera in 3D space.
[0044] Neural point-based rendering methods can be similar to image-to-image transformation methods, except that internal neural network features can be projected instead of raw RGB pixel values. Also, for each pixel in the generated image, the viewpoint rendering neural network can collate a "buffer" of projected features corresponding to different distances to the camera origin.
[0045] The online hypernetwork-based NeRF method can include performing online generation of an implicit model of the scene, which can be a mapping from image coordinates and incidence angles to color and transparency information.
[0046] Online image-based NeRF may differ from online hypervisor-based NeRF in that, instead of using an implicit model, the viewpoint-rendering neural network may perform lookup operations on feature maps of the source image during the ray tracing operations that generate the output image.
[0047] Geometry-free autoregressive modeling can involve using a large-capacity neural network to learn to directly associate image patches and transformation parameters, requiring zero manual geometric transformations as part of the neural network pipeline. Because geometry-free autoregressive modeling renders output images in patches, it can handle large viewpoint changes and ensure that subsequent output image patches are consistent with the previous image patches.
[0048] Neural viewpoint rendering techniques can use, for example, image transformation, point-based rendering, or autoregressive models. When using these neural viewpoint rendering techniques, the common coordinate frame can be the coordinate frame of the output image, i.e., the desired third-person viewpoint. This provides the advantage that no further processing is required to transform the common coordinate frame into the coordinate frame of the desired viewpoint.
[0049] The neural viewpoint rendering approach can use, for example, the online hypernetwork NeRF. When using the online hypernetwork NeRF, the common coordinate frame can be a bird's-eye view image.
[0050] According to an embodiment that can be combined with any of the above-described embodiments or further embodiments described below, the first and second projection steps can include re-rendering the image using a rendering neural network, such as at least one of learned denoising / filling, online neural radiance fields, neural point-based graphics, and autoregressive models. Re-rendering the image using the rendering neural network can create approximations of many different viewpoints that can be used as input to the machine learning model 106. These many different viewpoints can simplify the problem that the machine learning model 106 must solve; i.e., the machine learning model 106 can be trained to select the viewpoint that best matches the third-person viewpoint, instead of determining the depth details of objects in the input image 110 and then reconstructing a third-person viewpoint based on the depth details.
[0051] According to an embodiment that can be combined with any of the above-described embodiments or any further embodiments described below, the device 300 may further include performing a fill operation on the output image, i.e., the third-person perspective image 120. The device 100 may further include a fill neural network trained to perform the hierarchical fill operation. The method 300 may include downsampling the third-person perspective image 120 to a lower resolution. The method 300 may further include identifying a region of interest in the downsampled image where empty pixels are present due to occlusion. The method 300 may further include applying the fill neural network to the region of interest to generate image data for the empty pixels. The generated image data may be added to the downsampled image to obtain an intermediate filled image. The fill neural network may then be applied to the entire intermediate filled image to remove blur or add detail to obtain a processed filled image. The processed filled image may then be upsampled to a higher resolution to provide an improved third-person perspective image 120 without losing pixel data. By performing the fill operation, the third-person perspective image 120 may be completed, resulting in no empty pixels in the image.
[0052] According to an embodiment that can be combined with any of the above-described embodiments or any further embodiments described below, the filling process can include processing multiple planes of input data based on a vector quantization (VQ) layer in a VQ-VAE. The machine learning model 106 can include a vector quantization (VQ) layer specialized for processing data across multiple depth scales. Processing the multiple planes of input data based on a VQ layer can include building a codebook of feature vectors and determining the closest element of the codebook for each feature vector in each input layer feature map. The filling process can further include calculating weights for each of the input layers, for example, using a neural network. The filling process can further include quantizing the feature vectors, calculating a weighted average of the quantized vectors, and repeating the quantization process on the averaged vector. Advantageously, the VQ layer can stabilize image rendering with a smaller neural network. Information regarding an example of a filling neural network can be found in "Generating Diverse Structure for Image Inpainting with Hierarchical VQ-QAE" by Peng et al.
[0053] According to an embodiment that can be combined with any of the above-described embodiments or any further embodiments described below, the method 300 may further include generating a synthetic image of the object 202 based on a 3D virtual model of the object 202, and may further include adding the synthetic image of the object 202 to the third-person perspective image 120. The machine learning model 106 may be trained with training images that exclude the object 202, such that the machine learning model 106 can operate on any input image acquired relating to any type of object; for example, in the context of an automotive application, the machine learning model 106 can generate the third-person perspective image 120 without being unduly influenced by the type of vehicle. The synthetic image of the object 202 may be added to the third-person perspective image 120 to approximate a ground truth view in which the object 202 is present.
[0054] According to an embodiment that can be combined with any of the above-described embodiments or any of the further embodiments described below, the machine learning model 106 may include a fully convolutional neural network 400. The fully convolutional neural network 400 may include an encoder network and a decoder network. The encoder network may include multiple convolutional layers. When an image is input to the convolutional layer, a convolution operation may pass through the entire image with a kernel. The output of the convolution operation may include a value of the response to one or other kernel at each point in the image. The output may pass through additional convolutional layers. After the image has been processed by the required number of convolutional layers, the image may be provided to a pooling layer. The pooling layer may reduce the size of the input image. Similar to the convolutional layer, the pooling layer may include a moving kernel and calculate a unique value for each image region. Reducing the size of the image may speed up data processing in the neural network 400. When the image size is reduced, a convolutional layer of the same size may be able to capture most of the desired features in the image. The sequence of convolutional layers followed by pooling layers may be repeated multiple times until a minimum image size is achieved, which can be determined experimentally.
[0055] The decoder network may include upsampling layers and convolutional layers. The features highlighted by the encoder network may be expanded using upsampling layers, returning them to their initial size. The decoder network and the encoder network may be symmetrical. The decoder network may include convolutional layers between the upsampling layers, but the number of outputs from these convolutional layers may decrease as the image progresses through the decoder network. A repeated sequence of upsampling layers followed by convolutional layers may return the image to its initial size while reducing the number of possible image interpretations to the number of desired features.
[0056] One embodiment of neural network 400 is further described with respect to FIG.
[0057] According to an embodiment that can be combined with any of the above-described embodiments or any further embodiments described below, training of neural network 400 can be based, for example, on a standard backpropagation-based gradient descent method. As an example of how neural network 400 can be trained, a training data set can be provided to neural network 400, and the following training process can be performed:
[0058] The training data set can be generated by collecting ground truth third-person perspective images, for example, according to the methods described with respect to Figures 7A and 7B. Alternatively, an example of a suitable training data set for training the correction neural network 508 can be the publicly-based database RealEstate10k.
[0059] The ground truth third-person images can serve as a training signal for neural network 400. The training dataset can further include a plurality of surround-view images captured by a surround-view camera installed in the vehicle. These surround-view images can be processed to turn off masking of the vehicle, e.g., the body of the vehicle.
[0060] Before training neural network 400, the weights may be randomly initialized to numbers between 0.01 and 0.1, while the biases may be randomly initialized to numbers between 0.1 and 0.9.
[0061] The first observation of the dataset may then be loaded into the input layer of the neural network, and the output value(s) are generated by forward propagation of the input values through the input layer. A loss can then be calculated on the output value(s) using the following loss function:
[0062] TIFF2025526975000002.tif7170In the formula, n represents the number of neurons in the output layer, and y represents the actual output value. TIFF2025526975000003.tif5170 represents the predicted output. In other words, TIFF2025526975000004.tif5170 represents the difference between the actual output and the predicted output.
[0063] The weights and biases may then be updated by an AdamOptimizer with a learning rate of 0.001. Other parameters of the AdamOptimizer may be set to default values, for example: beta_1=0.9 beta_2=0.999 eps=1e-08 weight_decay=0
[0064] The above steps may be repeated with the next set of observations until all observations have been used for training. This represents the first training epoch. This may be repeated until 10 epochs have been completed.
[0065] According to an embodiment that can be combined with any of the above-described embodiments or any of the further embodiments described below, the fully convolutional neural network 400 may include a U-Net architecture. In a U-Net, the upsampling portion has multiple feature channels, allowing the network to propagate context information to higher-resolution layers. As a result, the dilation path is more or less symmetrical with respect to the erosion portion, resulting in a U-shaped architecture. The network uses only the significant portion of each convolution without fully connected layers. To predict pixels within the boundary regions of an image, missing context can be estimated by mirroring the input image. This tiling strategy can enable the neural network to be applied to large images, where resolution would otherwise be limited by the memory of the graphics processing unit or the processor of the device 100.
[0066] According to an embodiment that can be combined with any of the above-described embodiments or any further embodiments described below, the machine learning model 106 may include a deformation neural network 502. The deformation neural network 502 may be configured to generate a deformation map 510 and a composite map based on the transformed dataset 114. The machine learning model 106 may include a mapping module 504 configured to map each surface image of the set of surface images based on the deformation map 510 to generate a set of remapped images. The machine learning model 106 may further include a combiner module 506 configured to combine one or more of the remapped images based on the composite map to generate the third-person perspective image 120. The deformation neural network is further described with respect to FIG. 5. Directly generating the deformation map and the composite map instead of the third-person perspective image may provide the advantage of reducing the computational complexity required from the neural network. The deformation neural network 502 is further described with respect to FIG. 5.
[0067] According to an embodiment that can be combined with any of the above-described embodiments or any further embodiments described below, the training of the modified neural network 502 can be based on, for example, a standard backpropagation-based gradient descent method. As an example of how the modified neural network 502 can be trained, a training data set can be provided to the modified neural network 502, and the following training process can be performed.
[0068] An example of a suitable training data set for training the correction neural network 508 may be the publicly-based database RealEstate10k.
[0069] Before training the modified neural network 502, the weights may be randomly initialized to numbers between 0.01 and 0.1, while the biases may be randomly initialized to numbers between 0.1 and 0.9.
[0070] The first observation of the dataset may then be loaded into the input layer of the neural network, and the output value(s) are generated by forward propagation of the input values through the input layer. A loss can then be calculated on the output value(s) using the following loss function:
[0071] TIFF2025526975000005.tif7170In the formula, n represents the number of neurons in the output layer, and y represents the actual output value. TIFF2025526975000006.tif5170 represents the predicted output. In other words, TIFF2025526975000007.tif5170 represents the difference between the actual output and the predicted output.
[0072] The weights and biases may then be updated by an AdamOptimizer with a learning rate of 0.001. Other parameters of the AdamOptimizer may be set to default values, for example: beta_1=0.9 beta_2=0.999 eps=1e-08 weight_decay=0
[0073] The above steps may be repeated with the next set of observations until all observations have been used for training. This represents the first training epoch, and may be repeated until 10 epochs have been completed.
[0074] According to an embodiment that may be combined with any of the above-described embodiments or any further embodiments described below, the machine learning model 106 may further include a correction neural network 508 configured to correct for aberrations in the third-person perspective image 120. The correction neural network 508 may include, for example, a U-Net architecture. The correction neural network 508 may provide the advantage of removing distortions in the generated image.
[0075] According to an embodiment that can be combined with any of the above-described embodiments or any further embodiments described below, the correction neural network 508 may be further configured to perform inpainting on the third-person image 120. By doing so, the third-person image 120 may be completed by filling empty pixels, i.e., blank areas, with approximated content.
[0076] According to an embodiment that can be combined with any of the above-described embodiments or any further embodiments described below, the combiner module can be configured to combine one or more of the remapped images based on weights determined by the compositing map 520. As a result, the correction neural network 508 can receive a more complete image representing the scene as an intermediate step, thereby improving the accuracy of generating the third-person perspective image 120.
[0077] According to an embodiment that can be combined with any of the above-described embodiments or any further embodiments described below, the training of the correction neural network 508 can be based, for example, on a standard backpropagation-based gradient descent method. As an example of how the correction neural network 508 can be trained, a training data set can be provided to the correction neural network 508, and the following training process can be performed.
[0078] An example of a suitable training data set for training the correction neural network 508 may be the publicly-based database RealEstate10k.
[0079] Before training the compensation neural network 508, the weights may be randomly initialized to numbers between 0.01 and 0.1, while the biases may be randomly initialized to numbers between 0.1 and 0.9.
[0080] The first observation of the dataset may then be loaded into the input layer of the neural network, and the output value(s) are generated by forward propagation of the input values through the input layer. A loss can then be calculated on the output value(s) using the following loss function:
[0081] TIFF2025526975000008.tif7170In the formula, n represents the number of neurons in the output layer, and y represents the actual output value. TIFF2025526975000009.tif5170 represents the predicted output. In other words, TIFF2025526975000010.tif5170 represents the difference between the actual output and the predicted output.
[0082] The weights and biases may then be updated by an AdamOptimizer with a learning rate of 0.001. Other parameters of the AdamOptimizer may be set to default values, for example: beta_1=0.9 beta_2=0.999 eps=1e-08 weight_decay=0
[0083] The above steps may be repeated with the next set of observations until all observations have been used for training. This represents the first training epoch, and may be repeated until 10 epochs have been completed.
[0084] FIG. 4 illustrates a block diagram of an example neural network 400 according to various embodiments. The neural network 400 may include a modified architecture based on a fully convolutional network. The neural network 400 may include a U-shaped encoder-decoder network architecture. The neural network 400 may include multiple encoder blocks 422 forming an encoder network, also referred to as a contraction path. The neural network 400 may further include multiple decoder blocks 424 forming a decoder network, also referred to as a dilation path. The encoder network 402 may halve the spatial dimensions of its input while doubling the number of filters, or feature channels, in each encoder block 422. The decoder network 404 may double the spatial dimensions and halve the number of feature channels in each decoder block 424.
[0085] The encoder block 402 may be configured to extract features from input data and may further be configured to learn an abstract representation of the input data through a sequence of encoder blocks 422. Each encoder block 422 may include two 3x3 convolutional layers 430, each of which may be followed by a rectified linear unit (ReLU) activation function 432. The ReLU activation function 432 may introduce nonlinearity to better generalize the training data. The output of the ReLU activation function 432 may serve as a skip connection 406 for the corresponding decoder block 424. The skip connection 406 may provide additional information to help the decoder block 424 generate improved semantic features. The skip connection 406 may also facilitate the indirect flow of gradients to previous layers without degradation. In other words, the skip connection 406 may facilitate better flow of gradients in backpropagation, which helps the neural network 400 learn representations.
[0086] The neural network 400 may further include 2x2 max pooling 408, which reduces the spatial dimensions (height and width) of the feature maps generated by the preceding encoder block 422 by half. This reduces the computational cost by reducing the number of trainable parameters. The neural network 400 may further include a bridge 410, which connects the encoder network 402 to the decoder network 404. The bridge 410 may include two 3x3 convolutional layers 430, each followed by a ReLU activation function 432 similar to the encoder block 422.
[0087] The decoder network 404 may be configured to generate a semantic segmentation mask based on the abstract representation output by the encoder network 402. The decoder network 404 may include a 2x2 transposed convolutional layer 440 between every two consecutive decoder blocks 424. Each decoder block 424 may receive a feature map from the corresponding encoder block 422 via a skip connection 406. Each decoder block 424 may include two 3x3 convolutional layers 430. Each convolutional layer 430 may be followed by a ReLU activation function 432. The output of the last decoder block 424 may be passed through a 1x1 convolutional layer 440 with a sigmoid activation. The sigmoid activation function may provide a segmentation mask representing a pixel-by-pixel classification.
[0088] The number of input channels to neural network 400 may be three times the number of virtual surfaces 240. The number of output channels may be three.
[0089] The loss function used for neural network 400 may be mean squared error (MSE). As an example, neural network 400 may be trained using the Adam optimizer with a learning rate of 0.0001 and other parameters set to default, and trained for 10 epochs.
[0090] According to an embodiment that can be combined with any of the above-described embodiments or any further embodiments described below, neural network 400 may include a U-Net architecture. Information regarding the U-Net architecture can be found in "U-Net: Convolutional Networks for Biomedical Image Segmentation" by Ronneberger et al.
[0091] According to an embodiment that can be combined with any of the above-described embodiments or any further embodiments described below, the neural network 400 can be trained using an input training dataset that includes a set of re-rendered images or feature maps at desired poses in a common coordinate frame, a coordinate frame centered on the object 202. Ground truth third-person images taken around the object 202 can be provided as training signals to the neural network 400. Generating the re-rendered images or feature maps may include re-projecting input images captured by the SV camera using the first projection module 102 and the second projection module 104. The neural network 400 may be trained to correlate the ground truth third-person images with the re-rendered images or feature maps in the common coordinate frame.
[0092] According to an embodiment that can be combined with any of the above-described embodiments or further embodiments described below, a training method for training neural network 400 can include masking off pixels in a ground truth third-person image that correspond to a data collection vehicle and object 202. A data collection vehicle can refer to a tool used to carry a camera for capturing the ground truth third-person image. For example, a data collection vehicle can be movable arm 704 or movable device 750 shown in FIGS. 7A and 7B .
[0093] 5 shows a block diagram of machine learning model 106, according to various embodiments. In these embodiments, machine learning model 106 may include a deformation neural network 502 instead of neural network 400. Deformation neural network 502 may differ from neural network 400 in that, instead of directly generating third-person perspective images 120, deformation neural network 502 may generate deformation map 510 and composite map 520 based on transformed dataset 114. Advantageously, deformation neural network 502 may require lower computational resources than neural network 400.
[0094] The machine learning model 106 may further include a mapping module 504 and a combiner module 506. The deformation neural network 502 may generate a deformation map 510 of the input image 110 by correlating different projections of the input image 110 in the transformed dataset 114. The deformation neural network 502 may be trained using the same training input data as the neural network 400, but uses ground truth 3D information (instead of ground truth third-person images) as the training signal. The ground truth 3D information may include synthetic training data, such as that generated by a CARLA simulator. The mapping module 504 is configured to remap the images in the transformed dataset 114 based on the deformation map 510 to generate a set of remapped images. The machine learning model 106 may further include a combiner module 506 configured to combine one or more of the remapped images based on a composite map 520 to generate an initial third-person image 512. The combiner module 506 may combine the remapped images using a weighted average, where the weights may be determined by the composite map 520. In some embodiments, the third-person perspective image 120 may include the initial third-person perspective image 512.
[0095] According to an embodiment that can be combined with any of the above-described embodiments or any further embodiments described below, the deformation neural network 502 may include a deep learning network for deformable image registration (DIRNet). Information about DIRNet can be found in "End-to-end Unsupervised Deformable Image Registration with a Convolutional Neural Network" by Bob D. de Vos et al.
[0096] According to an embodiment that may be combined with any of the above-described embodiments or any further embodiments described below, the machine learning model 106 may further include a correction neural network 508. The correction neural network 508 may receive the initial third-person perspective image 512 and correct aberrations and inpainting on the initial third-person perspective image 512 to generate a final third-person perspective image 514. In some embodiments, the third-person perspective image 120 may include the final third-person perspective image 514.
[0097] FIG. 6 shows a flow diagram of a method 600 for training a machine learning model according to various embodiments. Method 600 may include processes 602, 604, and 606. Process 602 may include projecting each training image of a plurality of training images capturing the periphery of a training object 702 (shown in FIGS. 7A and 7B ) onto a respective virtual surface 240 of a set of virtual surfaces 240 around the training object 702 to generate a set of surface images. Process 604 may further include projecting each surface image of the set of surface images onto a common coordinate frame to generate a training dataset. Process 606 may include training the machine learning model using the training dataset as input to the machine learning model and further using a set of ground truth third-person perspective images corresponding to the plurality of training images as a desired output.
[0098] According to an embodiment that can be combined with any of the above-described embodiments or any further embodiments described below, the method 600 can further include a step of obtaining a training dataset including a plurality of training images and a set of ground truth third-person perspective images.
[0099] 7A and 7B show example equipment setups for collecting ground truth third-person images. The equipment may also be referred to herein as a data collection vehicle. Referring to FIG. 7A, the ground truth third-person images may be collected using a movable arm 704 attached to a training object 702. The training object 702 may be, for example, a car, and the movable arm 704 may be attached to the roof of the car. The movable arm 704 may have a first end 742 and a second end 744 opposite the first end. The first end 742 may be coupled to the training object 702, and the second end 744 may be coupled to an external sensor 706. The movable arm 704 may be rotatable at the first end 742 to move the external sensor 706 around the training object 702 to capture ground truth third-person images from various angles around the training object 702. The movable arm 704 can include multiple segments 722 connected by rotatable joints 724, whereby the position of the external sensor 706 can be adjusted by displacing any segment 722 relative to another segment 722 about the connected rotatable joints 724.
[0100] 7B, ground truth third-person perspective images may be collected using a mobile device 750 equipped with an external sensor 706. The mobile device 750 may be, for example, a robot or an unmanned ground vehicle. The mobile device 750 may be driven around the training object 702 such that the external sensor 706 captures third-person perspective images from various angles around the training object 702.
[0101] According to an embodiment that can be combined with any of the above-described embodiments or any further embodiments described below, method 600 may further include preprocessing the training dataset. Preprocessing the training dataset may include performing internal and / or external calibration of the surround view camera, e.g., sensor 204, and the third-party camera, e.g., external sensor 706. If the pose and temporal alignment of external sensor 706 are unknown, preprocessing may include jointly estimating the temporal offset and trajectory of external sensor 706 by formulating a calibration alignment problem. This may be necessary, for example, if the exact position of movable arm 704 or movable device 750 is unknown.
[0102] Preprocessing in method 600 can include removing segments of images collected by external sensor 706. These segments can include pixels that represent training object 702, movable arm 704, or movable device 750.
[0103] To remove the training objects 702 from the training images collected by the external sensor 706, a specially trained auxiliary neural network can be used to create a binary mask for each image. The binary mask can indicate which pixels in the collected image belong to the training object 702. By removing the training objects 702 from the training images, the machine learning model 106 can be trained for any general object, for example, any vehicle model.
[0104] Additionally, the collected images may also include a faintly visible movable arm 704 or movable device 750. A mask can be created to cover pixels that represent the movable arm 704 or movable device 750. A filler neural network can create reasonable image data in place of the covered pixels.
[0105] Optionally, method 600 may further include applying a depth estimation algorithm (e.g., stereo vision or monocular depth estimation) or additional sensors, such as radar or LiDAR sensors, to generate a 3D model of the scene. Method 600 may further include transforming the surface images into a common coordinate frame based on the 3D model. By doing so, the resulting transformed dataset can be used as a weak ground truth for training machine learning model 106 without using heuristics or neural networks to generate the transformed dataset.
[0106] The method 600 may further include creating an adaptive virtual surface based on a 3D model of the scene.
[0107] The training images may be captured by a sensor 204 attached to the training object 702. The method 600 may include projecting and reprojecting the training images by a first projection module and a second projection module in a process similar to that described with respect to FIG.
[0108] According to an embodiment, which may be combined with any of the above-described embodiments or any further embodiments described below, the machine learning model 106 shown in FIG. 1A may be trained using a method 600.
[0109] According to an embodiment that can be combined with any of the above-described embodiments or any further embodiments described below, the method 600 can further include masking off a training object from the set of ground truth third-person images before using the set of ground truth third-person images as a desired output for training the machine learning model, thereby allowing the machine learning model 106 to be trained without bias to the object type of the object 202, e.g., a vehicle model, thereby allowing the method to be applied to any type of object.
[0110] According to an embodiment that may be combined with any of the above-described embodiments or any further embodiments described below, method 600 may further include defining each virtual surface 240 of the set of virtual surfaces 240 as a mesh before projecting each training image of the plurality of training images onto the respective virtual surface 240. Method 600 may further include backpropagating the set of surface images onto the mesh. Method 600 may further include determining an output of a loss function based on the backpropagation to refine the set of virtual surfaces 240.
[0111] According to various embodiments, a data structure may be provided that is generated by the method 600. The data structure may include a trained machine learning model, such as the machine learning model 106.
[0112] While embodiments of the present invention have been particularly shown and described with reference to specific embodiments, it should be understood by those skilled in the art that various changes in form and detail can be made herein without departing from the spirit and scope of the invention as defined by the appended claims. The scope of the present invention is indicated by the appended claims, and therefore all changes that come within the meaning and range of equivalency of the claims are intended to be embraced. It will be understood that common numerals used in related drawings refer to components serving similar or the same purpose.
[0113] Those skilled in the art will understand that the terminology used herein is for the purpose of describing various embodiments only and is not intended to limit the invention. As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly dictates otherwise. It will be further understood that the terms "comprises" and / or "comprising," as used herein, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0114] It is understood that the specific order or hierarchy of blocks in the disclosed processes / flowcharts is illustrative of example approaches. Based on design preferences, it is understood that the specific order or hierarchy of blocks within a process / flowchart can be rearranged. Also, some blocks may be combined or omitted. The accompanying method claims present elements of the various blocks in a sample order and are not meant to be limited to the specific order or hierarchy presented.
[0115] The foregoing description is provided to enable those skilled in the art to practice the various aspects described herein. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects. Accordingly, the claims are not intended to be limited to the aspects set forth herein but are to be accorded the full scope consistent with the language of the claims, and references to elements in the singular do not mean "only" or "one or more" unless expressly stated otherwise. The term "exemplary" is used herein to mean "serving as an example, instance, or illustration." Any aspect described herein as "exemplary" should not necessarily be construed as preferred or advantageous over other aspects. Unless otherwise stated, the term "some" refers to one or more. Combinations such as "at least one of A, B, or C," "one or more of A, B, or C," "at least one of A, B, and C," "one or more of A, B, and C," and "A, B, C, or any combination thereof" include any combination of A, B, and / or C, and can include multiples of A, multiples of B, or multiples of C. Specifically, combinations such as "at least one of A, B, or C," "one or more of A, B, or C," "at least one of A, B, and C," "one or more of A, B, and C," and "A, B, C, or any combination thereof" can be A only, B only, C only, A and B, A and C, B and C, or A and B and C, and any such combination can include one or more elements of A, B, or C. All structural and functional equivalents to the elements of the various aspects described throughout this disclosure that are known, or that later become known, to those of ordinary skill in the art are expressly incorporated herein by reference and are intended to be encompassed by the claims.
Claims
1. A computer-implemented method (300) for generating a third-person perspective image, comprising: receiving a plurality of input images (110) capturing a surrounding of an object (202); projecting each input image (110) of the plurality of input images (110) onto a respective virtual surface (240) of a set of virtual surfaces (240) around the object (202) to generate a set of surface images (112); projecting each surface image (112) of the set of surface images (112) onto a common coordinate frame to generate a transformed data set (114); generating a third-person perspective image (120) based on the transformed dataset (114) by a machine learning model (106); The method (300).
2. generating a synthetic image of the object (202) based on a three-dimensional virtual model of the object (202); adding the composite image of the object (202) to the third-person perspective image (120); The method (300) of claim 1, further comprising:
3. The method (300) of any one of claims 1 to 2, wherein the machine learning model (106) comprises a fully convolutional neural network.
4. The machine learning model (106) a deformation neural network (502) configured to generate a deformation map (510) and a composite map (520) based on the transformed dataset (114); a mapping module (504) configured to map each surface image (112) of the set of surface images (112) based on the deformation map (510) to generate a set of remapped images; a combiner module (506) configured to combine one or more of the remapped images based on the composite map (520) to generate the third-person perspective image (120); The method (300) of any one of claims 1 to 3, comprising:
5. The method (300) of any one of claims 1 to 4, wherein the machine learning model (106) further comprises a correction neural network (508) configured to correct aberrations in the third-person perspective image (120).
6. The method (300) of any one of claims 1 to 5, wherein the correction neural network (508) is further configured to perform inpainting within the third-person perspective image (120).
7. 7. The method (300) of any one of claims 1 to 6, wherein the combiner module (506) is configured to combine the one or more of the remapped images based on an average of weights determined by the composite map (520).
8. A computer-implemented training method (600) for training a machine learning model (106), comprising: projecting each training image of a plurality of training images capturing a periphery of a training object onto a respective virtual surface of a set of virtual surfaces surrounding the training object to generate a set of surface images; projecting each surface image of the set of surface images onto a common coordinate frame to generate a training data set; training the machine learning model using the training dataset as an input to the machine learning model and further using a set of ground truth third-person images corresponding to the plurality of training images as a training signal; A training method (600) comprising:
9. masking off the training subject from the set of ground truth third-person images before using the set of ground truth third-person images as a desired output for training the machine learning model. The training method (600) of claim 8, further comprising:
10. defining each virtual surface of the set of virtual surfaces as a mesh before projecting each training image of the plurality of training images onto the respective virtual surface; backpropagating the set of surface images onto the mesh; determining an output of a loss function based on the backpropagation to improve the set of virtual surfaces; The training method (600) of any one of claims 8 to 9, further comprising:
12. The method (300) of any one of claims 1 to 7, wherein the machine learning model (106) is trained according to a training method (600) of any one of claims 8 to 10.
11. A data structure generated by the training method (600) of any one of claims 8 to 10.
13. A device (100) for generating a third-person perspective image (120), the device (100) comprising a processor (130) configured to perform the method (300) of any one of claims 1 to 7.
14. Multiple sensors 204 Furthermore, each sensor (204) of the plurality of sensors (204) configured to capture a respective input image (110) of the plurality of input images (110); The device (100) of claim 13.
15. vehicle Furthermore, the object (202) includes the vehicle; A device (100) according to any one of claims 13 to 15.
16. Use of the device according to claim 15 for remotely controlling the vehicle.
Citation Information
Patent Citations
Vehicular display device
JP2012138660A
Video synthesizing apparatus, program and method for synthesizing viewpoint video by projecting object information onto plural surfaces
JP2019046077A
Image complementary system
JP2022051700A
Neural rendering
US20210248811A1