Learning device, image processing device, learning method, image processing method, and computer program
The learning device and method address the challenge of generating color information outside the field of view by training a model with LiDAR data and neural radiance field techniques, enabling accurate object identification and visual annotation.
Patent Information
- Application Number
- JP2024542551
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-08-26
- Publication Date
- 2026-01-27
- Estimated Expiration
- 2042-08-26
AI Technical Summary
Existing methods struggle to generate color information for objects outside the field of view range and accurately identify stationary objects at night due to limitations in passive sensors and active sensors, respectively.
A learning device and method that utilizes three-dimensional coordinate values, gaze direction information, and point cloud data to train a model for generating images with RGB values and transparency, even outside the field of view range, by integrating LiDAR data with neural radiance field techniques.
Enables the generation of arbitrary viewpoint images with assigned RGB values outside the field of view range, enhancing object identification and visual annotation by combining shape and color information effectively.
Smart Images

Figure 0007806910000001 
Figure 0007806910000002 
Figure 0007806910000003
Abstract
Description
[Technical Field]
[0001] The disclosed technology relates to a learning device, an image processing method, a learning method, an image processing method, and a computer program. [Background technology]
[0002] Non-Patent Document 1 proposes "Neural Radiance Field (NeRF)," a volume representation using a Deep Neural Network (DNN) that synthesizes images from new viewpoints based on a set of images. NeRF represents one scene using a single DNN, and optimizes the DNN parameters based on images from multiple viewpoints so that it returns appropriate R (red), G (green), B (blue), and σ (transparency) based on input information on three-dimensional spatial coordinates and two-dimensional viewing direction (polar angle θ, azimuth angle φ). [Prior art documents] [Non-patent literature]
[0003] [Non-Patent Document 1] Mildenhall, B., Srinivasan, P.P., Tancik, M., Barron, J.T., Ramamoorthi, R., Ng, R., “NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis.”, Internet<URL:https: / / arxiv.org / pdf / 2003.08934.pdf> Summary of the Invention [Problem to be solved by the invention]
[0004] Creating a 3D map of a city requires obtaining location information for only stationary objects such as buildings and facilities, excluding moving objects such as pedestrians and cars. To obtain information on stationary objects, data can be collected at night, when there are fewer moving objects and fewer changes in the scene due to changes in the placement of billboards and other objects. However, due to the lack of sunlight at night, it is difficult to obtain color information using passive sensors such as visible light cameras. On the other hand, observations using active sensors such as LiDAR (Light Detection and Ranging) can efficiently obtain object shape information at night, when there are fewer moving objects, but they cannot obtain color information at wavelengths other than laser wavelengths. This can make it difficult to identify objects that are attached to road surfaces or walls, making visual annotation of objects more difficult. For this reason, it is conceivable to support identification by displaying shape information (referred to as point cloud data in this disclosure) visualized with a work tool during annotation and other tasks by assigning RGB based on an RGB image acquired during the day. However, simply assigning R, G, and B by superimposing has problems such as not being able to assign R, G, and B values outside the range of the image, and moving objects reflected in the daytime RGB image being transferred.
[0005] The disclosed technology has been made in consideration of the above points, and aims to provide a learning device, an image processing method, a learning method, an image processing method, and a computer program for generating arbitrary viewpoint images to which RGB is assigned even outside the field of view range. [Means for solving the problem]
[0006] A first aspect of the present disclosure is a learning device that includes an acquisition unit that receives three-dimensional coordinate values, gaze direction information, and point cloud data as input data and acquires images captured from multiple directions as training data, and a learning unit that uses the input data and the training data to learn a model for outputting an image from a specified gaze direction by outputting color and density for each pixel.
[0007] A second aspect of the present disclosure is an image processing device that includes: an estimation unit that inputs three-dimensional coordinate values, gaze direction information, and point cloud data as input data, and images captured from multiple directions as training data, inputs the gaze direction into a trained model for outputting an image from a specified gaze direction by outputting color and density for each pixel, and outputs the color and transparency for each pixel from the gaze direction from the model; and an image processing unit that generates an image from the gaze direction using the color and transparency output by the estimation unit.
[0008] A third aspect of the present disclosure is a learning method in which a processor receives three-dimensional coordinate values, gaze direction information, and point cloud data as input data, acquires images captured from multiple directions as training data, and executes a process of learning a model for outputting an image from a specified gaze direction by using the input data and the training data to output color and density for each pixel.
[0009] A fourth aspect of the present disclosure is an image processing method in which a processor inputs three-dimensional coordinate values, gaze direction information, and point cloud data as input data, and uses images captured from multiple directions as training data, inputs the gaze direction to a trained model for outputting an image from a specified gaze direction by outputting color and density for each pixel, outputs the color and transparency for each pixel from the gaze direction from the model, and executes a process to generate an image from the gaze direction using the color and transparency.
[0010] A fifth aspect of the present disclosure is a computer program that causes a computer to function as the learning device of the first aspect of the present disclosure or the image processing device described in the second aspect of the present disclosure. [Effects of the Invention]
[0011] According to the disclosed technology, it is possible to provide a learning device, an image processing method, a learning method, an image processing method, and a computer program for generating an arbitrary viewpoint image to which RGB is added even outside the field of view range. [Brief explanation of the drawings]
[0012] [Figure 1] FIG. 1 illustrates an example of an image processing system according to an embodiment. [Figure 2] FIG. 2 is a block diagram showing the hardware configuration of the learning device. [Figure 3] FIG. 2 is a block diagram illustrating an example of a functional configuration of a learning device. [Figure 4] FIG. 2 is a block diagram showing a hardware configuration of the image processing apparatus. [Figure 5] FIG. 2 is a block diagram illustrating an example of a functional configuration of the image processing apparatus. [Figure 6] FIG. 1 is a diagram illustrating an overview of a learning process in NeRF. [Figure 7] FIG. 2 is a diagram illustrating an overview of a learning process in a learning device. [Figure 8] FIG. 2 is a diagram illustrating an overview of a learning process in a learning device. [Figure 9] FIG. 2 is a diagram illustrating an overview of a learning process in a learning device. [Figure 10] 10 is a flowchart showing the flow of a learning process performed by the learning device. [Figure 11] 10 is a flowchart showing the flow of image processing by the image processing device. DETAILED DESCRIPTION OF THE INVENTION
[0013] An example of an embodiment of the disclosed technology will be described below with reference to the drawings. Note that the same or equivalent components and parts in each drawing are given the same reference numerals. Also, the dimensional proportions in the drawings are exaggerated for the convenience of explanation and may differ from the actual proportions.
[0014] 1 is a diagram showing an example of an image processing system according to this embodiment. The image processing system according to this embodiment includes a learning device 10 and an image processing device 20.
[0015] The learning device 10 is a device that performs learning processing on a model using images captured from multiple directions, point cloud data, and viewpoint information, and generates a trained model 1 that outputs information for generating an image from any viewpoint.
[0016] When training the trained model 1, the learning device 10 uses the three-dimensional space coordinates on the line of sight of each pixel in an image from a certain viewpoint, information on the line of sight direction, and point cloud data as input data, and uses an image captured from that viewpoint as training data, and trains the trained model 1 so that it outputs appropriate R (red), G (green), B (blue) values and σ (transparency) as output data to reduce errors with the training data. A specific example of the training process by the learning device 10 will be described in detail later. Furthermore, the input three-dimensional space coordinates, line of sight direction information, and coordinate system of the point cloud data are assumed to be the same. The point cloud data can be acquired using an active sensor such as LiDAR, for example.
[0017] The image processing device 20 is a device that inputs information on the field of view angle from the viewpoint from which an image is to be generated into the trained model 1, and generates an image from that viewpoint using the R, G, B values and σ (transparency) for each pixel output from the trained model 1.
[0018] By using not only coordinates in three-dimensional space and information on two-dimensional viewing angles from a certain viewpoint, but also point cloud data, the learning device 10 can perform a learning process for expressing three-dimensional information using a DNN, supplemented by three-dimensional shape information from the point cloud. By performing such a learning process, the learning device 10 can generate a trained model 1 for generating images from any viewpoint with R, G, and B assigned, even outside the range of the viewing angle.
[0019] In addition, the image processing device 20 can generate images from any viewpoint with R, G, and B assigned, even outside the range of the field of view, by inputting field of view angle information into the trained model 1 trained by the learning device 10.
[0020] 1, the learning device 10 and the image processing device 20 are separate devices, but the present disclosure is not limited to this example, and the learning device 10 and the image processing device 20 may be the same device. Furthermore, the learning device 10 may be composed of multiple devices.
[0021] Next, the configuration of the learning device 10 will be described.
[0022] FIG. 2 is a block diagram showing the hardware configuration of the learning device 10. As shown in FIG.
[0023] 2, the learning device 10 includes a CPU (Central Processing Unit) 11, a ROM (Read Only Memory) 12, a RAM (Random Access Memory) 13, a storage 14, an input unit 15, a display unit 16, and a communication interface (I / F) 17. Each component is connected to each other via a bus 19 so that they can communicate with each other.
[0024] The CPU 11 is a central processing unit that executes various programs and controls each component. That is, the CPU 11 reads a program from the ROM 12 or the storage 14 and executes the program using the RAM 13 as a work area. The CPU 11 controls the above components and performs various arithmetic processing in accordance with the program stored in the ROM 12 or the storage 14. In this embodiment, the ROM 12 or the storage 14 stores a learning processing program for generating a trained model 1 that executes a learning process and outputs information for generating an image from an arbitrary viewpoint.
[0025] The ROM 12 stores various programs and various data. The RAM 13 temporarily stores programs or data as a working area. The storage 14 is configured with a storage device such as an HDD (Hard Disk Drive) or an SSD (Solid State Drive) and stores various programs including the operating system and various data.
[0026] The input unit 15 includes a pointing device such as a mouse and a keyboard, and is used to perform various inputs.
[0027] The display unit 16 is, for example, a liquid crystal display, and displays various information. The display unit 16 may also function as the input unit 15 by adopting a touch panel system.
[0028] The communication interface 17 is an interface for communicating with other devices. For this communication, for example, a wired communication standard such as Ethernet (registered trademark) or FDDI, or a wireless communication standard such as 4G, 5G, or Wi-Fi (registered trademark) is used.
[0029] Next, the functional configuration of the learning device 10 will be described.
[0030] FIG. 3 is a block diagram showing an example of the functional configuration of the learning device 10.
[0031] 3, the learning device 10 has, as its functional components, an acquisition unit 101 and a learning unit 102. Each functional component is realized by the CPU 11 reading out a language processing program stored in the ROM 12 or storage 14, expanding it in the RAM 13, and executing it.
[0032] The acquisition unit 101 acquires data to be used for learning processing. In this embodiment, the acquisition unit 101 receives, as input data, three-dimensional spatial coordinates in the line of sight of each pixel in an image from a certain viewpoint, information on two-dimensional viewing angles, and point cloud data, and acquires, as training data, an image captured from the viewpoint.
[0033] The learning unit 102 uses the three-dimensional spatial coordinates and field of view angle information in the line of sight direction of each pixel in an image from a certain viewpoint acquired by the acquisition unit 101, and point cloud data as input data, and uses the image captured from that viewpoint as training data, and trains the trained model 1 to output appropriate R (red), G (green), B (blue) values and σ (transparency) as output data to minimize errors with the training data.
[0034] Next, the configuration of the image processing device 20 will be described.
[0035] FIG. 4 is a block diagram showing the hardware configuration of the image processing device 20. As shown in FIG.
[0036] 4, the image processing device 20 includes a CPU 21, a ROM 22, a RAM 23, a storage 24, an input unit 25, a display unit 26, and a communication interface (I / F) 27. Each component is connected to each other via a bus 29 so as to be able to communicate with each other.
[0037] The CPU 21 is a central processing unit that executes various programs and controls each component. That is, the CPU 21 reads a program from the ROM 22 or the storage 24 and executes the program using the RAM 23 as a work area. The CPU 21 controls each of the above components and performs various arithmetic processing in accordance with the program stored in the ROM 22 or the storage 24. In this embodiment, the ROM 12 or the storage 14 stores an image processing program for inputting information on the viewing angle of a certain viewpoint to the trained model 1 and generating an image from that viewpoint using information output by the trained model 1.
[0038] The ROM 22 stores various programs and various data. The RAM 23 serves as a working area and temporarily stores programs or data. The storage 24 is configured with a storage device such as an HDD or SSD, and stores various programs including the operating system, and various data.
[0039] The input unit 25 includes a pointing device such as a mouse and a keyboard, and is used to perform various inputs.
[0040] The display unit 26 is, for example, a liquid crystal display, and displays various information. The display unit 26 may function as the input unit 25 by adopting a touch panel system.
[0041] The communication interface 27 is an interface for communicating with other devices. For this communication, for example, a wired communication standard such as Ethernet (registered trademark) or FDDI, or a wireless communication standard such as 4G, 5G, or Wi-Fi (registered trademark) is used.
[0042] Next, the functional configuration of the image processing device 20 will be described.
[0043] FIG. 5 is a block diagram showing an example of the functional configuration of the image processing device 20. As shown in FIG.
[0044] 5, the image processing device 20 has, as functional components, an acquisition unit 201, an estimation unit 202, and an image generation unit 203. Each functional component is realized by the CPU 21 reading out an image processing program stored in the ROM 22 or the storage 24, expanding it in the RAM 23, and executing it.
[0045] The acquisition unit 201 acquires information on the line of sight direction of the viewpoint to be generated. The line of sight direction information, or the viewing angle information, is input by the user via a predetermined user interface that the image processing device 20 displays on the display unit 26, for example.
[0046] The estimation unit 202 inputs the gaze direction information acquired by the acquisition unit 201 into the trained model 1, and estimates the image from the gaze direction by outputting the color and transparency of each pixel from the gaze direction from the trained model 1.
[0047] The image generating unit 203 generates and outputs an image from the viewpoint based on the estimation result by the estimating unit 202 of the image from the viewpoint of the viewing angle acquired by the acquiring unit 201.
[0048] The image processing device 20 has such a configuration, and can generate an arbitrary viewpoint image to which RGB is assigned even outside the range of the angle of view by using the trained model 1.
[0049] Next, the operation of the learning device 10 will be described.
[0050] First, an overview of the learning process in NeRF will be described. Fig. 6 is a diagram illustrating an overview of the learning process in NeRF.
[0051] NeRF assumes an image at an arbitrary viewpoint and samples the spatial coordinate x on the line of sight corresponding to each pixel. During learning, the image at the arbitrary viewpoint is assumed to be the viewpoint of the correct image. NeRF also creates two patterns of learning: coarse sampling and fine sampling.
[0052] When the NeRF model receives input of spatial coordinates x(x,y,z) and gaze direction d(θ,φ), it outputs the R, G, and B values RGB(x) at said spatial coordinate x and the density value σ(x) at said spatial coordinate x. The model is configured as shown in Figure 6. For gaze direction d(θ,φ), parameters of the ground truth image are used during learning. The spatial coordinates x(x,y,z) in the gaze direction corresponding to each pixel are not included in the ground truth image acquired by a camera rather than by rendering, so are generated by sampling.
[0053] The spatial coordinate x is input to function γ, and then input to a five-layer neural network with 60, 256, 256, 256, and 256 nodes. The feature F after passing through the five-layer neural network is further combined with the spatial coordinate x input to function γ and input to a four-layer neural network with 256, 256, 256, and 256 nodes. The value after passing through the four-layer neural network is output as the density value σ(x). Furthermore, the value after passing through the four-layer neural network is combined with the gaze direction d input to function γ to become feature F', which is then input to the neural network. The value after passing through this neural network is output as RGB(x).
[0054] Once the NeRF model outputs RGB(x) and σ(x) for all pixels, an image at an arbitrary viewpoint is generated by volume rendering. The NeRF model is then trained to minimize the error between the image generated by the NeRF model and the ground truth image for that viewpoint.
[0055] In the NeRF model, when an image captured at night is used as the ground truth image, there is a problem in that R, G, and B values cannot be assigned outside the range of the image. Therefore, the learning device 10 according to this embodiment trains the trained model 1 using point cloud data in addition to the spatial coordinate x and the gaze direction d.
[0056] FIG. 7 is a diagram illustrating an overview of the learning process in the learning device 10. The learning process shown in FIG. 7 is configured to emphasize the use of point clouds to assist in learning three-dimensional shapes, and assigns R, G, and B based on the position in the scene. This configuration is effective, for example, in scenes where the color changes depending on the position (such as an indoor room where the floor, ceiling, and walls are all the same color). The framework for training the deep neural network is similar to the NeRF model training described in FIG. 6 in that the deep neural network is trained based on the generated image, which is the result of volume rendering, and the ground truth image, and two patterns of coarse sampling and fine sampling are created during training. However, a point cloud of an area corresponding to the ground truth image is added as an input to the deep neural network. In this case, the coordinate systems for the point cloud and the camera position coordinates are assumed to be the same. For example, if the point cloud is expressed in a Cartesian coordinate system and the camera position coordinates are expressed in a geographic coordinate system (latitude and longitude), they are aligned in advance to the same coordinate system using a corresponding coordinate system conversion method. Since point cloud processing and NeRF algorithms often use a Cartesian coordinate system, it is easier to implement a program using a Cartesian coordinate system rather than a geographic coordinate system.
[0057] The spatial coordinate x is input to a function γ, and then input to a four-layer third neural network 303 with 60, 256, 256, and 256 nodes. Furthermore, point cloud data consisting of point clouds and intensities is input to a model that captures the overall scene characteristics, such as PointNet. The output of this model is combined with the output from the four-layer neural network to generate the feature F.
[0058] The feature F is input to a predetermined first neural network. The value after passing through the first neural network 301 is output as a density value σ(x). The feature F is also combined with the gaze direction d input to the function γ to become a feature F', which is input to the second neural network 302. The value after passing through this second neural network 302 is output as RGB(x).
[0059] FIG. 8 is a diagram illustrating an overview of the learning process in the learning device 10. The learning process shown in FIG. 8 is configured to emphasize color estimation based on local shape information and brightness information from a point cloud, and assigns R, G, and B using the local shape as a clue. This configuration is effective for scenes where the color changes in response to the local shape (such as outdoor scenes with a mixture of trees and utility poles). The fact that two patterns of coarse sampling and fine sampling are created during learning and learning is the same as in the learning of the NeRF model described in FIG. 6.
[0060] Point cloud data consisting of point clouds and brightness is input to a model such as PointNet++ or KPConv that captures the peripheral features of each point. In addition, a point with spatial coordinate x is set as the center point and neighboring points are input to the model that captures the peripheral features. From the input to the model, local features are extracted and R, G, and B are assigned based on the local features. The output of the model is the feature F.
[0061] The feature F is input to a predetermined first neural network 301. The value after passing through the first neural network 301 is output as a density value σ(x). The feature F is also combined with the gaze direction d input to the function γ to become a feature F′, which is input to a predetermined second neural network 302. The value after passing through this second neural network 302 is output as RGB(x).
[0062] The learning device 10 trains the trained model 1 so as to reduce the error between an image from an arbitrary viewpoint generated from RGB(x) and σ(x) output by the trained model 1 and the correct image. Here, when training the trained model 1, the learning device 10 calculates the error only at coordinates that overlap with the correct image. Areas that do not overlap with the correct image are colored in accordance with the area to be trained.
[0063] FIG. 9 is a diagram illustrating an overview of the learning process in the learning device 10. The learning process shown in FIG. 9 is configured to emphasize local shape information, brightness information, and color estimation based on coordinates from a point cloud, and is configured to assign R, G, and B based on both the position in the scene and the local shape. This configuration is effective for outdoor scenes, for example, where the roads and sidewalks have a consistent color and trees and utility poles are mixed. During learning, two patterns of sampling, coarse and fine, are created and learned, similar to the learning of the NeRF model described in FIG. 6.
[0064] 9 adds to the learning process shown in Fig. 8 by combining a feature associated with a spatial position obtained by nonlinearly transforming the spatial coordinate x using a neural network with the feature F to generate a feature F'. By adding information about the spatial coordinate x when generating the feature F', the learning device 10 can train a trained model 1 that performs color estimation taking into account the relative position within the target area as well as local shape features.
[0065] 10 is a flowchart showing the flow of the learning process by the learning device 10. The CPU 11 reads out the learning process program from the ROM 12 or the storage 14, loads it into the RAM 13, and executes it, thereby performing the learning process.
[0066] In step S101, the CPU 11 acquires three-dimensional coordinate values, line-of-sight direction information, point cloud data, and a correct image that is an image captured from the line-of-sight direction, which are used in the learning process.
[0067] Following step S101, in step S102, CPU 11 uses three-dimensional coordinate values, gaze direction information, and point cloud data as input data and a correct image as training data to optimize the model parameters of trained model 1. CPU 11 optimizes the model parameters of trained model 1 by executing, for example, any of the learning processes shown in FIGS.
[0068] Following step S102, in step S103, the CPU 11 saves the optimized model parameters of the trained model 1.
[0069] 11 is a flowchart showing the flow of image processing by the image processing device 20. The CPU 21 reads out an image processing program from the ROM 22 or storage 24, loads it into the RAM 23, and executes it, thereby performing image processing.
[0070] In step S201, the CPU 21 acquires information on a viewpoint to be generated when generating an image using the trained model 1.
[0071] Following step S201, in step S202, the CPU 21 reads the model parameters of the trained model 1.
[0072] Following step S202, in step S203, CPU 21 inputs information about the viewpoint to be generated into trained model 1 into which the model parameters have been read, and generates an image from the target viewpoint using the color and transparency of each pixel output from trained model 1.
[0073] In the above embodiments, the learning process and image processing executed by the CPU by reading software (programs) may be executed by various processors other than the CPU. Examples of such processors include programmable logic devices (PLDs) whose circuit configuration can be changed after fabrication, such as field-programmable gate arrays (FPGAs), and dedicated electrical circuits, such as application-specific integrated circuits (ASICs), which are processors with circuit configurations specifically designed to execute specific processes. The learning process and image processing may be executed by one of these various processors, or by a combination of two or more processors of the same or different types (e.g., multiple FPGAs, or a combination of a CPU and an FPGA). The hardware structure of these various processors is, more specifically, an electrical circuit that combines circuit elements such as semiconductor devices.
[0074] In addition, in each of the above embodiments, the learning processing program is pre-stored (installed) in the storage 14, and the image processing program is pre-stored (installed) in the storage 24, but this is not limiting. The programs may be provided in a form stored in a non-transitory storage medium such as a CD-ROM (Compact Disk Read Only Memory), a DVD-ROM (Digital Versatile Disk Read Only Memory), or a USB (Universal Serial Bus) memory. The programs may also be downloaded from an external device via a network.
[0075] The following additional notes are provided regarding the above-described embodiments.
[0076] (Additional note 1) Memory and at least one processor coupled to said memory; Including, The processor: Three-dimensional coordinate values, line-of-sight direction information, and point cloud data are used as input data, and images captured from multiple directions are acquired as training data. Using the input data and the training data, a model is learned for outputting an image from a specified line of sight by outputting color and density for each pixel. A learning device configured as follows.
[0077] (Additional note 2) Memory and at least one processor coupled to said memory; Including, The processor: The gaze direction is input to a trained model for outputting an image from a specified gaze direction by using three-dimensional coordinate values, gaze direction information, and point cloud data as input data and images captured from a plurality of directions as training data, and outputting the color and density for each pixel, and outputting the color and transparency for each pixel from the gaze direction from the model. Generate an image from the viewing direction using the color and the transparency The image processing device is configured as follows.
[0078] (Additional note 3) A non-transitory storage medium storing a program executable by a computer to perform a learning process, The learning process includes: Three-dimensional coordinate values, line-of-sight direction information, and point cloud data are used as input data, and images captured from multiple directions are acquired as training data. Using the input data and the training data, a model is learned for outputting an image from a specified gaze direction by outputting color and density for each pixel. Non-transitory storage medium.
[0079] (Additional note 4) A non-transitory storage medium storing a program executable by a computer to perform image processing, The image processing The gaze direction is input to a trained model for outputting an image from a specified gaze direction by using three-dimensional coordinate values, gaze direction information, and point cloud data as input data and images captured from a plurality of directions as training data, and outputting the color and density for each pixel, and outputting the color and transparency for each pixel from the gaze direction from the model. Generate an image from the viewing direction using the color and the transparency Non-transitory storage medium. [Explanation of symbols]
[0080] 1. Pre-trained model 10 Learning Device 20 Image processing device
Claims
1. an acquisition unit that receives three-dimensional coordinate values, line-of-sight direction information, and point cloud data as input data and acquires images captured from multiple directions as training data; a learning unit that uses the input data and the training data to learn a model for outputting an image from a specified line of sight by outputting a color and density for each pixel; Equipped with the learning unit inputs a first feature amount obtained from the point cloud data and the three-dimensional coordinate values into a predetermined first neural network to output a density for each pixel, and inputs information on the line of sight direction and a feature amount obtained from the first feature amount into a predetermined second neural network to learn the model so as to output a color for each pixel; the first feature is obtained from a feature obtained by inputting the three-dimensional coordinate values into a predetermined third neural network and a feature obtained by inputting the point cloud data into a predetermined model.
2. The learning device according to claim 1 , wherein the first feature is obtained from a neighboring point set with the three-dimensional coordinate value as a center point and a feature obtained by inputting the point cloud data into a predetermined model.
3. an estimation unit that inputs a gaze direction into a trained model that uses three-dimensional coordinate values, gaze direction information, and point cloud data as input data, and images captured from a plurality of directions as training data, thereby outputting a color and density for each pixel, thereby outputting the color and transparency for each pixel from the gaze direction; an image processing unit that generates an image from the line of sight direction using the color and the transparency output by the estimation unit; Equipped with the trained model is trained to output a density for each pixel by inputting a first feature amount obtained from the point cloud data and the three-dimensional coordinate values into a predetermined first neural network, and to output a color for each pixel by inputting information on the gaze direction and a feature amount obtained from the first feature amount into a predetermined second neural network; The image processing device wherein the first feature is obtained from a feature obtained by inputting the three-dimensional coordinate values into a predetermined third neural network and a feature obtained by inputting the point cloud data into a predetermined model.
4. The processor: Three-dimensional coordinate values, line-of-sight direction information, and point cloud data are used as input data, and images captured from multiple directions are acquired as training data. Using the input data and the training data, a model is learned for outputting an image from a specified line of sight by outputting color and density for each pixel. Execute the process, the processor inputs the point cloud data and a first feature amount obtained from the three-dimensional coordinate values into a predetermined first neural network to output a density for each pixel, and inputs the gaze direction information and a feature amount obtained from the first feature amount into a predetermined second neural network to train the model so as to output a color for each pixel; A learning method in which the first feature is obtained from a feature obtained by inputting the three-dimensional coordinate values into a predetermined third neural network and a feature obtained by inputting the point cloud data into a predetermined model.
5. The processor: The gaze direction is input to a trained model for outputting an image from a specified gaze direction by using three-dimensional coordinate values, gaze direction information, and point cloud data as input data and images captured from a plurality of directions as training data, and outputting the color and density for each pixel, and outputting the color and transparency for each pixel from the gaze direction from the model. Generate an image from the viewing direction using the color and the transparency Execute the process, the trained model is trained to output a density for each pixel by inputting a first feature amount obtained from the point cloud data and the three-dimensional coordinate values into a predetermined first neural network, and to output a color for each pixel by inputting information on the gaze direction and a feature amount obtained from the first feature amount into a predetermined second neural network; an image processing method in which the first feature is obtained from a feature obtained by inputting the three-dimensional coordinate values into a predetermined third neural network and a feature obtained by inputting the point cloud data into a predetermined model.
6. A computer is a learning device according to claim 1 or A computer program for causing the image processing device according to claim 3 to function.
Citation Information
Patent Citations
Three-dimensional nail arm modelling method
JP2017018158A
Methods and systems for generating and using localization reference data
JP2018533721A