Image processing method, neural network learning method, three-dimensional image display method, image processing system, neural network learning system, and three-dimensional image display system

The image processing method addresses distortion issues in free viewpoint image synthesis by estimating residual information through machine learning, eliminating the need for costly three-dimensional sensing devices and achieving high-quality images.

JP7761134B2Active Publication Date: 2025-10-28SOCIONEXT INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024513585
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-04-04
Publication Date
2025-10-28
Estimated Expiration
2042-04-04

AI Technical Summary

Technical Problem

Existing image processing systems using predefined projection planes in synthesizing free viewpoint images suffer from distortion due to mismatches between the predefined projection plane and the actual three-dimensional structure, and adding three-dimensional sensing devices like LiDAR increases costs.

Method used

An image processing method that estimates residual information using machine learning to map captured images onto a display projection surface, eliminating the need for three-dimensional sensing devices by calculating the difference between a predefined default projection surface and the actual display projection surface.

Benefits of technology

Synthesizes free viewpoint images with minimal distortion without the use of three-dimensional sensing devices, reducing costs and improving image quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007761134000001
    Figure 0007761134000001
  • Figure 0007761134000002
    Figure 0007761134000002
  • Figure 0007761134000003
    Figure 0007761134000003
Patent Text Reader

Abstract

In this image processing method for synthesizing a free viewpoint image on a display projection plane using a plurality of captured images on the basis of viewpoint information, a computer performs: an image acquisition step for acquiring the plurality of captured images using each of a plurality of cameras; a residual estimation step for inputting the plurality of captured images and the viewpoint information, and estimating projection plane residual information indicating the difference between a predefined bowl-shaped default projection plane and the display projection plane by machine learning; and a mapping step for obtaining the free viewpoint image by mapping the plurality of captured images onto the display projection plane using information about the default projection plane, the residual information, and the viewpoint information.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an image processing method, a neural network learning method, a three-dimensional image display method, an image processing system, a neural network learning system, and a three-dimensional image display system. [Background technology]

[0002] 2. Description of the Related Art Image processing systems are known that use a plurality of captured images captured by a plurality of cameras to synthesize a free viewpoint image, which is a three-dimensional image that can be displayed by freely moving the viewpoint.

[0003] For example, a technique is known in which a bowl-shaped (cone-shaped) projection surface is defined in advance, and images captured by multiple cameras are mapped onto the projection surface to synthesize a free-viewpoint image (see Non-Patent Document 1). Also known is a technique in which a projection surface is calculated using distance information measured by a three-dimensional sensing device such as LiDAR (Laser Imaging Detection and Ranging), and images captured by multiple cameras are mapped onto the projection surface to synthesize a free-viewpoint image (see Non-Patent Document 2). [Prior art documents] [Non-patent literature]

[0004] [Non-Patent Document 1] Seiya Shimizu, et.al, "Wraparound View System for Motor Vehicles", Fujitsu Scientific & Technical Journal 46(1):95-102, 2010. [Non-patent document 2] "World's first! Development of in-vehicle 3D image synthesis technology that displays people and objects around the vehicle without distortion and clearly indicates the risk of contact," Fujitsu Press Release, 2013, [online], Fujitsu, Internet<URL: https: / / pr.fujitsu.com / jp / news / 2013 / 10 / 9-2.html> ,[Retrieved March 24, 1992] Summary of the Invention [Problem to be solved by the invention]

[0005] The technique disclosed in Non-Patent Document 1 has a problem in that a mismatch between a predefined projection plane and the actual three-dimensional structure causes distortion in the composite image projected onto the projection plane.

[0006] In addition, as in the technology disclosed in Non-Patent Document 2, distortion of the composite image can be suppressed by adding a three-dimensional sensing device such as LiDAR, but there is a problem in that adding a three-dimensional sensing device increases costs.

[0007] One embodiment of the present invention has been made in consideration of the above-mentioned problems, and enables an image processing system that synthesizes a free-viewpoint image using a plurality of captured images to synthesize a free-viewpoint image with little distortion without using a three-dimensional sensing device. [Means for solving the problem]

[0008] In order to solve the above-mentioned problems, an image processing method according to one embodiment of the present invention is an image processing method for synthesizing a free viewpoint image using a plurality of captured images on a display projection surface based on viewpoint information, wherein the image processing method includes an image acquisition step of acquiring the plurality of captured images by each of a plurality of cameras, a residual estimation step of inputting the plurality of captured images and the viewpoint information and estimating, by machine learning, residual information on the projection surface that indicates the difference between a predefined bowl-shaped default projection surface and the display projection surface, and a mapping step of mapping the plurality of captured images on the display projection surface using information on the default projection surface, the residual information, and the viewpoint information to obtain the free viewpoint image. [Effects of the Invention]

[0009] According to one embodiment of the present invention, in an image processing system that synthesizes a free viewpoint image using a plurality of captured images, it becomes possible to synthesize a free viewpoint image with little distortion without using a three-dimensional sensing device. [Brief explanation of the drawings]

[0010] [Figure 1] FIG. 1 is a diagram illustrating an example of a system configuration of an image processing system according to an embodiment. [Figure 2] FIG. 1 is a diagram illustrating an overview of image processing according to an embodiment. [Figure 3] FIG. 2 is a diagram illustrating an example of a hardware configuration of a computer according to an embodiment. [Figure 4] FIG. 1 is a diagram illustrating an example of a functional configuration of an image processing apparatus according to an embodiment. [Figure 5] 10 is a flowchart illustrating an example of image processing according to an embodiment. [Figure 6] FIG. 2 is a diagram for explaining an overview of a learning process according to the first embodiment. [Figure 7] FIG. 2 is a diagram illustrating an example of a functional configuration of the image processing device (during learning) according to the first embodiment. [Figure 8] 5 is a flowchart illustrating an example of a learning process according to the first embodiment. [Figure 9] FIG. 10 is a diagram illustrating an outline of a learning process according to a second embodiment. [Figure 10] FIG. 10 is a diagram illustrating an example of a functional configuration of an image processing device (during learning) according to a second embodiment. [Figure 11] 10 is a flowchart illustrating an example of a learning process according to the second embodiment. [Figure 12] 10 is a flowchart illustrating an example of a residual information calculation process according to the second embodiment. [Figure 13] FIG. 11 is a diagram illustrating an example of the configuration of a residual estimation model according to the third embodiment. [Figure 14] FIG. 10 is a diagram illustrating an example of a functional configuration of an image processing apparatus according to a third embodiment. [Figure 15] FIG. 10 is a diagram illustrating an outline of a learning process according to the third embodiment. [Figure 16] 10 is a flowchart (1) illustrating an example of a learning process according to the third embodiment. [Figure 17] 13 is a flowchart (2) illustrating an example of a learning process according to the third embodiment. [Figure 18] FIG. 10 is a diagram illustrating an example of a system configuration of a three-dimensional image display system according to a fourth embodiment. [Figure 19] FIG. 10 is a diagram illustrating an example of a hardware configuration of an edge device according to a fourth embodiment. [Figure 20] FIG. 10 is a diagram illustrating an example of the functional configuration of a three-dimensional image display system according to a fourth embodiment. [Figure 21] FIG. 13 is a sequence diagram illustrating an example of a three-dimensional image display process according to the fourth embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0011] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings.

[0012] The image processing system according to this embodiment is a system that uses a plurality of captured images captured by a plurality of cameras to synthesize a free-viewpoint image, which is a three-dimensional image that can be displayed by freely moving the viewpoint. The image processing system according to this embodiment can be applied to, for example, periphery monitoring of a moving object such as an automobile, a robot, or a drone, or to AR (Augmented Reality) / VR (Virtual Reality) technology. Here, as an example, a case where the image processing system according to this embodiment is installed in a vehicle such as an automobile will be described.

[0013] <System configuration> Fig. 1 is a diagram showing an example of the system configuration of an image processing system according to an embodiment. In the example of Fig. 1, the image processing system 100 includes an image processing device 10 mounted on a vehicle 1 such as an automobile, a plurality of cameras 12, and a display device 16. The above components are connected to each other so as to be able to communicate with each other, for example, via an in-vehicle network, a wired cable, or wireless communication.

[0014] The vehicle 1 is an example of a moving body that is equipped with the image processing system 100 according to this embodiment. The moving body is not limited to the vehicle 1, and may be, for example, a robot that moves by legs or the like, a manned or unmanned aircraft, or any other device or machine that has a moving function.

[0015] Camera 12 is an imaging device that captures images of the periphery of vehicle 1 and acquires the captured images. In the example of FIG. 1, four cameras 12A to 12D are provided on vehicle 1, facing different imaging areas E1 to E4. In the following description, "camera 12" will be used to refer to any of the four cameras 12A to 12D. Furthermore, "imaging area E" will be used to refer to any of the four imaging areas E1 to E4. The numbers of cameras 12 and imaging areas E shown in FIG. 1 are merely examples, and may be any number greater than or equal to two.

[0016] 1, for example, camera 12A is provided facing an imaging area E1 in front of vehicle 1, and camera 12B is provided facing an imaging area E2 on one side of vehicle 1. Camera 12C is provided facing an imaging area E3 on another side of vehicle 1, and camera 12D is provided facing an imaging area E4 behind vehicle 1.

[0017] The display device 16 is, for example, a display device such as an LCD (Liquid Crystal Display) or an organic EL (Electro-Luminescence) display, or various other devices or apparatuses having a display function for displaying various types of information.

[0018] The image processing device 10 is a computer that executes a predetermined program to perform image processing for synthesizing a free viewpoint image on a display projection surface using a plurality of captured images taken by the cameras 12A to 12D. The free viewpoint image is a three-dimensional image that can be displayed by freely moving the viewpoint using a plurality of captured images taken by a plurality of cameras.

[0019] (Processing Overview) 2 is a diagram for explaining an overview of image processing according to an embodiment. The image processing device 10 has projection surface information 230, which is information on a bowl-shaped (or cone-shaped) projection surface (hereinafter referred to as a default projection surface 231) predefined around the vehicle 1.

[0020] Furthermore, the image processing device 10 inputs a plurality of captured images 210 of the periphery of the vehicle 1 captured by the plurality of cameras 12 and viewpoint information 240 indicating the viewpoint of the free viewpoint image into a residual estimation model, and estimates residual information 220 of the projection plane by machine learning (step S1). Here, the residual estimation model is a trained neural network that receives the plurality of captured images 210 and the viewpoint information 240 as input data and outputs residual information 220 of the projection plane indicating the difference between a projection plane onto which the free viewpoint image is projected (hereinafter referred to as a display projection plane) and a default projection plane 231.

[0021] Furthermore, the image processing device 10 generates a free viewpoint image 250 by mapping the multiple captured images 210 onto the display projection surface, using the projection surface information 230, which is information on the default projection surface 231, the residual information 220, and the viewpoint information 240 (step S2). Here, as described above, the residual information 220 is information indicating the difference between the display projection surface and the default projection surface 231, so the image processing device 10 can calculate the display projection surface from the projection surface information 230 and the residual information 220.

[0022] The residual estimation model has been machine-trained in advance to estimate the difference between the default projection plane 231 and the display projection plane from multiple learning images, viewpoint information 240, and three-dimensional information of one or more three-dimensional objects captured in the multiple learning images.

[0023] Therefore, according to this embodiment, in the image processing system 100 that synthesizes a free viewpoint image 250 using a plurality of captured images 210, it is possible to synthesize a free viewpoint image with little distortion without using a three-dimensional sensing device.

[0024] The system configuration of the image processing system 100 shown in FIG. 1 is an example. For example, The image processing system 100 includes a plurality of cameras 12, a display device 16, and the like, and may be a wearable device such as AR goggles or VR goggles worn by a user.

[0025] <Hardware configuration> The image processing device 10 includes, for example, the hardware configuration of a computer 300 as shown in FIG.

[0026] 3 is a diagram illustrating an example of the hardware configuration of a computer according to an embodiment. The computer 300 includes, for example, a processor 301, a memory 302, a storage device 303, an I / F (Interface) 304, an input device 305, an output device 306, a communication device 307, and a bus 308.

[0027] The processor 301 is an arithmetic device such as a CPU (Central Processing Unit) or GPU (Graphics Processing Unit) that executes predetermined processing by executing a program stored in a storage medium such as the storage device 303. The memory 302 includes, for example, a RAM (Random Access Memory), which is a volatile memory used as a work area or the like for the processor 301, and a ROM (Read Only Memory), which is a nonvolatile memory that stores programs for starting up the processor 301. The storage device 303 is, for example, a large-capacity nonvolatile storage device such as an SSD (Solid State Drive) or an HDD (Hard Disk Drive). The I / F 304 includes, for example, various interfaces for connecting external devices such as the camera 12 and the display device 16 to the computer 300.

[0028] The input device 305 includes various devices that accept external input (e.g., a keyboard, a touch panel, a pointing device, a microphone, a switch, a button, a sensor, etc.). The output device 206 includes various devices that perform output to the outside (e.g., a display, a speaker, an indicator, etc.). The communication device 307 includes various communication devices for communicating with other devices via a wired or wireless network. The bus 308 is connected to each of the above components and transmits, for example, address signals, data signals, and various control signals.

[0029] <Functional configuration> Fig. 4 is a diagram showing an example of the functional configuration of an image processing device according to an embodiment. The image processing device 10 implements an image acquisition unit 401, a residual estimation unit 402, a mapping unit 403, a display control unit 404, a setting unit 405, a storage unit 406, and the like, for example, by executing an image processing program in the processor 301 shown in Fig. 3. Note that at least a part of the above functional configurations may be implemented by hardware.

[0030] The image acquisition unit 401 executes an image acquisition process for acquiring a plurality of captured images 210 by each of the plurality of cameras 12. For example, the image acquisition unit 401 acquires a plurality of captured images 210 of the periphery of the vehicle 1 by the plurality of cameras 12A, 12B, 12C, and 12D.

[0031] The residual estimation unit 402 inputs the multiple captured images 210 acquired by the image acquisition unit 401 and the viewpoint information 240, and performs a residual estimation process to estimate residual information 220 of the projection surface that indicates the difference between a predefined bowl-shaped default projection surface 231 and the display projection surface using machine learning.

[0032] Preferably, the residual estimation unit 402 has a residual estimation model 410 that has been trained to infer a difference between the default projection plane 231 and the display projection plane from a plurality of learning captured images, viewpoint information 240, and three-dimensional information of one or more three-dimensional objects captured in the plurality of learning captured images. The residual estimation model 410 is a trained neural network (hereinafter referred to as NN) that receives the plurality of captured images 210 and the viewpoint information 240 as input data, and outputs residual information 220 of the projection plane that indicates the difference between the display projection plane onto which a free viewpoint image is projected and the default projection plane 231. In this embodiment, of the NNs that receive the plurality of captured images 210 and the viewpoint information 240 as input data and output the residual information 220, the trained NN is referred to as the residual estimation model 410.

[0033] The residual estimation unit 402 inputs the multiple captured images 210 acquired by the image acquisition unit 401 and viewpoint information 240 to the residual estimation model 410, and acquires the residual information 220 output by the residual estimation model 410. Here, the viewpoint information 240 is coordinate information indicating the viewpoint of the free viewpoint image generated by the image processing device 10, and is expressed, for example, in Cartesian coordinates or polar coordinates. The three-dimensional information is, for example, data including three-dimensional point cloud information measured by a three-dimensional sensing device such as LiDAR (Laser Imaging Detection and Ranging), or three-dimensional distance information of objects around the vehicle 1, such as a depth image including depth information.

[0034] The mapping unit 403 performs a mapping process using the projection surface information 230 related to the default projection surface 231 and the residual information 220 estimated by the residual estimation unit 402 to map the acquired multiple captured images 210 onto the display projection surface to obtain a free viewpoint image 250. As described above, the residual information 220 is information indicating the difference between the display projection surface and the default projection surface 231, so the mapping unit 403 can calculate the display projection surface from the projection surface information 230, which is information about the default projection surface 231, and the residual information 220. Furthermore, the process of mapping the multiple captured images onto the calculated display projection surface to obtain the free viewpoint image 250 can be performed using known techniques such as those shown in Non-Patent Documents 1 and 2.

[0035] The display control unit 404 executes a display control process for displaying the free viewpoint image 250 etc. generated by the mapping unit 403 on the display device 16 etc.

[0036] The setting unit 405 executes a setting process for setting information such as the projection surface information 230 and the viewpoint information 240 in the image processing device 10 .

[0037] The memory unit 406 is realized, for example, by a program executed by the processor 301, a storage device 303, a memory 302, etc., and performs storage processing to store various information (or data) including the captured image 210, projection surface information 230, and viewpoint information 240, etc.

[0038] 4 is an example of the functional configuration of the image processing device 10. For example, the functional configurations included in the image processing device 10 may be distributed among a plurality of computers 300.

[0039] <Processing flow> Next, the processing flow of the image processing method according to this embodiment will be described.

[0040] (Image Processing) 5 is a flowchart showing an example of image processing according to an embodiment. This processing shows a specific example of the image processing described in FIG. 2, which is executed by the image processing device 10 described in FIG.

[0041] In step S501, the image acquisition unit 401 acquires a plurality of captured images 210 of the periphery of the vehicle 1 using, for example, a plurality of cameras 12.

[0042] In step S502, the residual estimation unit 402 inputs the multiple captured images 210 acquired by the image acquisition unit 401 and viewpoint information 240 indicating the viewpoint of the free viewpoint image 250 into the residual estimation model 410, and estimates residual information 220 of the projection surface.

[0043] 2 and the residual information 220 estimated by the residual estimation unit 402. For example, the mapping unit 403 calculates the display projection surface by reflecting the residual information 220 on the default projection surface 231.

[0044] In step 504 , the mapping unit 403 maps the plurality of captured images 210 acquired by the image acquisition unit 401 onto the display projection surface to generate a free viewpoint image 250 .

[0045] In step S505, the display control unit 404 displays the free viewpoint image 250 generated by the mapping unit 403 on the display device 16 or the like.

[0046] By the processing of Figure 5, the image processing device 10 can synthesize a free viewpoint image 250 with little distortion in the image processing system 100 that synthesizes a free viewpoint image 250 using multiple captured images 210, without using a three-dimensional sensing device.

[0047] <About the learning process> Next, a learning process for machine learning the residual estimation model 410 will be described.

[0048] [First embodiment] (Processing Overview) 6 is a diagram for explaining an overview of the learning process according to the first embodiment. This process shows an example of a learning process in which the image processing device 10 performs machine learning on a residual learning model, which is an NN that receives a plurality of captured images 210 and viewpoint information 240 as input data and outputs residual information 220. In this embodiment, of the NNs that receive a plurality of captured images 210 and viewpoint information 240 as input data and output residual information 220, an NN before learning or currently learning is referred to as a residual learning model. Furthermore, the image processing device 10 that executes the learning process may be the same computer as the computer 300 that executes the image processing described with reference to FIGS. 2 to 4, or may be a different computer.

[0049] The image processing device 10 acquires a plurality of captured images 210 captured by a plurality of cameras 12 and three-dimensional information (for example, three-dimensional point cloud information, depth image, etc.) acquired by a three-dimensional sensor such as LiDAR. Furthermore, the image processing device 10 reconstructs a three-dimensional image of the plurality of captured images 210 using the acquired three-dimensional information, and generates (renders) a teacher image 602, which is a free viewpoint image for the teacher, based on the input viewpoint information 240 (step S11).

[0050] Furthermore, the image processing device 10 inputs the plurality of captured images 210 and the viewpoint information 240 to a residual learning model, and acquires residual information (hereinafter referred to as training residual information 601) output by the residual learning model (step S12). Subsequently, the image processing device 10 uses the acquired training residual information, the projection surface information 230, and the viewpoint information 240 to map the plurality of captured images 210 onto a display projection surface, thereby generating free viewpoint images for training (hereinafter referred to as training images 603) (step S13).

[0051] Furthermore, the image processing apparatus 10 trains a residual learning model (NN) so as to reduce the error between the generated teacher image 602 and the learning image 603 (step S14).

[0052] <Functional configuration> 7 is a diagram showing an example of the functional configuration of the image processing device (during learning) according to the first embodiment. The image processing device 10, for example, implements a captured image preparation unit 701, a three-dimensional information preparation unit 702, a teacher image preparation unit 703, a learning unit 704, a setting unit 705, a storage unit 706, etc. by executing a program for learning processing in the processor 301 of FIG. 3. Note that at least a part of the above functional configurations may be implemented by hardware.

[0053] The captured image preparation unit 701 executes a captured image preparation process to prepare a plurality of captured images 210 for learning captured by each of the plurality of cameras 12. Note that the captured image preparation unit 701 may acquire the plurality of captured images 210 in real time using the plurality of cameras 12, or may acquire the plurality of captured images 210 required for the learning process from captured images 711 captured in advance and stored in the storage unit 706 or the like.

[0054] The three-dimensional information preparation unit 702 executes a three-dimensional information preparation process to acquire three-dimensional information (e.g., three-dimensional point cloud information or depth images) corresponding to the plurality of captured images 210 for learning prepared by the captured image preparation unit 701. As an example, the three-dimensional information preparation unit 702 uses a three-dimensional sensor such as a LiDAR 707 to acquire three-dimensional point cloud information of the periphery of the vehicle 1 at the same timing (synchronization) as the capture of the plurality of captured images 210. Note that the three-dimensional information preparation unit 702 may acquire the three-dimensional information using another three-dimensional sensor such as a stereo camera, a depth camera that captures depth images, or a wireless sensing device.

[0055] As another example, the three-dimensional information preparation unit 702 may acquire three-dimensional information indicating the positions of surrounding three-dimensional objects by using a technique such as Visual SLAM (Simultaneous Localization and Mapping) on ​​the plurality of captured images 210 for learning stored in the storage unit 706, etc. In short, the three-dimensional information preparation unit 702 may prepare three-dimensional information indicating the positions of three-dimensional objects around the vehicle 1, synchronized with the plurality of captured images 210 for learning prepared by the captured image preparation unit 701, using any method.

[0056] The teacher image preparation unit 703 performs a teacher image preparation process to restore a three-dimensional image of the multiple captured images 210 using the three-dimensional information prepared by the three-dimensional information preparation unit 702, and to generate (render) a teacher image 602, which is a free viewpoint image for the teacher, based on the viewpoint information 240.

[0057] The learning unit 704 executes a learning process to learn a residual learning model (NN) 710 using a plurality of captured images 210, viewpoint information 240, projection surface information 230, and a teacher image 602. For example, the learning unit 704 inputs the plurality of captured images 210 and viewpoint information 240 into the residual learning model (NN) 710 to acquire training residual information 601, and calculates a display projection surface for training using the training residual information 601 and the projection surface information 230. Furthermore, the learning unit 704 maps the plurality of captured images 210 onto the calculated display projection surface using the viewpoint information 240, to generate a training image 603, which is a free viewpoint image for training. Furthermore, the learning unit 704 trains the residual learning model (NN) 710 so as to reduce an error between the generated teacher image 602 and the training image 603.

[0058] The setting unit 705 executes a setting process for setting various pieces of information such as the projection surface information 230 and the viewpoint information 240 in the image processing device 10 .

[0059] The memory unit 706 is realized, for example, by a program executed by the processor 301, a storage device 303, a memory 302, etc., and stores various information (or data) such as a captured image 711, three-dimensional information 712, projection surface information 230, and viewpoint information 240.

[0060] 7 is an example. For example, the functional components included in the image processing device 10 may be distributed among a plurality of computers 300.

[0061] <Processing flow> Next, the processing flow of the neural network learning method according to the first embodiment will be described.

[0062] Fig. 8 is a flowchart showing an example of the learning process according to the first embodiment. This process shows a specific example of the learning process described in Fig. 6, which is executed by the image processing device 10 described in Fig. 7.

[0063] In step S801a, the captured image preparation unit 701 prepares a plurality of captured images 210 for learning captured by each of the plurality of cameras 12. For example, the captured image preparation unit 701 acquires a plurality of captured images 210 capturing the periphery of the vehicle 1 using the plurality of cameras 12.

[0064] In step S801b, the three-dimensional information preparation unit 702 acquires three-dimensional information corresponding to the plurality of captured images 210 for learning prepared by the captured image preparation unit 701. For example, the three-dimensional information preparation unit 702 synchronizes with the captured image preparation unit 701 and acquires three-dimensional information (e.g., three-dimensional point cloud information) of the surroundings of the vehicle 1 at the same time.

[0065] In step S802, the image processing apparatus 10 prepares viewpoint information 240 indicating a viewpoint to be learned. For example, the setting unit 705 sets the coordinates of the viewpoint to be learned in the residual learning model 710 in the viewpoint information 240.

[0066] In step S803, the teacher image preparation unit 703 reconstructs three-dimensional images of the multiple captured images 210 using the three-dimensional information prepared by the three-dimensional information preparation unit 702, and generates (renders) a teacher image 602 based on the viewpoint information 240.

[0067] In step S804, the learning unit 704 inputs the plurality of captured images 210 and the viewpoint information 240 to the residual learning model (NN) 710 to acquire the residual information for learning 601 in parallel with the processing of step S803.

[0068] In step S805, the learning unit 704 calculates the display projection surface from the learning residual information 601 and the projection surface information 230, and generates learning images 603 by mapping the multiple captured images 210 onto the display projection surface based on the viewpoint information 240.

[0069] In step S806, the learning unit 704 learns the residual learning model 710 so as to minimize the difference between the generated teacher image 602 and the learning image 603. For example, the learning unit 704 calculates weights for the residual learning model 710 that minimize the difference between the two images (for example, the sum of the differences in pixel values ​​of all pixels), and sets the calculated weights to the residual learning model 710.

[0070] In step S807, the learning unit 704 determines whether learning has ended. For example, the learning unit 704 may determine that learning has ended when the processes of steps S801 to S806 have been executed a predetermined number of times. Alternatively, the learning unit 704 may determine that learning has ended when the difference between the teacher image 602 and the learning image 603 becomes equal to or less than a predetermined value.

[0071] If the learning has not finished, the learning unit 704 returns the process to steps S801a and S801b. On the other hand, if the learning has finished, the learning unit 704 ends the process of FIG.

[0072] The image processing device 10 described in FIG. 4 can execute the image processing described in FIG. 5 using the NN (residual learning model 710) trained in the processing of FIG. 8 as the residual estimation model 410.

[0073] [Second embodiment] (Processing Overview) 9 is a diagram for explaining an outline of the learning process according to the second embodiment. This process shows another example of a learning process in which the image processing device 10 performs machine learning on a residual learning model, which is a NN that receives a plurality of captured images 210 and viewpoint information 240 as input data and outputs residual information 220. Note that detailed description of the process similar to that of the first embodiment will be omitted here.

[0074] The image processing device 10 uses a plurality of captured images 210, projection plane information 230, and viewpoint information 240 to generate an unmodified free viewpoint image (hereinafter referred to as an unmodified image 901) by mapping the plurality of captured images 210 onto a default projection plane 231 (step S21).

[0075] In addition, the image processing device 10 acquires three-dimensional information, restores the multiple captured images 210 using the acquired three-dimensional information, and generates (renders) a teacher image 602, which is a free viewpoint image for the teacher, based on the viewpoint information 240 (step S22).

[0076] The image processing device 10 also compares the generated uncorrected image 901 with the teacher image 602 to obtain residual information of the two images (step S23). Furthermore, the image processing device 10 inputs the multiple captured images 210 and the viewpoint information 240 to the residual learning model 710 to obtain the training residual information 601 (step S24). Subsequently, the image processing device 10 trains the residual learning model 710 so as to minimize the difference between the residual information of the two images and the training residual information 601 (step S25).

[0077] <Functional configuration> Fig. 10 is a diagram showing an example of the functional configuration of an image processing device (during learning) according to the second embodiment. As shown in Fig. 10, the image processing device 10 according to the second embodiment has an uncorrected image preparation unit 1001 and a residual calculation unit 1002 in addition to the functional configuration of the image processing device 10 according to the first embodiment described in Fig. 7. Furthermore, the learning unit 704 executes learning processing different from that of the first embodiment, as described in Fig. 9.

[0078] The unedited image preparation unit 1001 is realized by, for example, a program executed by the processor 301. The unedited image preparation unit 1001 executes an unedited image preparation process to generate an unedited image 901 by mapping the plurality of captured images 210 onto a default projection plane 231, using the plurality of captured images 210, projection plane information 230, and viewpoint information 240.

[0079] The residual calculation unit 1002 is realized by, for example, a program executed by the processor 301, and executes a residual calculation process that compares the generated uncorrected image 901 with the training image 602 and calculates residual information of the two images.

[0080] The learning unit 704 of the second embodiment performs a learning process to learn the residual learning model 710 so that the difference between the residual information of two images calculated by the residual calculation unit 1002 and the learning residual information 601 output by the residual learning model 710 is minimized.

[0081] The functional configuration other than that described above is the same as the functional configuration of the image processing device 10 according to the first embodiment described with reference to FIG. 7, and therefore will not be described here.

[0082] <Processing flow> Next, the processing flow of the neural network learning method according to the second embodiment will be described.

[0083] Fig. 11 is a flowchart showing an example of the learning process according to the second embodiment. This process shows a specific example of the learning process described in Fig. 9, which is executed by the image processing device 10 described in Fig. 10. Of the processes shown in Fig. 11, the processes of steps S801a, S801b, and S802 are the same as the learning process according to the first embodiment described in Fig. 8, and therefore will not be described here.

[0084] In step S1101, the unedited image preparation unit 1001 generates an unedited image 901 by mapping the plurality of captured images 210 onto the default projection plane 231 using the plurality of captured images 210, the projection plane information 230, and the viewpoint information 240.

[0085] In step S1102, the teacher image preparation unit 703 reconstructs three-dimensional images of the multiple captured images 210 using the three-dimensional information prepared by the three-dimensional information preparation unit 702, and generates (renders) a teacher image 602 based on the viewpoint information 240.

[0086] In step S1103, the residual calculation unit 1002 compares the generated uncorrected image 901 with the training image 602, and calculates residual information between the two images.

[0087] In step S1104, the learning unit 704 inputs the plurality of captured images 210 and the viewpoint information 240 to the residual learning model (NN) 710 to acquire the residual information for learning 601, for example, in parallel with the processing of steps S1101 to S1103.

[0088] In step S1105, the learning unit 704 learns the residual learning model 710 so that the difference between the residual information of the two images calculated by the residual calculation unit 1002 and the residual information for learning is minimized. For example, the learning unit 704 calculates weights for the residual learning model 710 that minimize the difference between the two pieces of residual information, and sets the calculated weights to the residual learning model 710.

[0089] In step S1106, the learning unit 704 determines whether learning has ended. If learning has not ended, the learning unit 704 returns the process to steps S801a and S801b. On the other hand, if learning has ended, the learning unit 704 ends the process of FIG. 11.

[0090] The image processing device 10 described in FIG. 4 can execute the image processing described in FIG. 5 using the NN (residual learning model 710) trained in the processing of FIG. 11 as the residual estimation model 410.

[0091] (Calculation of residual information) Fig. 12 is a flowchart showing an example of residual information calculation processing according to the second embodiment. This processing shows an example of residual information calculation processing executed by the residual calculation unit 1002 in step S1103 in Fig. 11, for example.

[0092] In step S1201, the residual calculation unit 1002 calculates the difference between each pixel of the uncorrected image 901 and the teacher image 602 (for example, the difference between the pixel values ​​of each pixel).

[0093] In step S1202, the residual calculation unit 1002 determines whether the calculated difference is equal to or less than a predetermined value. If the difference is equal to or less than the predetermined value, the residual calculation unit 1002 proceeds to step S1207, where it sets the current projection plane residual as the residual information of the two images. On the other hand, if the difference is not equal to or less than the predetermined value, the residual calculation unit 1002 proceeds to step S1203.

[0094] In step S1203, the residual calculation unit 1002 obtains the location on the image where the difference is large, and obtains the coordinates of the corresponding projection plane.

[0095] In step S1204, the residual calculation unit 1002 sets residual information of the projection plane so that the difference becomes small near the acquired coordinates.

[0096] In step S1205, the residual calculation unit 1002 generates a free viewpoint image by reflecting the set residual information.

[0097] In step S1206, the residual calculation unit 1002 calculates the difference between the pixel values ​​of the generated free viewpoint image and the teacher image, and the process returns to step S1202.

[0098] The residual calculation unit 1002 can calculate the residual information of the two images by repeatedly executing the process of Fig. 12 until the difference between the pixel values ​​of the two images becomes equal to or less than a predetermined value. However, the method by which the residual calculation unit 1002 calculates the residual information of the two images is not limited to this.

[0099] The learning process according to the first embodiment uses backpropagation based on image errors to learn (update weights) the residual learning model 710. As a prerequisite, each calculation in the residual calculation process must be a differentiable procedure in the learning system.

[0100] On the other hand, the learning process according to the second embodiment directly uses the residual information of two images to learn the residual learning model 710, which has the advantage that the residual calculation process does not have to be a differentiable procedure.

[0101] Developers and the like can select the learning process according to the first embodiment or the second embodiment depending on whether they want the advantage of being able to use direct images as training data or the advantage of being able to relax the conditions of the backpropagation method.

[0102] [Third embodiment] In the third embodiment, a preferred configuration example of a residual estimation model 410 and a residual learning model will be described.

[0103] 13 is a diagram showing an example of the configuration of a residual estimation model according to the third embodiment. The residual estimation model 410 may be configured in a form separated into a plurality of camera characteristic correction models 1301-1, 1301-2, 1301-3, ... and a base model 1302 that is common to each of the camera characteristic correction models. In the following description, when referring to any of the plurality of camera characteristic correction models 1301-1, 1301-2, 1301-3, ..., the term "camera characteristic correction model 1301" is used.

[0104] In this case, the image processing device 10 switches the camera characteristic correction model 1301 in accordance with, for example, a setting made by a user. Specifically, a user API (Application Programming Interface) is provided with an argument for specifying the camera characteristic correction model 1301, and a user SDK is provided with a database in which multiple camera characteristic correction models 1301 that are referenced in conjunction with the argument are defined.

[0105] The camera characteristic correction model 1301 is a network portion of the residual estimation model 410 that is mainly close to the image input, and learns weight data that is sensitive to camera characteristic parameters (focal length, etc.). The camera characteristic correction model 1301 is an example of a camera model inference engine that is trained to infer feature map information of feature points of one or more three-dimensional objects from a plurality of learning captured images and three-dimensional information of the three-dimensional objects captured in the plurality of learning captured images.

[0106] The base model 1302 learns weight data that is unlikely to be affected by the camera characteristic parameters and is common to each camera characteristic correction model 1301. The base model 1302 is an example of a base model inference engine that is trained to infer the difference between the default projection plane 231 and the display projection plane from the feature map information output by the camera characteristic correction model 1301 and the viewpoint information 240.

[0107] The residual estimation model 410 is an example of an inference engine that is separated into multiple camera model inference engines and a base model inference engine. The weight data after learning of the camera model inference engine is more influenced by the characteristic parameters of multiple cameras than the weight data after learning of the base model inference engine.

[0108] <Functional configuration> Fig. 14 is a diagram showing an example of the functional configuration of an image processing device according to the third embodiment. As shown in Fig. 14, the image processing device 10 according to the third embodiment stores a correction model DB (Database) 1401 in a storage unit 406 in addition to the functional configuration of the image processing device 10 according to the embodiment described in Fig. 4.

[0109] The correction model DB 1401 is a database in which a plurality of camera characteristic correction models 1301-1, 1301-2, 1301-3, . . . are defined.

[0110] For example, when the setting unit 405 displays a setting screen for a camera set and receives a setting of the camera set by the user, the setting unit 405 acquires the camera characteristic correction model 1301 corresponding to the received camera set from the correction model DB 1401. Furthermore, the setting unit 405 sets the acquired camera characteristic correction model 1301 in the residual estimation model 410.

[0111] As a result, when a user sets a first camera set, for example, the image processing device 10 executes the image processing described in Fig. 5 using the residual estimation model 410 including the camera characteristic correction model 1301-1 corresponding to the first camera set and the base model 1302. Similarly, when a user sets a second camera set, for example, the image processing device 10 executes the image processing described in Fig. 5 using the residual estimation model 410 including the camera characteristic correction model 1301-2 corresponding to the second camera set and the base model 1302.

[0112] <Learning process> Next, the learning process according to the third embodiment will be described.

[0113] (Processing Overview) 15 is a diagram for explaining an outline of the learning process according to the third embodiment. As the first learning process, the image processing device 10 learns a first camera characteristic correction model 1301-1 and a base model 1302 (step S31).

[0114] Furthermore, in the second learning process, the image processing device 10 combines the second camera characteristic correction model 1301-2 with the base model 1302 learned in the first learning process to learn a camera characteristic correction model 1302-2 (step S32).

[0115] Similarly, as the nth learning process, the image processing device 10 can learn the camera characteristic correction model 1301-n by combining the nth camera characteristic correction model 1301-n with the base model 1302 learned in the first learning process.

[0116] (Learning process 1) Fig. 16 is a flowchart (1) showing an example of learning processing according to the third embodiment. This processing shows an example of learning processing when the third embodiment is applied to the image processing device 10 according to the first embodiment described in Fig. 7. Note that detailed description of processing similar to that of the first embodiment will be omitted here.

[0117] In step S1601, the image processing apparatus 10 initializes a counter n to 1, and then executes the process of step S1602.

[0118] In step S1602, the image processing device 10 trains the residual training model 710 including the first camera characteristic correction model 1301-1 and the base model 1302 by the training process according to the first embodiment described with reference to FIG.

[0119] For example, referring to the flowchart of FIG. 8, in step S801a, the captured image preparation unit 701 executes a first captured image preparation process to prepare a first plurality of captured images by each of the first plurality of cameras (first camera set) 12.

[0120] In step S801b, the three-dimensional information preparation unit 702 executes a first three-dimensional information preparation process to prepare first three-dimensional information of one or more solid objects captured in the first plurality of captured images.

[0121] In step S803, the teacher image preparation unit 703 performs a first teacher image preparation process to restore a three-dimensional image of the first plurality of captured images using the first three-dimensional information and generate a first teacher image based on the input viewpoint information.

[0122] In step S804, the learning unit 704 inputs the first plurality of captured images and characteristic parameters for at least one camera 12 among the first plurality of cameras 12 into the first camera characteristic correction model 1301-1 to obtain first learning residual information.

[0123] In step S805, the learning unit 704 generates first learning residual information, projection surface information 23, and first learning image.

[0124] In step S806, the learning unit 704 trains both the camera characteristic correction model 1301-1 and the base model 1302 so as to reduce the error between the first teacher image and the first learning image.

[0125] Returning to Fig. 16, the processing from step S1603 onwards will now be described. In step S1603, the image processing device 10 determines whether n≦N (N is the number of camera characteristic correction models 1301). If n≦N is not satisfied, the image processing device 10 shifts the processing to step S1604. On the other hand, if n≦N is satisfied, the image processing device 10 ends the learning processing of Fig. 16.

[0126] In step S1604, the image processing apparatus 10 adds 1 to n and executes the process of step S1605.

[0127] In step S1605, the image processing apparatus 10 fixes the base model 1302, learns the n-th camera characteristic correction model through the learning process according to the first embodiment described with reference to FIG. 8, and returns the process to step S1603.

[0128] For example, when n=2, referring to the flowchart of FIG. 8, in step S801a, the captured image preparation unit 701 executes a second captured image preparation process to prepare a second plurality of captured images by each of the second plurality of cameras (second camera set) 12.

[0129] In step S801b, the three-dimensional information preparation unit 702 executes a second three-dimensional information preparation process to prepare second three-dimensional information of one or more solid objects captured in the second plurality of captured images.

[0130] In step S803, the teacher image preparation unit 703 performs a second teacher image preparation process to restore a three-dimensional image of the second plurality of captured images using the second three-dimensional information and generate a second teacher image based on the input viewpoint information.

[0131] In step S804, the learning unit 704 inputs the second plurality of captured images and characteristic parameters for at least one camera 12 among the second plurality of cameras 12 into the second camera characteristic correction model 1301-2 to obtain second learning residual information.

[0132] In step S805, the learning unit 704 uses the second learning residual information, the projection surface information 230, and the viewpoint information 240 to map the second plurality of captured images onto the display projection surface to generate second learning images.

[0133] In step S806, the learning unit 704 fixes the base model 1302 and trains the camera characteristic correction model 1301-2 so as to reduce the error between the second teacher image and the second learning image.

[0134] 16, the image processing device 10 can obtain a residual estimation model 410 including a plurality of camera characteristic correction models 1301-1, 1301-2, 1302-3, . . . and the base model 1302, as shown in FIG.

[0135] (Learning process 2) Fig. 17 is a flowchart (2) showing an example of learning processing according to the third embodiment. This processing shows an example of learning processing when the third embodiment is applied to the image processing device 10 according to the second embodiment described in Fig. 10. Note that detailed description of processing similar to that of the second embodiment will be omitted here.

[0136] In step S1701, the image processing apparatus 10 initializes a counter n to 1, and then executes the process of step S1702.

[0137] In step S1702, the image processing device 10 trains the residual training model 710 including the first camera characteristic correction model 1301-1 and the base model 1302 by the training process according to the second embodiment described with reference to FIG.

[0138] For example, referring to the flowchart of FIG. 11, in step S801a, the captured image preparation unit 701 executes a first captured image preparation process to prepare a first plurality of captured images by each of the first plurality of cameras (first camera set) 12.

[0139] In step S801b, the three-dimensional information preparation unit 702 executes a first three-dimensional information preparation process to prepare first three-dimensional information of one or more solid objects captured in the first plurality of captured images.

[0140] In step S1101, the unedited image preparation unit 1101 performs a first unedited image preparation process in which a first plurality of captured images are mapped onto a default projection plane and a first unedited image is generated based on input viewpoint information.

[0141] In step S1102, the teacher image preparation unit 703 performs a first teacher image preparation process to restore a three-dimensional image of the first plurality of captured images using the first three-dimensional information and generate a first teacher image based on the input viewpoint information 240.

[0142] In step S1103, the residual calculation unit 1002 executes a first residual calculation process to compare the first uncorrected image with the first training image and prepare first residual information.

[0143] In step S1104, the learning unit 704 inputs the first plurality of captured images and the viewpoint information 240 to the residual learning model 710 under training to acquire first residual information for training. Specifically, the learning unit 704 inputs the first plurality of captured images and characteristic parameters related to at least one camera among the first plurality of cameras to the first camera characteristic correction model 1301-1 to acquire first feature map information. The learning unit 704 also inputs the acquired first feature map information and the viewpoint information 240 to the base model 1302 to acquire first residual information for training.

[0144] In step S1105, the learning unit 704 learns the residual learning model 710 so as to minimize the difference between the first residual information calculated by the residual calculation unit 1002 and the first residual information for learning. In this way, the learning unit 704 uses the first residual information as training data to simultaneously learn both the first camera characteristic correction model 1301-1 and the base model 1302.

[0145] Returning to Fig. 17, the processing from step S1703 onwards will now be described. In step S1703, the image processing device 10 determines whether n≦N (N is the number of camera characteristic correction models 1301). If n≦N is not satisfied, the image processing device 10 shifts the processing to step S1704. On the other hand, if n≦N is satisfied, the image processing device 10 ends the learning processing of Fig. 17.

[0146] In step S1704, the image processing apparatus 10 adds 1 to n and executes the process of step S1705.

[0147] In step S1705, the image processing apparatus 10 fixes the base model 1302, learns the n-th camera characteristic correction model through the learning process according to the second embodiment described with reference to FIG. 11, and returns the process to step S1703.

[0148] For example, when n=2, referring to the flowchart of FIG. 11, in step S801a, the captured image preparation unit 701 executes a second captured image preparation process to prepare a second plurality of captured images by each of the second plurality of cameras (second camera set) 12.

[0149] In step S801b, the three-dimensional information preparation unit 702 executes a second three-dimensional information preparation process to prepare second three-dimensional information of one or more solid objects captured in the second plurality of captured images.

[0150] In step S1101, the unedited image preparation unit 1101 performs a second unedited image preparation process in which the second plurality of captured images are mapped onto the default projection plane 231 and a second unedited image is generated based on the input viewpoint information 240.

[0151] In step S1102, the teacher image preparation unit 703 performs a second teacher image preparation process to restore a three-dimensional image of the second plurality of captured images using the second three-dimensional information and generate a second teacher image based on the input viewpoint information 240.

[0152] In step S1103, the residual calculation unit 1002 performs a second residual calculation process to compare the second uncorrected image with the 21 training images and prepare second residual information.

[0153] In step S1104, the learning unit 704 inputs the second plurality of captured images and the viewpoint information 240 to the residual learning model 710 under training to acquire second residual information for learning. Specifically, the learning unit 704 inputs the second plurality of captured images and characteristic parameters related to at least one camera among the second plurality of cameras to the second camera characteristic correction model 1301-2 to acquire second feature map information. The learning unit 704 also inputs the acquired second feature map information and the viewpoint information 240 to the base model 1302 to acquire second residual information for learning.

[0154] In step S1105, the learning unit 704 fixes the base model 1302 and trains the second camera characteristic correction model 1301-2 so as to minimize the difference between the second residual information calculated by the residual calculation unit 1002 and the second residual information for training. In this way, the learning unit 704 trains the second camera characteristic correction model 1301-2 using the second residual information as training data.

[0155] By the learning process shown in FIG. 17, the image processing device 10 can obtain a residual estimation model 410 including a plurality of camera characteristic correction models 1301-1, 1301-2, 1302-3, . . . and the base model 1302, as shown in FIG.

[0156] [Fourth embodiment] In the above embodiments, an example has been described in which the image processing system 100 is mounted on a vehicle 1 such as an automobile. In the fourth embodiment, an example will be described in which the image processing system 100 is applied to a 3D image display system that displays 3D images on an edge device such as AR goggles.

[0157] 18 is a diagram showing an example of the system configuration of a 3D image display system according to the fourth embodiment. The 3D image display system 1800 includes an edge device 1801 such as AR goggles, and a server 1802 that can communicate with the edge device 1801 via a communication network N such as the Internet or a LAN (Local Area Network).

[0158] The edge device 1801 is equipped with, for example, one or more peripheral cameras, a three-dimensional sensor, a display device, a communication I / F, etc., and transmits images captured by the peripheral cameras and three-dimensional information acquired by the three-dimensional sensor to the server 1802.

[0159] The server 1802 includes one or more computers 300, and executes a predetermined program to generate a three-dimensional image using the captured image and three-dimensional information received from the edge device 1801, and transmits the generated three-dimensional image to the edge device 1801. The server 1802 is an example of a remote processing means.

[0160] The edge device 1801 displays the three-dimensional image received from the server 1802 on the display device, thereby displaying a three-dimensional image of the surroundings.

[0161] However, in conventional systems, there is a problem that after the edge device 1801 transmits the captured image and three-dimensional information to the server 1802, the three-dimensional image cannot be displayed until the three-dimensional image is received from the server 1802.

[0162] Therefore, the edge device 1801 according to this embodiment displays a free viewpoint image generated using, for example, the image processing described with reference to Fig. 5 after transmitting the captured image and the three-dimensional information to the server 1802 and before receiving the three-dimensional image from the server 1802. As a result, according to the three-dimensional image display system 1800 according to this embodiment, the edge device 1801 can display a virtual space before receiving the three-dimensional image from the server 1802.

[0163] <Hardware configuration> 19 is a diagram illustrating an example of the hardware configuration of an edge device according to the fourth embodiment. The edge device 1801 has a computer configuration, and includes, for example, a processor 1901, a memory 1902, a storage device 1903, a communication I / F 1904, a display device 1905, multiple peripheral cameras 1906, an IMU 1907, a three-dimensional sensor 1908, and a bus 1909.

[0164] The processor 1901 is, for example, an arithmetic device such as a CPU or GPU that executes predetermined processing by executing a program stored in a storage medium such as a storage device 1903. The memory 1902 includes, for example, a RAM, which is a volatile memory used as a work area or the like for the processor 1901, and a ROM, which is a nonvolatile memory that stores programs for starting up the processor 1901. The storage device 1903 is, for example, a large-capacity nonvolatile storage device such as an SSD or HDD.

[0165] The communication I / F 1904 is a communication device such as a WAN (Width Area Network) or a LAN (Local Area Network) that connects the edge device 1801 to a communication network N and communicates with the server 1802. The display device 1905 is, for example, a display means such as an LCD or an organic EL. The multiple peripheral cameras 1906 are cameras that capture images of the periphery of the edge device 1801.

[0166] The IMU (Inertial Measurement Unit) 1907 is an inertial measurement device that detects three-dimensional angular velocity and acceleration using, for example, a gyro sensor and an acceleration sensor. The three-dimensional sensor 1908 is a sensor that acquires three-dimensional information, such as a LiDAR, a stereo camera, a depth camera, or a wireless sensing device. The bus 1909 is connected to the above components and transmits, for example, address signals, data signals, and various control signals.

[0167] <Functional configuration> FIG. 20 is a diagram showing the functional configuration of a 3D image display system according to the fourth embodiment.

[0168] (Edge device functional configuration) The processor 1901 executes a predetermined program, and the edge device 1801 has a three-dimensional information acquisition unit 2001, a transmission unit 2002, a reception unit 2003, and the like in addition to the functional configuration of the image processing device 10 described in Fig. 4. The edge device 1801 also has a display control unit 2004 instead of the display control unit 404. Note that the image acquisition unit 401, the residual estimation unit 402, the mapping unit 403, the setting unit 405, and the storage unit 406 are similar to the functional configurations of the image processing device 10 described in Fig. 4, and therefore description thereof will be omitted here.

[0169] The three-dimensional information acquisition unit 2001 acquires three-dimensional information of the periphery of the edge device 1801 using the three-dimensional sensor 1908. The transmission unit 2002 transmits the three-dimensional information acquired by the three-dimensional information acquisition unit 2001 and the multiple captured images acquired by the image acquisition unit 401 to the server 1802.

[0170] The receiving unit 2003 receives the three-dimensional image transmitted by the server 1802 in accordance with the three-dimensional information and the plurality of captured images transmitted by the transmitting unit 2002. The display control unit 2004 displays the free viewpoint image 250 generated by the mapping unit 403 on the display device 16 or the like before the receiving unit 2003 completes reception of the three-dimensional image. Furthermore, the display control unit 2004 displays the received three-dimensional image on the display device 16 or the like after the receiving unit 2003 completes reception of the three-dimensional image.

[0171] (Server functional configuration) The server 1802 executes predetermined programs on one or more computers 300 to realize a receiving unit 2011, a three-dimensional image generating unit 2012, a transmitting unit 2013, and the like.

[0172] The receiving unit 2011 receives, for example, the three-dimensional information and the plurality of captured images transmitted by the edge device 1801 using the communication device 307.

[0173] The three-dimensional image generating unit 2012 uses the three-dimensional information received by the receiving unit 2011 and the plurality of captured images to render the plurality of captured images in a three-dimensional space, thereby generating a three-dimensional image of the periphery of the edge device 1801. Note that in this embodiment, the method for generating the three-dimensional image by the server 1802 may be any method.

[0174] The transmission unit 2013 transmits the three-dimensional image generated by the three-dimensional image generation unit 2012 to the edge device using, for example, the communication device 307.

[0175] <Processing flow> FIG. 21 is a sequence diagram illustrating an example of a three-dimensional image display process according to the fourth embodiment.

[0176] In step S2101, the image acquisition unit 401 of the edge device 1801 acquires a plurality of captured images of the periphery of the edge device 1801 by each of the plurality of peripheral cameras 1906.

[0177] In step S2102, the three-dimensional information acquisition unit 2001 of the edge device 1801 acquires three-dimensional information of one or more three-dimensional objects captured in multiple captured images using the three-dimensional sensor 1908. For example, the three-dimensional information acquisition unit 2001 acquires three-dimensional point cloud information around the edge device 1801.

[0178] In step S2103, the transmitting unit 2002 of the edge device 1801 transmits the plurality of captured images acquired by the image acquiring unit 401 and the three-dimensional information acquired by the three-dimensional information acquiring unit 2001 to the server 1802.

[0179] In step S2104, the three-dimensional image generation unit 2012 of the server 1802 executes a three-dimensional image generation process to generate a three-dimensional image by rendering the multiple captured images in three-dimensional space, using the multiple captured images and three-dimensional information received from the edge device 1801. However, this process takes time, and the processing time may vary depending on the communication state with the edge device 1801, the load on the server 1802, etc.

[0180] 5 in parallel with the processing of step S2104, the edge device 1801 generates a free viewpoint image by mapping a plurality of captured images onto a display projection surface, and displays the image on the display device 1905. Note that this processing can be completed in a shorter time than the 3D image generation processing executed by the server 1802, and is not affected by communication information with the server 1802, the load on the server 1802, etc., and therefore, the image around the edge device 1801 can be displayed in a shorter time.

[0181] In step S2106, when the three-dimensional image generating unit 2012 of the server 1802 completes the generation of the three-dimensional image, the transmitting unit 2013 of the server 1802 transmits the generated three-dimensional image to the edge device 1801.

[0182] In step S2107, upon receiving the three-dimensional image from the server 1802, the display control unit 2004 of the edge device 1801 displays the received three-dimensional image on the display device 1905.

[0183] Through the processing of FIG. 21, the three-dimensional image display system 1800 can display a virtual space after transmitting a plurality of captured images and three-dimensional information to the server 1802 and before receiving a three-dimensional image from the server 1802.

[0184] As described above, according to each embodiment of the present invention, in an image processing system that synthesizes a free viewpoint image using a plurality of captured images, it becomes possible to synthesize a free viewpoint image with little distortion without using a three-dimensional sensing device. [Explanation of symbols]

[0185] 1. Image processing system 10 Image processing device 12, 12A~12D Camera 16 Display device 210 Captured Images 130 Projection surface information 220 Residual information 230 Projection surface information (information about the default projection surface) 231 Default projection plane 240 Viewpoint Information 250 free viewpoint images 300 Computers 401 Image acquisition unit 402 Residual Estimator 403 Mapping Department 404, 2004 Display control unit 410 Residual Estimation Model (Inference Engine) 601 Residual information for training 602 Teacher Images 603 training images 701 Captured image preparation unit 702 Three-dimensional information preparation department 703 Teacher Image Preparation Department 704 Learning Department 405, 705 Settings 710 Residual Learning Model (Neural Network) 901 Uncensored Images 1001 Uncensored Image Preparation Department 1002 Residual calculation section 1301, 1301-1 to 1301-3 Camera characteristic correction model (camera model inference engine) 1302 Base Model (Base Model Inference Engine) 1800 3D image display system 2001 3D Information Acquisition Department 2002 Transmitter 2003 Receiving section

Claims

1. An image processing method for synthesizing a free viewpoint image using a plurality of captured images on a display projection surface, comprising: an image acquisition step of acquiring the plurality of captured images by each of a plurality of cameras; a residual estimation step of inputting the plurality of captured images and viewpoint information and estimating residual information of a projection plane indicating a difference between a predefined bowl-shaped default projection plane and the display projection plane by machine learning; a mapping step of mapping the plurality of captured images onto the display projection surface using information about the default projection surface, the residual information, and the viewpoint information to obtain the free viewpoint image; A computer-implemented image processing method.

2. The residual estimation step includes: The residual information is estimated using an inference engine that has been trained to infer a difference between the default projection plane and the display projection plane from a plurality of learning captured images, the viewpoint information, and three-dimensional information of one or more three-dimensional objects captured in the plurality of learning captured images. The image processing method according to claim 1 .

3. The residual estimation step includes: a camera model inference engine that is trained to infer feature map information of feature points of one or more three-dimensional objects from a plurality of captured images for learning and three-dimensional information of the one or more three-dimensional objects captured in the plurality of captured images for learning; a base model inference engine that is trained to infer a difference between the default projection plane and the display projection plane from the feature map information and the viewpoint information output by the camera model inference engine; and estimating the residual information using The image processing method according to claim 1 .

4. The weight data after learning of the camera model inference engine is Compared with the weight data after learning of the base model inference engine, The characteristic parameters of the plurality of cameras have a large influence. The image processing method according to claim 3 .

5. The camera model inference engine receives characteristic parameters relating to at least one camera among the plurality of cameras.

5. The image processing method according to claim 3 or 4.

6. the camera model inference engine is capable of selecting the camera model inference engine that infers the feature map information from a plurality of candidate camera model inference engines, each candidate having been trained based on different characteristic parameters; 5. The image processing method according to claim 3 or 4.

7. A neural network learning method for inferring residual information of a projection plane that reflects three-dimensional information of one or more three-dimensional objects captured in a plurality of captured images with respect to a predefined bowl-shaped default projection plane based on the plurality of captured images, comprising: a captured image preparation step of preparing the plurality of captured images captured by each of the plurality of cameras; a three-dimensional information preparation step of preparing the three-dimensional information; a teacher image preparation step of restoring a three-dimensional image of the plurality of captured images using the three-dimensional information and generating a teacher image, which is a free viewpoint image for the teacher, based on the input viewpoint information; a learning step of inputting the plurality of captured images and the viewpoint information into the neural network to obtain residual information for learning, mapping the plurality of captured images on a display projection surface using the residual information for learning, information about the default projection surface, and the viewpoint information to generate training images that are free viewpoint images for learning, and training the neural network so that an error between the teacher image and the training image is reduced; A neural network training method in which a computer executes the following.

8. A neural network learning method for inferring residual information of a projection plane that reflects three-dimensional information of one or more three-dimensional objects captured in a plurality of captured images with respect to a predefined bowl-shaped default projection plane based on the plurality of captured images, comprising: an image preparation step of preparing the plurality of captured images by each of a plurality of cameras; an unedited image preparation step of mapping the plurality of captured images onto the predetermined projection plane and generating unedited images that are unedited free viewpoint images based on input viewpoint information; a three-dimensional information preparation step of preparing the three-dimensional information; a teacher image preparation step of restoring a three-dimensional image of the plurality of captured images using the three-dimensional information and generating a teacher image, which is a free viewpoint image for the teacher, based on the input viewpoint information; a residual calculation step of comparing the unmodified free viewpoint image with the training image to prepare the residual information; a learning step of inputting the plurality of captured images and viewpoint information into the neural network and causing the neural network to learn using the residual information prepared in the residual calculation step as training data; A neural network training method in which a computer executes the following.

9. The neural network a camera model inference network to which the plurality of captured images and characteristic parameters related to at least one of the plurality of cameras are input; a base model inference network to which the output of the camera model inference network and the viewpoint information are input; 9. The neural network training method according to claim 7 or 8.

10. A neural network learning method for inferring residual information of a projection plane that reflects three-dimensional information of one or more three-dimensional objects captured in a plurality of captured images with respect to a predefined bowl-shaped default projection plane based on the plurality of captured images, comprising: the neural network is composed of a first camera model inference network or a second camera model inference network and a base model inference network; a first captured image preparation step of preparing a first plurality of captured images by each of a first plurality of cameras; a second captured image preparation step of preparing a second plurality of captured images by each of a second plurality of cameras; a first three-dimensional information preparation step of preparing first three-dimensional information of one or more three-dimensional objects captured in the first plurality of captured images; a second three-dimensional preparation step of preparing second three-dimensional information of one or more three-dimensional objects captured in the second plurality of captured images; a first teacher image preparation step of restoring a three-dimensional image of the first plurality of captured images using the first three-dimensional information and generating a first teacher image, which is a free viewpoint image for a first teacher, based on input viewpoint information; a second teacher image preparation step of restoring a three-dimensional image of the second plurality of captured images using the second three-dimensional information and generating a second teacher image, which is a free viewpoint image for a second teacher, based on the viewpoint information; inputting the first plurality of captured images and characteristic parameters related to at least one camera of the first plurality of cameras into the first camera model inference network; inputting the output of the first camera model inference network and viewpoint information into the base model inference network to obtain first learning residual information, information related to the default projection plane, and the viewpoint information to map the first plurality of captured images onto a display projection plane to generate first learning images that are first free viewpoint images for learning; and simultaneously training both the first camera model inference network and the base model inference network so as to reduce an error between the first teacher image and the first learning image; a learning step of inputting the second plurality of captured images and characteristic parameters related to at least one camera of the second plurality of cameras into the second camera model inference network, mapping the second plurality of captured images onto a display projection plane using second learning residual information obtained by inputting an output of the second camera model inference network and viewpoint information into the trained base model inference network, information related to the default projection plane, and the viewpoint information to generate second learning images which are second learning free viewpoint images, and training the second camera model inference network so as to reduce an error between the second teacher image and the second learning image. How neural networks are trained.

11. A neural network learning method for inferring residual information of a projection plane that reflects three-dimensional position information of one or more three-dimensional objects captured in a plurality of captured images with respect to a predefined bowl-shaped default projection plane based on the plurality of captured images, comprising: the neural network is composed of a first camera model inference network or a second camera model inference network and a base model inference network; a first captured image preparation step of preparing a first plurality of captured images by each of a first plurality of cameras; a second captured image preparation step of preparing a second plurality of captured images by each of a second plurality of cameras; a first unedited image preparation step of mapping the first plurality of captured images onto the default projection plane and generating a first unedited image, which is a first unedited free viewpoint image, based on input viewpoint information; a second unedited image preparation step of mapping the second plurality of captured images onto the default projection plane and generating a second unedited image that is a second unedited free viewpoint image based on input viewpoint information; a first three-dimensional information preparation step of preparing first three-dimensional information of one or more three-dimensional objects captured in the first plurality of captured images; a second three-dimensional information preparation step of preparing second three-dimensional information of one or more three-dimensional objects captured in the second plurality of captured images; a first teacher image preparation step of restoring a three-dimensional image of the first plurality of captured images using the first three-dimensional information and generating a first teacher image, which is a free viewpoint image for a first teacher, based on input viewpoint information; a second teacher image preparation step of restoring a three-dimensional image of the second plurality of captured images using the second three-dimensional information and generating a second teacher image, which is a free viewpoint image for a second teacher, based on the viewpoint information; a first residual calculation step of comparing the first uncorrected image with the first training image to prepare first residual information; a second residual calculation step of comparing the second uncorrected image with the second training image to prepare second residual information; inputting the first plurality of captured images and characteristic parameters related to at least one camera of the first plurality of cameras into the first camera model inference network, inputting an output of the first camera model inference network and viewpoint information into the base model inference network, and simultaneously training both the first camera model inference network and the base model inference network using the first residual information prepared in the first residual calculation step as training data; a learning step of inputting the second plurality of captured images and characteristic parameters related to at least one camera of the second plurality of cameras into the second camera model inference network, inputting an output of the second camera model inference network and viewpoint information into the trained base model inference network, and training the second camera model inference network using the second residual information prepared in the second residual calculation step as training data. How neural networks are trained.

12. A 3D image display method for displaying a free viewpoint image synthesized using a plurality of captured images on a display projection surface, comprising: a captured image acquisition step of acquiring the plurality of captured images by each of a plurality of cameras; a three-dimensional information acquisition step of acquiring three-dimensional information of one or more three-dimensional objects captured in the plurality of captured images; a residual estimation step of inputting the plurality of captured images and viewpoint information and estimating residual information of a projection plane indicating a difference between a predefined bowl-shaped default projection plane and the display projection plane by machine learning; a mapping step of mapping the plurality of captured images onto a display projection surface using information about the default projection surface, the residual information, and the viewpoint information to obtain a free viewpoint image; a transmitting step of transmitting the plurality of captured images and the three-dimensional information to a remote processing means; a receiving step of receiving, from the remote processing means, three-dimensional images of the plurality of captured images reconstructed based on the three-dimensional information in the remote processing means; a display control step of displaying the free viewpoint image on a display means before completing reception of the three-dimensional image in the receiving step, and displaying the three-dimensional image on the display means after completing reception of the three-dimensional image in the receiving step; A computer-implemented, three-dimensional image display method.

13. An image processing system that synthesizes a free viewpoint image using a plurality of captured images on a display projection screen, an image acquisition unit configured to acquire the plurality of captured images by each of a plurality of cameras; a residual estimation unit configured to input the plurality of captured images and viewpoint information, and estimate residual information of a projection plane that indicates a difference between a predefined bowl-shaped default projection plane and the display projection plane by machine learning; a mapping unit configured to map the plurality of captured images onto the display projection surface using information about the default projection surface, the residual information, and the viewpoint information to obtain the free viewpoint image; An image processing system comprising:

14. A neural network learning system that infers residual information of a projection plane that reflects three-dimensional information of one or more three-dimensional objects captured in a plurality of captured images with respect to a predefined bowl-shaped default projection plane based on the plurality of captured images, comprising: a captured image preparation unit configured to prepare the plurality of captured images by each of a plurality of cameras; a three-dimensional information preparation unit configured to prepare the three-dimensional information; a teacher image preparation unit configured to restore a three-dimensional image of the plurality of captured images using the three-dimensional information and to generate a teacher image, which is a free viewpoint image for a teacher, based on input viewpoint information; a learning unit configured to input the plurality of captured images and the viewpoint information into the neural network to acquire learning residual information, map the plurality of captured images onto a display projection surface using the learning residual information, information related to the default projection surface, and the viewpoint information to generate learning images that are free viewpoint images for learning, and train the neural network so that an error between the teacher image and the learning image is reduced; A neural network learning system having:

15. A neural network learning system that infers residual information of a projection plane that reflects three-dimensional information of one or more three-dimensional objects captured in a plurality of captured images with respect to a predefined bowl-shaped default projection plane based on the plurality of captured images, comprising: an image preparation unit configured to prepare the plurality of captured images by each of a plurality of cameras; an unedited image preparation unit configured to map the plurality of captured images to the predetermined projection plane and generate unedited images that are unedited free viewpoint images based on input viewpoint information; a three-dimensional information preparation unit configured to prepare the three-dimensional information; a teacher image preparation unit configured to restore a three-dimensional image of the plurality of captured images using the three-dimensional information and to generate a teacher image, which is a free viewpoint image for a teacher, based on input viewpoint information; a residual calculation unit configured to compare the unmodified free viewpoint image with the training image to prepare the residual information; a learning unit configured to input the plurality of captured images and viewpoint information to the neural network and to cause the neural network to learn using the residual information prepared by the residual calculation unit as training data; A neural network learning system having:

16. A 3D image display system for displaying a free viewpoint image synthesized using a plurality of captured images on a display projection surface, comprising: an image acquisition unit configured to acquire the plurality of captured images by each of a plurality of cameras; a three-dimensional information acquisition unit configured to acquire three-dimensional information of one or more three-dimensional objects captured in the plurality of captured images; a residual estimation unit configured to input the plurality of captured images and viewpoint information, and estimate residual information of a projection plane that indicates a difference between a predefined bowl-shaped default projection plane and the display projection plane by machine learning; a mapping unit configured to map the plurality of captured images onto a display projection surface using information about the default projection surface, the residual information, and the viewpoint information to obtain a free viewpoint image; a transmitting unit configured to transmit the plurality of captured images and the three-dimensional information to a remote processing means; a receiving unit configured to receive, from the remote processing means, three-dimensional images of the plurality of captured images reconstructed based on the three-dimensional information in the remote processing means; a display control unit configured to display the free viewpoint image on a display means before the receiving unit completes reception of the 3D image, and to display the 3D image on the display means after the receiving unit completes reception of the 3D image; A three-dimensional image display system comprising:

Citation Information

Patent Citations

  • Surroundings monitoring apparatus, image processing method, and image processing program

    WO2017122294A1

  • Image processing device

    WO2019053922A1