Image generation device, method, and program

The image generation apparatus addresses the issue of incomplete learning in NeRF by using density and color distribution calculation units with interpolation, ensuring high-quality images from arbitrary viewpoints.

WO2025158662A1PCT designated stage Publication Date: 2025-07-31NT T INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/002490
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-26
Publication Date
2025-07-31

AI Technical Summary

Technical Problem

Existing image generation methods using Neural Radiance Fields (NeRF) produce blurred or cloud-like images when learning has not sufficiently progressed or is incomplete for certain coordinates and directions.

Method used

An image generation apparatus and method that includes a density calculation unit, color distribution calculation unit, training degree calculation unit, and interpolation unit to perform interpolation processing based on the training degree of the model's parameters, ensuring accurate image generation from arbitrary viewpoints even when learning is insufficient.

Benefits of technology

Enables high-quality image generation from arbitrary viewpoints by aggregating density and color distribution calculations, effectively handling incomplete learning scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024002490_31072025_PF_FP_ABST
    Figure JP2024002490_31072025_PF_FP_ABST
Patent Text Reader

Abstract

An image generation device according to one embodiment comprises: a density calculation unit that calculates the density of a drawing space in an image; a color distribution calculation unit that calculates the color distribution of the drawing space; a training degree calculation unit that calculates a training degree indicating the degree of progress of learning of parameters for a model used in calculation performed by the density calculation unit and a model used in calculation performed by the color distribution calculation unit; an interpolation unit that performs an interpolation process on at least one of a calculation result pertaining to the density and a calculation result pertaining to the color distribution in accordance with the magnitude of the training degree calculated by the training degree calculation unit; and a luminance field aggregation unit that aggregates, along coordinates of a line of sight leading to a drawing object in the drawing space, a calculation result produced by the density calculation unit and a calculation result produced by the color distribution calculation unit after the interpolation process performed by the interpolation unit, thereby estimating the color of an image of the drawing object.
Need to check novelty before this filing date? Find Prior Art

Description

Image generation device, method, and program

[0001] FIELD Embodiments of the present invention relate to an image generation device, a method, and a program.

[0002] Neural Radiance Fields (NeRF) is a free viewpoint image generation method that generates an image from an arbitrary viewpoint different from the shooting viewpoint from a group of images captured from multiple viewpoints (see, for example, Non-Patent Document 1).

[0003] NeRF sets the region containing the object to be drawn, called the drawing space, as a rectangular parallelepiped or sphere, and models the luminance field in the drawing space with a neural network (sometimes simply called a neural network or NN). The luminance field has a color distribution c and a density σ.

[0004] Figure 1 shows an example of image generation using NeRF. For example, assume that the rendering space is set to a sphere with center (0,0,0) and radius R. When rendering using NeRF, camera parameters (symbol a in Figure 1) are specified, and an image (symbol b in Figure 1) is calculated from the luminance field for each pixel. The camera parameters include a projection function K, a rotation matrix, and a translation vector. The rotation matrix and translation vector are expressed as (1) and (2) below.

[0005]

[0006] The direction vector (symbol e in FIG. 1) of the line of sight (symbol d in FIG. 1), which is a straight line passing from this translation vector through the pixel coordinate (symbol c in FIG. 1) of the image, is expressed as in (3) and (4) below, and this direction vector can be calculated using (5) and (6) below.

[0007]

[0008] K -1 is the inverse matrix of K. On the above line of sight, the starting point t n and the end point t f By defining the color distribution and density of the luminance field (symbol f in Fig. 1), an estimate of the color of the pixel coordinate can be calculated. Symbol g in Fig. 1 indicates the coordinate where the color is estimated. The estimate of the color of the pixel image is given by (7) below.

[0009]

[0010] The above estimated values ​​are obtained according to the following (8) to (14) and (15): (14) indicates the coordinates of the point on the line of sight.

[0011]

[0012] Here, the above PE and DE are functions that convert the pixel coordinate position and direction vector into a high-dimensional space one-to-one to generate a sharp image. For example, the following functions (16), (17), and (18) are used.

[0013]

[0014] Also, t i , i=1,...,N are sampled to satisfy the following (19), where N is the number of samples. n ≦t 1 ≦…≦t N ≦t f …(19)

[0015] Neural network NN for calculating color distribution of luminance field c , and a neural network NN that calculates the density of the brightness field. σ The parameters are learned from a set of images whose camera parameters are known. The configuration and specific learning method of NeRF are widely known, so a detailed explanation will be omitted here.

[0016] NeRF: Representing Scenes as Neural Radiance Fields for View 0Synthesis”, Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi and Ren Ng,ECCV, 2020

[0017] The above-mentioned conventional technology can generate high-quality images when the neural network is sufficiently trained with the coordinates and directions of the new viewpoint. However, if the training is not sufficiently advanced or if the coordinates and directions that have not been trained are included, blurred images or images containing objects such as clouds or fog may be generated.

[0018] This occurs because computing the color of a pixel aggregates density and color distributions at coordinates and directions that are not well trained or have not been trained.

[0019] This invention has been made in light of the above circumstances, and its purpose is to provide an image generation device, method, and program that are capable of generating appropriate images from any viewpoint, even when the model that calculates the color distribution and density of the image drawing space is not sufficiently trained.

[0020] An image generating device according to one aspect of the present invention comprises a density calculation unit that calculates the density of a drawing space of an image; a color distribution calculation unit that calculates the color distribution of the drawing space; a training level calculation unit that calculates a training level that indicates the degree of progress of learning of the parameters of the model used for calculations by the density calculation unit and the model used for calculations by the color distribution calculation unit; an interpolation unit that performs interpolation processing of at least one of the density calculation results and the color distribution calculation results depending on the level of training calculated by the training level calculation unit; and a luminance field aggregation unit that estimates the color of the image of the drawing object by aggregating the calculation results by the density calculation unit and the calculation results by the color distribution calculation unit after the interpolation processing by the interpolation unit along the coordinates of the line of sight to the drawing object in the drawing space.

[0021] An image generation method according to one aspect of the present invention is a method performed by an image generation device, and includes: a density calculation unit of the image generation device calculating the density of a drawing space of an image; a color distribution calculation unit of the image generation device calculating the color distribution of the drawing space; a training level calculation unit of the image generation device calculating a training level indicating the degree of progress of learning of the parameters of the model used for the calculation by the density calculation unit and the model used for the calculation by the color distribution calculation unit; an interpolation unit of the image generation device performing interpolation processing on at least one of the density calculation results and the color distribution calculation results depending on the level of training level calculated by the training level calculation unit; and a luminance field aggregation unit of the image generation device aggregating the calculation results by the density calculation unit and the calculation results by the color distribution calculation unit after the interpolation processing by the interpolation unit along the coordinates of the line of sight to the drawing object in the drawing space, thereby estimating the color of the image of the drawing object.

[0022] According to the present invention, even when the model for calculating the color distribution and density of the rendering space of the image is not sufficiently trained, it is possible to generate an appropriate image at any viewpoint.

[0023] FIG. 1 is a diagram showing an example of image generation using NeRF. FIG. 2 is a diagram showing an application example of an image generation device according to an embodiment of the present invention. FIG. 3 is a flowchart showing an example of the procedure of processing operations of the image generation device during learning. FIG. 4 is a flowchart showing an example of the procedure of processing operations of the image generation device during drawing. FIG. 5 is a block diagram showing an example of the hardware configuration of an image generation device according to an embodiment of the present invention.

[0024] An embodiment of the present invention will be described below with reference to the drawings. Fig. 2 is a diagram showing an application example of an image generating device according to an embodiment of the present invention.

[0025] As shown in FIG. 2, an image generating device 100 according to one embodiment of the present invention includes an input unit 10, a storage unit 20, a control unit 30, a luminance field aggregation unit 40, a density calculation unit 50, a color distribution calculation unit 60, a training level calculation unit 70, an interpolation unit 80, and an output unit 90.

[0026] The luminance field aggregation unit 40 aggregates the luminance field in the rendering space to determine the color of a pixel. The density calculation unit 50 calculates the density in the rendering space. The color distribution calculation unit 60 calculates the color distribution in the rendering space. The training level calculation unit 70 calculates the training level of the model used to calculate the density and color distribution. The interpolation unit 80 performs interpolation processing on the calculation results of the density and color distribution.

[0027] The training level calculation unit 70 of this practical embodiment can determine whether or not the calculation results of the density and color distribution should be used as is by estimating whether or not the learning of the model parameters used in the calculations by the density calculation unit 50 and the color distribution calculation unit 60 has progressed sufficiently. Furthermore, the interpolation unit 80 can find alternative values ​​when the calculation results of the density and color distribution should not be used as is.

[0028] The functions of each unit of the image generating device 100 will be described below in order. (First embodiment) First, the first embodiment will be described. (Input unit 10) During training of the neural network used to calculate the color distribution and density (hereinafter, sometimes simply referred to as training), the input unit 10 inputs the radius R of the drawing space, the number of samples N, the number of steps S, the image group, and the training loss coefficient (λ P , λ D ), and the training threshold (δ P , δ D ) input is accepted. P is the loss factor related to the training degree of the model used to calculate the density, and λ D is a loss coefficient related to the training degree of the model used to calculate the color distribution, and δ P is the threshold for the training level of the model used to calculate the density, and δ D is a threshold for the training level of the model used to calculate the color distribution.

[0029] The image group is image I m and the camera parameters are expressed as follows (20).

[0030]

[0031] m is expressed as follows: m=1,...,M

[0032] When drawing the image after the learning, the input unit 10 receives input of the camera parameters and the width W and height H of the image to be generated. The camera parameters are expressed as shown in (21) below.

[0033]

[0034] (Storage Unit 20) The storage unit 20 includes a storage device such as a non-volatile memory. The storage device stores the number of samples N, the number of steps S, the image set, and the training loss coefficient (λ P , λ D ), training threshold (δ P , δ D ), and neural network NN σ , N.N. c , N.N. P , and N.N. D The parameters are stored. σ is the neural network used to calculate the density of the brightness field, c is a neural network used to calculate the color distribution of the luminance field. P is a neural network that calculates the training level of a model used in density calculation (hereinafter, sometimes referred to as the training level of density calculation or the training level of density, etc.), and NN D is a neural network that calculates the training level of a model used to calculate a color distribution (hereinafter, this may be referred to as the training level of the color distribution calculation or the training level of the color distribution).

[0035] The image group is image I m and the camera parameters are expressed as follows (22).

[0036]

[0037] (Controller 30) During learning, the controller 30 randomly repeats image selection and pixel selection for the number of steps S. The controller 30 calls the luminance field aggregator 40 using the translation vector, direction vector, and radius of the rendering space related to the selected image as arguments. Here, the direction vector is calculated according to the following (23) and (24).

[0038]

[0039] Furthermore, the control unit 30 calculates the following loss function from the color estimates and the colors of the original image returned from the luminance field aggregating unit 40, and trains the accumulated neural network parameters so as to reduce this loss function. The colors of the original image are expressed as in (25) below, and the formula for calculating the loss function is expressed as in (26) below.

[0040]

[0041] The configuration of the neural network and the specific learning method are generally known, so details will be omitted here. Furthermore, when drawing, the control unit 30 calls the luminance field summarizing unit 40 the number of times equal to the image size (W × H) consisting of the image width W and height H, and generates an image using this luminance field summarizing unit 40.

[0042] (Luminance field summarizing unit 40) The luminance field summarizing unit 40 calculates the start point t of the line of sight sample from the input translation vector, direction vector, and radius R of the rendering space. n and the end point t f is calculated as follows (27) and (28).

[0043]

[0044] Also, t i , i=1,...,N are sampled so that the following (29) is satisfied: n ≦t 1 ≦…≦t N ≦t f …(29)

[0045] During learning, the luminance field summarization unit 40 i, the interpolation unit 80 is called to obtain the density and color distribution values. As a result, the interpolation unit 80 calls the density calculation unit 50 to calculate the density, the color distribution calculation unit 60 to calculate the color distribution, and the training degree calculation unit 70 to update the training degree of the density and the training degree of the color distribution. The calculation results of the density and color distribution are shown as (30) and (31) below.

[0046]

[0047] The luminance field summarizing unit 40 then summarises the obtained density and color distribution values ​​as shown in (32), (33), and (34) below to obtain an estimated color value.

[0048]

[0049] During drawing, the luminance field summarizing unit 40 calculates each t i , the interpolation unit 80 is called to obtain the density and color distribution values. As a result, the interpolation unit 80 calls the training level calculation unit 70 to calculate the training level of the density and the training level of the color distribution, and performs calculation or interpolation of the density and calculation or interpolation of the color distribution according to the calculated training levels.

[0050] The luminance field summarizing unit 40 then summarises the density and color distribution values ​​in the same way as during learning, to obtain an estimated color value.

[0051] (Density Calculation Unit 50) The density calculation unit 50 calculates the density according to the following (35) and (36).

[0052]

[0053] In the first embodiment, the following functions (37), (38), and (39), i.e., the above functions (16), (17), and (18), are used to transform the pixel coordinate position and direction vector into a high-dimensional space.

[0054]

[0055] (Color Distribution Calculation Unit 60) The color distribution calculation unit 60 calculates the color distribution as shown in the following (40) and (41).

[0056]

[0057] (Training Degree Calculation Unit 70) During learning and drawing, the training degree calculation unit 70 calculates the density training degree and the color distribution training degree shown in the following (42) in accordance with the following (43), (44), (45), and (46) and the above (41). (45) is a loss function used to calculate the density training degree, and (46) is a loss function used to calculate the color distribution training degree.

[0058]

[0059] Here, PKF and DKF are known functions and are used to calculate the training degree of the neural network NN. P and N.N. D These parameters are used to determine whether or not each parameter has been learned. If this learning has progressed sufficiently, the training level indicated by (43) and (44) above approaches 1; if not, the training level approaches 0. Furthermore, the above (43) and (44) are assumed to be (41).

[0060] The training degree calculation unit 70 calculates the density and color distribution using the neural network NN σ and N.N. c The neural network NN for training degree calculation is trained using the position and direction vectors shown in (47) below. P and N.N. D The parameters of are also learned.

[0061]

[0062] Therefore, as the learning of the neural network for training level calculation progresses, it is considered that the learning of the neural network for density and color distribution also progresses.

[0063] In the first embodiment, functions that return position and direction vectors to PKF and DKF are used, as shown in (48) and (49) below.

[0064]

[0065] During learning, the training level calculation unit 70 updates the training level. The training level is updated by using a loss function (lossP , loss D ) is small, the neural network NN used to calculate the training degree P and N.N. D Each of these is trained.

[0066] (Interpolation unit 80) During learning, the interpolation unit 80 calls the density calculation unit 50 to calculate the density, calls the color distribution calculation unit 60 to calculate the color distribution, and calls the training level calculation unit 70 to update the training level of the density and the training level of the color distribution.

[0067] When drawing, if the density at the pixel coordinate position shown in (50) below is highly trained, the interpolation unit 80 calls the density calculation unit 50 and calculates the density at that position as shown in (51) and (52) below.

[0068] Furthermore, if the density at the position is not sufficiently trained, the interpolation unit 80 does not perform density calculation at that position but performs interpolation.

[0069]

[0070] In the first embodiment, the density interpolation function is set to be identically zero as shown in the following (53).

[0071]

[0072] When drawing, if the color distribution at the pixel coordinate position and direction vector shown in (47) above is highly trained, the interpolation unit 80 further calls the color distribution calculation unit 60 and calculates the color distribution at that position and direction vector as shown in (55) and (56) below.

[0073] Furthermore, if the degree of training of the color distribution is low, the interpolation unit 80 does not perform color distribution calculation at that position and direction vector, but performs interpolation.

[0074]

[0075] Here, the color distribution interpolation function erp cis the weighted average of the color distribution in the vicinity of the pixel coordinate position and direction vector shown in (47) above, weighted by the orientation and training level, as shown in (56), (57), and (58) below. In the first embodiment, (56) and (57) above are assumed to be (59) below.

[0076]

[0077] Furthermore, when j is 1 to 9, the value of (59) above (sometimes referred to as value q) is expressed as (60) to (68) below.

[0078]

[0079] Furthermore, when j is 10 to 20, the value q in the above (59) is expressed as in the following (69) to (79).

[0080]

[0081] Here, it is assumed that the following (80) is true.

[0082]

[0083] The value q in (59) above may be a randomly specified number of points on the unit sphere instead of the vertices of a regular dodecahedron. The interpolation unit 80 returns density and color distribution values ​​both during learning and during drawing.

[0084] (Output Unit 90) The output unit 90 outputs the image generated during drawing.

[0085] (Operations During Learning) Next, operations during learning in the first embodiment will be described with reference to Fig. 3. Fig. 3 is a flowchart showing an example of the procedure of processing operations of the image generating device during learning. First, the input unit 10 accepts input such as a group of images and stores the input results in the storage unit 20 (S11). If there are remaining steps (Yes in S12), the control unit 30 randomly selects an image (S13) and randomly selects pixels from the selected image (S14).

[0086] Furthermore, the control unit 30 calls the luminance field summarizing unit 40 with the image selected in S13, the camera parameters of this image, that is, the projection function, rotation matrix, and translation vector, and the pixel position selected in S14 as input.

[0087] Next, the luminance field summarizing unit 40 generates gaze samples based on the given gaze (S15). Next, if there are remaining samples (Yes in S16), the luminance field summarizing unit 40 calls the interpolation unit 80 with the coordinates of the gaze samples generated in S15 as input.

[0088] The interpolation unit 80 calls the density calculation unit 50 to calculate the density, and calls the training level calculation unit 70 to update the neural network parameters for the training level of the density (S17, S18). The interpolation unit 80 also calls the color distribution calculation unit 60 to calculate an estimate of the color distribution, and calls the training level calculation unit 70 to update the neural network parameters for the training level of the color distribution (S19, S20).

[0089] The luminance field summarizing unit 40 repeats steps S17 to S20 until there are no more samples remaining. When there are no more samples remaining (No in S16), it aggregates the results of the processing by the interpolation unit 80 and calculates an estimated value of the pixel color (S21). Finally, the control unit 30 calculates the difference between the original value of the pixel color and the estimated value calculated in S21, and updates the parameters of the neural network for density calculation and the neural network for color distribution calculation so as to reduce this difference (S22). The processing from S13 to S22 is repeated until there are no more steps remaining. When the repetition of the number of steps has finished (No in S12), the control unit 30 stores the learned parameters of the various neural networks in the storage unit 20, and the learning process ends.

[0090] In the first embodiment, the neural networks for density calculation and training are trained on the same position vector, and the neural networks for color distribution calculation and training are trained on the same position and direction vectors. Therefore, the training levels of the neural networks for density and color distribution calculation can be estimated by examining the loss of the training level of the neural network.

[0091] (Operations during drawing) Next, operations during image drawing in the first embodiment will be described with reference to Fig. 4. Fig. 4 is a flowchart showing an example of the procedure of processing operations of the image generating device during drawing. First, the input unit 10 inputs camera parameters and the width W and height H of the image to be generated (S31).

[0092] When the control unit 30 has selected all pixels (width W x height H) of the image, i.e., when there are pixels remaining (Yes in S32), it sequentially selects pixel coordinates and calls the luminance field aggregation unit 40 with the camera parameters, i.e., the projection function, rotation matrix, and translation vector, as input.

[0093] The luminance field summarizing unit 40 generates gaze samples based on the given gaze (S33). Next, if there are remaining samples (Yes in S34), the luminance field summarizing unit 40 calls the interpolation unit 80 using the coordinates of the gaze samples generated in S33 as input.

[0094] The interpolation unit 80 calls the training level calculation unit 70 to calculate the training level of the density and color distribution neural network, and determines whether to use the density and color distribution calculation values ​​or perform interpolation depending on the calculated training level.

[0095] In detail, the interpolation unit 80 calls the training level calculation unit 70 to calculate the training level of the density neural network (S35), and if the calculated training level is high (Yes in S36), it calls the density calculation unit 50 to calculate the density (S37), and if the calculated training level is low (No in S36), it performs interpolation processing of the density calculation result (S38).

[0096] After S37 or S38, the interpolation unit 80 calls the training level calculation unit 70 to calculate the training level of the color distribution neural network (S39). If the calculated training level is high (Yes in S40), the interpolation unit 80 calls the color distribution calculation unit 60 to calculate the color distribution (S41). If the calculated training level is low (No in S40), the interpolation unit 80 performs an interpolation process on the color distribution calculation results (S42).

[0097] The luminance field aggregation unit 40 repeats the processes from S35 to S42 until there are no more samples remaining, and when there are no more samples remaining (No in S34), it aggregates the processing results from the interpolation unit 80 and calculates an estimated value of the pixel color (S43).

[0098] The processes from S32 to S43 are repeated until there are no more pixels remaining. When there are no more pixels remaining (No in S32), the control unit 30 generates an image based on the pixel positions and the estimated color values ​​at these positions. The output unit 90 outputs the generated image. This completes the drawing process.

[0099] Second Embodiment Next, a second embodiment will be described. In the following embodiments, detailed descriptions of the same parts as those in the first embodiment may be omitted. In this second embodiment, the color distribution interpolation function erp c The value w related to the calculation is calculated as follows (81), (82), and (83).

[0100]

[0101] Third Embodiment In a third embodiment, the training level calculation unit 70 calculates the color distribution interpolation function erp c is calculated by selecting the point that maximizes the product of the orientation and the training degree, as shown in (84) to (88) below.

[0102]

[0103] Fourth Embodiment Because neural networks are initialized randomly, values ​​close to the arguments may change even before training. Therefore, in the fourth embodiment, the training level calculation unit 70 uses functions that repeat the position and direction n times for the PKF and DKF, as shown in the following (89) and (90), so that the training level value when the neural network has not yet trained differs more from the value when training has been performed.

[0104]

[0105] Fifth Embodiment For the same reasons as those described in the fourth embodiment, in the fifth embodiment, the functions PE and DE that transform position and direction vectors into a high-dimensional space are used in the PKF and DKF as shown in (91) and (92) below, so that the training level value when not yet learned becomes more different from the value when learning has been performed.

[0106]

[0107] Sixth Embodiment In the sixth embodiment, the training level calculation unit 70 calculates the density interpolation function described in the first embodiment by selecting points that maximize the product of the position vector and the training level, as shown in (93), (94), and (95) below. 1 is a positive integer, and λ 1 and λ 2 is a constant greater than or equal to zero.

[0108]

[0109] Seventh Embodiment In the seventh embodiment, the training level calculation unit 70 calculates the training level by using the training level calculation unit 70 described in the sixth embodiment. P is calculated as follows (96) and (97): 1 and λ 2 is a constant greater than or equal to zero.

[0110]

[0111] Eighth Embodiment In the eighth embodiment, the training level calculation unit 70 uses the following functions (98), (99), (100), and (101) as the density interpolation functions described in the first embodiment. N 1 is a positive integer, and λ 1 and λ 2 is a constant greater than or equal to zero.

[0112]

[0113] Ninth Embodiment In the ninth embodiment, the functions PE and DE that convert a position and direction vector into a high-dimensional space on a one-to-one basis use the following functions (102) and (103) on a one-to-one basis.

[0114]

[0115] The "MultiresolutionHashEncoding" shown in (102) above is described in detail in the following document: "Instant Neural Graphics Primitives with a Multiresolution Hash Encoding", Thomas Muller, Alex Evans, Christoph Schied, Alexander Keller, SIGGRAPH, 2022" The "SphericalHarmonicFunction" shown in (103) above is described in detail in the following document: "Ref-NeRF: Structured View-Dependent Appearance for Neural Radiance Fields", Dor Verbin, Peter Hedman, Ben Mildenhall, Todd Zickler, Jonathan T. Barron, Pratul P. Srinivasan, CVPR 2022" (Effects of Each Embodiment) As described above, each of the above embodiments can generate a high-resolution free-viewpoint image even when coordinates and directions for which learning has not progressed sufficiently or has not been performed are included.

[0116] Fig. 5 is a block diagram showing an example of the hardware configuration of an image generating apparatus according to an embodiment of the present invention. In the example shown in Fig. 5, the image generating apparatus 100 according to the embodiment is configured, for example, as a server computer or a personal computer, and has a hardware processor 111A such as a CPU. A program memory 111B, a data memory 112, an input / output interface 113, and a communication interface 114 are connected to this hardware processor 111A via a bus 115.

[0117] The communication interface 114 includes, for example, one or more wireless communication interface units, and enables transmission and reception of information to and from a communication network. As the wireless interface, for example, an interface that adopts a low-power wireless data communication standard such as a wireless LAN (Local Area Network) is used.

[0118] An input device 200 and an output device 300, which are attached to the image generating device 100 and used by a user or the like, are connected to the input / output interface 113. The input / output interface 113 can take in operation data input by a user or the like via the input device 200, such as a keyboard, touch panel, touchpad, or mouse, and can output output data to an output device 300, which includes a display device using a liquid crystal or organic electroluminescence (EL) display, for display. The input device 200 and the output device 300 may be devices built into the image generating device 100, or may be input devices and output devices of other information terminals that can communicate with the image generating device 100 via a network.

[0119] The program memory 111B is a non-transitory tangible storage medium that is a combination of a non-volatile memory that can be written to and read from at any time, such as a hard disk drive (HDD) or a solid state drive (SSD), and a non-volatile memory such as a read only memory (ROM), and can store programs necessary to execute various control processes, etc., according to one embodiment.

[0120] The data memory 112 is a tangible storage medium that is, for example, a combination of the above-mentioned nonvolatile memory and a volatile memory such as RAM (Random Access Memory), and can be used to store various data or information acquired and created during various processes.

[0121] The image generating device 100 according to one embodiment of the present invention can be configured as a data processing device having the units shown in FIG. 2 as software-based processing functional units.

[0122] The information storage unit used as a work memory or the like by each unit of image generating device 100 can be configured using data memory 112 shown in Fig. 5. However, these configured storage areas are not essential components within image generating device 100, and may be areas provided in, for example, an external storage medium such as a USB (Universal Serial Bus) memory, or a storage device such as a database server located in the cloud.

[0123] The processing function units in each of the above units can be realized by reading and executing a program stored in the program memory 111B by the hardware processor 111A. Note that some or all of these processing function units may be realized in various other forms, including integrated circuits such as an application specific integrated circuit (ASIC) or a field-programmable gate array (FPGA).

[0124] The methods described in each embodiment may be stored as a program (software means) that can be executed by a computer on a recording medium such as a magnetic disk (e.g., a floppy disk, a hard disk, etc.), an optical disk (e.g., a CD-ROM, a DVD, an MO, etc.), or a semiconductor memory (e.g., a ROM, a RAM, a flash memory, etc.), or may be transmitted and distributed via a communication medium. The program stored on the medium also includes a configuration program that configures the software means (including not only executable programs but also tables and data structures) that the computer executes. The computer that realizes this device reads the program stored on the recording medium and, in some cases, configures the software means using the configuration program, and executes the above-described processing by having the operation controlled by this software means. The term "recording medium" as used herein is not limited to a storage medium for distribution, but also includes a storage medium such as a magnetic disk or semiconductor memory installed inside the computer or in a device connected via a network.

[0125] The present invention is not limited to the above-described embodiments, and various modifications can be made in the implementation stage without departing from the spirit of the invention. Furthermore, the embodiments may be implemented in appropriate combinations, in which case the combined effects can be obtained. Furthermore, the above-described embodiments include various inventions, and various inventions can be extracted by combining selected elements from the disclosed elements. For example, if the problem can be solved and the desired effect can be obtained even if some elements are deleted from all elements shown in the embodiments, the configuration from which these elements are deleted can be extracted as an invention.

[0126] REFERENCE SIGNS LIST 100... image generating device 10... input unit 20... storage unit 30... control unit 40... luminance field summarizing unit 50... density calculation unit 60... color distribution calculation unit 70... training level calculation unit 80... interpolation unit 90... output unit

Claims

1. A density calculation unit that calculates the density of the drawing space of an image, a color distribution calculation unit that calculates the color distribution of the drawing space, a training degree calculation unit that calculates the degree of progress of learning of the parameters of the model used in the calculation by the density calculation unit and the model used in the calculation by the color distribution calculation unit, an interpolation unit that performs interpolation processing on at least one of the calculation result of the density and the calculation result of the color distribution according to the magnitude of the training degree calculated by the training degree calculation unit, and a luminance field aggregation unit that estimates the color of the image of the drawing target by aggregating the calculation result by the density calculation unit and the calculation result by the color distribution calculation unit after the interpolation processing by the interpolation unit along the coordinates of the line of sight to the drawing target in the drawing space. An image generation device comprising:

2. An input unit that receives an input of an image used for learning the parameters of the model used in the calculation by the density calculation unit, the model used in the calculation by the color distribution calculation unit, and the model used in the calculation by the training degree calculation unit, a first learning unit that learns the parameters of the model used in the calculation by the training degree calculation unit based on the coordinates of the line of sight to the drawing target in the drawing space, and a second learning unit that learns the parameters of the model used in the calculation by the density calculation unit and the model used in the calculation by the color distribution calculation unit based on the difference between the estimated color and the color of the image used for learning. The image generation device according to claim 1, further comprising:

3. A method performed by an image generation device, comprising: calculating, by a density calculation unit of the image generation device, the density of a drawing space of an image; calculating, by a color distribution calculation unit of the image generation device, the color distribution of the drawing space; calculating, by a training degree calculation unit of the image generation device, a training degree indicating the degree of progress of learning of parameters of a model used in the calculation by the density calculation unit and a model used in the calculation by the color distribution calculation unit; performing, by an interpolation unit of the image generation device, interpolation processing on at least one of the calculation result of the density and the calculation result of the color distribution according to the magnitude of the training degree calculated by the training degree calculation unit; and estimating the color of an image of the drawing object by aggregating, by a luminance field aggregation unit of the image generation device, the calculation result by the density calculation unit and the calculation result by the color distribution calculation unit after the interpolation processing by the interpolation unit along the coordinates of the line of sight to the drawing object in the drawing space. An image generation method comprising the steps of:

4. An image generation processing program that causes a processor to function as each unit of the image generation device according to claim 1 or 2.

Citation Information

Patent Citations

  • Deformable neural radiance fields

    JP2023549821A

  • Rendering new images of scenes using geometry-aware neural networks conditioned on latent variables

    WO2022167602A2