Information processing device
By performing learning and quality improvement processes on rendering images generated from two-dimensional images captured from multiple viewpoints, the method addresses inconsistent quality issues in conventional methods, enhancing the quality and reducing noise in three-dimensional image generation.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-26
- Publication Date
- 2026-04-02
AI Technical Summary
Conventional methods for generating three-dimensional images from two-dimensional images captured from multiple viewpoints suffer from inconsistent quality improvement, leading to blurring and noise due to noise during imaging, which affects the consistency and quality of the generated three-dimensional images.
Perform learning based on Gaussian Splatting using two-dimensional images before quality improvement processing, generate rendering images, and then apply quality improvement processes to these rendering images, repeatedly updating the learning model to ensure image consistency and reduce noise.
The proposed method improves the quality of three-dimensional images by maintaining image consistency and reducing blurring and noise, resulting in higher-quality three-dimensional image generation.
Smart Images

Figure JP2024034316_02042026_PF_FP_ABST
Abstract
Description
Information processing apparatus
[0001] The present invention relates to a technique for generating a three-dimensional image using two-dimensional images captured from a plurality of viewpoints.
[0002] As one technique for generating a three-dimensional image from two-dimensional images captured from a plurality of viewpoints (imaging positions), Gaussian Splatting is known. Gaussian Splatting represents a three-dimensional space using a point cloud that spreads in a Gaussian shape, and is realized by learning the position, color, diffusion direction, scale, and transparency of the point cloud using two-dimensional image data as teacher data. For example, in Patent Document 1, when a self-propelled inspection robot stops at an inspection target imaging location and the camera is directed at the inspection target, the imaging difficulty is evaluated based on the amount of the three-dimensional point cloud in front of the inspection target among the three-dimensional point clouds that fit within the imaging angle.
[0003] Japanese Patent Application Laid-Open No. 2023-72514
[0004] In Gaussian Splatting, the quality of the generated three-dimensional image depends on the quality of the two-dimensional images input to the learning model. Therefore, for example, it is conceivable to perform quality improvement processing such as super-resolution processing on the two-dimensional images before inputting them to the learning model and then input them to the learning model.
[0005] However, the two-dimensional images input to the learning model may be affected by noise generated during imaging. Due to this influence, in the two-dimensional images after quality improvement processing, the consistency of quality improvement between a plurality of viewpoints for the same subject cannot be maintained, and the quality of the image deteriorates, such as blurring and noise appearing in the generated three-dimensional image.
[0006] Therefore, an object of the present invention is to improve the quality of a three-dimensional image generated using two-dimensional images captured from a plurality of viewpoints.
[0007] To solve the above problems, the present invention provides an information processing device comprising: a generation unit that generates rendering image data showing two-dimensional images taken from multiple viewpoints using a learning model generated through learning to generate three-dimensional images using two-dimensional images taken from multiple viewpoints; a quality improvement processing unit that performs quality improvement processing on the generated rendering image data to improve the image quality; and a learning unit that updates the learning model by repeatedly performing the learning using the rendering image data after the quality improvement processing.
[0008] According to the present invention, the quality of a three-dimensional image generated using two-dimensional images taken from multiple viewpoints can be improved.
[0009] This is a block diagram showing an example of the hardware configuration of an information processing device 10 according to one embodiment of the present invention. This is a diagram illustrating an example of a conventional quality improvement method. This is a diagram illustrating the reason why the quality of a 3D image generated through a conventional quality improvement method deteriorates. This is a diagram illustrating an example of the quality improvement method in this embodiment. This is a diagram illustrating the reason why the quality of a 3D image generated through the quality improvement method in this embodiment improves. This is a block diagram showing an example of the functional configuration of the information processing device 10. This is a flowchart showing an example of the operation of the information processing device 10.
[0010] [Embodiment] [Configuration] An embodiment of the present invention will be described below. Figure 1 is a diagram showing the hardware configuration of an information processing device 10 according to an embodiment of the present invention. Physically, the information processing device 10 is configured as a computer including a processor 1001, memory 1002, storage 1003, communication device 1004, and a bus connecting these. Each of these devices operates on power supplied from a battery (not shown). In the following description, the word "device" can be read as a circuit, device, unit, etc. The hardware configuration of the information processing device 10 may be configured to include one or more of the devices shown in Figure 1, or it may be configured to omit some of the devices. Alternatively, multiple devices with different housings may be connected to each other via communication to constitute the information processing device 10.
[0011] Each function in the information processing device 10 is realized by loading predetermined software (programs) onto hardware such as the processor 1001 and memory 1002, which allows the processor 1001 to perform calculations, control communication by the communication device 1004, and control at least one of data reading and writing in the memory 1002 and storage 1003.
[0012] The processor 1001 controls the entire computer, for example, by running an operating system. The processor 1001 may consist of a central processing unit (CPU) that includes interfaces with peripheral devices, control units, arithmetic units, registers, and so on.
[0013] The processor 1001 reads programs (program code), software modules, data, etc., from at least one of the storage 1003 and the communication device 1004 into the memory 1002 and executes various processes accordingly. The program used is one that causes the computer to execute at least a part of the operations described later. Functional blocks of the information processing device 10 may be stored in the memory 1002 and implemented by control programs that run on the processor 1001. Various processes may be executed by one processor 1001, but may also be executed simultaneously or sequentially by two or more processors 1001. The processor 1001 may be implemented by one or more chips. The program may also be transmitted to the information processing device 10 via a telecommunications line.
[0014] The memory 1002 is a computer-readable recording medium and may consist of at least one of the following: ROM (Read Only Memory), EPROM (Erasable Programmable ROM), EEPROM (Electrically Erasable Programmable ROM), RAM (Random Access Memory), etc. The memory 1002 may also be called a register, cache, main memory, etc. The memory 1002 can store executable programs (program code), software modules, etc., for carrying out the method according to this embodiment.
[0015] The storage 1003 is a computer-readable recording medium and may consist of at least one of the following: an optical disc such as a CD-ROM (Compact Disc ROM), a hard disk drive, a flexible disk, a magneto-optical disk (e.g., a compact disc, a digital multipurpose disc, a Blu-ray® disc), a smart card, flash memory (e.g., a card, a stick, a key drive), a floppy® disk, a magnetic strip, etc. The storage 1003 may also be called an auxiliary storage device.
[0016] The communication device 1004 is hardware (transmitting / receiving device) for communication between computers, and is also called a network device, network controller, network card, or communication module. Data showing two-dimensional images captured from multiple viewpoints (shooting positions) by a shooting device (not shown) is input to the information processing device 10 via the communication device 1004.
[0017] Each device, such as the processor 1001 and the memory 1002, is connected by a bus for communicating information. The bus may be configured using a single bus, or different buses may be used for each device.
[0018] The information processing device 10 may include hardware such as a microprocessor, a digital signal processor (DSP), an ASIC (Application Specific Integrated Circuit), a PLD (Programmable Logic Device), and an FPGA (Field Programmable Gate Array), and some or all of each functional block may be realized by this hardware. For example, the processor 1001 may be implemented using at least one of these hardware components.
[0019] Here, Figure 2 illustrates an example of a conventional quality improvement method, and Figure 3 illustrates the reason why the quality of 3D images generated through conventional quality improvement methods deteriorates.
[0020] In Figure 2, images i1 and i2 are images of the same subject (in this example, an apple is shown) taken from multiple different viewpoints (viewpoint 1 and viewpoint 2). Although these images i1 and i2 are of the same subject, they contain minute noise caused by various factors during shooting.
[0021] Therefore, to improve the quality of the captured images, quality improvement processes such as well-known super-resolution processing and noise reduction processing are performed. As a result, the captured images i1 and i2 become improved quality images I1 and I2, respectively.
[0022] Then, the data of the quality-improved images after the quality improvement process is used as training data, and learning is performed using a technique called Gaussian Splatting (represented as GS in the figure). Note that in Figure 2, only the captured images i1 and i2 and the quality-improved images I1 and I2 corresponding to viewpoints 1 and 2 are shown, but in reality, learning is performed using a large number of images obtained by capturing images from a large number of viewpoints.
[0023] By inputting the data of images captured from viewpoints 1 and 2 into the learned model obtained through this learning process, 3D image data (referred to as a 3D model image in the diagram) can be obtained.
[0024] In Figure 3, captured image i01 is a representation of a portion of captured image i1 in Figure 2 using a 4x4 pixel array, and captured image i02 is a representation of a portion of captured image i2 in Figure 2 (the same portion as the portion of captured image i1) using a 4x4 pixel array. Captured images i01 and i02 capture the same portion of the subject, but due to the effects of noise during shooting, their pixel values differ slightly when compared at the pixel level. In Figure 3, the difference in pixel values is represented by the difference in the color of each pixel.
[0025] When image quality improvement processing is applied to these captured images i01 and i02, the pixel values of the improved images I01 and I02 will also differ due to the difference in pixel values as described above. In other words, there is no image consistency between improved image I01 and improved image I02, even though they depict the same part of the subject.
[0026] As a result of performing training based on Gaussian Splatting using the inconsistent quality improvement images I01 and I02, the 3D model image D generated by the trained model may contain blur and noise, resulting in an image that differs significantly from the captured images i01 and i02.
[0027] Thus, with conventional methods, the quality improvement processing performed to improve the quality of captured images does not guarantee image consistency between viewpoints, resulting in blurring and noise in the 3D images generated by the learning model.
[0028] In contrast, Figure 4 is a diagram illustrating an example of a quality improvement method in this embodiment, and Figure 5 is a diagram illustrating the reason why the quality of the 3D image generated by the quality improvement method in this embodiment is improved. In Figure 4, the captured images i1 and i2 are images of the same subject taken from multiple different viewpoints (viewpoint 1, viewpoint 2), similar to Figure 2. As mentioned above, these captured images i1 and i2 are of the same subject, but they contain minute noise caused by various factors during shooting.
[0029] In this embodiment, instead of performing image quality improvement processing first as in conventional methods, the captured images i1 and i2 (2D images) before quality improvement processing are used as training data, and learning based on Gaussian Splatting is performed.
[0030] Next, using the learned model obtained through the above learning process, rendering images r1 and r2 are generated, showing two-dimensional images viewed from multiple viewpoints (viewpoints 1 and 2).
[0031] Then, quality improvement processes such as well-known super-resolution processing and noise reduction processing are performed on these rendered images r1 and r2. The resulting rendered images rh1 and rh2 are used as training data, and training based on Gaussian Splatting is performed again. In other words, the trained model is updated based on Gaussian Splatting.
[0032] In Figure 5, rendering image r01 is a representation of a portion of rendering image r1 in Figure 4 using a 4x4 pixel array, and rendering image r02 is a representation of a portion of rendering image r2 in Figure 4 (the same portion as the portion of rendering image r1) using a 4x4 pixel array. Since rendering image r01 and rendering image r02 are images rendered using the same learning model, when compared pixel by pixel, the pixel values of pixels corresponding to the same part of the subject are the same. Note that in Figure 5, the difference in pixel values is represented by the difference in the color of each pixel.
[0033] When rendering images r01 and r02 are subjected to quality improvement processing, the pixel values are the same as described above, and therefore the pixel values of the rendering quality-improved images rh01 and rh02, which have undergone quality improvement processing, are also the same. In other words, it can be said that there is image consistency between rendering quality-improved image rh01 and rendering quality-improved image rh02.
[0034] As a result of performing training based on Gaussian Splatting using the rendering quality improvement images rh01 and rh02, which have consistent image quality, the 3D model image D generated by the trained model no longer contains the blurring and noise described in Figure 3.
[0035] In this embodiment, the learning model is updated one or more times by repeatedly performing the set of learning and quality improvement processes described above.
[0036] Next, Figure 6 is a block diagram showing the functional configuration of the information processing device 10. In the information processing device 10, the processor 1001 reads programs and the like from the storage 1003 into the memory 1002 and executes them, thereby realizing the functions of the acquisition unit 11, the storage unit 12, the learning unit 13, the generation unit 14, the quality improvement processing unit 15, and the output unit 16.
[0037] The acquisition unit 11 acquires data showing captured images (2D images) taken from multiple viewpoints using a photography device (not shown).
[0038] The learning unit 13 generates a learning model by performing learning based on Gaussian Splatting using the captured image data (2D image) as training data. The generated learning model is stored in the storage unit 12.
[0039] The generation unit 14 uses the above-mentioned learning model to render two-dimensional images viewed from multiple viewpoints and generates rendered image data representing those two-dimensional images.
[0040] The quality improvement processing unit 15 performs quality improvement processing on the rendered image data generated by the generation unit 14 to improve the image quality. This quality improvement processing includes well-known quality improvement processing such as image super-resolution processing or image noise reduction processing. The learning unit 13 then repeatedly performs learning based on Gaussian Splatting using the rendered image data after the quality improvement processing to update the learning model stored in the storage unit 12. By repeatedly performing the above set of learning processing and quality improvement processing, the learning model is updated one or more times.
[0041] The output unit 16 outputs a three-dimensional image generated by the learning model as viewed from an arbitrary viewpoint.
[0042] [Operation] Next, the operation of this embodiment will be described. FIG. 7 is a flowchart showing an example of the operation of the information processing apparatus 10. In FIG. 7, the acquisition unit 11 acquires data indicating captured images (two-dimensional images) captured from a plurality of viewpoints by a capturing device not shown (step S11).
[0043] Next, the learning unit 13 performs learning based on Gaussian Splatting using the data of the captured images (two-dimensional images) as teacher data to generate a learning model (step S12).
[0044] Next, the generation unit 14 uses the above learning model to render two-dimensional images viewed from a plurality of viewpoints and generates rendering image data indicating the two-dimensional images (step S13).
[0045] The quality improvement processing unit 15 performs quality improvement processing to improve the quality of the rendering image data generated by the generation unit 14 (step S14).
[0046] Then, the learning unit 13 uses the rendering image data after the quality improvement processing to perform learning based on Gaussian Splatting again, and if it is necessary to update the learning model stored in the storage unit 12 (step S15; YES), learning is performed again using the rendering image data after the quality improvement processing (step S12). By repeatedly performing the processing from the above learning processing to the quality improvement processing (steps S12 to S14), the learning model is updated one or more times. The number of times of repeating the processing from this learning processing to the quality improvement processing (steps S12 to S14) can be arbitrarily determined.
[0047] When the update of the learning model is completed (step S15; NO), the output unit 16 outputs a three-dimensional image generated by the learning model as viewed from the specified arbitrary viewpoint (step S16).
[0048] According to the embodiment described above, instead of first performing the image quality improvement process, by performing learning based on Gaussian Splatting using the data of the two-dimensional image before passing through the quality improvement process as teacher data, it becomes possible to improve the quality of the three-dimensional image generated using the two-dimensional images taken from a plurality of viewpoints.
[0049] [Modification Example] The present invention is not limited to the above-described embodiment. The above-described embodiment may be modified as follows. Also, two or more of the following modification examples may be combined and implemented.
[0050] [Modification Example 1] Instead of performing the quality improvement process on all of the captured images, a quality improvement process may be performed by selecting a partial area of the captured image. That is, the quality improvement unit 15 may select a partial image area from the generated rendering image data and perform the quality improvement process. This contributes to speeding up the quality improvement process.
[0051] For example, if the quality improvement process is super-resolution processing, an area corresponding to a subject that is relatively far away in the captured image is selected. For example, since the rendering image data generated by the generation unit 14 includes depth data corresponding to the depth in the three-dimensional space, in that case, the quality improvement unit 15 selects a partial image area that satisfies the condition that the depth data is far from the viewpoint at the time of shooting in the three-dimensional space (for example, the condition that the distance from the viewpoint is greater than or equal to a threshold value) and performs the quality improvement process. This is because the resolution of a subject that is far from the viewpoint at the time of shooting is more degraded and thus needs to be preferentially improved. By thus selecting a partial image area from the rendering image data and performing the quality improvement process, it contributes to speeding up the quality improvement process.
[0052] Also, it is generally known that an image area with many high-frequency components in the captured image has a lot of noise. Therefore, the quality improvement unit 15 may perform frequency conversion processing (for example, two-dimensional Fourier transform) on the generated rendering image data and select a partial image area that includes a determined high-frequency component greater than or equal to a threshold value and perform the quality improvement process. This contributes to speeding up the quality improvement process.
[0053] [Modification 2] In the above embodiment, a set of one learning process and one quality improvement process was repeated, but the number of times each process is repeated when the learning process and quality improvement process are repeated is not limited to this example. For example, if a rendering quality improvement image that has undergone one quality improvement process is used for learning twice, the quality improvement process may be performed again. In other words, the quality improvement processing unit 15 may perform one quality improvement process on the generated rendering image data each time the learning unit 13 updates the learning model by performing learning multiple times using the rendering image data after the quality improvement process has been performed. This will contribute to speeding up the processing.
[0054] [Variation 3] Quality improvement processing may be performed in stages. For example, if you want to quadruple the resolution with super-resolution processing, instead of immediately performing super-resolution processing to quadruple the resolution, you can perform super-resolution processing to double the resolution to learn the model, and then perform super-resolution processing to double the resolution. In other words, the quality improvement processing unit 15 may divide the quality improvement processing into multiple stages, perform the first stage of quality improvement processing on the generated rendering image data, and then perform the second stage of quality improvement processing on the rendering image data generated using the learning model updated with the rendering image data after the first stage of quality improvement processing.
[0055] [Modification 4] When the quality improvement process includes image super-resolution processing and noise reduction, the quality improvement processing unit 15 may first perform image noise reduction on the generated rendering image data, the learning unit 13 may repeatedly perform learning using the rendering image data after noise reduction to update the learning model, and then the quality improvement processing unit 15 may perform image super-resolution processing on the generated rendering image data, and the learning unit 13 may perform learning using the rendering image data after super-resolution processing.
[0056] [Modification 5] The quality improvement processing unit 15 may perform quality improvement processing with different intensities for each image region of the generated rendering image data. For example, it is possible to set a higher level of noise reduction or a higher super-resolution magnification for image regions that contain high-frequency components above a threshold compared to other image regions.
[0057] [Other Modifications] The block diagram used in the description of the above embodiment shows functional units. These functional blocks (components) are realized by any combination of at least one of hardware and software. Furthermore, the method of realizing each functional block is not particularly limited. That is, each functional block may be realized using one device that is physically or logically coupled, or it may be realized using two or more physically or logically separated devices that are directly or indirectly connected (for example, using wired or wireless connections). A functional block may be realized by combining the above one device or the above multiple devices with software.
[0058] Functions include, but are not limited to, judgment, decision, determination, calculation, calculation, processing, derivation, investigation, exploration, confirmation, reception, transmission, output, access, resolution, selection, selection, establishment, comparison, assumption, expectation, assumption, broadcasting, notifying, communicating, forwarding, configuring, reconfiguring, allocating (mapping), and assigning. For example, a functional block (configuration part) that enables transmission is called a transmission unit or transmitter. In all cases, as mentioned above, the method of implementation is not particularly limited.
[0059] For example, the information processing device 10 in one embodiment of the present disclosure may function as a computer that performs the processing of the present disclosure.
[0060] Each aspect / embodiment described in this disclosure may be applied to at least one of the following: LTE (Long Term Evolution), LTE-A (LTE-Advanced), SUPER 3G, IMT-Advanced, 4G (4th generation mobile communication system), 5G (5th generation mobile communication system), FRA (Future Radio Access), NR (new Radio), W-CDMA®, GSM®, CDMA2000, UMB (ULTRA Mobile Broadband), IEEE 802.11 (Wi-Fi®), IEEE 802.16 (WiMAX®), IEEE 802.20, UWB (Ultra-WideBand), Bluetooth®, and other appropriate systems, as well as next-generation systems extended based thereon. Furthermore, multiple systems may be applied in combination (for example, a combination of at least one of LTE and LTE-A with 5G).
[0061] The processing procedures, sequences, flowcharts, etc., of each aspect / embodiment described in this disclosure may be reordered, provided they do not contradict each other. For example, the methods described in this disclosure present various step elements using exemplary order and are not limited to the specific order presented.
[0062] Input and output information may be stored in a specific location (e.g., memory) or managed using a management table. Input and output information may be overwritten, updated, or appended to. Output information may be deleted. Input information may be transmitted to other devices.
[0063] The determination may be made by a value represented by one bit (0 or 1), by a boolean value (true or false), or by a numerical comparison (for example, a comparison with a predetermined value).
[0064] Although the present disclosure has been described in detail above, it will be clear to those skilled in the art that the present disclosure is not limited to the embodiments described herein. The present disclosure can be implemented in modified and altered forms without departing from the intent and scope of the present disclosure as defined by the claims. Therefore, the descriptions in the present disclosure are illustrative and not intended to be restrictive in any way.
[0065] Software should be broadly interpreted to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software modules, applications, software applications, software packages, routines, subroutines, objects, executable files, execution threads, procedures, functions, etc., whether they are called software, firmware, middleware, microcode, hardware description languages, or by any other name. Furthermore, software, instructions, information, etc., may be transmitted and received via a transmission medium. For example, if software is transmitted from a website, server, or other remote source using at least one of wired technologies (such as coaxial cable, fiber optic cable, twisted pair, or digital subscriber line (DSL)) and wireless technologies (such as infrared or microwave), at least one of these wired and wireless technologies is included in the definition of a transmission medium.
[0066] The information, signals, etc., described herein may be represented using any of the following different technologies. For example, data, instructions, commands, information, signals, bits, symbols, chips, etc., which may be referred to throughout the above description, may be represented by voltage, current, electromagnetic waves, magnetic fields or magnetic particles, optical fields or photons, or any combination thereof. Terms used herein and terms necessary for understanding this disclosure may be replaced with terms having the same or similar meaning.
[0067] Furthermore, the information, parameters, etc., described in this disclosure may be expressed using absolute values, relative values from a predetermined value, or corresponding other information.
[0068] In this disclosure, the phrase "based on" does not mean "based solely on" unless otherwise specified. In other words, the phrase "based on" means both "based solely on" and "based at least on."
[0069] Any reference to elements using the designations “First,” “Second,” etc., as used in this disclosure does not generally limit the quantity or order of those elements. These designations may be used in this disclosure as a convenient way to distinguish between two or more elements. Accordingly, references to the First and Second elements do not imply that only two elements may be employed, or that the First element must precede the Second element in any way.
[0070] In the above-described configuration of each device, the term "part" may be replaced with "means," "circuit," "device," etc.
[0071] Where the terms “include,” “including,” and variations thereof are used in this disclosure, these terms are intended to be inclusive, as is the term “comprising.” Furthermore, the term “or” as used in this disclosure is not intended to mean exclusive OR.
[0072] In this disclosure, if articles are added through translation, such as a, an, and the in English, this disclosure may include the fact that the noun following these articles is plural.
[0073] In this disclosure, the term "A and B are different" may mean "A and B are different from each other." The term may also mean "A and B are each different from C." Terms such as "separate" and "combine" may be interpreted similarly to "different."
[0074] 10: Information processing device, 11: Acquisition unit, 12: Memory unit, 13: Learning unit, 14: Generation unit, 15: Quality improvement processing unit, 16: Output unit, 1001: Processor, 1002: Memory, 1003: Storage, 1004: Communication device, i1, i2, i01, i02: Captured images, I1, I2, I01, I02: Quality improvement images, d, D: 3D model images, r1, r2, r01, r02: Rendering images, rh1, rh2, rh01, rh02: Rendering quality improvement images.
Claims
1. An information processing device comprising: a generation unit that generates rendering image data showing two-dimensional images viewed from multiple viewpoints using a learning model generated through learning to generate three-dimensional images using two-dimensional images taken from multiple viewpoints; a quality improvement processing unit that performs quality improvement processing on the generated rendering image data to improve the image quality; and a learning unit that updates the learning model by repeatedly performing the learning process using the rendering image data after the quality improvement processing.
2. The information processing apparatus according to claim 1, characterized in that the quality improvement process includes super-resolution processing of images or noise reduction processing of images.
3. The information processing apparatus according to claim 1, characterized in that the learning is based on Gaussian Splatting.
4. The information processing apparatus according to claim 1, characterized in that the quality improvement processing unit selects a portion of the generated rendering image data and performs quality improvement processing.
5. The information processing apparatus according to claim 4, characterized in that, if the generated rendering image data includes depth data corresponding to depth in three-dimensional space, the quality improvement processing apparatus selects a portion of the image region that satisfies the condition that the depth data is far from the viewpoint in three-dimensional space and performs quality improvement processing.
6. The information processing apparatus according to claim 4, characterized in that the quality improvement processing unit performs frequency conversion processing on the generated rendering image data, selects a portion of the image region containing a predetermined high-frequency component above a threshold, and performs quality improvement processing on that region.
7. The information processing apparatus according to claim 1, wherein the quality improvement processing unit performs one quality improvement process on the generated rendering image data each time the learning unit updates the learning model by performing the learning process multiple times using the rendering image data after the quality improvement process has been performed.
8. The information processing apparatus according to claim 1, characterized in that the quality improvement processing unit divides the quality improvement processing into multiple stages, performs a first-stage quality improvement processing on the generated rendering image data, and performs a second-stage quality improvement processing on the rendering image data generated using the learning model updated with the rendering image data after the first-stage quality improvement processing.
9. When the quality improvement process includes image super-resolution processing and noise reduction, the information processing apparatus according to claim 1, characterized in that, first, the quality improvement processing unit performs image noise reduction on the generated rendering image data, the learning unit repeatedly performs the learning process using the rendering image data after noise reduction to update the learning model, and then the quality improvement processing unit performs image super-resolution processing on the generated rendering image data, and the learning unit performs the learning process using the rendering image data after super-resolution processing.
10. The information processing apparatus according to claim 1, characterized in that the quality improvement processing unit performs quality improvement processing on a portion of the generated rendering image data with different intensities for each image region.