Information processing device, information processing method, and information processing program
The model enhances image quality by using a diffusion model with classifier-free guidance to decompose and train super-resolution tasks, addressing the challenges of noise and blur in low-resolution images, resulting in improved high-resolution image generation.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- LY CORP
- Filing Date
- 2023-11-20
- Publication Date
- 2026-05-20
AI Technical Summary
Existing image processing techniques struggle to effectively super-resolve low-resolution images degraded by noise, blur, and compression, resulting in suboptimal image quality improvement.
A model is developed that utilizes a diffusion model combined with classifier-free guidance (CFG) to decompose super-resolution tasks into multiple classes, training the model to generate high-resolution images by gradually adding and removing noise and blur, while using class labels with a certain probability to enhance fidelity without reducing diversity.
The model achieves improved image quality in super-resolving low-resolution images by increasing fidelity and reducing noise, effectively converting noisy or low-resolution images into high-resolution images within the same domain.
Smart Images

Figure 0007863079000001 
Figure 0007863079000002 
Figure 0007863079000003
Abstract
Description
Technical Field
[0001] The present invention relates to an information processing apparatus, an information processing method, and an information processing program.
Background Art
[0002] A technique related to an image processing circuit that executes noise reduction processing for reducing noise in an image is disclosed (see Patent Document 1).
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] However, there is room for improvement in the above prior art. For example, in the above prior art, there is room for improvement in terms of super-resolution (high image quality improvement) of a low-resolution image that complexly includes deterioration due to noise, blur (blur: blurring, defocusing), compression, etc. Therefore, a method for more appropriately super-resolving (high image quality improvement) a low-resolution image is required.
[0005] The present application has been made in view of the above, and an object thereof is to realize a model that more appropriately super-resolves (high image quality improvement) a low-resolution image.
Means for Solving the Problems
[0006] The information processing apparatus according to the present application includes a generation unit that generates a set of a reference image and an application image to which a predetermined task is applied to the reference image for each of a plurality of different tasks, and a learning unit that causes a diffusion model to learn to generate a reference image when an application image and task information indicating the task applied to the application image are input. The generation unit then decomposes the original task into multiple tasks and generates a set of reference image and application image for each task. It is characterized by the above. [Effects of the Invention]
[0007] According to one embodiment, a model can be realized that more appropriately super-resolutions (enhances image quality) low-resolution images. [Brief explanation of the drawing]
[0008] [Figure 1] Figure 1 is an explanatory diagram showing an overview of the information processing system according to the embodiment. [Figure 2] Figure 2 shows an image of CFG (classifier-free guidance). [Figure 3] Figure 3 is an explanatory diagram illustrating the overview of real-world super-resolution using CFG. [Figure 4] Figure 4 shows the evaluation of the results of conducting CFG. [Figure 5] Figure 5 shows an example of the configuration of a terminal device according to this embodiment. [Figure 6] Figure 6 shows an example of the configuration of a server device according to the embodiment. [Figure 7] Figure 7 is a flowchart showing the processing procedure according to the embodiment. [Figure 8] Figure 8 shows an example of a hardware configuration. [Modes for carrying out the invention]
[0009] The following describes in detail, with reference to the drawings, embodiments for implementing the information processing device, information processing method, and information processing program according to the present application (hereinafter referred to as "embodiments"). Note that these embodiments do not limit the information processing device, information processing method, and information processing program according to the present application. Furthermore, the same parts are denoted by the same reference numerals in the following embodiments, and redundant descriptions are omitted.
[0010] [1. Overview of the Information Processing System] First, with reference to Figure 1, an overview of the information processing system according to the embodiment will be described. Figure 1 is an explanatory diagram showing an overview of the information processing system according to the embodiment. As shown in Figure 1, the information processing system 1 according to the embodiment includes a terminal device 10 and a server device 100. These various devices are connected to each other via a network N, either by wire or wireless means, enabling communication. As a result, the terminal device 10 can cooperate with the server device 100. The network N is, for example, a LAN (Local Area Network) or a WAN (Wide Area Network) such as the Internet.
[0011] Terminal device 10 is an information processing device used by user U. For example, terminal device 10 may be a smart device such as a smartphone or tablet, a mobile phone such as a feature phone, a PC (Personal Computer), a PDA (Personal Digital Assistant), a game console or AV equipment with communication functions, an information appliance or digital appliance, a car navigation system, a wearable device such as a smartwatch, head-mounted display, or smart glasses. Alternatively, terminal device 10 may be a house or building compatible with the Internet of Things (IoT), a car, a home appliance, or an electronic device.
[0012] In this embodiment, the terminal device 10 is a smart device such as a smartphone or tablet used by user U, and is a mobile terminal device capable of communicating with any server device via a wireless communication network such as LTE (Long Term Evolution), 4G (4th Generation), or 5G (5th Generation). The terminal device 10 also has a screen such as a liquid crystal display with touch panel functionality, and accepts various operations on displayed data such as content from user U using a finger or stylus, such as tapping, sliding, and scrolling. Operations performed on the area of the screen where content is displayed may also be considered as operations on the content. Furthermore, the terminal device 10 may be an information processing device such as a desktop PC (Personal Computer) or a notebook PC, not just a smart device.
[0013] Furthermore, the terminal device 10 can connect to the network N via wireless communication networks such as LTE, 4G, 5G, Bluetooth®, or wireless LAN (Local Area Network), and communicate with the server device 100.
[0014] The server device 100 is, for example, a computer such as a PC or blade server, or a mainframe or workstation. The server device 100 may also be implemented through cloud computing.
[0015] In this embodiment, the server device 100 is an information processing device that works in conjunction with each user U's terminal device 10 and provides each user U's terminal device 10 with API (Application Programming Interface) services for various applications (hereinafter referred to as "apps") and various data, and is implemented by a computer or cloud system.
[0016] Further, the server device 100 may be an information processing device that provides some kind of web service online to the terminal device 10 of each user U. For example, as a web service, the server device 100 may provide services such as Internet connection, search service, SNS (Social Networking Service), e-commerce (EC: Electronic Commerce), electronic payment, online game, online banking, online trading, accommodation and ticket reservation, video and music distribution, news, map, route search, route guidance, route information, operation information, weather forecast, etc. In fact, the server device 100 may cooperate with various servers that provide the above web services and mediate the web services, or be responsible for the processing of the web services.
[0017] In addition, the server device 100 can acquire user information regarding the user U. For example, as user information, the server device 100 acquires information (attribute information) regarding the attributes of the user U, such as the gender, age, and residential area of the user U. Further, the server device 100 can acquire information regarding attributes such as the demographics (demographic attributes), psychographics (psychological attributes), geographics (geographical attributes), and behavioral (behavioral attributes) of the user U. Also, the server device 100 may acquire, as user information, the segment or persona (persona) to which the user U belongs in the field of marketing. Then, the server device 100 stores and manages the information (attribute information) regarding the attributes of the user U together with the identification information (such as user ID) indicating the user U.
[0018] In addition, the server device 100 acquires various types of history information (log data) indicating the actions of the user U from the terminal device 10 of the user U or from various servers, etc. based on the user ID, etc. For example, the server device 100 acquires a position history, which is a history of the position and time of the user U, from the terminal device 10. Further, the server device 100 acquires a search history, which is a history of search queries input by the user U, from a search server (search engine). Further, the server device 100 acquires a browsing history, which is a history of the content browsed by the user U, from a content server. Further, the server device 100 acquires a purchase history (settlement history), which is a history of the user U's product purchases and settlement processes, from an e-commerce server or a settlement processing server. Further, the server device 100 may acquire a listing history and a sales history, which are histories of the user U's listings on the marketplace, from an e-commerce server or a settlement processing server. Further, the server device 100 acquires a posting history, which is a history of the user U's posts, from a posting server or an SNS server that provides a word-of-mouth posting service. Note that each of the above-mentioned various servers, etc. may be the server device 100 itself. That is, the server device 100 may function as each of the above-mentioned various servers, etc.
[0019] Also, the number of each device included in the information processing system 1 shown in FIG. 1 is not limited to that shown. For example, in FIG. 1, only one terminal device 10 is shown for simplicity of illustration, but this is merely an example and not limiting, and two or more may be used.
[0020] [2. Real-world Super-Resolution] Real-world super-resolution is a super-resolution (image quality improvement) task for images that are complexly degraded by noise, blur, and compression. For example, when not related to the real world, it refers to a task of restoring an image reduced by a factor of 2 or 4 to its original resolution.
[0021] [2-1. Mechanism of Super-Resolution by Diffusion Model] A diffusion model is a type of generative AI (Artificial Intelligence) that learns the process of gradually adding noise to an image to degrade it (diffusion process) and the process of gradually removing noise (denoicing) to reconstruct the image (de-diffusion process).
[0022] In the super-resolution mechanism using the diffusion model, the server device 100 sets the condition for conditional generation by the diffusion model to a low-resolution image and generates a high-resolution image from a noisy image. Furthermore, during model training, the server device 100 generates a low-resolution image by degrading the original high-resolution image data through blurring, JPEG compression, etc.
[0023] For example, the server device 100 generates a noise-free image by overlaying a low-resolution image as a condition onto a noisy image and inputting it into a diffusion model. Then, the server device 100 overlays the low-resolution image onto the noise-free image (the generated image) and inputs it into the diffusion model to generate yet another noise-free image. The server device 100 repeats this process to generate a high-resolution image.
[0024] Images utilize a 3-channel image representation (representation method), such as RGB (red, green, blue) or YcbCr (luminance signal (Y), difference between luminance signal and blue component (Cb), and difference between luminance signal and red component (Cr)). In the case of super-resolution using a diffusion model, a total of 6 channels are input—3 channels of a noisy image plus 3 channels of a low-resolution image—to produce a 3-channel image with reduced noise. For example, when using a 3-channel image representation of YCbCr, a low-resolution image in YCbCr format is superimposed on the noisy image.
[0025] At this time, the server device 100 inputs further time information t to the diffusion model and repeats the noise reduction process according to the time information t input to the diffusion model. In other words, the server device 100 performs noise reduction according to the input time information t.
[0026] The time information t indicates the current level of noise. For example, if the noise level is around 1000, inputting 1000 as the time information t will cause the process of generating an image with progressively reduced noise to repeat 1000 times. Each time the process is executed, the time information t is decreased by 1, and with each iteration, the noise gradually decreases until it reaches 0, at which point a completely noise-free, clean image (high-resolution image) is output.
[0027] During training, the noise reduction process described above is reversed. For example, the server device 100 starts with a high-resolution image and gradually adds noise from noise level 0 (no noise) to 1000 (complete noise) to generate an image with only noise. Then, by reversing this process, the diffusion model is trained to gradually denoise (de-noise) the noisy image to ultimately generate a high-resolution image.
[0028] In this way, the server device 100 generates high-resolution images by repeatedly denoising (denoising) from complete noise using a diffusion model.
[0029] For example, as shown in Figure 1, the server device 100 inputs a noisy image and a low-resolution image into the diffusion model (step S1).
[0030] Furthermore, the server device 100 inputs time information t into the diffusion model (step S2).
[0031] The server device 100 uses a diffusion model to generate an image from the input noisy image and low-resolution image with one level of noise removed, according to the time information t (step S3).
[0032] Then, each time the above process is executed, the server device 100 decreases the value of time information t by one, inputs the denoised image and the low-resolution image into the diffusion model, and repeats the denoising process until the value of time information t becomes 0 (step S4).
[0033] Finally, the server device 100 generates a high-resolution image (step S5).
[0034] However, currently, the generation (reconstruction) of high-resolution images using diffusion models is far from perfect, and there is room for improvement. In this embodiment, quality is improved by combining super-resolution using diffusion models with CFG (classifier-free guidance).
[0035] [2-2. Explanation of CFG (classifier-free guidance)] CFG (classifier-free guidance) is a technique primarily used in Text2Image generative models. Originally, there was a technique called classifier guidance, which improved the generative results by separately training a special classifier. Classifier guidance involves using a classifier trained on noisy images to guide the model to generate images that match the labels. CFG (classifier-free guidance) is so named because it can be implemented without training a separate classifier.
[0036] In CFG (classifier-free guidance), the model is trained while removing (dropping) text conditions with a certain probability. During generation, the generation results with and without text conditions are calculated separately, and the calculations are performed as follows.
[0037] The generated result (CFG (classifier-free guidance)) = scale × (generated result (with text conditions) - generated result (without text conditions)) + generated result (without text conditions)
[0038] Here, `scale` indicates the strength of the guidance (guidance scale). In other words, the value of `scale` corresponds to the strength of the guidance. The larger the value of `Scale`, the stronger the guidance. When actually running the program, `scale` is specified as a parameter, so the user can specify any value. Note that if `Scale=1`, no guidance is provided. If `scale=1`, the "No text conditions" option is canceled and disappears. Therefore, if you want guidance, the value of `scale` should be greater than 1.
[0039] Figure 2 shows an image of CFG (classifier-free guidance) (in the 2D case) (scale=2). Of the lines shown in Figure 2, A represents the generation result (with text conditions), B represents the generation result (without text conditions), and C represents the generation result (CFG (classifier-free guidance)). The coordinates at the end of the arrow in C reflect the text conditions more accurately than the coordinates at the end of the arrow in A.
[0040] Applying CFG (classifier-free guidance) increases fidelity but reduces diversity. For example, applying CFG increases the likelihood of generating images from the same domain and decreases the likelihood of generating images from different domains. However, for real-world super-resolution, which converts noisy or low-resolution images from the same domain into high-resolution images, increased fidelity is preferable, and reduced diversity is not a problem. In other words, applying CFG has the secondary effect of increasing fidelity in real-world super-resolution.
[0041] [2-3. Class-conditional super-resolution] Referring to Figure 3, we will explain real-world super-resolution using CFG. Figure 3 is an explanatory diagram showing an overview of real-world super-resolution using CFG. The low-resolution image, which is a condition for conditional generation using the diffusion model described above, is an essential condition (an indispensable condition) for super-resolution because without it, it would be impossible to indicate the direction of generating a high-resolution image from a noisy image, and a completely different image might be generated. Therefore, in order to add new conditions or manipulate conditions, it is necessary to specify the conditions in a different form.
[0042] In this embodiment, the server device 100 decomposes the task of performing real-world super-resolution (real-world super-resolution task) into multiple classes (tasks) and performs class-conditional super-resolution. That is, the server device 100 decomposes the original task (reference task) into multiple tasks (constituent tasks that make up the reference task) and changes it into a class-conditional super-resolution model.
[0043] Specifically, the server device 100 inputs class labels c (task labels) along with time information t to the diffusion model. For example, the class labels c may be "Class 0: Super Resolution", "Class 1: Denoising (Denoising) Only", and "Class 2: Real-World Super Resolution". Alternatively, the class labels c may be "Class 0: 2x Super Resolution", "Class 1: 3x Super Resolution", and "Class 2: 4x Super Resolution". In this case, the original task (reference task) may be designated as Class 2, and the original task may be decomposed into multiple tasks (constituent tasks) such as Class 0 and Class 1.
[0044] Furthermore, during model training, the server device 100 generates training samples by adding degradation that simulates the real world to high-resolution images. At this time, the server device 100 applies degradation to the high-resolution images according to the class label c (task label). For example, if "Class 0: Super Resolution," "Class 1: Denoising (Denoising) Only," and "Class 2: Real-World Super Resolution," the server device 100 applies only scaling (enlargement or reduction) for Class 0, only noise addition for Class 1, and both scaling and noise addition for Class 2.
[0045] The server device 100 applies a stepwise degradation to the high-resolution image according to the class label c (task label). The server device 100 then trains the diffusion model stepwise using the training samples generated for each class (task) along with the class label c (task label). For example, in the case of class 1, the server device 100 starts with a high-resolution image and gradually adds noise from noise level 0 (no noise) to 1000 (complete noise) to generate an image with only noise. By reversing this process, the diffusion model is trained to gradually denoise (de-noise) the noisy image to generate a high-resolution image. The same applies to classes 0 and 2. For example, in the case of image reduction, the server device 100 starts with a high-resolution image and gradually reduces it from reduction level 0 (original image) to 1000 (image reduced by 2x or 4x) to generate an image reduced by 2x or 4x.
[0046] At this time, the server device 100 drops the class label c with a certain probability (for example, a probability of about 10%) during model training, making CFG (classifier-free guidance) available. The diffusion model operates according to the class label c, such as class 0, class 1, class 2, etc., but in the case of no label, it does not know which class (task) it is, but it will operate on average so that it does not cause problems regardless of the class. In this way, the diffusion model operates safely in the case of no label. Therefore, in the case of no label, the performance will be worse than at least in the case of class 2.
[0047] Furthermore, when generating high-resolution images, the server device 100 can improve the quality of real-world super-resolution by specifying class 2 as the class label c input to the diffusion model and by applying CFG (classifier-free guidance). In this case, simply specifying class 2 improves the quality of real-world super-resolution, and applying guidance further improves the quality of real-world super-resolution.
[0048] For example, as shown in Figure 3, the server device 100 generates a low-resolution image by degrading the original high-resolution image data through blurring, JPEG compression, etc. (step S11).
[0049] Furthermore, the server device 100 decomposes the super-resolution task into multiple classes (tasks) and sets class labels c (task labels) (step S12).
[0050] Next, the server device 100 generates training samples by gradually adding degradation to the high-resolution image according to the class label c (step S13).
[0051] Next, the server device 100 trains the diffusion model step by step using the generated training samples along with the class label c (step S14).
[0052] At this time, the server device 100 drops the class label c with a certain probability (for example, a probability of about 10%), making CFG (classifier-free guidance) available (step S15).
[0053] Subsequently, the server device 100 superimposes a low-resolution image as a condition onto the input image (noised image, etc.) and inputs it into the diffusion model. It also inputs time information t and class label c into the diffusion model, thereby performing real-world super-resolution on the input image and ultimately generating a high-resolution image (step S16).
[0054] [2-4. Effects of CFG (classifier-free guidance)] The results of classifier-free guidance (CFG) performed with specified classes were evaluated using the RealSRv3 and DrealSR datasets, which are widely used for evaluating real-world super-resolution. A lower FID10k value indicates higher quality.
[0055] Figure 4 shows the evaluation of the results of performing CFG. As shown in Figure 4, specifying Class 2 (real-world super-resolution) and strengthening the guidance by setting the CFG scale to 1.5 or 2.0 results in a lower FID10k value and improved quality, compared to specifying Class 0 or 1. Here, the scale value corresponds to the strength of the guidance, meaning that the guidance is stronger when scale=2.0 than when scale=1.5.
[0056] [2-5. Supplement] During model training, the server device 100 generates pairs of reference images and applied images obtained by applying a predetermined task to the reference image, for each of several different tasks. The reference image is, for example, a high-resolution image. The applied image is, for example, an image obtained by adding one level of noise to the reference image. The multiple tasks include a reference task performed using the trained model and constituent tasks that make up the reference task. Note that the multiple tasks include a task to generate a real-world super-resolution image (high-resolution image).
[0057] The noise is applied in stages. For example, the server device 100 uses an image with n levels of noise applied to a high-resolution image as the reference image, and an image with one level of noise applied to the reference image (an image with n+1 levels of noise applied to the high-resolution image) as the applied image, thereby generating a pair of the reference image and the applied image. The server device 100 then uses the applied image from this stage as the reference image for the next stage, and repeats this process a predetermined number of times (for example, 1000 times) corresponding to the time information t mentioned above.
[0058] The server device 100 then trains its model to generate a reference image when it receives an application image and task information indicating the task applied to that application image as input. In other words, the server device 100 trains its model on each of the reference image and application image pairs that have been generated a predetermined number of times.
[0059] As a result, the server device 100 trains its model to generate an image (reference image) from a noisy image (application image) and a low-resolution image (task-specific image) by gradually (stepwise) removing the noise according to the time information t. Then, using the trained model, the server device 100 generates and outputs a real-world super-resolution image (high-resolution image) of the input image according to the time information t.
[0060] In this embodiment, classifier-free guidance (CFG) is applied to the super-resolution using a diffusion model, and the level of processing to be performed is input. Since CFG is a technique that makes the model more faithful to the original as the level is increased, the server device 100 applies this to receive a label indicating a task and trains the model to perform the task corresponding to that label.
[0061] The server device 100 performs a certain amount of unlabeled learning in parallel with labeled learning. The model will start to perform tasks by estimating labels, and will therefore perform average operations. At this time, it is sufficient to include at least the final objective in the labels.
[0062] The label combinations should be set to tasks that construct real-world super-resolution. Alternatively, they could be set to level differences of real-world super-resolution with only the magnification being different.
[0063] Applying CFG (classifier-free guidance) reduces diversity, but since the main purpose of super-resolution is to improve the resolution of images within the same domain, reduced diversity is actually an advantage.
[0064] [3. Example of terminal device configuration] Next, the configuration of the terminal device 10 will be explained using Figure 5. Figure 5 is a diagram showing an example of the configuration of the terminal device 10. As shown in Figure 5, the terminal device 10 comprises a communication unit 11, a display unit 12, an input unit 13, a positioning unit 14, a sensor unit 20, a control unit 30 (controller), and a storage unit 40.
[0065] (Communications Section 11) The communication unit 11 is connected to the network N by wire or wireless connection and transmits and receives information to and from the server device 100 via the network N. For example, the communication unit 11 can be implemented using a NIC (Network Interface Card) or an antenna.
[0066] (Display section 12) The display unit 12 is a display device that displays various information such as location information. For example, the display unit 12 may be a liquid crystal display (LCD) or an organic electro-luminescent display (OLED). The display unit 12 may also be a touch panel display, but is not limited to this.
[0067] (Input section 13) The input unit 13 is an input device that receives various operations from the user U. For example, the input unit 13 has buttons for inputting characters, numbers, etc. The input unit 13 may also be an input / output port (I / O port) or a USB (Universal Serial Bus) port. If the display unit 12 is a touch panel display, a part of the display unit 12 functions as the input unit 13. The input unit 13 may also be a microphone that receives voice input from the user U. The microphone may be wireless.
[0068] (Positioning unit 14) The positioning unit 14 receives signals (radio waves) transmitted from GPS (Global Positioning System) satellites and, based on the received signals, acquires position information (e.g., latitude and longitude) indicating the current position of the terminal device 10. In other words, the positioning unit 14 determines the position of the terminal device 10. Note that GPS is just one example of a GNSS (Global Navigation Satellite System).
[0069] Furthermore, the positioning unit 14 can determine its position using various methods other than GPS. For example, the positioning unit 14 may use various communication functions of the terminal device 10 to determine its position as an auxiliary positioning means for position correction, etc., as described below.
[0070] (Wi-Fi positioning) For example, the positioning unit 14 determines the location of the terminal device 10 by utilizing the Wi-Fi® communication function of the terminal device 10 and the communication network provided by each telecommunications company. Specifically, the positioning unit 14 determines the location of the terminal device 10 by performing Wi-Fi communication, etc., and determining the distance to nearby base stations and access points.
[0071] (Beacon positioning) Furthermore, the positioning unit 14 may determine the location using the Bluetooth® function of the terminal device 10. For example, the positioning unit 14 determines the location of the terminal device 10 by connecting to a beacon transmitter connected via the Bluetooth® function.
[0072] (Geomagnetic positioning) Furthermore, the positioning unit 14 determines the position of the terminal device 10 based on the geomagnetic pattern of the structure, which has been measured in advance, and the geomagnetic sensor provided by the terminal device 10.
[0073] (RFID positioning) Furthermore, if, for example, the terminal device 10 is equipped with an RFID (Radio Frequency Identification) tag function equivalent to that of a contactless IC card used at a train station ticket gate or in a store, or if it is equipped with a function to read RFID tags, the location where it was used will be recorded along with the information on the payment or other transactions made by the terminal device 10. The positioning unit 14 may determine the location of the terminal device 10 by acquiring such information. Alternatively, the location may be determined by an optical sensor or infrared sensor equipped in the terminal device 10.
[0074] The positioning unit 14 may, if necessary, determine the position of the terminal device 10 using one or a combination of the positioning means described above.
[0075] (Sensor unit 20) The sensor unit 20 includes various sensors mounted on or connected to the terminal device 10. The connection can be wired or wireless. For example, the sensors may be detection devices other than the terminal device 10, such as wearable devices or wireless devices. In the example shown in Figure 5, the sensor unit 20 includes an acceleration sensor 21, a gyro sensor 22, a barometric pressure sensor 23, a temperature sensor 24, a sound sensor 25, a light sensor 26, a magnetic sensor 27, and an image sensor (camera) 28.
[0076] The sensors 21-28 described above are merely examples and not limiting. In other words, the sensor unit 20 may be configured to include some of the sensors 21-28, or it may include other sensors such as humidity sensors in addition to or instead of the sensors 21-28.
[0077] The acceleration sensor 21 is, for example, a 3-axis acceleration sensor and detects the physical movement of the terminal device 10, such as its direction of movement, velocity, and acceleration. The gyro sensor 22 detects the physical movement of the terminal device 10, such as its tilt in the three axes, based on its angular velocity. The barometric pressure sensor 23 detects the atmospheric pressure around the terminal device 10, for example.
[0078] Since the terminal device 10 is equipped with the acceleration sensor 21, gyroscope 22, barometric pressure sensor 23, etc., it becomes possible to determine the position of the terminal device 10 using technologies such as pedestrian dead-reckoning (PDR) that utilize these sensors 21 to 23. This makes it possible to obtain indoor location information that is difficult to obtain with positioning systems such as GPS.
[0079] For example, a pedometer using an accelerometer 21 can calculate the number of steps, walking speed, and distance walked. Additionally, a gyroscope 22 can be used to determine the user U's direction of movement, gaze direction, and body tilt. Furthermore, the barometric pressure detected by the barometric pressure sensor 23 can be used to determine the altitude and floor number of the user U's terminal device 10.
[0080] The temperature sensor 24 detects, for example, the ambient temperature around the terminal device 10. The sound sensor 25 detects, for example, the ambient sound around the terminal device 10. The light sensor 26 detects the ambient illumination around the terminal device 10. The magnetic sensor 27 detects, for example, the Earth's magnetic field around the terminal device 10. The image sensor 28 captures an image of the area around the terminal device 10.
[0081] The aforementioned pressure sensor 23, temperature sensor 24, sound sensor 25, light sensor 26, and image sensor 28 can detect the surrounding environment and conditions of the terminal device 10 by detecting atmospheric pressure, temperature, sound, and illuminance, respectively, and by capturing images of the surroundings. Furthermore, it becomes possible to improve the accuracy of the location information of the terminal device 10 based on the surrounding environment and conditions.
[0082] (Control Unit 30) The control unit 30 includes, for example, a microcomputer having a CPU (Central Processing Unit), ROM (Read Only Memory), RAM, input / output ports, and various circuits. Alternatively, the control unit 30 may be composed of hardware such as an integrated circuit (ASIC) or FPGA (Field Programmable Gate Array). The control unit 30 includes a transmission unit 31, a reception unit 32, and a processing unit 33.
[0083] (Transmitter 31) The transmission unit 31 can transmit various information, such as information input by the user U using the input unit 13, various information detected by sensors 21-28 mounted on or connected to the terminal device 10, and location information of the terminal device 10 determined by the positioning unit 14, to the server device 100 via the communication unit 11.
[0084] (Receiving unit 32) The receiving unit 32 can receive various information provided by the server device 100, as well as requests for various information from the server device 100, via the communication unit 11.
[0085] (Processing 33) The processing unit 33 controls the entire terminal device 10, including the display unit 12. For example, the processing unit 33 can output and display various information transmitted by the transmission unit 31 and various information received from the server device 100 by the reception unit 32 to the display unit 12.
[0086] (Storage unit 40) The storage unit 40 is implemented by, for example, semiconductor memory elements such as RAM (Random Access Memory) and flash memory, or by storage devices such as HDD (Hard Disk Drive), SSD (Solid State Drive), and optical discs. Various programs and various data are stored in this storage unit 40.
[0087] [4. Example of Server Device Configuration] Next, the configuration of the server device 100 according to the embodiment will be described using Figure 6. Figure 6 is a diagram showing an example of the configuration of the server device 100 according to the embodiment. As shown in Figure 6, the server device 100 includes a communication unit 110, a storage unit 120, and a control unit 130.
[0088] (Communications Department 110) The communication unit 110 is implemented, for example, by a NIC (Network Interface Card). The communication unit 110 is connected to the network N by wire or wireless connection.
[0089] (Storage unit 120) The storage unit 120 is implemented by, for example, semiconductor memory elements such as RAM (Random Access Memory) and flash memory, or by storage devices such as HDDs, SSDs, and optical discs. The storage unit 120 may store identification information (such as a user ID) indicating user U, as well as attribute information and history information (log data) of user U.
[0090] (Control unit 130) The control unit 130 is a controller, and is realized by executing various programs (corresponding to an example of an information processing program) stored in the internal memory of the server device 100 using a memory area such as RAM as a working area, for example, by a CPU (Central Processing Unit), MPU (Micro Processing Unit), ASIC (Application Specific Integrated Circuit), or FPGA (Field Programmable Gate Array). In the example shown in Figure 6, the control unit 130 has an acquisition unit 131, a generation unit 132, a learning unit 133, a super-resolution unit 134, a management unit 135, and a provision unit 136.
[0091] (Acquisition part 131) The acquisition unit 131 acquires the search query entered by the user U. For example, when the user U enters a search query into a search engine or the like and performs a keyword search, the acquisition unit 131 acquires the search query via the communication unit 110. In other words, the acquisition unit 131 acquires the keyword entered by the user U into the search box of a search engine, website, or application via the communication unit 110.
[0092] Furthermore, the acquisition unit 131 acquires user information about user U via the communication unit 110. For example, the acquisition unit 131 acquires identification information (such as user ID), location information, and attribute information of user U from user U's terminal device 10. The acquisition unit 131 may also acquire identification information and attribute information of user U when user U is registered. The acquisition unit 131 then stores the user information in the storage unit 120.
[0093] Furthermore, the acquisition unit 131 acquires various historical information (log data) indicating the user U's actions via the communication unit 110. For example, the acquisition unit 131 acquires various historical information indicating the user U's actions from the user U's terminal device 10, or from various servers based on the user ID, etc. The acquisition unit 131 then stores the various historical information in the storage unit 120.
[0094] Furthermore, the acquisition unit 131 acquires image data via the communication unit 110 from the user U's terminal device 10, another server device 100, or an external storage device or storage medium. The acquisition unit 131 also acquires image data stored in the storage unit 120. For example, the acquisition unit 131 acquires a high-resolution image. The domain of the high-resolution image is not restricted; that is, the subject of the high-resolution image can be arbitrary.
[0095] (Generation unit 132) The generation unit 132 generates pairs of a reference image and an applied image obtained by applying a predetermined task to the reference image, for each of several different tasks. For example, if the reference image is a high-resolution image and the predetermined task is noise reduction, the generation unit 132 generates a noisy image as the applied image by adding noise to the high-resolution image.
[0096] At this time, the generation unit 132 decomposes the original task into multiple tasks and generates a set of a reference image and an applied image for each task. That is, the generation unit 132 generates a set of a reference image and an applied image for a reference task performed using a trained diffusion model, and further generates a set of a reference image and an applied image for the constituent tasks that make up the reference task. For example, if the reference task is a real-world super-resolution task, the generation unit 132 decomposes the real-world super-resolution task into a super-resolution task and a denoising task as constituent tasks. Then, the class labels as task information are set as follows: super-resolution task is class 0, denoising task is class 1, and real-world super-resolution task is class 2.
[0097] For example, the generation unit 132 generates a set of a high-resolution image as a reference image and an applied image obtained by gradually scaling the high-resolution image, as a training sample when the class label of the task information is class 0. The generation unit 132 also generates a set of a high-resolution image as a reference image and an applied image obtained by gradually adding noise to the high-resolution image, as a training sample when the class label of the task information is class 1. The generation unit 132 also generates a set of a high-resolution image as a reference image and an applied image obtained by gradually scaling and adding noise to the high-resolution image, as a training sample when the class label of the task information is class 2.
[0098] (Learning Section 133) The learning unit 133 trains a diffusion model to generate a reference image when it receives an application image and task information indicating the task applied to the application image. For example, the learning unit 133 trains the diffusion model to generate an application image by gradually degrading the reference image according to a predetermined task, and to reconstruct the reference image from the application image step by step by reversing the degradation process.
[0099] Furthermore, when the learning unit 133 receives a class label as task information indicating a task, it trains the diffusion model to generate a reference image. At this time, the learning unit 133 trains the diffusion model for each class label using the generated training samples.
[0100] Furthermore, the learning unit 133 removes class labels with a certain probability (for example, a probability of about 10%), and trains the diffusion model using the generated training samples without class labels, thereby making it possible to apply CFG (classifier-free guidance) to super-resolution by the diffusion model.
[0101] From another perspective, the learning unit 133 decomposes the super-resolution task into multiple classes and trains the diffusion model to become a class-conditional super-resolution model.
[0102] From another perspective, the learning unit 133 trains the diffusion model so that CFG (classifier-free guidance) can be used in super-resolution using the diffusion model.
[0103] (Super-resolution section 134) The super-resolution unit 134 inputs the application image and task information indicating a real-world super-resolution task into a trained diffusion model, performs real-world super-resolution on the application image, and generates a reference image (in reality, a real-world super-resolution image equivalent to the reference image).
[0104] For example, the super-resolution unit 134 inputs a noisy image as the application image and a low-resolution image into a trained diffusion model, and further inputs time information corresponding to the noise level and a class label as task information. It then performs real-world super-resolution on the noisy image and generates a high-resolution image as a real-world super-resolution image equivalent to the reference image.
[0105] From another perspective, the super-resolution unit 134 performs real-world super-resolution of the input image using a pre-trained diffusion model that becomes a class-conditional super-resolution model.
[0106] From another perspective, the super-resolution unit 134 performs real-world super-resolution of the input image using a diffusion model that makes CFG (classifier-free guidance) available.
[0107] (Management Department 135) The management unit 135 stores the reference image (actually, a real-world super-resolution image equivalent to the reference image) generated by the super-resolution unit 134. For example, the management unit 135 stores the generated reference image (actually, a real-world super-resolution image equivalent to the reference image) in the storage unit 120. Alternatively, the management unit 135 may store the generated reference image (actually, a real-world super-resolution image equivalent to the reference image) in an external storage device or storage medium via the communication unit 110.
[0108] (Provider 136) The provisioning unit 136 provides, via the communication unit 110, a reference image (actually a real-world super-resolution image equivalent to the reference image) generated by the super-resolution unit 134 or stored by the management unit 135 to the user U's terminal device 10 or another server device 100, or an external storage device or storage medium.
[0109] [5. Processing Procedure] Next, the processing procedure by the server device 100 according to the embodiment will be described using Figure 7. Figure 7 is a flowchart of the processing procedure according to the embodiment. Note that the processing procedure shown below is repeatedly executed by the control unit 130 of the server device 100.
[0110] For example, as shown in Figure 7, the acquisition unit 131 of the server device 100 acquires a high-resolution image that will serve as a reference image (step S101).
[0111] Next, the generation unit 132 of the server device 100 decomposes the original task into multiple tasks, and for each task, including the original task and the multiple tasks after decomposition, generates a pair of a high-resolution image that serves as a reference image and an applied image obtained by applying the task to the reference image (step S102).
[0112] Next, the learning unit 133 of the server device 100, for each task, specifies a class label indicating the task and trains the diffusion model on the process of generating an applied image by gradually degrading the reference image according to the task, and the process of reconstructing the reference image from the applied image step by step by reversing the degradation process (step S103).
[0113] At this time, the learning unit 133 of the server device 100 removes class labels with a certain probability (for example, a probability of about 10%), and trains the diffusion model using the generated training samples without class labels, thereby training the diffusion model so that CFG (classifier-free guidance) can be used in super-resolution by the diffusion model (step S104).
[0114] Next, the super-resolution unit 134 of the server device 100 inputs the application image and task information indicating a real-world super-resolution task to a trained diffusion model for which CFG (classifier-free guidance) is available, performs real-world super-resolution on the application image, and generates a reference image (actually, a real-world super-resolution image equivalent to the reference image) (step S105).
[0115] Next, the management unit 135 of the server device 100 stores the reference image (actually, a real-world super-resolution image equivalent to the reference image) generated by the super-resolution unit 134 (step S106).
[0116] Next, the supply unit 136 of the server device 100 provides the reference image (actually, a real-world super-resolution image equivalent to the reference image) stored by the management unit 135 to an external party (step S107).
[0117] [6. Variant Example] The terminal device 10 and server device 100 described above may be implemented in various other forms besides those of the embodiment described above. Therefore, the following describes modifications of the embodiment.
[0118] In the above embodiment, some or all of the processing performed by the server device 100 may actually be performed by the terminal device 10 (or an application running on the terminal). For example, the processing may be completed in a standalone manner (by the terminal device 10 alone). In this case, the terminal device 10 is assumed to have the functions of the server device 100 in the above embodiment. Furthermore, in the above embodiment, since the terminal device 10 is in cooperation with the server device 100, from the perspective of the user U, it appears as if the processing of the server device 100 is also being performed by the terminal device 10. In other words, from another perspective, it can be said that the terminal device 10 is equipped with the server device 100.
[0119] Furthermore, in the above embodiment, the high-resolution image may be a high-resolution aerial photograph. The low-resolution image may be a satellite image from a map application. Note that the high-resolution image and the low-resolution image are images from the same domain.
[0120] Furthermore, in the above embodiment, the number of tasks when decomposing the original task (reference task) into multiple tasks (configuration tasks that make up the reference task) is arbitrary. That is, the number of configuration tasks may be equal to the number of tasks that can be decomposed from the reference task.
[0121] Furthermore, in the above embodiment, the number of steps in which degradation according to the class label c (task label) is applied to the high-resolution image is arbitrary. For example, the number of noise levels when degrading the target image by applying noise in stages is arbitrary.
[0122] [7. Effects] As described above, the information processing device (terminal device 10 and server device 100) according to the present invention is characterized by comprising: a generation unit 132 that generates pairs of a reference image (e.g., a high-resolution image) and an applied image (e.g., a noisy image) obtained by applying a predetermined task to the reference image, for each of several different tasks; and a learning unit 133 that causes a diffusion model to learn to generate a reference image when an applied image and task information (e.g., a class label) indicating the task (class) applied to the applied image are input.
[0123] The generation unit 132 decomposes the original task into multiple tasks and generates a pair of reference image and application image for each task.
[0124] The generation unit 132 generates pairs of reference images and applied images for a reference task performed using a trained diffusion model, and further generates pairs of reference images and applied images for constituent tasks that make up the reference task.
[0125] The learning unit 133 trains a diffusion model to perform the process of generating an applied image by gradually degrading a reference image according to a predetermined task, and the process of reconstructing the reference image from the applied image step by step, in reverse order of the degradation process.
[0126] The learning unit 133 trains a diffusion model to generate a reference image when it receives a class label as task information indicating the task.
[0127] The generation unit 132 generates a pair of a high-resolution image as a reference image and an applied image obtained by gradually scaling the high-resolution image, as a training sample when the class label as task information is class 0; a pair of a high-resolution image as a reference image and an applied image obtained by gradually adding noise to the high-resolution image, as a training sample when the class label as task information is class 1; and a pair of a high-resolution image as a reference image and an applied image obtained by gradually scaling and adding noise to the high-resolution image, as a training sample when the class label as task information is class 2. The learning unit 133 trains the diffusion model using the generated training samples for each class label.
[0128] The learning unit 133 removes class labels with a certain probability and trains the diffusion model using the generated training samples without class labels, thereby making it possible to apply CFG (classifier-free guidance) to super-resolution by the diffusion model.
[0129] Furthermore, the information processing device according to the present invention is characterized by further comprising a super-resolution unit 134 that performs real-world super-resolution on the applied image and generates a real-world super-resolution image corresponding to a reference image by inputting the applied image and task information indicating a real-world super-resolution task to a trained diffusion model.
[0130] From another perspective, the information processing device according to the present invention is characterized by comprising a learning unit 133 that decomposes a super-resolution task into multiple classes and trains a diffusion model to become a class-conditional super-resolution model, and a super-resolution unit 134 that performs real-world super-resolution of an input image using the trained diffusion model.
[0131] From another perspective, the information processing device according to the present invention is characterized by comprising a learning unit 133 that trains a diffusion model so that CFG (classifier-free guidance) can be used in super-resolution using a diffusion model, and a super-resolution unit 134 that performs real-world super-resolution of an input image using the diffusion model in which CFG (classifier-free guidance) has become available.
[0132] By any or a combination of the above-described processes, the information processing device according to the present invention can realize a model that more appropriately super-resolutions (enhances image quality) low-resolution images.
[0133] [8. Hardware Configuration] Furthermore, the terminal device 10 and server device 100 according to the above-described embodiment are realized by a computer 1000 having a configuration such as that shown in Figure 8. The following explanation will use the server device 100 as an example. Figure 8 is a diagram showing an example of the hardware configuration. The computer 1000 is connected to an output device 1010 and an input device 1020, and has a configuration in which an arithmetic unit 1030, a primary storage device 1040, a secondary storage device 1050, an output interface 1060, an input interface 1070, and a network interface 1080 are connected by a bus 1090.
[0134] The arithmetic unit 1030 operates based on programs stored in the primary storage device 1040 and the secondary storage device 1050, as well as programs read from the input device 1020, and executes various processes. The arithmetic unit 1030 can be implemented using, for example, a CPU (Central Processing Unit), an MPU (Micro Processing Unit), an ASIC (Application Specific Integrated Circuit), or an FPGA (Field Programmable Gate Array).
[0135] The primary storage device 1040 is a memory device, such as RAM (Random Access Memory), that temporarily stores data used by the arithmetic unit 1030 for various calculations. The secondary storage device 1050 is a storage device where data used by the arithmetic unit 1030 for various calculations and various databases are registered, and can be implemented using ROM (Read Only Memory), HDD (Hard Disk Drive), SSD (Solid State Drive), flash memory, etc. The secondary storage device 1050 may be internal storage or external storage. The secondary storage device 1050 may also be a removable storage medium such as USB (Universal Serial Bus) memory or SD (Secure Digital) memory card. The secondary storage device 1050 may also be cloud storage (online storage), NAS (Network Attached Storage), file server, etc.
[0136] The output I / F 1060 is an interface for transmitting information to be output to output devices 1010, such as displays, projectors, and printers, and is implemented using connectors of standards such as USB (Universal Serial Bus), DVI (Digital Visual Interface), and HDMI (High Definition Multimedia Interface). The input I / F 1070 is an interface for receiving information from various input devices 1020, such as mice, keyboards, keypads, buttons, and scanners, and is implemented using, for example, USB.
[0137] Furthermore, the output interface 1060 and input interface 1070 may be wirelessly connected to the output device 1010 and input device 1020, respectively. In other words, the output device 1010 and input device 1020 may be wireless devices.
[0138] Furthermore, the output device 1010 and the input device 1020 may be integrated as a touch panel. In this case, the output I / F 1060 and the input I / F 1070 may also be integrated as an input / output I / F.
[0139] The input device 1020 may also be a device that reads information from, for example, an optical recording medium such as a CD (Compact Disc), DVD (Digital Versatile Disc), or PD (Phase Change Rewritable Disk), a magneto-optical recording medium such as an MO (Magneto-Optical disk), a tape medium, a magnetic recording medium, or a semiconductor memory.
[0140] The network interface 1080 receives data from other devices via network N and sends it to the computing unit 1030, and also transmits data generated by the computing unit 1030 to other devices via network N.
[0141] The arithmetic unit 1030 controls the output device 1010 and the input device 1020 via the output interface 1060 and the input interface 1070. For example, the arithmetic unit 1030 loads a program from the input device 1020 or the secondary storage device 1050 onto the primary storage device 1040 and executes the loaded program.
[0142] For example, when computer 1000 functions as a server device 100, the arithmetic unit 1030 of computer 1000 realizes the functions of the control unit 130 by executing a program loaded onto the primary storage device 1040. Alternatively, the arithmetic unit 1030 of computer 1000 may load a program obtained from another device via the network interface 1080 onto the primary storage device 1040 and execute the loaded program. Furthermore, the arithmetic unit 1030 of computer 1000 may cooperate with other devices via the network interface 1080 and call and use program functions, data, etc., from other programs on other devices.
[0143] [9. Other] Although embodiments of the present invention have been described above, the present invention is not limited by the content of these embodiments. Furthermore, the aforementioned components include those that can be easily conceived by those skilled in the art, those that are substantially the same, and those that fall within the so-called equivalent range. Moreover, the aforementioned components can be combined as appropriate. Furthermore, various omissions, substitutions, or modifications of the components can be made without departing from the gist of the embodiments described above.
[0144] Furthermore, among the processes described in the above embodiments, all or part of the processes described as being performed automatically can be performed manually, or all or part of the processes described as being performed manually can be performed automatically by known methods. In addition, the processing procedures, specific names, and information including various data and parameters shown in the above document and drawings can be arbitrarily changed unless otherwise specified. For example, the various information shown in each figure is not limited to the information shown.
[0145] Furthermore, the components of each illustrated device are functionally conceptual and do not necessarily need to be physically configured as shown. In other words, the specific forms of distribution and integration of each device are not limited to those shown, and all or part of them can be functionally or physically distributed and integrated in any unit according to various loads and usage conditions.
[0146] For example, the server device 100 described above may be implemented using multiple server computers, and the configuration can be flexibly changed, such as by calling external platforms via APIs (Application Programming Interfaces) or network computing depending on the function.
[0147] Furthermore, the embodiments and modifications described above can be combined as appropriate, provided that the processing content is not inconsistent.
[0148] Furthermore, the terms "section, module, unit" mentioned above can be replaced with "means" or "circuit." For example, the acquisition unit can be replaced with acquisition means or acquisition circuit. [Explanation of Symbols]
[0149] 1. Information Processing System 10 Terminal devices 100 Server Devices 110 Communications Department 120 Storage section 130 Control Unit 131 Acquisition Department 132 Generation part 133 Learning Department 134 Super-resolution section 135 Management Department 136 Provision Department
Claims
1. A generation unit that generates pairs of reference images and applied images obtained by applying a predetermined task to the reference image, for each of several different tasks, A learning unit that, upon inputting an application image and task information indicating the task applied to that application image, trains a diffusion model to generate a reference image. Equipped with, The generation unit decomposes the original task into multiple tasks and generates a set of reference image and application image for each task. An information processing device characterized by the following:
2. A generation unit that generates pairs of a reference image and an applied image obtained by applying a predetermined task to the reference image, each for a plurality of different tasks, A learning unit that, upon inputting an application image and task information indicating the task applied to that application image, trains a diffusion model to generate a reference image. Equipped with, The generation unit generates pairs of reference images and applied images for a reference task performed using a trained diffusion model, and further generates pairs of reference images and applied images for constituent tasks that make up the reference task. An information processing device characterized by the following:
3. The learning unit trains a diffusion model to perform the process of generating an applied image by gradually degrading a reference image according to a predetermined task, and the process of reconstructing the reference image from the applied image step by step, reversing the degradation process. The information processing apparatus according to claim 1 or 2.
4. A generation unit that generates pairs of a reference image and an applied image obtained by applying a predetermined task to the reference image, each for a plurality of different tasks, A learning unit that, upon inputting an application image and task information indicating the task applied to that application image, trains a diffusion model to generate a reference image. Equipped with, The learning unit, upon receiving a class label as task information indicating a task, trains a diffusion model to generate a reference image. An information processing device characterized by the following:
5. A generation unit that generates pairs of a reference image and an applied image obtained by applying a predetermined task to the reference image, each for a plurality of different tasks, A learning unit that, upon inputting an application image and task information indicating the task applied to that application image, trains a diffusion model to generate a reference image. Equipped with, The generating unit is A pair of a high-resolution image as a reference image and an applied image obtained by gradually scaling that high-resolution image is generated as a training sample when the class label as task information is class 0. A pair of a high-resolution image as a reference image and an applied image obtained by gradually adding noise to that high-resolution image is generated as a training sample when the class label as task information is class 1. A pair of a high-resolution image as a reference image and an applied image obtained by gradually scaling and adding noise to that high-resolution image is generated as a training sample when the class label as task information is class 2. The learning unit trains the diffusion model using the generated training samples for each class label. An information processing device characterized by the following:
6. The learning unit removes class labels with a certain probability and trains a diffusion model using the generated training samples without class labels, thereby making CFG usable for super-resolution by the diffusion model. The information processing apparatus according to feature 5.
7. A generation unit that generates pairs of a reference image and an applied image obtained by applying a predetermined task to the reference image, each for a plurality of different tasks, A learning unit that, upon inputting an application image and task information indicating the task applied to that application image, trains a diffusion model to generate a reference image. A super-resolution unit that takes an application image and task information indicating a real-world super-resolution task as input to a trained diffusion model, performs real-world super-resolution on the application image, and generates a real-world super-resolution image equivalent to a reference image. An information processing device characterized by comprising:
8. A learning unit that decomposes the super-resolution task into multiple classes and trains a diffusion model to become a class-conditional super-resolution model, A super-resolution unit that performs real-world super-resolution of an input image using a pre-trained diffusion model, An information processing device characterized by comprising:
9. A learning unit that trains the diffusion model so that CFG (classifier-free guidance) can be used in super-resolution using the diffusion model, A super-resolution unit that performs real-world super-resolution of an input image using a diffusion model for which classifier-free guidance (CFG) is available, An information processing device characterized by comprising:
10. An information processing method performed by an information processing device, A generation process that generates pairs of reference images and applied images obtained by applying a predetermined task to the reference image, each for multiple different tasks, A learning process in which a diffusion model is trained to generate a reference image when an application image and task information indicating the task applied to that application image are input, Includes, In the generation process described above, the original task is broken down into multiple tasks, and a pair of a reference image and an applied image is generated for each task. An information processing method characterized by the following:
11. A generation procedure for generating pairs of reference images and applied images obtained by applying a predetermined task to the reference image, each for multiple different tasks, A learning procedure for training a diffusion model to generate a reference image when given an application image and task information indicating the task applied to that application image as input, An information processing program that causes a computer to execute, In the aforementioned generation procedure, the original task is broken down into multiple tasks, and a pair of a reference image and an applied image is generated for each task. An information processing program characterized by the following features.