Information processing apparatus, information processing method, and information processing program
The information processing apparatus addresses the challenge of super-resolving low-resolution images by using a diffusion model to generate high-resolution images from low-resolution images and task information, effectively improving image quality.
Patent Information
- Application Number
- JP2023196802
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-20
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2043-11-20
AI Technical Summary
Existing image processing techniques struggle to effectively super-resolve low-resolution images degraded by noise, blur, and compression in a complex manner.
An information processing apparatus that includes a generation unit for creating reference and applied images for various tasks, and a learning unit that utilizes a diffusion model to generate high-resolution images from low-resolution images and task information.
The solution enables more appropriate super-resolution of low-resolution images, improving image quality by effectively addressing noise, blur, and compression degradation.
Smart Images

Figure 2025083111000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an information processing apparatus, an information processing method, and an information processing program.
Background Art
[0002] A technique related to an image processing circuit that executes noise reduction processing for reducing noise in an image is disclosed (see Patent Document 1).
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] However, there is room for improvement in the above prior art. For example, in the above prior art, there is room for improvement in terms of super-resolving (enhancing the image quality) a low-resolution image that contains degradation due to noise, blur, compression, etc. in a complex manner. Therefore, a method for more appropriately super-resolving (enhancing the image quality) a low-resolution image is required.
[0005] The present application has been made in view of the above, and an object thereof is to realize a model that more appropriately super-resolves (enhances the image quality) a low-resolution image.
Means for Solving the Problems
[0006] The information processing apparatus according to the present application includes a generation unit that generates a set of a reference image and an applied image obtained by applying a predetermined task to the reference image for each of a plurality of different tasks, and a learning unit that causes a diffusion model to learn to generate a reference image when an applied image and task information indicating the task applied to the applied image are input.
Effects of the Invention
[0007] According to one aspect of the embodiment, a model for more appropriately super-resolving (enhancing the image quality of) a low-resolution image can be realized.
Brief Description of the Drawings
[0008]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Modes for Carrying Out the Invention
[0009] Hereinafter, embodiments for implementing the information processing apparatus, information processing method, and information processing program according to the present application (hereinafter referred to as "embodiments") will be described in detail with reference to the drawings. Note that the information processing apparatus, information processing method, and information processing program according to the present application are not limited by this embodiment. Also, in the following embodiments, the same parts are denoted by the same reference numerals, and duplicate explanations are omitted.
[0010] 〔1. Overview of the Information Processing System〕 First, referring to FIG. 1, the outline of the information processing system according to the embodiment will be described. FIG. 1 is an explanatory diagram showing the outline of the information processing system according to the embodiment. As shown in FIG. 1, the information processing system 1 according to the embodiment includes a terminal device 10 and a server device 100. These various devices are connected to each other via a network N so as to be communicable with each other by wire or wirelessly. Thereby, the terminal device 10 can cooperate with the server device 100. The network N is, for example, a LAN (Local Area Network) or a WAN (Wide Area Network) such as the Internet.
[0011] The terminal device 10 is an information processing device used by a user U. For example, the terminal device 10 is a smart device such as a smartphone (smartphone) or a tablet terminal, a mobile phone such as a feature phone (galaxy phone), a PC (Personal Computer), a PDA (Personal Digital Assistant), a game machine or an AV device having a communication function, an information home appliance or a digital home appliance, a car navigation system, a wearable device (Wearable Device) such as a smart watch or a head-mounted display, or smart glasses. Further, the terminal device 10 may be a house, building, vehicle, home appliance, electronic device, etc. corresponding to the IoT (Internet of Things).
[0012] In this embodiment, the terminal device 10 is a smart device such as a smartphone or a tablet terminal used by the user U, and is a portable terminal device capable of communicating with an arbitrary server device via a wireless communication network such as LTE (Long Term Evolution), 4G (4th Generation), or 5G (5th Generation: 5th Generation Mobile Communication System). Further, the terminal device 10 has a screen such as a liquid crystal display and has a screen having a touch panel function, and receives various operations on display data such as content, such as a tap operation, a slide operation, and a scroll operation, from the user U using a finger, a stylus, or the like. Note that, among the operations performed on the screen, an operation performed on the area where the content is displayed may be regarded as an operation on the content. Further, the terminal device 10 may be not only a smart device but also an information processing device such as a desktop PC (Personal Computer) or a notebook PC.
[0013] Further, such a terminal device 10 can be connected to the network N via a wireless communication network such as LTE, 4G, or 5G, or a short-range wireless communication such as Bluetooth (registered trademark) or wireless LAN (Local Area Network), and communicate with the server device 100.
[0014] The server device 100 is, for example, a computer such as a PC or a blade server, or a mainframe or a workstation. Note that the server device 100 may be realized by cloud computing.
[0015] In this embodiment, the server device 100 is an information processing device that cooperates with the terminal device 10 of each user U and provides an API (Application Programming Interface) service and various data for various applications (hereinafter referred to as apps) and the like to the terminal device 10 of each user U, and is realized by a computer, a cloud system, or the like.
[0016] Further, the server device 100 may be an information processing device that provides some kind of web service online to the terminal device 10 of each user U. For example, as a web service, the server device 100 may provide services such as Internet connection, search service, SNS (Social Networking Service), electronic commerce (EC: Electronic Commerce), electronic payment, online game, online banking, online trading, accommodation and ticket reservation, video and music distribution, news, map, route search, route guidance, route information, operation information, weather forecast, etc. In fact, the server device 100 may cooperate with various servers that provide the above web services and mediate the web service, or be responsible for the processing of the web service.
[0017] In addition, the server device 100 can acquire user information about the user U. For example, the server device 100 acquires information (attribute information) regarding the attributes of the user U, such as the gender, age, and residential area of the user U, as user information. Further, the server device 100 can acquire information regarding attributes such as the demographics (demographic attributes), psychographics (psychological attributes), geographics (geographical attributes), and behavioral (behavioral attributes) of the user U. Also, the server device 100 may acquire, as user information, the segment or persona (persona) to which the user U belongs in the field of marketing. Then, the server device 100 stores and manages the information (attribute information) regarding the attributes of the user U together with the identification information (such as user ID) indicating the user U.
[0018] In addition, the server device 100 acquires various types of history information (log data) indicating the actions of the user U from the terminal device 10 of the user U or from various servers etc. based on the user ID etc. For example, the server device 100 acquires a location history, which is a history of the location and date / time of the user U, from the terminal device 10. Also, the server device 100 acquires a search history, which is a history of the search queries input by the user U, from a search server (search engine). Further, the server device 100 acquires a browsing history, which is a history of the content browsed by the user U, from a content server. Additionally, the server device 100 acquires a purchase history (payment history), which is a history of the user U's product purchases and payment processing, from an e-commerce server or a payment processing server. Moreover, the server device 100 may acquire a listing history or a sales history, which is a history of the user U's listings on a marketplace, from an e-commerce server or a payment processing server. Also, the server device 100 acquires a posting history, which is a history of the user U's posts, from a posting server or an SNS server that provides a word-of-mouth posting service. Note that each of the above-mentioned various servers etc. may be the server device 100 itself. That is, the server device 100 may function as each of the above-mentioned various servers etc.
[0019] Also, the number of each device included in the information processing system 1 shown in FIG. 1 is not limited to that shown. For example, in FIG. 1, for the sake of simplicity of illustration, only one terminal device 10 is shown, but this is merely an example and is not limiting, and there may be two or more.
[0020] [2. Real-world Super-Resolution] Real-world super-resolution is a super-resolution (image quality improvement) task for images that contain a complex combination of noise, blur, and degradation due to compression. For example, when not related to the real world, it refers to a task of restoring an image reduced by a factor of 2 or 4 to its original resolution.
[0021] [2-1. Mechanism of Super-Resolution by Diffusion Model] The diffusion model is a type of generative AI. It is a generative model that learns a process of gradually adding noise to a target image to degrade it (diffusion process) and a process of gradually removing noise (denoising) to reconstruct the image by tracing back the degradation process (reverse diffusion process).
[0022] In the super-resolution mechanism using the diffusion model, the server device 100 uses the low-resolution image as the condition for conditional generation by the diffusion model and generates a high-resolution image from the noisy image. Note that when training the model, the server device 100 degrades the high-resolution image that is the original image data by blurring, JPEG compression, etc. to generate a low-resolution image.
[0023] For example, the server device 100 generates an image with noise removed by inputting a low-resolution image as a condition to the noisy image and then inputting it to the diffusion model. Then, the server device 100 inputs the low-resolution image to the image with noise removed (the generated image) and inputs it to the diffusion model to generate an image with further noise removed. The server device 100 repeats this process to generate a high-resolution image.
[0024] For images, for example, image representations (representation methods) with 3 channels such as RGB (red, green, blue) and YcbCr (luminance signal (Y), difference between luminance signal and blue component (Cb), difference between luminance signal and red component (Cr)) are used. In the case of super-resolution by the diffusion model, by inputting a total of 6 channels obtained by adding the 3 channels of the low-resolution image to the 3 channels of the noisy image, an image with 3 channels with reduced noise is output. For example, when using an image representation with 3 channels of YCbCr, a low-resolution image in YCbCr format is used by overlapping it with the noisy image.
[0025] At this time, the server device 100 further inputs the time information t to the diffusion model and repeats the above noise removal process according to the time information t input to the diffusion model. That is, the server device 100 performs noise removal according to the input time information t.
[0026] The time information t indicates the degree of current noise (noise level). For example, if there are about 1000 levels of noise intensity, 1000 is input as the time information t, and the generation of an image with gradually removed noise is repeated 1000 times. Each time the process is executed, the value of the time information t is decreased by 1, and as the number of repetitions increases, the noise gradually (step by step) decreases, and finally, when it reaches 0, a clean image without noise (high-quality image) is output.
[0027] When learning, the above noise removal process is reversed. For example, the server device 100 generates an image with only noise by gradually adding noise from a noise level of 0 (no noise) to 1000 (complete noise) based on a high-quality image, and then makes the diffusion model learn to gradually remove noise (denoise) from the noisy image to finally generate a high-quality image by following the reverse process.
[0028] In this way, the server device 100 uses the diffusion model to repeatedly remove noise (denoise) from complete noise to generate a high-resolution image.
[0029] For example, as shown in FIG. 1, the server device 100 inputs a noisy image and a low-resolution image into the diffusion model (step S1).
[0030] Also, the server device 100 inputs the time information t into the diffusion model (step S2).
[0031] The server device 100 uses the diffusion model to generate an image with one-step removed noise from the input noisy image and low-resolution image according to the time information t (step S3).
[0032] Then, each time the server device 100 executes the above process, the value of the time information t is decreased by one, the image with removed noise and the low-resolution image are input into the diffusion model, and the above noise removal process is repeated until the value of the time information t becomes 0 (step S4).
[0033] Finally, the server device 100 generates a high-resolution image (step S5).
[0034] However, currently, it is difficult to say that the generation (reconstruction) of high-resolution images by the diffusion model is perfect, and there is room for improvement. In this embodiment, classifier-free guidance (CFG) is combined with super-resolution by the diffusion model to improve the quality.
[0035] [2-2. Explanation of classifier-free guidance (CFG)] Classifier-free guidance (CFG) is a technique mainly used in the Text2Image generation model. Originally, there was a technique called classifier guidance, which separately trained a special classifier to improve the generation result. Classifier guidance is to use a classifier trained with a noisy image to guide the generation of an image as per the label. CFG (classifier-free guidance) is called so because it can be realized without separately training a classifier.
[0036] In CFG (classifier-free guidance), during the learning of the model, the text condition is dropped with a certain probability while learning. During generation, the generation result with the text condition and the generation result without the text condition are calculated respectively, and calculated as follows.
[0037] Generation result (CFG (classifier-free guidance)) = scale × (generation result (with text condition) - generation result (without text condition)) + generation result (without text condition)
[0038] Here, scale indicates the strength of the guidance (guidance scale). That is, the value of scale corresponds to the strength of the guidance. The larger the value of Scale, the stronger the guidance. When actually running the program, since scale is specified as a parameter, the user can specify any value. Note that when Scale = 1, no guidance is performed. When scale = 1, the item of "no text condition" is canceled and disappears. Therefore, when performing guidance, the value of scale should be greater than 1.
[0039] Figure 2 is a diagram showing an image of CFG (classifier-free guidance) in the 2D case (scale = 2). Among the lines shown in Figure 2, A represents the generation result (with text condition), B represents the generation result (without text condition), and C represents the generation result (CFG (classifier-free guidance)). The coordinate at the tip of the arrow of C represents a generation result that reflects the text condition more than the coordinate at the tip of the arrow of A.
[0040] Note that when applying CFG (classifier-free guidance), the fidelity increases while the diversity decreases. For example, by applying CFG (classifier-free guidance), images of the same domain are more likely to be generated, and images of different domains are less likely to be generated. However, for real-world super-resolution, since it is to convert noisy or low-resolution images of the same domain into high-resolution images, it is preferable that the fidelity increases, and it is not a problem even if the diversity decreases. That is, by applying CFG (classifier-free guidance), a secondary effect of increasing the fidelity in real-world super-resolution can be obtained.
[0041] [2-3. Super-resolution with class conditions] Referring to FIG. 3, the real-world super-resolution applying CFG will be described. FIG. 3 is an explanatory diagram showing an overview of the real-world super-resolution applying CFG. The low-resolution image as a condition for conditional generation by the above diffusion model is an essential condition (indispensable condition) for performing super-resolution and cannot be changed because without it, it is impossible to indicate the direction of generating a high-resolution image from a noisy image and there is a possibility of generating a completely different image. Therefore, in order to add new conditions or manipulate conditions, it is necessary to specify conditions in another form.
[0042] In this embodiment, the server device 100 decomposes a task of performing real-world super-resolution (real-world super-resolution task) into a plurality of classes (tasks) to perform class-conditioned super-resolution. That is, the server device 100 decomposes the original task (reference task) into a plurality of tasks (constituent tasks constituting the reference task) and changes it to a class-conditioned super-resolution model.
[0043] Specifically, the server device 100 inputs a class label c (task label) together with time information t to the diffusion model. For example, the class label c is set to "Class 0: Super-resolution", "Class 1: Denoising only", "Class 2: Real-world super-resolution". Also, the class label c may be set to "Class 0: 2x super-resolution", "Class 1: 3x super-resolution", "Class 2: 4x super-resolution". At this time, the original task (reference task) may be set to Class 2 and decomposed into a plurality of tasks (constituent tasks) such as Class 0 and Class 1.
[0044] In addition, when learning the model, the server device 100 adds degradation assumed in the real world to the high-resolution image to generate learning samples. At this time, the server device 100 adds degradation according to the class label c (task label) to the high-resolution image. For example, when set to "Class 0: Super-resolution", "Class 1: Denoising only", "Class 2: Real-world super-resolution", the server device 100 adds only scaling (enlarging or reducing) in the case of Class 0, adds only noise addition in the case of Class 1, and adds both scaling and noise addition in the case of Class 2.
[0045] The server device 100 gradually adds degradation according to the class label c (task label) to the high-resolution image. Then, the server device 100 causes the diffusion model to learn the training samples generated for each class (task) step by step together with the class label c (task label). For example, in the case of class 1, the server device 100 gradually adds noise from noise level 0 (no noise) to 1000 (complete noise) based on the high-quality image to finally generate an image consisting only of noise, and by tracing the reverse process, it causes the diffusion model to learn to gradually remove noise (denoise) from the noisy image to finally generate a high-quality image. The same applies to the cases of class 0 and class 2. For example, in the case of image reduction, the server device 100 gradually reduces the high-quality image from reduction level 0 (original image) to 1000 (an image reduced by 2 or 4 times) to finally generate an image reduced by 2 or 4 times.
[0046] At this time, the server device 100 deletes (drops) the class label c with a certain probability (for example, a probability of about 10%) during model learning to enable CFG (classifier-free guidance). The diffusion model operates according to class labels c such as class 0, class 1, and class 2. When there is no label, it is not known which class (task) it is, but it will perform an average operation so that there is no problem regardless of which class it is. In this way, the diffusion model will perform a safe operation when there is no label. Therefore, when there is no label, the performance is at least worse than in the case of class 2.
[0047] Also, when generating a high-resolution image, the server device 100 can improve the quality of real-world super-resolution by specifying class 2 as the class label c input to the diffusion model and applying the guidance of CFG (classifier-free guidance). At this time, just specifying class 2 can improve the quality of real-world super-resolution, and applying the guidance can further improve the quality of real-world super-resolution.
[0048] For example, as shown in FIG. 3, the server device 100 generates a low-resolution image by degrading a high-resolution image that is the original image data through blurring, JPEG compression, or the like (step S11).
[0049] Further, the server device 100 decomposes the super-resolution task into a plurality of classes (tasks) and sets a class label c (task label) (step S12).
[0050] Subsequently, the server device 100 gradually applies degradation according to the class label c to the high-resolution image to generate a training sample (step S13).
[0051] Subsequently, the server device 100 gradually causes the generated training sample to be learned by the diffusion model together with the class label c (step S14).
[0052] At this time, the server device 100 deletes (drops) the class label c with a certain probability (for example, a probability of about 10%) to enable classifier-free guidance (CFG) (step S15).
[0053] After that, the server device 100 superimposes the low-resolution image as a condition on the input image (noisy image, etc.) and inputs it to the diffusion model, and inputs the time information t and the class label c to the diffusion model, thereby performing real-world super-resolution on the input image and finally generating a high-resolution image (step S16).
[0054] 〔2-4. Effect of CFG (classifier-free guidance)〕 Regarding the result of performing classifier-free guidance (CFG) by specifying a class, it was evaluated using the datasets of RealSRv3 and DrealSR, which are widely used for the evaluation of real-world super-resolution. Note that the lower the FID10k value, the higher the quality.
[0055] Figure 4 is a diagram showing the evaluation of the results of implementing CFG. As shown in Figure 4, rather than specifying classes 0 or 1, it is better to strengthen the guidance by specifying class 2 (real-world super-resolution) and setting the scale of CFG to 1.5 or 2.0, as this results in a lower FID10k value and improved quality. Here, since the value of the scale corresponds to the strength of the guidance, it means that the guidance is stronger when scale = 2.0 than when scale = 1.5.
[0056] 〔2-5. Supplementary〕 During the learning of the model, the server device 100 generates a set of a reference image and an applied image obtained by applying a predetermined task to the reference image for each of a plurality of different tasks. The reference image is, for example, a high-resolution image. Also, the applied image is, for example, an image obtained by adding noise to the reference image in one step. The plurality of tasks include a reference task performed using a learned model and constituent tasks that make up the reference task. Note that the plurality of tasks include a task of generating a real-world super-resolution image (high-resolution image).
[0057] Note that the noise is added step by step. For example, the server device 100 uses an image obtained by adding noise to a high-resolution image in n steps as the reference image, and an image obtained by adding noise to the reference image in one step (an image obtained by adding noise to the high-resolution image in n + 1 steps) as the applied image, and generates a set of the reference image and the applied image. Then, the server device 100 repeats this for a predetermined number of times (for example, 1000 times) corresponding to the above-mentioned time information t, using the current applied image as the next reference image.
[0058] Then, when the server device 100 inputs an applied image and task information indicating the task applied to the applied image, it causes the model to learn to generate a reference image. That is, the server device 100 causes the model to learn each set of the reference image and the applied image generated the above-mentioned predetermined number of times.
[0059] As a result, the server device 100 causes the model to learn to generate an image (reference image) with gradually (step by step) removed noise from the noisy image (applied image) and the low-resolution image (image according to the task) according to the time information t. Then, the server device 100 uses the learned model to generate and output a real-world super-resolution image (high-resolution image) of the input image according to the time information t.
[0060] In the present embodiment, classifier-free guidance (CFG) is applied to super-resolution by a diffusion model, and the level of how much processing is to be performed is input. Since CFG is a technique that makes it more faithful to revert to the original as the level is increased, by applying this, the server device 100 inputs a label indicating a task and causes the model to learn to execute a task according to the label.
[0061] The server device 100 performs a certain amount of learning without labels in parallel with the learning with labels. Since the model comes to estimate labels and execute tasks, it comes to perform average operations. At this time, at least the ultimate goal may be included in the labels.
[0062] As the combination of labels, it is preferably set to a task that constitutes real-world super-resolution. Or, it may be the level difference of real-world super-resolution that only differs in magnification.
[0063] Applying classifier-free guidance (CFG) reduces the diversity, but since the main purpose of super-resolution is to improve the resolution of images in the same domain, the reduction in diversity is one of the advantages.
[0064] [3. Configuration Example of Terminal Device] Next, the configuration of the terminal device 10 will be described with reference to FIG. 5. FIG. 5 is a diagram showing a configuration example of the terminal device 10. As shown in FIG. 5, the terminal device 10 includes a communication unit 11, a display unit 12, an input unit 13, a positioning unit 14, a sensor unit 20, a control unit 30 (controller), and a storage unit 40.
[0065] (Communication unit 11) The communication unit 11 is connected to the network N either by wire or wirelessly, and transmits and receives information to and from the server device 100 via the network N. For example, the communication unit 11 is realized by a NIC (Network Interface Card), an antenna, or the like.
[0066] (Display unit 12) The display unit 12 is a display device that displays various types of information such as location information. For example, the display unit 12 is a liquid crystal display (LCD) or an organic electro-luminescent display (Organic Electro-Luminescent Display). Further, the display unit 12 is a touch panel type display, but is not limited thereto.
[0067] (Input unit 13) The input unit 13 is an input device that receives various operations from the user U. For example, the input unit 13 has buttons or the like for inputting characters, numbers, etc. Note that the input unit 13 may be an input / output port (I / O port), a USB (Universal Serial Bus) port, or the like. Further, when the display unit 12 is a touch panel type display, a part of the display unit 12 functions as the input unit 13. Also, the input unit 13 may be a microphone or the like that receives voice input from the user U. The microphone may be wireless.
[0068] (Positioning unit 14) The positioning unit 14 receives signals (radio waves) transmitted from GPS (Global Positioning System) satellites, and based on the received signals, acquires position information (for example, latitude and longitude) indicating the current position of the terminal device 10, which is the own device. That is, the positioning unit 14 measures the position of the terminal device 10. Note that GPS is only an example of GNSS (Global Navigation Satellite System).
[0069] In addition to GPS, the positioning unit 14 can also measure the position by various methods. For example, the positioning unit 14 may measure the position by using various communication functions of the terminal device 10 as follows as auxiliary positioning means for position correction and the like.
[0070] (Wi-Fi positioning) For example, the positioning unit 14 measures the position of the terminal device 10 by using the Wi-Fi (registered trademark) communication function of the terminal device 10 and the communication network provided by each communication company. Specifically, the positioning unit 14 performs Wi-Fi communication or the like, and measures the distance from a nearby base station or access point to measure the position of the terminal device 10.
[0071] (Beacon positioning) In addition, the positioning unit 14 may measure the position by using the Bluetooth (registered trademark) function of the terminal device 10. For example, the positioning unit 14 measures the position of the terminal device 10 by connecting to a beacon transmitter connected by the Bluetooth (registered trademark) function.
[0072] (Geomagnetic positioning) In addition, the positioning unit 14 measures the position of the terminal device 10 based on the geomagnetic pattern of the structure measured in advance and the geomagnetic sensor provided in the terminal device 10.
[0073] (RFID positioning) Also, for example, when the terminal device 10 has a function equivalent to a non-contact IC card used at a station ticket gate or a store, or has a function of reading an RFID (Radio Frequency Identification) tag, the used position is recorded together with the information on which settlement or the like is performed by the terminal device 10. The positioning unit 14 may measure the position of the terminal device 10 by acquiring such information. Also, the position may be measured by an optical sensor, an infrared sensor, or the like provided in the terminal device 10.
[0074] The positioning unit 14 may, if necessary, measure the position of the terminal device 10 using one or a combination of the above-described positioning means.
[0075] (Sensor unit 20) The sensor unit 20 includes various sensors mounted on or connected to the terminal device 10. Note that the connection may be either wired or wireless. For example, the sensors may be detection devices other than the terminal device 10, such as wearable devices or wireless devices. In the example shown in FIG. 5, the sensor unit 20 includes an acceleration sensor 21, a gyro sensor 22, a pressure sensor 23, a temperature sensor 24, a sound sensor 25, a light sensor 26, a magnetic sensor 27, and an image sensor (camera) 28.
[0076] Note that the above-described sensors 21 to 28 are merely examples and are not limited thereto. That is, the sensor unit 20 may be configured to include some of the sensors 21 to 28, or may include other sensors such as a humidity sensor in addition to or instead of the sensors 21 to 28.
[0077] The acceleration sensor 21 is, for example, a three-axis acceleration sensor, and detects physical movements of the terminal device 10 such as the moving direction, speed, and acceleration of the terminal device 10. The gyro sensor 22 detects physical movements of the terminal device 10 such as the inclination in three-axis directions based on the angular velocity of the terminal device 10. The pressure sensor 23 detects, for example, the atmospheric pressure around the terminal device 10.
[0078] Since the terminal device 10 includes the above-described acceleration sensor 21, gyro sensor 22, pressure sensor 23, etc., it becomes possible to measure the position of the terminal device 10 using technologies such as pedestrian dead reckoning (PDR) that utilize these sensors 21 to 23, etc. As a result, it becomes possible to obtain position information indoors, which is difficult to obtain with a positioning system such as GPS.
[0079] For example, a pedometer using the acceleration sensor 21 can calculate the number of steps, the walking speed, and the distance walked. Also, by using the gyro sensor 22, the traveling direction, the direction of the line of sight, and the body inclination of the user U can be known. Further, from the atmospheric pressure detected by the atmospheric pressure sensor 23, the altitude at which the terminal device 10 of the user U exists and the floor number can also be known.
[0080] The temperature sensor 24 detects, for example, the temperature around the terminal device 10. The sound sensor 25 detects, for example, the sound around the terminal device 10. The light sensor 26 detects the illuminance around the terminal device 10. The magnetic sensor 27 detects, for example, the geomagnetism around the terminal device 10. The image sensor 28 captures an image around the terminal device 10.
[0081] The above-described atmospheric pressure sensor 23, temperature sensor 24, sound sensor 25, light sensor 26, and image sensor 28 can detect the environment and situation around the terminal device 10 by detecting the atmospheric pressure, temperature, sound, illuminance, or capturing an image of the surroundings, respectively. Also, it becomes possible to improve the accuracy of the position information of the terminal device 10 from the environment and situation around the terminal device 10.
[0082] (Control unit 30) The control unit 30 includes, for example, a microcomputer having a CPU (Central Processing Unit), ROM (Read Only Memory), RAM, input / output ports, etc., and various circuits. Also, the control unit 30 may be configured by hardware such as an integrated circuit such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array). The control unit 30 has a transmission unit 31, a reception unit 32, and a processing unit 33.
[0083] (Transmission unit 31) The transmission unit 31 can transmit various information input by the user U using, for example, the input unit 13, various information detected by the sensors 21 to 28 mounted on or connected to the terminal device 10, the position information of the terminal device 10 positioned by the positioning unit 14, etc. to the server device 100 via the communication unit 11.
[0084] (Receiving unit 32) The receiving unit 32 can receive various information provided from the server device 100 and requests for various information from the server device 100 via the communication unit 11.
[0085] (Processing unit 33) The processing unit 33 controls the entire terminal device 10 including the display unit 12 etc. For example, the processing unit 33 can output and display various information transmitted by the transmission unit 31 and various information from the server device 100 received by the receiving unit 32 on the display unit 12.
[0086] (Storage unit 40) The storage unit 40 is realized by, for example, semiconductor memory elements such as RAM (Random Access Memory), flash memory, or storage devices such as HDD (Hard Disk Drive), SSD (Solid State Drive), optical disks. Various programs and various data etc. are stored in such a storage unit 40.
[0087] [4. Configuration example of server device] Next, with reference to FIG. 6, the configuration of the server device 100 according to the embodiment will be described. FIG. 6 is a diagram showing a configuration example of the server device 100 according to the embodiment. As shown in FIG. 6, the server device 100 includes a communication unit 110, a storage unit 120, and a control unit 130.
[0088] (Communication unit 110) The communication unit 110 is realized by, for example, a NIC (Network Interface Card) etc. Also, the communication unit 110 is connected to the network N by wire or wirelessly.
[0089] (Memory unit 120) The memory unit 120 is realized by, for example, a semiconductor memory element such as a RAM (Random Access Memory) or a flash memory, or a storage device such as an HDD, an SSD, or an optical disk. The memory unit 120 may store the attribute information and history information (log data) of the user U together with the identification information (such as user ID) indicating the user U.
[0090] (Control unit 130) The control unit 130 is a controller, and is realized, for example, by a CPU (Central Processing Unit), an MPU (Micro Processing Unit), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), etc., when various programs (corresponding to an example of an information processing program) stored in the internal storage device of the server device 100 are executed with a storage area such as a RAM as a work area. In the example shown in FIG. 6, the control unit 130 includes an acquisition unit 131, a generation unit 132, a learning unit 133, a super-resolution unit 134, a management unit 135, and a provision unit 136.
[0091] (Acquisition unit 131) The acquisition unit 131 acquires the search query input by the user U. For example, when the user U inputs a search query to a search engine or the like and performs a keyword search, the acquisition unit 131 acquires the search query via the communication unit 110. That is, the acquisition unit 131 acquires the keyword input by the user U to the search window of the search engine, site, or application via the communication unit 110.
[0092] Further, the acquisition unit 131 acquires user information regarding the user U via the communication unit 110. For example, the acquisition unit 131 acquires identification information (such as a user ID) indicating the user U, the location information of the user U, the attribute information of the user U, etc. from the terminal device 10 of the user U. Further, the acquisition unit 131 may acquire identification information indicating the user U, the attribute information of the user U, etc. at the time of user registration of the user U. Then, the acquisition unit 131 stores the user information in the storage unit 120.
[0093] Further, the acquisition unit 131 acquires various history information (log data) indicating the actions of the user U via the communication unit 110. For example, the acquisition unit 131 acquires various history information indicating the actions of the user U from the terminal device 10 of the user U or from various servers, etc. based on the user ID, etc. Then, the acquisition unit 131 stores the various history information in the storage unit 120.
[0094] Further, the acquisition unit 131 acquires image data from the terminal device 10 of the user U, another server device 100, or an external storage device or storage medium, etc. via the communication unit 110. Further, the acquisition unit 131 acquires the image data stored in the storage unit 120. For example, the acquisition unit 131 acquires a high-resolution image. Note that the domain of the high-resolution image is not limited. That is, the subject of the high-resolution image is arbitrary.
[0095] (Generation unit 132) The generation unit 132 generates a pair of a reference image and an applied image obtained by applying a predetermined task to the reference image for each of a plurality of different tasks. For example, when the reference image is a high-resolution image and the predetermined task is noise removal, the generation unit 132 generates a noisy image obtained by adding noise to the high-resolution image as the applied image.
[0096] At this time, the generation unit 132 decomposes the original task into a plurality of tasks, and generates a pair of a reference image and an application image for each task. That is, the generation unit 132 generates a pair of a reference image and an application image for a reference task performed using a learned diffusion model, and further generates a pair of a reference image and an application image for a constituent task constituting the reference task. For example, when the reference task is a real-world super-resolution task, the generation unit 132 decomposes the real-world super-resolution task into a super-resolution task and a noise removal task as constituent tasks. And the class label as task information is set such that the super-resolution task is class 0, the noise removal task is class 1, and the real-world super-resolution task is class 2.
[0097] For example, when the class label as task information is class 0, the generation unit 132 generates, as a training sample, a pair of a high-resolution image as a reference image and an application image obtained by gradually enlarging and shrinking the high-resolution image. Also, when the class label as task information is class 1, the generation unit 132 generates, as a training sample, a pair of a high-resolution image as a reference image and an application image obtained by gradually adding noise to the high-resolution image. Further, when the class label as task information is class 2, the generation unit 132 generates, as a training sample, a pair of a high-resolution image as a reference image and an application image obtained by gradually performing enlargement / shrinkage and noise addition on the high-resolution image.
[0098] (Learning unit 133) When the learning unit 133 inputs an application image and task information indicating the task applied to the application image, it causes the diffusion model to learn to generate a reference image. For example, the learning unit 133 causes the diffusion model to learn a process of generating an application image by gradually applying deterioration according to a predetermined task to a reference image, and a process of reconstructing the reference image from the application image step by step so as to reverse the deterioration process.
[0099] Further, when the learning unit 133 inputs a class label as task information indicating a task, it causes the diffusion model to learn to generate a reference image. At this time, the learning unit 133 causes the diffusion model to learn for each class label using the generated learning samples.
[0100] Also, the learning unit 133 deletes the class label with a certain probability (for example, a probability of about 10%), and causes the diffusion model to learn using the generated learning samples in a state without a class label, thereby applying classifier-free guidance (CFG) to super-resolution by the diffusion model and making it available.
[0101] From another perspective, the learning unit 133 causes the diffusion model to learn so as to decompose the super-resolution task into a plurality of classes and become a class-conditional super-resolution model.
[0102] From yet another perspective, the learning unit 133 causes the diffusion model to learn so that classifier-free guidance (CFG) can be used for super-resolution by the diffusion model.
[0103] (Super-resolution unit 134) The super-resolution unit 134 inputs an application image and task information indicating a real-world super-resolution task to the learned diffusion model, performs real-world super-resolution on the application image, and generates a reference image (actually, a real-world super-resolution image corresponding to the reference image).
[0104] For example, the super-resolution unit 134 inputs a noisy image as an application image and a low-resolution image to the learned diffusion model, and further inputs time information corresponding to the noise level and a class label as task information, thereby performing real-world super-resolution on the noisy image and generating a high-resolution image as a real-world super-resolution image corresponding to the reference image.
[0105] From another perspective, the super-resolution unit 134 performs real-world super-resolution of the input image using the learned diffusion model so as to become a class-conditional super-resolution model.
[0106] From another perspective, the super-resolution unit 134 performs real-world super-resolution of the input image using a diffusion model for which classifier-free guidance (CFG) has become available.
[0107] (Management unit 135) The management unit 135 stores the reference image (actually, the real-world super-resolution image corresponding to the reference image) generated by the super-resolution unit 134. For example, the management unit 135 stores (records) the generated reference image (actually, the real-world super-resolution image corresponding to the reference image) in the storage unit 120. Note that the management unit 135 may store the generated reference image (actually, the real-world super-resolution image corresponding to the reference image) in an external storage device or storage medium via the communication unit 110.
[0108] (Provision unit 136) The provision unit 136 provides, via the communication unit 110, the reference image (actually, the real-world super-resolution image corresponding to the reference image) generated by the super-resolution unit 134 or stored by the management unit 135 to the terminal device 10 of the user U, another server device 100, or an external storage device or storage medium.
[0109] [5. Processing procedure] Next, the processing procedure by the server device 100 according to the embodiment will be described with reference to FIG. 7. FIG. 7 is a flowchart showing the processing procedure according to the embodiment. Note that the following processing procedure is repeatedly executed by the control unit 130 of the server device 100.
[0110] For example, as shown in FIG. 7, the acquisition unit 131 of the server device 100 acquires a high-resolution image to be a reference image (step S101).
[0111] Subsequently, the generation unit 132 of the server device 100 decomposes the original task into a plurality of tasks, and for each task of the original task and the plurality of decomposed tasks, generates a pair of a high-resolution image to be a reference image and an applied image obtained by applying the task to the reference image (step S102).
[0112] Subsequently, for each task, the learning unit 133 of the server device 100 specifies a class label indicating the task, and causes the diffusion model to learn a process of generating an application image by gradually applying deterioration corresponding to the task to the reference image, and a process of reconstructing the reference image from the application image step by step so as to reverse the deterioration process (step S103).
[0113] At this time, the learning unit 133 of the server device 100 deletes the class label with a certain probability (for example, a probability of about 10%), and causes the diffusion model to learn using the generated training sample in a state without a class label, so that classifier-free guidance (CFG) can be used in super-resolution by the diffusion model. The diffusion model is trained (step S104).
[0114] Subsequently, the super-resolution unit 134 of the server device 100 inputs the application image and task information indicating the real-world super-resolution task into the trained diffusion model in which classifier-free guidance (CFG) is available, and performs real-world super-resolution on the application image to generate a reference image (actually, a real-world super-resolution image corresponding to the reference image) (step S105).
[0115] Subsequently, the management unit 135 of the server device 100 stores the reference image (actually, a real-world super-resolution image corresponding to the reference image) generated by the super-resolution unit 134 (step S106).
[0116] Subsequently, the providing unit 136 of the server device 100 provides the reference image (actually, a real-world super-resolution image corresponding to the reference image) stored by the management unit 135 to the outside (step S107).
[0117] 〔6. Modification Example〕 The above-described terminal device 10 and server device 100 may be implemented in various different forms other than the above-described embodiment. Therefore, below, modification examples of the embodiment will be described.
[0118] In the above embodiment, part or all of the processing executed by the server device 100 may actually be executed by the terminal device 10 (or an application operating on the terminal). For example, the processing may be completed in a stand-alone manner (by the terminal device 10 alone). In this case, it is assumed that the terminal device 10 has the functions of the server device 100 in the above embodiment. Further, in the above embodiment, since the terminal device 10 is in cooperation with the server device 100, from the perspective of the user U, the processing of the server device 100 also seems to be executed by the terminal device 10. That is, from another perspective, it can be said that the terminal device 10 includes the server device 100.
[0119] Further, in the above embodiment, the high-resolution image may be a high-resolution aerial photograph. Also, the low-resolution image may be a satellite photograph of a map application. Note that the high-resolution image and the low-resolution image are images in the same domain.
[0120] Also, in the above embodiment, the number of tasks when decomposing the original task (reference task) into a plurality of tasks (constituent tasks that constitute the reference task) is arbitrary. That is, the constituent tasks may be set according to the number of tasks that can be decomposed from the reference task.
[0121] Also, in the above embodiment, the number of steps when gradually adding degradation corresponding to the class label c (task label) to the high-resolution image is arbitrary. For example, the number of noise levels when gradually adding noise to the target image to degrade it is arbitrary.
[0122] 〔7. Effects〕 As described above, the information processing apparatus (terminal device 10 and server device 100) according to the present application includes a generation unit 132 that generates a set of a reference image (e.g., a high-resolution image) and an applied image (e.g., an image with noise) obtained by applying a predetermined task to the reference image for each of a plurality of different tasks, and a learning unit 133 that causes a diffusion model to learn to generate a reference image when an applied image and task information (e.g., a class label) indicating a task (class) applied to the applied image are input.
[0123] The generation unit 132 decomposes the original task into a plurality of tasks and generates a set of a reference image and an applied image for each task.
[0124] The generation unit 132 generates a set of a reference image and an applied image for a reference task performed using a learned diffusion model, and further generates a set of a reference image and an applied image for a constituent task that constitutes the reference task.
[0125] The learning unit 133 causes the diffusion model to learn a process of generating an applied image by gradually applying degradation corresponding to a predetermined task to a reference image and a process of gradually reconstructing the reference image from the applied image so as to reverse the degradation process.
[0126] When the learning unit 133 inputs a class label as task information indicating a task, it causes the diffusion model to learn to generate a reference image.
[0127] The generation unit 132 generates a pair of a high-resolution image as a reference image and an applied image obtained by gradually magnifying and shrinking the high-resolution image as a learning sample when the class label as task information is class 0, and generates a pair of a high-resolution image as a reference image and an applied image obtained by gradually adding noise to the high-resolution image as a learning sample when the class label as task information is class 1, and generates a pair of a high-resolution image as a reference image and an applied image obtained by gradually magnifying and shrinking and adding noise to the high-resolution image as a learning sample when the class label as task information is class 2. The learning unit 133 causes the diffusion model to be trained using the generated learning samples for each class label.
[0128] The learning unit 133 deletes the class label with a certain probability, and by causing the diffusion model to be trained using the generated learning samples in a state without a class label, applies classifier-free guidance (CFG) to super-resolution by the diffusion model to make it available.
[0129] Further, the information processing apparatus according to the present application further includes a super-resolution unit 134 that inputs an applied image and task information indicating a real-world super-resolution task to a trained diffusion model, performs real-world super-resolution on the applied image, and generates a real-world super-resolution image corresponding to the reference image.
[0130] From another perspective, the information processing apparatus according to the present application includes a learning unit 133 that decomposes a super-resolution task into a plurality of classes and causes the diffusion model to be trained so as to be a class-conditional super-resolution model, and a super-resolution unit 134 that performs real-world super-resolution of an input image using the trained diffusion model.
[0131] From yet another perspective, the information processing apparatus according to the present application includes a learning unit 133 that causes a diffusion model to learn so that classifier-free guidance (CFG) can be used in super-resolution by the diffusion model, and a super-resolution unit 134 that performs real-world super-resolution of an input image using the diffusion model in which classifier-free guidance (CFG) has become available.
[0132] By any one or combination of the above-described processes, the information processing apparatus according to the present application can realize a model that more appropriately super-resolves (improves the image quality) a low-resolution image.
[0133] 〔8. Hardware Configuration〕 Also, the terminal device 10 and the server device 100 according to the above-described embodiments are realized by a computer 1000 having a configuration as shown in FIG. 8, for example. Hereinafter, the server device 100 will be described as an example. FIG. 8 is a diagram showing an example of a hardware configuration. The computer 1000 has a form in which an output device 1010, an input device 1020 are connected, and an arithmetic device 1030, a primary storage device 1040, a secondary storage device 1050, an output I / F (Interface) 1060, an input I / F 1070, and a network I / F 1080 are connected by a bus 1090.
[0134] The arithmetic device 1030 operates based on programs stored in the primary storage device 1040 and the secondary storage device 1050, programs read from the input device 1020, etc., and executes various processes. The arithmetic device 1030 is realized by, for example, a CPU (Central Processing Unit), an MPU (Micro Processing Unit), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or the like.
[0135] The primary storage device 1040 is a memory device that primarily stores data used by the arithmetic unit 1030 for various calculations, such as a RAM (Random Access Memory). Also, the secondary storage device 1050 is a storage device in which data used by the arithmetic unit 1030 for various calculations and various databases are registered, and is realized by a ROM (Read Only Memory), HDD (Hard Disk Drive), SSD (Solid State Drive), flash memory, etc. The secondary storage device 1050 may be an internal storage or an external storage. Also, the secondary storage device 1050 may be a removable storage medium such as a USB (Universal Serial Bus) memory or an SD (Secure Digital) memory card. Also, the secondary storage device 1050 may be a cloud storage (online storage), NAS (Network Attached Storage), file server, etc.
[0136] The output I / F 1060 is an interface for transmitting information to be output to an output device 1010 that outputs various types of information, such as a display, a projector, and a printer, and is realized by a connector of a standard such as USB (Universal Serial Bus), DVI (Digital Visual Interface), or HDMI (registered trademark) (High Definition Multimedia Interface). Also, the input I / F 1070 is an interface for receiving information from various input devices 1020 such as a mouse, a keyboard, a keypad, a button, and a scanner, and is realized by, for example, USB or the like.
[0137] Also, the output I / F 1060 and the input I / F 1070 may be wirelessly connected to the output device 1010 and the input device 1020, respectively. That is, the output device 1010 and the input device 1020 may be wireless devices.
[0138] Further, the output device 1010 and the input device 1020 may be integrated like a touch panel. In this case, the output I / F 1060 and the input I / F 1070 may also be integrated as an input / output I / F.
[0139] Note that the input device 1020 may be a device that reads information from an optical recording medium such as a CD (Compact Disc), a DVD (Digital Versatile Disc), or a PD (Phase change rewritable Disk), a magneto-optical recording medium such as an MO (Magneto-Optical disk), a tape medium, a magnetic recording medium, or a semiconductor memory.
[0140] The network I / F 1080 receives data from other devices via the network N and sends it to the arithmetic unit 1030, and also sends the data generated by the arithmetic unit 1030 via the network N to other devices.
[0141] The arithmetic unit 1030 controls the output device 1010 and the input device 1020 via the output I / F 1060 and the input I / F 1070. For example, the arithmetic unit 1030 loads a program from the input device 1020 or the secondary storage device 1050 onto the primary storage device 1040 and executes the loaded program.
[0142] For example, when the computer 1000 functions as the server device 100, the arithmetic unit 1030 of the computer 1000 realizes the function of the control unit 130 by executing the program loaded onto the primary storage device 1040. Further, the arithmetic unit 1030 of the computer 1000 may load a program acquired from other devices via the network I / F 1080 onto the primary storage device 1040 and execute the loaded program. Further, the arithmetic unit 1030 of the computer 1000 may cooperate with other devices via the network I / F 1080 and call and use the functions and data of the program from other programs of other devices.
[0143] 〔9. Others〕 The embodiments of the present application have been described above, but the present invention is not limited by the contents of these embodiments. Further, the components described above include those that can be easily assumed by those skilled in the art, those that are substantially the same, and those within the so-called equivalent range. Furthermore, the above-described components can be combined as appropriate. Moreover, various omissions, substitutions, or changes of the components can be made without departing from the gist of the above-described embodiments.
[0144] Also, among the respective processes described in the above embodiments, all or part of the processes described as being automatically performed can be manually performed, or all or part of the processes described as being manually performed can be automatically performed by a known method. In addition, regarding the processing procedures, specific names, and information including various data and parameters shown in the above documents and drawings, they can be arbitrarily changed unless otherwise specified. For example, the various information shown in each figure is not limited to the illustrated information.
[0145] Also, each component of each illustrated device is a functional concept, and it is not necessarily physically configured as shown in the figure. That is, the specific form of the distribution and integration of each device is not limited to that shown in the figure, and all or part of it can be functionally or physically distributed and integrated in any unit according to various loads and usage situations.
[0146] For example, the above-described server device 100 may be realized by a plurality of server computers, and depending on the function, the configuration can be flexibly changed, such as by calling an external platform or the like through an API (Application Programming Interface) or network computing.
[0147] Also, the above-described embodiments and modification examples can be appropriately combined as long as the processing contents do not conflict.
[0148] Also, the "section (section, module, unit)" described above can be read as "means", "circuit", etc. For example, the acquisition section can be read as an acquisition means or an acquisition circuit.
Explanation of Signs
[0149] 1 Information processing system 10 Terminal device 100 Server device 110 Communication section 120 Storage section 130 Control section 131 Acquisition section 132 Generation section 133 Learning section 134 Super-resolution section 135 Management section 136 Provision section
Claims
1. A generation unit that generates a set of a reference image and an applied image obtained by applying a predetermined task to the reference image for each of a plurality of different tasks; A learning unit that causes a diffusion model to learn to generate a reference image when an applied image and task information indicating the task applied to the applied image are input; An information processing apparatus comprising the above.
2. The generation unit decomposes the original task into a plurality of tasks, and generates a set of a reference image and an applied image for each task. The information processing apparatus according to claim 1, characterized in that.
3. The generation unit generates a set of a reference image and an applied image for a reference task performed using a learned diffusion model, and further generates a set of a reference image and an applied image for a constituent task constituting the reference task. The information processing apparatus according to claim 1, characterized in that.
4. The learning unit causes the diffusion model to learn a process of generating an applied image by gradually applying deterioration corresponding to a predetermined task to a reference image, and a process of reconstructing the reference image from the applied image step by step so as to reverse the deterioration process. The information processing apparatus according to claim 1, characterized in that.
5. When the learning unit inputs a class label as task information indicating a task, it causes the diffusion model to learn to generate a reference image. The information processing apparatus according to claim 1, characterized in that.
6. The generation unit generates, as a learning sample when the class label as task information is class 0, a set of a high-resolution image as a reference image and an applied image obtained by gradually enlarging and reducing the high-resolution image; generates, as a learning sample when the class label as task information is class 1, a set of a high-resolution image as a reference image and an applied image obtained by gradually adding noise to the high-resolution image; generates, as a learning sample when the class label as task information is class 2, a set of a high-resolution image as a reference image and an applied image obtained by gradually performing enlargement / reduction and noise addition on the high-resolution image; The learning unit causes the diffusion model to learn using the generated learning samples for each class label. The information processing apparatus according to claim 1, characterized in that.
7. The learning unit deletes class labels with a certain probability and uses the generated training samples in a state without class labels to train the diffusion model, so that classifier-free guidance (CFG) can be applied to super-resolution by the diffusion model and be made available for use. The information processing apparatus according to claim 6, characterized in that.
8. A super-resolution unit that inputs an application image and task information indicating a real-world super-resolution task into a trained diffusion model, performs real-world super-resolution on the application image, and generates a real-world super-resolution image corresponding to a reference image. The information processing apparatus according to claim 1, further comprising.
9. A learning unit that decomposes a super-resolution task into multiple classes and trains the diffusion model to be a class-conditional super-resolution model, A super-resolution unit that performs real-world super-resolution of an input image using a trained diffusion model, An information processing apparatus, characterized by comprising.
10. A learning unit that trains the diffusion model so that classifier-free guidance (CFG) can be used in super-resolution by the diffusion model, A super-resolution unit that performs real-world super-resolution of an input image using a diffusion model in which classifier-free guidance (CFG) has become available, An information processing apparatus, characterized by comprising.
11. An information processing method executed by an information processing apparatus, A generation step of generating a pair of a reference image and an application image to which a predetermined task is applied to the reference image for each of a plurality of different tasks, A learning step of training the diffusion model to generate a reference image when an application image and task information indicating the task applied to the application image are input, An information processing method, characterized by including.
12. A generation procedure for generating a pair of a reference image and an application image to which a predetermined task is applied to the reference image for each of a plurality of different tasks, A learning procedure for training the diffusion model to generate a reference image when an application image and task information indicating the task applied to the application image are input, An information processing program, characterized by causing a computer to execute.
Citation Information
Patent Citations
Soybean stalk internal structure identification method and system based on CT image
CN117011316A
Deep learning super resolution of medical images
WO2023183504A1
Image processing apparatus, image processing method, and program
JP2020191046A