Information processing apparatus and information processing method

JP2024134427A5Pending Publication Date: 2026-03-31CANON KK
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-03-20
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing neural networks for noise reduction in videos suffer from image quality deterioration due to frame-to-frame fluctuations.

Method used

A two-stage learning approach for a machine learning model, where the first learning focuses on noise suppression and improving resolution, and the second learning addresses image quality deterioration caused by inter-frame fluctuations, using datasets tailored for each purpose with different motion and brightness conditions.

Benefits of technology

The approach effectively suppresses noise in videos while minimizing image quality deterioration, achieving stable and high-quality image processing results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To provide an information processing apparatus and an information processing method that achieve a machine learning model that can reduce noise in moving images, while preventing a reduction in image quality caused by variations in luminance between frames.SOLUTION: An information processing apparatus applies a first data set for learning to a machine learning model for reducing noise in moving images to perform first learning, and applies a second data set for learning to the machine learning model after the completion of the first learning to perform second learning. The machine learning model outputs, for an input image with a plurality of frames including a target frame in which noise is reduced, an image as a result of processing on the target frame. The first learning is learning intended to reduce noise, and the second learning is learning intended to prevent a reduction in image quality caused by variations in luminance between the plurality of frames.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to an information processing device and an information processing method. [Background technology]

[0002] In recent years, image processing using machine learning (more specifically, trained neural networks) has been widely used (Patent Document 1). In Patent Document 1, the learning accuracy is improved by training a neural network that predicts images in a specific order.

[0003] Specifically, in Patent Document 1, a neural network is trained first using video with large subject movements, whose movements are easy to predict, and thereafter, the neural network is trained using video with smaller subject movements. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Patent Publication No. 2021-189857 Summary of the Invention [Problem to be solved by the invention]

[0005] However, when a neural network for reducing noise is trained in a similar manner, there is a problem that image quality may be degraded due to effects caused by inter-frame variations.

[0006] In view of such problems with the conventional technology, in one embodiment, the present invention provides an information processing device and an information processing method that realize a machine learning model that can suppress noise in moving images while suppressing degradation of image quality caused by fluctuations between frames. [Means for solving the problem]

[0007] In one aspect, the present invention provides an information processing device that trains a machine learning model that suppresses noise in video, the information processing device having a learning means that performs a first learning by applying a first learning data set to the machine learning model and performs a second learning by applying a second learning data set to the machine learning model after the first learning has been completed, the machine learning model outputs an image as a processing result for a target frame for input images of multiple frames including a target frame for which noise is to be suppressed, the first learning is learning aimed at suppressing noise, and the second learning is learning aimed at suppressing degradation of image quality caused by fluctuations between multiple frames. Effect of the Invention

[0008] According to the present invention, it is possible to provide an information processing device and an information processing method that realize a machine learning model that can suppress noise in moving images while suppressing degradation of image quality caused by fluctuations between frames. [Brief description of the drawings]

[0009] [Figure 1] FIG. 1 is a block diagram showing an example of a hardware configuration of an information processing device according to an embodiment; [Diagram 2] FIG. 1 is a block diagram showing an example of a functional configuration of an information processing device according to an embodiment; [Diagram 3] A flowchart showing an overview of a learning method for a machine learning model according to an embodiment. [Figure 4] A flowchart showing details of a learning method for a machine learning model according to an embodiment. [Diagram 5] FIG. 1 is a diagram illustrating input and output of a machine learning model according to an embodiment. [Figure 6] FIG. 1 is a diagram showing a method for generating a learning dataset in the first embodiment. [Figure 7] FIG. 1 is a diagram comparing the first and second learning according to the first embodiment; [Figure 8] FIG. 10 is a diagram showing a learning data set in the second embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0010] ●(First embodiment) The present invention will be described in detail below based on its exemplary embodiments with reference to the accompanying drawings. Note that the following embodiments do not limit the invention according to the claims. In addition, although multiple features are described in the embodiments, not all of them are necessarily essential to the invention, and multiple features may be combined arbitrarily. Furthermore, in the accompanying drawings, the same reference numbers are used for the same or similar configurations, and duplicated explanations are omitted.

[0011] In the following embodiments, the present invention will be described with respect to a case where the present invention is implemented in a computer device (personal computer, tablet computer, media player, PDA, etc.). However, the present invention can be implemented in any electronic device that uses a processor. Such electronic devices include imaging devices (digital cameras), smartphones, game consoles, robots, drones, and drive recorders. These are merely examples, and the present invention can also be implemented in other electronic devices.

[0012] FIG. 1 is a block diagram illustrating an example of a hardware configuration of an information processing device according to an embodiment. The information processing device 100 includes a CPU 101, a RAM 102, a ROM 103, a secondary storage device 104, an input interface 105, and an output interface 106. The components of the information processing device 100 are connected to each other via a system bus 107 so as to be able to communicate with each other.

[0013] The information processing device 100 is connected to an external storage device 108 and an operation unit 110 via an input interface 105. The information processing device 100 is also connected to an external storage device 108 and a display device 109 via an output interface 106. Note that although the operation unit 110, the external storage device 108, and the display device 109 are described as external devices, they may be included in the information processing device 100.

[0014] CPU 101 uses RAM 102 as a work memory, executes programs (such as an OS and applications) stored in ROM 103 and secondary storage device 104, and controls the operation of each component of information processing device 100 via system bus 107. The operation of information processing device 100, which will be described later, is realized by CPU 101 executing an appropriate program for realizing the operation.

[0015] In the following description, some of the operations executed by the CPU 101 may be executed by another processor instead of the CPU 101 or in cooperation with the CPU 101. The other processor may be, for example, an NPU or a GPU configured to be able to execute calculations related to machine learning at high speed. The programs executed by the other processor may also be stored in the ROM 103 or the secondary storage device 104.

[0016] The secondary storage device 104 stores programs executed by the CPU 101, user data, various data handled by the information processing device 100, and the like. The secondary storage device 104 may be a storage device having a larger capacity than the ROM 103, such as an SSD or HDD. The CPU 101 accesses the secondary storage device 104 via a system bus 107.

[0017] The input interface 105 is, for example, a serial bus interface such as a USB. The information processing device 100 can communicate with an external device via the input interface 105. In this embodiment, the operation unit 110 and the external storage device 108 are shown as examples of external devices that can be connected to the input interface 105, but other external devices may be connected. In addition, there is no particular limit to the type and number of the input interface 105.

[0018] The external storage device 108 may be a storage device that uses a removable storage medium such as a memory card.

[0019] The operation unit 110 is an input device for a user of the information processing device 100 to input instructions to the information processing device 100, and may have one or more of a keyboard, a pointing device, a touch pad, a touch panel, a switch, a button, and the like.

[0020] The output interface 106 is, for example, a serial bus interface such as USB, similar to the input interface 105. The output interface 106 may be, for example, a video output terminal such as DVI or HDMI (registered trademark). The information processing device 100 outputs data and the like to an external device via the output interface 106. In this embodiment, the display device 109 and the external storage device 108 are shown as examples of external devices that can be connected to the output interface 106, but other external devices may also be connected. Furthermore, there is no particular limitation on the type and number of the output interface 106.

[0021] Although the input interface 105 and the output interface 106 are described separately, they may be a single input / output interface in practice. The input interface 105 and the output interface 106 may be one or more types of wired or wireless communication interfaces to which external devices can be connected.

[0022] 2 is a block diagram showing an example of a functional configuration of the information processing device 100. A storage unit 201 corresponds to the RAM 102, the ROM 103, the secondary storage device 104, and the external storage device 108. The other blocks 202 to 207 correspond to functions realized by the CPU 101 and / or other processors executing programs.

[0023] 3 is a flowchart simply illustrating a learning method of a machine learning model used to reduce noise in a moving image by the information processing device 100. In this embodiment, learning is performed in two stages.

[0024] For ease of explanation and understanding, it is assumed here that a machine learning model using a convolutional neural network (CNN) is learned, but there is no particular limitation on the method of realizing the machine learning model. For example, a recurrent neural network (RNN), a transformer, etc. may be used. It is assumed here that the machine learning model to be learned is coded in advance in an appropriate program language and stored in the secondary storage device 104, for example.

[0025] The machine learning model takes as input images a frame (target frame) in which noise is to be suppressed and a certain number of frames before and after it out of multiple frames constituting a video, and takes as an output image an image of the target frame in which noise is suppressed, which is inferred from the input image. The configuration of the machine learning model (filters, activation functions, loss functions, etc. used in the convolution layer) is appropriately determined according to the resolution of the input video, etc.

[0026] In S301, the CPU 101 and / or another processor (hereinafter simply referred to as the CPU 101) applies a first learning to a machine learning model. The first learning is learning for the purpose of removing noise and improving the resolution of an input image.

[0027] When the first learning is completed, in S302, the CPU 101 applies the second learning to the machine learning model. The second learning is intended to suppress degradation of image quality caused by fluctuations between frames of an input image, specifically, afterimages that occur around moving object regions due to movement between frames.

[0028] By inputting the target frame and the frames before and after it, it is possible to reduce the variation in processing for each frame in terms of noise suppression and resolution improvement, and obtain a stable processing result, compared to when only the target frame is input. However, on the other hand, it becomes necessary to suppress afterimages that occur due to the movement between the input frames. In this embodiment, after performing learning suitable for noise suppression and resolution improvement, learning suitable for suppressing afterimages is performed, making it possible to fully enjoy the effect of inputting multiple frames.

[0029] As described later, the first learning and the second learning differ in the data set used for learning and the initial learning rate. In order to perform learning efficiently, the data set used for learning generates multiple images from the same still image by varying the cut-out positions. Noise is added to these multiple images, and they are used as training data for the machine learning model as pseudo video frames. In this case, one of the multiple images is set as a target frame for noise suppression, and the image of the target frame before noise is added is set as teacher data (correct answer data). The target frame is treated as the frame located chronologically in the center of the multiple images.

[0030] Next, a learning method of the machine learning model will be described with reference to the block diagram shown in Fig. 2 and the flowchart shown in Fig. 4. The learning method is common to the first learning and the second learning.

[0031] In S401, the learning unit 207 sets parameters of the machine learning model. The parameters are mainly weight parameters of the neural network, and may also include settings of a learning rate, a loss function, and an optimizer (optimization algorithm).

[0032] In this embodiment, the first learning uses weight parameters having initial values ​​determined by random numbers, and the second learning uses weight parameters stored in the storage unit 201 as the results of the first learning in S412 described later.

[0033] Moreover, the initial learning rate of the second learning is set to a value smaller than the initial learning rate of the first learning, in order to prevent the weight parameters reflecting the first learning from being significantly changed by the second learning.

[0034] In S402, the parameter acquisition unit 202 acquires parameters related to the input image data included in the training data to be input to the machine learning model. The parameters related to the input image data include, for example, the number of frames of the input image, the amount of motion between frames, and the like.

[0035] In this embodiment, the number of frames simultaneously input to the machine learning model is an odd number equal to or greater than 3. Here, as an example, it is set to 5 as shown in FIG. 5. The frame located in the center in the time series is set as the target frame, and the remaining frames are set as reference frames. The machine learning model uses the target frame and the reference frame as input images, infers the image of the target frame with noise suppressed, and sets it as an output image.

[0036] In this embodiment, the amount of motion between frames is the maximum value of the horizontal and vertical amounts of motion between adjacent frames for multiple frames used as input images. The amount of motion between frames is a parameter used to suppress afterimages that occur around moving object areas in the frames.

[0037] Learning with a data set with a large amount of motion between frames can improve the effect of suppressing afterimages around moving object regions, but the perceived resolution of the image decreases.On the other hand, learning with a data set with a small amount of motion between frames can improve the perceived resolution of the image while suppressing noise, but afterimages are more likely to occur around moving object regions.

[0038] Therefore, the first learning uses a data set with a small amount of motion between frames, and the second learning uses a data set with a larger amount of motion between frames than that of the first learning. Also, by making the initial learning rate of the second learning smaller than that of the first learning, the influence of the second learning on the results of the first learning is suppressed.

[0039] By applying this two-stage learning process, it is possible to combine weighting parameters (results of the first learning) that suppress noise and improve resolution with the effect of suppressing afterimages that occur around moving object areas (results of the second learning).

[0040] The inter-frame motion amount of the data set used in the first learning may be, for example, the motion amount caused by typical camera shake that occurs when shooting a video while holding it in the horizontal and vertical directions. Although it depends on the pixel pitch of the image sensor, it can be, for example, 10 pixels in both the horizontal and vertical directions.

[0041] The inter-frame motion amount of the data set used in the second learning is set to a value larger in both the horizontal and vertical directions than the inter-frame motion amount of the data set used in the first learning (a value larger than 20 pixels, and here, 30 pixels as an example). A specific inter-frame motion amount may be obtained, for example, experimentally, but an inter-frame motion amount suitable for learning to suppress afterimages in moving object areas is a value larger than an inter-frame motion amount suitable for learning to suppress noise and improve resolution. Figure 7 shows an overview of the first learning and the second learning.

[0042] In S403, the image acquisition unit 204 randomly selects one of the multiple still image data stored in the storage unit 201. The still image data is assumed to have been captured under imaging conditions (e.g., imaging sensitivity ISO 100) that provide an image with little noise.

[0043] In S404, the parameter acquisition unit 202 copies the still image data selected in S403 and stores the copies in the storage unit 201 so as to match the number of frames acquired in S402.

[0044] In S405, the parameter processing unit 203 determines the amount of horizontal and vertical motion to be applied to each of the still image data stored in S404. Here, the parameter processing unit 203 determines the amount of horizontal and vertical motion to be 0 for the still image data used as the target frame.

[0045] Furthermore, for still image data used as a reference frame, the parameter processing unit 203 randomly determines the horizontal and vertical motion amounts within a range of inter-frame motion amounts according to whether the training data set to be generated is for the first training or the second training. Specifically, if the inter-frame motion amount is X, the parameter processing unit 203 determines the horizontal motion amount nx and the vertical motion amount ny within the ranges of -X≦nx≦X, -X≦ny≦X. The motion amounts are integers in pixel units.

[0046] In S406, the image acquisition unit 204 extracts image (image patch) data of a predetermined size from each still image data based on the amount of motion determined in S405. FIG. 6 is a diagram showing a schematic diagram of the image patch acquisition operation in S406. The image patch to be used as a reference frame is acquired by moving the extraction position of the image patch to be used as a target frame by the amount of motion determined in S405. The image acquisition unit 204 stores the generated image patches (target frame and reference frame) in the storage unit 201.

[0047] The cut-out position of the image patch used as the target frame is set near the center of the still image before cut-out. The cut-out position used as the reference may be set so that the center of the image and the center of the image patch coincide with each other, or may be set in another way, such as to include a feature region (e.g., a face region, a human body region, etc.) or a main subject region included in the still image before cut-out. The feature region or the main subject region can be detected or determined by a known method.

[0048] In this way, image patches generated by changing the cut-out position from the same still image are used as pseudo video frames for training the machine learning model.

[0049] In S407, the image acquisition unit 204 stores the image patch to be used as the target frame in the storage unit 201 as training data.

[0050] In S408, the image processing unit 205 applies a predetermined image processing to each of the image patches generated in S406 to generate training data. A combination of training data and corresponding teacher data is a unit of learning data.

[0051] In this embodiment, noise suppression is one of the first learning objectives, but since the still image data that is the basis of the training data contains little noise, image processing can be applied to add artificial luminance noise that occurs during high-sensitivity shooting. When the imaging sensitivity of the video to be applied to the machine learning model after learning is known in advance, the image processing unit 205 can apply image processing to the training data that simulates noise that occurs at that imaging sensitivity. In addition, image processing such as reducing the image size to reduce the learning load can also be applied in S408.

[0052] In S409, the learning unit 207 applies the training data generated in S408 to the machine learning model, and acquires image data from the machine learning model.

[0053] In S410, the error calculation unit 206 calculates the error between the image data acquired from the machine learning model in S409 and the teacher data stored in the storage unit 201 in S407. The error can be calculated using a known loss function, such as the sum of absolute values ​​of differences between corresponding pixel values ​​of the image data.

[0054] In S411, the learning unit 207 updates the parameters of the machine learning model so as to minimize the error calculated in S410. Specifically, the learning unit 207 can update the parameters by, for example, an error backpropagation method. In this embodiment, the learning is terminated when the number of learning times is equal to or greater than a predetermined number and the calculation error is equal to or less than a predetermined value.

[0055] In S412, the learning unit 207 stores in the storage unit 201 the parameters of the machine learning model upon completion of learning.

[0056] When the first learning is performed, the parameters of the machine learning model for which the learning has been completed are used in the model parameter setting in S401 during the second learning. Also, the parameters of the machine learning model for which the second learning has been completed are used when applying the dest data to the machine learning model.

[0057] According to this embodiment, the learning of the machine learning model for suppressing noise in a video is divided into a first learning for noise suppression and improvement of resolution, and a second learning for suppressing afterimages that occur around moving object regions. This makes it possible to perform learning using a learning data set suitable for each purpose, and it is possible to suppress noise and improve resolution while suppressing afterimages caused by movement between frames.

[0058] Although the learning data set is generated by changing the cut-out position of the same still image, it may be generated by other methods. For example, the learning data set may be generated by extracting images corresponding to the inter-frame motion amount from still images continuously shot while panning, still images continuously shot of a scene including a moving object, frame images extracted from a video, etc. In this case, the operations after generating the learning data set (operations after S407) may be as described above.

[0059] Also, the number of reference frames before and after the target frame need not be equal, in which case the number of frames in the training data need not be an odd number.

[0060] Although the amount of movement of the cutout position relative to the reference position is randomly determined within the range of inter-frame movement, the inter-frame movement may be restricted to a predetermined direction, for example, to simulate the movement of a subject. The inter-frame movement may include scaling and / or rotation in addition to horizontal and vertical directions.

[0061] ●(Second embodiment) The second embodiment of the present invention will be described below. This embodiment can also be implemented by the information processing device 100 described with reference to Figures 1 and 2. Therefore, a description of the configuration common to the first embodiment will be omitted.

[0062] In this embodiment, the second learning is aimed at suppressing image quality degradation caused by changes between frames, specifically color unevenness caused by luminance changes between frames. The luminance changes targeted here are mainly caused by flickering light sources (fluorescent lights, LEDs, etc.) contained in ambient light. Since the luminance of a flickering light source changes periodically, a video captured under a flickering light source may have luminance fluctuations between frames. Luminance changes between frames constituting an input image of a machine learning model, particularly between a target frame and a reference frame, may cause color unevenness in an output image of the machine learning model.

[0063] The following will focus on the steps in the flow chart of FIG. 4 that perform operations different from those in the first embodiment.

[0064] In S402, the parameter acquisition unit 202 acquires parameters related to input image data included in the training data to be input to the machine learning model. The parameters related to the input image data are, for example, the number of frames of the input image, the brightness fluctuation rate between frames, etc. The brightness fluctuation rate between frames is, for example, the maximum value of the ratio of the difference in average brightness value between the target frame and the reference frame.

[0065] Since the first learning aims to improve the perceived resolution and stability in the time direction, the luminance fluctuation rate between frames is set to a sufficiently small value greater than 0, for example, 2% in this embodiment. On the other hand, the second learning aims to suppress adverse effects caused by luminance fluctuations between frames, so the rate of luminance fluctuation between frames is set to be larger than that in the first learning, for example, 20% in this embodiment.

[0066] In S405, the parameter processing unit 203 determines the brightness variation rate to be applied to each of the still image data stored in S404. Here, the parameter processing unit 203 determines the brightness variation rate to be 0% for the still image data used as the target frame.

[0067] Furthermore, for still image data used as reference frames, the parameter processing unit 203 determines, for example, randomly, a luminance variation rate within a range of luminance variation rates between frames depending on whether the learning data set to be generated is for the first learning or the second learning. Specifically, if the luminance variation rate between frames is Y%, the parameter processing unit 203 determines the luminance variation rate m for each still image data within the range of -Y≦m≦Y.

[0068] In S406, the image acquisition unit 204 cuts out image (image patch) data of a predetermined size from each still image data. Here, it is assumed that the cut-out position is fixed. If the cut-out position is fixed, the image patch cut out in S406 may be duplicated instead of performing the duplication process in S404. The image acquisition unit 204 stores the generated image patches (target frame and reference frame) in the storage unit 201.

[0069] In S407, the image acquisition unit 204 stores the image patch to be used as the target frame in the storage unit 201 as training data.

[0070] In S408, the image processing unit 205 applies the fluctuation rate determined in S405 to the pixel values ​​(luminance values) of the image patches of the reference frame among the image patches generated in S406, and generates training data. The image processing unit 205 also applies image processing to add noise to each of the image patches of the target frame and the reference frame in the same manner as in the first embodiment.

[0071] Although the cut-out position is fixed in S406, it may be changed. In this case, since the images in the pixel patch do not match, in S408, the image processing unit 205 applies image processing to change the pixel values ​​so that the average luminance value of the pixel patch satisfies the luminance fluctuation rate determined in S405.

[0072] According to this embodiment, the learning of the machine learning model for suppressing noise in a video is divided into a first learning for noise suppression and improvement of resolution, and a second learning for suppressing the influence of luminance fluctuations between frames. This enables learning using a learning data set suitable for each purpose, and it is possible to suppress noise and improve resolution while suppressing the influence of luminance changes between frames caused by, for example, a flickering light source.

[0073] (Other embodiments) The present invention can also be realized by a process in which a program for implementing one or more of the functions of the above-described embodiments is supplied to a system or device via a network or a storage medium, and one or more processors in a computer of the system or device read and execute the program. The present invention can also be realized by a circuit (e.g., ASIC) that implements one or more of the functions.

[0074] The disclosure of the present embodiment includes the following information processing device, image processing device, information processing method, and program. (Item 1) An information processing device that learns a machine learning model to suppress noise in a video, a learning means for performing a first learning by applying a first learning data set to the machine learning model, and performing a second learning by applying a second learning data set to the machine learning model after the first learning has been completed; The machine learning model outputs an image as a processing result for a plurality of input images of frames including a target frame for noise suppression, The first learning is learning for the purpose of suppressing noise, and the second learning is learning for the purpose of suppressing degradation of image quality caused by fluctuations between the multiple frames. 23. An information processing apparatus comprising: (Item 2) 2. The information processing device according to item 1, wherein the learning means sets an initial learning rate of the second learning to be smaller than an initial learning rate of the first learning. (Item 3) 3. The information processing device according to item 1 or 2, wherein the second learning is learning aimed at suppressing afterimages caused by movement between the multiple frames. (Item 4) 4. The information processing device according to item 3, further comprising a generating means for generating the first training data set and the second training data set based on a still image. (Item 5) Further comprising a generating means for generating the first training data set and the second training data set, Item 4. The information processing device according to item 3, wherein the generating means generates the first learning data set and the second learning data set such that a maximum value of the amount of motion between multiple frames used as the input image in the second learning is greater than a maximum value of the amount of motion between multiple frames used as the input image in the first learning. (Item 6) 3. The information processing device according to item 1 or 2, wherein the second learning is learning aimed at suppressing an influence caused by a change in luminance between the plurality of frames. (Item 7) Further comprising a generating means for generating the first training data set and the second training data set, 7. The information processing device according to item 6, wherein the generating means generates the first learning data set and the second learning data set such that a rate of variation in luminance between multiple frames used as the input image in the second learning is greater than a rate of variation in luminance between multiple frames used as the input image in the first learning. (Item 8) 8. The information processing device according to any one of items 1 to 7, wherein the machine learning model uses a neural network. (Item 9) A machine learning model trained by the information processing device according to any one of items 1 to 8; An acquisition means for inputting a video to the machine learning model to acquire the video with noise suppression; 13. An image processing device comprising: (Item 10) An information processing method executed by an information processing device, Applying a first training dataset to a machine learning model that suppresses noise in a video to perform first training; and performing second learning by applying a second learning dataset to the machine learning model after the first learning has been completed, The machine learning model outputs an image as a processing result for a plurality of input images of frames including a target frame for noise suppression, The first learning is learning for the purpose of suppressing noise, and the second learning is learning for the purpose of suppressing degradation of image quality caused by fluctuations between the multiple frames. 23. An information processing method comprising: (Item 11) A program for causing a computer to function as each of the means possessed by the information processing device according to any one of items 1 to 8.

[0075] The present invention is not limited to the above-described embodiments, and various modifications and variations are possible without departing from the spirit and scope of the invention. Therefore, the following claims are appended to disclose the scope of the invention. [Explanation of symbols]

[0076] 101...CPU, 102...RAM, 103...ROM, 104...secondary storage device, 105...input interface, 106...output interface, 107...bus, 108...external storage device, 109...display device, 110 operation unit

Claims

1. An information processing device that learns a machine learning model to suppress noise in a video, a learning means for performing a first learning by applying a first learning data set to the machine learning model, and performing a second learning by applying a second learning data set to the machine learning model after the first learning has been completed; The machine learning model outputs an image as a processing result for a plurality of input images of frames including a target frame for noise suppression, The first learning is learning for the purpose of suppressing noise, and the second learning is learning for the purpose of suppressing degradation of image quality caused by fluctuations between the multiple frames.

23. An information processing apparatus comprising:

2. 2. The information processing apparatus according to claim 1, wherein said learning means sets an initial learning rate of said second learning to be smaller than an initial learning rate of said first learning.

3. 2 . The information processing apparatus according to claim 1 , wherein the second learning is aimed at suppressing afterimages caused by movements between the plurality of frames.

4. The information processing apparatus according to claim 3 , further comprising a generating means for generating the first training data set and the second training data set based on a still image.

5. The method further includes generating means for generating the first training data set and the second training data set, The information processing device according to claim 3, characterized in that the generation means generates the first learning data set and the second learning data set so that a maximum value of the amount of motion between multiple frames used as the input images in the second learning is greater than a maximum value of the amount of motion between multiple frames used as the input images in the first learning.

6. 2. The information processing apparatus according to claim 1, wherein the second learning is aimed at suppressing an influence caused by a change in luminance between the plurality of frames.

7. The method further includes generating means for generating the first training data set and the second training data set, The information processing device according to claim 6, characterized in that the generation means generates the first learning data set and the second learning data set so that the rate of variation in luminance between multiple frames used as the input image in the second learning is greater than the rate of variation in luminance between multiple frames used as the input image in the first learning.

8. The information processing device according to claim 1 , wherein the machine learning model uses a neural network.

9. A machine learning model trained by the information processing device according to any one of claims 1 to 8; An acquisition means for inputting a video to the machine learning model to acquire the video with noise suppression; 13. An image processing device comprising:

10. An information processing method executed by an information processing device, Applying a first training dataset to a machine learning model that suppresses noise in a video to perform first training; and performing second learning by applying a second learning dataset to the machine learning model after the first learning is completed, The machine learning model outputs an image as a processing result for a plurality of input images of frames including a target frame for noise suppression, The first learning is learning for the purpose of suppressing noise, and the second learning is learning for the purpose of suppressing degradation of image quality caused by fluctuations between the multiple frames.

23. An information processing method comprising:

11. A program for causing a computer to function as each of the means included in the information processing device according to any one of claims 1 to 8.