Method and apparatus for reducing noise included in real-time image data

A neural network and motion correction method with unsupervised learning and edge enhancement techniques enhance low-dose fluoroscopy images, improving image quality and diagnostic accuracy in clinical environments.

WO2026095691A1PCT designated stage Publication Date: 2026-05-07EWHA UNIV IND COLLABORATION FOUND
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
EWHA UNIV IND COLLABORATION FOUND
Filing Date
2025-10-30
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

Low-dose X-ray fluoroscopy images in medical procedures suffer from increased noise and distortion, and existing deep learning-based methods struggle to perform consistently in clinical environments with data scarcity.

Method used

A method involving a neural network model and motion correction network to reduce noise, using unsupervised learning and SWT-based edge enhancement, which includes training a noise correction model with first and second output data to improve image quality.

Benefits of technology

Significantly improves image quality and diagnostic accuracy by reducing noise and preserving image detail, effectively addressing data scarcity in clinical settings.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025017633_07052026_PF_FP_ABST
    Figure KR2025017633_07052026_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to a method and apparatus for reducing noise included in real-time image data. The method for reducing noise included in real-time image data, according to an aspect, may comprise the steps of: acquiring first output data by inputting at least one frame before a time point t and at least one frame after the time point t, from among a plurality of frames included in real-time image data, into a neural network model; acquiring second output data by inputting a frame at the time point t, at least one frame before the time point t, and at least one frame after the time point t, from among the plurality of frames, into a motion correction network and a recursive filter; training a noise correction model by using at least one of the first output data and the second output data; and reducing noise included in the real-time image data by using the trained noise correction model.
Need to check novelty before this filing date? Find Prior Art

Description

Method and apparatus for reducing noise included in real-time video data

[0001] The present disclosure relates to a method and apparatus for reducing noise included in real-time video data.

[0002] Fluoroscopy is a technique that generates and visualizes real-time X-ray images and plays an important role in various medical procedures, such as catheter insertion or monitoring during orthopedic surgery. However, this technique generates radiation, which can pose a potential risk to patients and medical staff. Accordingly, the use of low-dose X-ray fluoroscopy is recommended to reduce radiation, but there is a problem in that the images generated by low-dose X-ray fluoroscopy are prone to increased noise or distortion.

[0003] To overcome these problems, methods such as restoring images by modeling noise and reducing noise using deep learning / machine learning-based artificial intelligence models are being attempted. However, these approaches must be consistently applicable in clinical environments where various fluoroscopy devices are used, and must be able to guarantee performance despite the problem of data scarcity in clinical settings.

[0004] The aforementioned background technology is technical information that the inventor possessed for the derivation of the present invention or acquired during the process of deriving the present invention, and it cannot be considered as prior art disclosed to the general public prior to the filing of the present invention.

[0005] The technical problem that the present invention aims to solve is to provide a 'method and apparatus for reducing noise included in real-time image data.' Additionally, the invention aims to provide a computer-readable recording medium storing a program for executing the method on a computer. The technical problems to be solved are not limited to those described above, and other technical problems may exist.

[0006] As a technical means for achieving the technical problem described above, the first aspect of the present disclosure may provide a method for reducing noise included in real-time video data, comprising: a step of obtaining first output data by inputting at least one frame prior to time t and at least one frame after time t among a plurality of frames included in real-time video data into a neural network model; a step of obtaining second output data by inputting the frame at time t, at least one frame prior to time t, and at least one frame after time t among the plurality of frames into a motion correction network and a recursive filter; a step of training the noise correction model using at least one of the first output data and the second output data; and a step of reducing noise included in the real-time video data using the trained noise correction model.

[0007] A second aspect of the present disclosure may provide a computer-readable recording medium that records a program for executing the method of the first aspect on a computer.

[0008] A third aspect of the present disclosure may provide an apparatus for reducing noise included in real-time video data, comprising at least one memory; and at least one processor, wherein the processor inputs at least one frame prior to time t and at least one frame after time t among a plurality of frames included in the real-time video data into a neural network model to obtain first output data, inputs the frame at time t, at least one frame prior to time t and at least one frame after time t among the plurality of frames into a motion correction network and a recursive filter to obtain second output data, learns the noise correction model using at least one of the first output data and the second output data, and reduces the noise included in the real-time video data using the learned noise correction model.

[0009] In addition, other methods for implementing the present invention, other systems, and computer-readable recording media storing a computer program for executing said methods may be further provided.

[0010] Other aspects, features, and advantages other than those described above will become clear from the following drawings, claims, and detailed description of the invention.

[0011] According to the means for solving the problem of the present disclosure described above, the image quality of low-dose fluoroscopy images can be significantly improved to increase diagnostic accuracy. In particular, a noise correction model can be trained without high-quality ground images through an unsupervised learning-based neural network model, and accordingly, the problem of data scarcity in clinical environments can be solved.

[0012] In addition, according to the means for solving the problem of the present disclosure described above, blur caused by movement between consecutive frames can be effectively removed through motion correction, and the detailed structure of the image can be clearly preserved by using an SWT-based edge enhancement technique.

[0013] FIG. 1 is a diagram illustrating an example of a system for reducing noise included in real-time image data according to one embodiment.

[0014] FIG. 2 is a flowchart illustrating a method for reducing noise included in real-time video data according to one embodiment.

[0015] FIG. 3 is a block diagram illustrating the learning of a noise correction model according to one embodiment.

[0016] FIG. 4 is a flowchart illustrating a method for obtaining first output data according to one embodiment.

[0017] FIG. 5 is a schematic diagram illustrating a method for obtaining second output data according to one embodiment.

[0018] FIG. 6 is a block diagram illustrating a method for training a noise correction model using edge-enhanced second output data and third output data according to one embodiment.

[0019] FIG. 7 is a block diagram illustrating a method for reinforcing the edges of a frame using an edge reinforcing module according to one embodiment.

[0020] FIG. 8 is a block diagram of a user device according to one embodiment.

[0021] A method for reducing noise included in real-time video data according to one aspect may include: a step of obtaining first output data by inputting at least one frame prior to time t and at least one frame after time t among a plurality of frames included in real-time video data into a neural network model; a step of obtaining second output data by inputting the frame at time t, at least one frame prior to time t, and at least one frame after time t among the plurality of frames into a motion correction network and a recursive filter; a step of training the noise correction model using at least one of the first output data and the second output data; and a step of reducing noise included in the real-time video data using the trained noise correction model.

[0022] The advantages and features of the present invention, and the methods for achieving them, will become clear by referring to the embodiments described in detail together with the accompanying drawings. However, the present invention is not limited to the embodiments presented below, but can be implemented in various different forms and should be understood to include all modifications, equivalents, and substitutions that fall within the spirit and scope of the present invention. The embodiments presented below are provided to ensure that the disclosure of the present invention is complete and to fully inform those skilled in the art of the scope of the invention. In describing the present invention, detailed descriptions of related known technologies are omitted if it is determined that such detailed descriptions may obscure the essence of the present invention.

[0023] The terms used in this application are used merely to describe specific embodiments and are not intended to limit the invention. The singular expression includes the plural expression unless the context clearly indicates otherwise. In this application, terms such as 'comprising' or 'having' are intended to specify the existence of the features, numbers, steps, actions, components, parts, or combinations thereof described in the specification, and should be understood as not precluding the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.

[0024] Additionally, terms such as 'unit' and 'module' described in the specification refer to a unit that processes at least one function or operation, and this may be implemented in hardware or software, or as a combination of hardware and software.

[0025] Additionally, terms including ordinal numbers, such as 'first' or 'second' used in the specification, may be used to describe various components, but said components shall not be limited by said terms. Such terms may be used for the purpose of distinguishing one component from another.

[0026] Some embodiments of the present disclosure may be represented by functional block configurations and various processing steps. Some or all of these functional blocks may be implemented by various numbers of hardware and / or software configurations that execute specific functions. For example, the functional blocks of the present disclosure may be implemented by one or more microprocessors or by circuit configurations for a specific function. Additionally, for example, the functional blocks of the present disclosure may be implemented in various programming or scripting languages. The functional blocks may be implemented as algorithms executed on one or more processors. Furthermore, the present disclosure may employ prior art for electronic configuration, signal processing, and / or data processing, etc. Terms such as 'mechanism', 'element', 'means', and 'configuration' may be used broadly and are not limited to mechanical and physical configurations.

[0027] Hereinafter, the present disclosure will be described in detail with reference to the attached drawings.

[0028] FIG. 1 is a diagram illustrating an example of a system for reducing noise included in real-time image data according to one embodiment.

[0029] Referring to FIG. 1, the system (1) may include a server (110), a user device (120), and a shooting device (130).

[0030] For convenience of explanation, FIG. 1 is illustrated as including a server (110), a user device (120), and a shooting device (130) in the system (1), but the present disclosure is not limited thereto. In other words, the system (1) may include any component necessary to reduce noise included in real-time video data.

[0031] Additionally, the operation of the server (110), user device (120), and shooting device (130) described below may be implemented by a single device or by more devices.

[0032] The server (110) refers to a server that comprehensively processes the collection, storage, reading, search, and management of real-time video data.

[0033] For example, the server (110) can receive real-time video data from the shooting device (130), store it in a mass storage device, and transmit the video data to a user device (120) or an external device. At this time, the server (110) can transmit the real-time video data received from the shooting device (130) to the user device (120) or an external device in a form divided into multiple frames.

[0034] In one embodiment, the server (110) may be a computing device including a processor. For example, the server (110) according to one embodiment may be a cloud server. If the server (110) is a computing device, the processor may perform at least some of the operations of the user device (120) described below with reference to FIGS. 1 to 7.

[0035] The user device (120) can train a noise correction model to reduce noise included in real-time video data and reduce noise included in real-time video data using the trained noise correction model.

[0036] In one embodiment, the user device (120) may be a computing device comprising a display device and a device for receiving user input (e.g., a keyboard, a mouse, etc.), and including memory and a processor. In this case, the function of receiving user input may be performed by a display device including a touch screen. For example, the user device (120) may include, but is not limited to, a notebook PC, a desktop PC, a laptop, a tablet computer, a smartphone, etc.

[0037] In the present disclosure, the shooting device (130) may refer to any device used to acquire image data or images.

[0038] In one embodiment, the imaging device (130) may refer to a device that generates image data for various purposes such as diagnosis, treatment, and monitoring, and outputs or stores the data.

[0039] In one embodiment, the imaging device (130) may refer to a concept including medical imaging devices such as an X-ray device, a CT device, an MRI device, an ultrasound device, a PET device, and a SPECT device. For example, the imaging device (130) may be an X-ray imaging device, in which case the imaging device (130) may be a device that transmits X-rays to a patient's body and detects the transmitted X-rays to generate digital image data.

[0040] In one embodiment, the image data may include medical image data such as fluoroscopic images, CT scan images, MRI images, and ultrasound images.

[0041] Meanwhile, the server (110), user device (120), and shooting device (130) can be connected via a network to transmit and receive data to and from each other. For example, real-time video data can be acquired by the shooting device (130), transmitted to the server (110), and transmitted from the server (110) to the user device (120). Alternatively, real-time video data can be acquired by the shooting device (130) and transmitted to the user device (120). Here, the network may refer to a comprehensive data communication network that enables each constituent entity of the system illustrated in FIG. 1 to communicate smoothly with one another.

[0042] In one embodiment, the network includes a Local Area Network (LAN), a Wide Area Network (WAN), a Value Added Network (VAN), a mobile radio communication network, a satellite communication network, and combinations thereof, and may include wired internet and wireless internet. Here, wireless communication may include, for example, Wi-Fi, Bluetooth, Bluetooth Low Energy, Zigbee, Wi-Fi Direct (WFD), Ultra Wideband (UWB), Infrared Data Association (IrDA), Near Field Communication (NFC), but is not limited thereto.

[0043] FIG. 2 is a flowchart illustrating a method for reducing noise included in real-time video data according to one embodiment.

[0044] A method for reducing noise included in real-time video data described with reference to FIG. 2 can be performed by the user device of FIG. 1, and specifically, can be performed by a processor included in the user device.

[0045] In step 210, the processor may input at least one frame prior to time t and at least one frame after time t among a plurality of frames included in real-time video data into a neural network model to obtain first output data. Here, the first output data according to one embodiment refers to the frame at time t predicted using the neural network model.

[0046] In one embodiment, the neural network model may receive at least one frame prior to time t and at least one frame after time t as inputs, and output a prediction result for the frame at time t. At this time, the processor may change time t in correspondence with the time flow of the image data and output a prediction result for the frame at all times.

[0047] Specific details regarding a method for a processor according to one embodiment to predict a frame at time t using a neural network model will be described later in FIG. 4.

[0048] In step 220, the processor may input a frame at time t, at least one frame prior to time t, and at least one frame after time t among a plurality of frames into a motion correction network and a recursive filter to obtain second output data. Here, according to one embodiment, the second output data refers to the frame at time t in which the processor has reduced noise using a motion correction network and a recursive filter.

[0049] In one embodiment, a motion correction network executed by a processor can receive a frame at time t, a frame before time t, and a frame after time t as inputs, and output motion-corrected frames. Here, the motion correction network executed by the processor can calculate information regarding the flow of movement between each input frame and perform motion correction based on the calculated information.

[0050] At this time, the processor can change the time point t in correspondence with the time flow of the video data and output motion-corrected frames for all time points.

[0051] Specific details regarding how a processor according to one embodiment performs motion correction using a motion correction network will be described later in FIG. 5.

[0052] In one embodiment, the processor can input a frame at time t, motion-corrected by a motion correction network, at least one frame before time t, and at least one frame after time t into a recursive filter to output a frame at time t with reduced noise.

[0053] At this time, the processor can change the time point t in correspondence with the time flow of the image data and output a frame with reduced noise at all time points.

[0054] In step 230, the processor can train the noise correction model using at least one of the first output data and the second output data.

[0055] In one embodiment, the noise correction model can receive a frame as input and output a corrected frame, and in this case, the processor can calculate an error between the corrected frame and the first and second output data and train the noise correction model based on the error.

[0056] Specific details regarding the training of a noise correction model according to one embodiment will be described later in FIG. 3.

[0057] In step 240, the processor can reduce noise contained in real-time image data by using a noise correction model learned according to steps 210 to 230.

[0058] In one embodiment, the processor can reduce noise in real time for input image data. At this time, the image data with noise reduced by the processor can be output to the user through a display included in the user device.

[0059] FIG. 3 is a block diagram illustrating the learning of a noise correction model according to one embodiment.

[0060] The learning of the noise correction model described with reference to FIG. 3 can be performed by the user device of FIG. 1, and specifically, by a processor included in the user device.

[0061] Referring to FIG. 3, a noise correction model (300), a neural network model (310), a motion correction network, and a recursive filter (320) are illustrated. Also, referring to FIG. 3, input data (301) and output data (302) of the noise correction model (300), input data (311) and output data (312) of the neural network model (310), and input data (321) and output data (322) of the motion correction network and recursive filter (320) are illustrated.

[0062] In one embodiment, a noise correction model (300) executed by a processor is trained according to a method described below and can receive at least one frame included in the image data as input and perform noise correction.

[0063] According to one embodiment, the input data (301) of the noise correction model (300) may include each frame included in the image data to be corrected for noise, and the output data (302) of the noise correction model (300) may include each frame with noise corrected.

[0064] For example, as illustrated in FIG. 3, the input data (301) of the noise correction model (300) may include a frame at time t that is the target of noise correction, and the output data (302) of the noise correction model (300) may include a frame at time t that has noise corrected.

[0065] In one embodiment, a neural network model (310) executed by a processor can predict a frame to be predicted based on a frame prior to the frame to be predicted and a frame following the frame to be predicted.

[0066] The input data (311) of the neural network model (310) according to one embodiment may include at least one frame prior to the frame to be predicted and at least one frame after the frame to be predicted. Additionally, the output data (312) of the neural network model (310) according to one embodiment may include the predicted frame.

[0067] For example, as illustrated in FIG. 3, the input data (301) of the neural network model (300) may include frames at time t-2, time t-1, time t+1, and time t+2, and the output data of the neural network model (300) according to one embodiment may include a frame predicted for time t.

[0068] In one embodiment, the processor may calculate the motion flow between multiple frames using a motion correction network and perform motion correction on multiple frames based thereon. At this time, the processor may output a frame with reduced noise by adjusting specific feature or pixel values, etc., according to the motion-corrected information of multiple frames using a recursive filter.

[0069] According to one embodiment, the input data of the motion correction network and the recursive filter (320) may include a frame to be reduced in noise, a frame prior to the frame to be reduced in noise, and a frame following the frame to be reduced in noise. Additionally, according to one embodiment, the output data of the motion correction network and the recursive filter (320) may include a frame with reduced noise.

[0070] For example, as illustrated in FIG. 3, the processor can input frames at time points t-2, t-1, t, t+1, and t+2 into a motion correction network to perform motion correction for each frame. Additionally, the processor can output a frame at time point t with reduced noise by adjusting specific feature values ​​or pixel values, etc., for the frame at time point t according to the information of the motion-corrected frames at time points t-2, t-1, t, t+1, and t+2 using a recursive filter.

[0071] Meanwhile, in the present disclosure, the first output data refers to the output data (302) of the neural network model (300), the second output data refers to the output data (322) of the motion correction network and the recursive filter (320), and the third output data refers to the output data (302) of the noise correction model (300).

[0072] In one embodiment, the neural network model (310) and the motion correction network may be trained in advance prior to the training of the noise correction model in order to output the first output data and the second data, respectively.

[0073] Referring again to FIG. 3, a first error (331) between the third output data (302) and the first output data (312) and a second error (332) between the third output data (302) and the second output data (322) are shown.

[0074] The processor can calculate a first error (331) based on the third output data (302) and the first output data (312), and can calculate a second error (332) based on the third output data (302) and the second output data (322).

[0075] Here, the first error (331) or the second error (332) according to one embodiment may mean the Mean Squared Error (MSE), defined as the squared mean of the difference in pixel values ​​between each frame, or the Mean Absolute Error (MAE), defined as the average of the difference in absolute pixel values ​​between each frame. Alternatively, it may mean the Structural Similarity (SSIM) index between frames.

[0076] According to one embodiment, the processor can train a noise correction model (300) based on at least one of a first error (331) and a second error (332). For example, the processor can train the noise correction model (300) in a direction that minimizes at least one of the first error (331) and the second error (332).

[0077] Meanwhile, the timing and number of input / output data described with reference to FIG. 3 are merely examples and the present disclosure is not limited thereto. In other words, the timing or number of input / output data of each architecture can be changed by a person skilled in the art by applying the present disclosure by analogy.

[0078] FIG. 4 is a flowchart illustrating a method for obtaining first output data according to one embodiment.

[0079] The method of a neural network model predicting a frame described with reference to FIG. 4 can be performed by the user device of FIG. 1, and specifically, can be performed by a processor included in the user device.

[0080] In addition, the process of a neural network model predicting a frame, as described with reference to FIG. 4, may correspond to an embodiment of step 210 of FIG. 2.

[0081] In step 410, the processor can extract features of the input frame.

[0082] In one embodiment, the processor can extract features of an input frame using a convolution filter. Here, the convolution filter according to one embodiment may include a recursive residual convolution filter and may be configured with various sizes, such as 3x3, 5x5, and 7x7, to adjust the sensitivity of feature extraction.

[0083] Specifically, the processor can extract features by scanning multiple regions included in an input frame using a convolution filter, multiplying the input pixel values ​​by weights, and summing the resulting values. In this case, the processor can perform recursive operations by accepting the output of the previous step as the current input. In this scenario, the processor can extract features more precisely by training a neural network model based on the difference between input and output data through residuals.

[0084] In step 420, the processor can calculate attention weights.

[0085] In one embodiment, the processor can calculate an attention weight that reflects importance for each feature extracted according to step 410.

[0086] Specifically, the processor can assign higher weights to regions containing more important features or information among multiple regions included in the input frame. Here, the attention weights can be repeatedly updated by the processor during the learning process.

[0087] In step 430, the processor can output a prediction result for a frame at time t in which a specific feature is emphasized, based on the attention weights calculated according to step 420.

[0088] According to one embodiment, the processor can output a prediction result for a frame at time t as first output data by emphasizing a feature with a high attention weight calculated according to step 420.

[0089] Meanwhile, according to one embodiment, a neural network model executed by a processor may be configured to include at least one architecture among a convolution filter, a recursive residual convolution, and an attention gate to perform the aforementioned steps 410 to 430. Specifically, a neural network model according to one embodiment may include MSAR2U-NET (Multi-scale attention recurrent U-Net).

[0090] According to the above-described embodiment, by adopting a neural network structure including recursive residual convolution, context information is integrated across the entire network without setting additional network parameters, thereby improving the prediction performance of the neural network model and minimizing the computational burden.

[0091] FIG. 5 is a schematic diagram illustrating a method for obtaining second output data according to one embodiment.

[0092] Referring to Fig. 5, , , A process is illustrated in which a frame is input into a motion correction network to be motion corrected, and the motion-corrected frames are input into a recursive filter to output second output data.

[0093] According to one embodiment, a motion correction network executed by a processor can calculate the optical flow of preceding and succeeding frames and perform motion correction for each frame based on the optical flow. In this case, the motion correction network according to one embodiment may include a pre-trained spynet.

[0094] An optical flow according to one embodiment may include a forward flow and a backward flow. Here, the forward flow refers to an optical flow over time, and the backward flow refers to an optical flow that runs counter to the flow of time. For example, In the frame Optical flow into the frame refers to forward flow, and In the frame Optical flow into the frame refers to reverse flow.

[0095] In one embodiment, the processor may perform motion correction using a forward warping method based on forward flow and a backward warping method based on backward flow. Here, the forward warping method refers to a method of correcting the subsequent frame based on optical flow information of pixels moving from the previous frame to the subsequent frame, and the backward warping method refers to a method of correcting the previous frame based on optical flow information of pixels moving from the subsequent frame to the previous frame. For example, the processor In the frame Based on optical flow information of pixels into the frame Motion correction can be performed on the frame, and In the frame Based on optical flow information of pixels into the frame Motion correction can be performed on the frame.

[0096] In one embodiment, the processor may output a frame with reduced noise as second output data by adjusting specific feature or pixel values, etc., using a recursive filter for a plurality of motion-corrected frames. At this time, the processor may combine the output data generated for the previous time point with the input data for the next time point and input it back into the recursive filter.

[0097] In addition, the processor can adjust the application ratio of frames from previous / subsequent time points in the process of calculating second output data at a specific time point by setting weights within a recursive filter. A mathematical formula for weight adjustment according to one embodiment can be expressed as follows.

[0098]

[0099] In mathematical formula 1, is the i-th frame motion-corrected by the motion correction network, is the second output data output for the i-th frame, means weight.

[0100] According to the above-described embodiment, the inter-frame discontinuity of the second output data output for multiple time points can be reduced by an architecture combining a motion correction network and a recursive filter.

[0101] Meanwhile, the timing and number of input / output data and the movement flow calculated between frames described with reference to FIG. 5 are merely examples and the present disclosure is not limited thereto. In other words, the timing or number of input / output data and the method of calculating the movement flow between frames can be modified by a person skilled in the art by applying the present disclosure by analogy.

[0102] FIG. 6 is a block diagram illustrating a method for training a noise correction model using edge-enhanced second output data and third output data according to one embodiment.

[0103] The learning of the noise correction model described with reference to FIG. 6 can be performed by the user device of FIG. 1, and specifically, by a processor included in the user device.

[0104] Referring to FIG. 6, the third output data (601) and the second output data (611) may correspond to the output data (302) of the noise correction model (300) of FIG. 3 and the output data (322) of the motion correction network and recursive filter (320), respectively. Additionally, the second output data (611) may correspond to the second output data obtained according to the method described above in FIG. 5.

[0105] In one embodiment, the processor can output edge-enhanced third output data (602) and edge-enhanced second output data (612) by inputting third output data (601) and second output data (611) to an edge enhancement module. Specific details regarding a method for outputting an edge-enhanced frame using an edge enhancement module according to one embodiment will be described later in FIG. 7.

[0106] In one embodiment, the processor may calculate an error (620) between edge-enhanced third output data (602) and edge-enhanced second output data (612), and train a noise correction model based on the error (620). Here, the process of calculating the error (620) and the process of training the noise correction model based on the error (620) are as described above in FIG. 3.

[0107] Meanwhile, according to the above-described embodiment, by training a noise correction model to output a frame in which edge information and detailed structure are maintained, the correction accuracy of the noise correction model and the visual quality of the correction result can be improved.

[0108] FIG. 7 is a block diagram illustrating a method for reinforcing the edges of a frame using an edge reinforcing module according to one embodiment.

[0109] The edge reinforcement of the frame described with reference to FIG. 7 can be performed by the user device of FIG. 1, and specifically, can be performed by a processor included in the user device.

[0110] Referring to FIG. 7, a process of decomposing an input frame (701) into a plurality of subbands (710 to 740) and obtaining an output frame (702) with enhanced high-frequency components is illustrated.

[0111] The processor can decompose an input frame (701) into four sub-bands using a wavelet transform (SWT) module. In one embodiment, the processor can decompose the input frame (701) into a low-frequency component (LL component, 710), a vertical edge component (LH component, 720), a horizontal edge component (HL component, 730), and a diagonal edge component (HH component, 740) based on row and column components.

[0112] The processor can adjust the intensity of each subband according to preset conditions. For example, among the multiple subbands of an input frame, a predetermined standard may be provided to emphasize the LH component (720), HL component (730), and HH component (740) containing high-frequency components, and to exclude the LL component (710).

[0113] The processor can output an output frame (702) by reconstructing the frame using an inverse wavelet transform method.

[0114] FIG. 8 is a block diagram of a user device according to one embodiment.

[0115] Referring to FIG. 8, the user device (800) may correspond to the user device (120) of FIG. 1.

[0116] Referring to FIG. 8, the user device (800) may include a processor (810) and a memory (820). Referring to FIG. 8, only components related to the embodiment are illustrated in the user device (800). Therefore, a person skilled in the art will understand that other general-purpose components may be included in addition to the components illustrated in FIG. 8. For example, although not illustrated in FIG. 8, the user device (800) may include a communication unit that enables communication with an external server or external device and an interface unit that enables interaction with a user.

[0117] The processor (810) controls the overall operation of the user device (800). For example, the processor (810) can control the components included in the user device (800) by executing programs stored in memory (820). Additionally, the processor (810) can control the operation of the user device (800) by executing programs stored in memory (820).

[0118] The processor (810) can control at least some of the operations of the user device (800) described in FIGS. 1 to 7. For example, the processor (810) can obtain first output data by inputting at least one frame before time t and at least one frame after time t among a plurality of frames included in real-time video data into a neural network model, obtain second output data by inputting a frame at time t, at least one frame before time t and at least one frame after time t among a plurality of frames into a motion correction network and a recursive filter, train the noise correction model using at least one of the first output data and the second output data, and reduce noise included in the real-time video data using the trained noise correction model.

[0119] Meanwhile, a specific example of the operation of the processor (810) is the same as described above with reference to FIGS. 1 to 7. Therefore, a specific description of the operation of the processor (810) will be omitted below.

[0120] The processor (810) may be implemented using at least one of ASICs (application specific integrated circuits), DSPs (digital signal processors), DSPDs (digital signal processing devices), PLDs (programmable logic devices), FPGAs (field programmable gate arrays), controllers, microcontrollers, microprocessors, and other electrical units for performing functions.

[0121] In one embodiment, the user device (800) may be a mobile electronic device. For example, the user device (800) may be implemented as a smartphone, tablet PC, PC, smart TV, PDA (personal digital assistant), laptop, media player, navigation, a device equipped with a camera, and other mobile electronic devices. Additionally, the user device (800) may be implemented as a wearable device such as a watch, glasses, a hair band, and a ring equipped with communication functions and data processing functions.

[0122] In another embodiment, the user device (800) may be an electronic device embedded within a medical device. For example, the user device (800) may be an electronic device inserted into a medical imaging device through tuning after the production process.

[0123] In another embodiment, the user device (800) may be a server. The server may be implemented as a computer device or a plurality of computer devices that communicate through a network to provide commands, code, files, content, services, etc.

[0124] An embodiment according to the present invention may be implemented in the form of a computer program that can be executed through various components on a computer, and such a computer program may be recorded on a computer-readable medium. In this case, the medium may include a magnetic medium such as a hard disk, a floppy disk, and a magnetic tape, an optical recording medium such as a CD-ROM and a DVD, a magneto-optical medium such as a floptical disk, and a hardware device specifically configured to store and execute program instructions, such as a ROM, RAM, or flash memory.

[0125] Meanwhile, the above-mentioned computer program may be one specifically designed and configured for the present invention, or one known and available to those skilled in the art of computer software. Examples of computer programs may include machine code, such as that generated by a compiler, as well as high-level language code that can be executed by a computer using an interpreter, etc.

[0126] According to one embodiment, the method according to various embodiments of the present disclosure may be provided by being included in a computer program product. The computer program product may be traded between a seller and a buyer as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or distributed online (e.g., download or upload) through an application store (e.g., Play Store™) or directly between two user devices. In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily created in a device-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.

[0127] Unless explicitly stated or contrary to the order of the steps constituting the method according to the present invention, said steps may be performed in a suitable order. The present invention is not necessarily limited by the order in which said steps are described. The use of all examples or exemplary terms in the present invention is merely for the purpose of describing the present invention in detail, and the scope of the present invention is not limited by said examples or exemplary terms unless limited by the claims. Furthermore, those skilled in the art will understand that various modifications, combinations, and changes may be made according to design conditions and factors within the scope of the claims or equivalents to which they are added.

[0128] Accordingly, the scope of the present invention should not be limited to the embodiments described above, and all scopes equivalent to or equivalently modified from the claims set forth below, as well as the claims set forth below, shall be considered to fall within the scope of the concept of the present invention.

Claims

1. A method for reducing noise included in real-time video data, A step of obtaining first output data by inputting at least one frame prior to time t and at least one frame after time t into a neural network model among a plurality of frames included in real-time video data; A step of obtaining second output data by inputting the frame at time t, at least one frame prior to time t, and at least one frame after time t among the plurality of frames into a motion correction network and a recursive filter; A step of training the noise correction model using at least one of the first output data and the second output data; and A step of reducing noise included in the real-time image data using the above-mentioned learned noise correction model; A method including 2. In Paragraph 1, The above-mentioned learning step is, A step of obtaining third output data by inputting the frame at time t into the noise correction model; and A step of training the noise correction model based on the error between the first output data and the third output data; A method including 3. In Paragraph 1, The step of obtaining the first output data above is, A step of extracting features of the above-mentioned input frame; A step of calculating the residual of each frame and, based on the residual, calculating the attention weight for each feature; and A step of obtaining the first output data based on the attention weights above; A method including 4. In Paragraph 1, The above neural network model is, A method comprising at least one architecture among a convolution filter, a recursive residual convolution, and an attention gate.

5. In Paragraph 1, The above-mentioned learning step is, A step of obtaining third output data by inputting the frame at time t into the noise correction model; A step of performing motion correction on the frame at time t, at least one frame prior to time t, and at least one frame after time t; A step of obtaining the second output data by inputting the motion-corrected frames into a recursive filter; and A step of training the noise correction model based on the error between the second output data and the third output data; A method including 6. In Paragraph 5, The step of performing the above motion correction is, A step of calculating an optical flow between a frame at time t, at least one frame prior to time t, and at least one frame after time t, and performing motion correction based on the optical flow; A method including 7. In Paragraph 5, The step of obtaining the second output data above is, Step of setting weights between the motion-corrected frames; and A step of obtaining the second output data based on the above weights; A method including 8. In Paragraph 5, A step of strengthening the edges of the second output data and the third output data; and A step of training the noise correction model based on the error between the edge-enhanced second output data and the edge-enhanced third output data; A method that further includes.

9. In Paragraph 8, The step of reinforcing the above edge is, A step of decomposing the second output data and the third output data into a plurality of sub-bands through wavelet transformation; A step of adjusting the intensity of a sub-band according to a preset condition among the plurality of sub-bands; and A step of strengthening edges by reconstructing the second output data and the third output data through an inverse wavelet transform; A method including 10. In Paragraph 1, A method in which the above time point t is updated over time.

11. A computer-readable recording medium storing a program for executing the method of claim 1 on a computer.

12. At least one memory; and It includes at least one processor, The above processor is, Among a plurality of frames included in real-time video data, at least one frame prior to time t and at least one frame after time t are input into a neural network model to obtain first output data, and Among the plurality of frames, the frame at time t, at least one frame prior to time t, and at least one frame after time t are input into a motion correction network and a recursive filter to obtain second output data, and The noise correction model is trained using at least one of the first output data and the second output data, and Reducing noise included in the real-time video data using the above-mentioned learned noise correction model, A device for reducing noise included in real-time video data.