Image denoising method and system based on physical information guidance
By using a speckle noise estimation network guided by physical information and employing still images of a fake eye as guidance, the generalization ability and computational cost issues of speckle noise suppression in retinal images in OCT systems are addressed. This achieves a high signal-to-noise ratio denoising effect and improves the denoising robustness and cross-platform applicability of OCT systems.
Patent Information
- Application Number
- CN202610476123.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-13
- Publication Date
- 2026-07-21
- Estimated Expiration
- 2046-04-13
AI Technical Summary
Existing methods for suppressing speckle noise in retinal images in OCT systems suffer from limited generalization ability, high computational cost, and denoising performance limited by registration or label generation methods, which affect the accuracy of fundus image layer structure observation and diagnosis.
A physical information-guided image denoising method is adopted. By constructing a speckle noise estimation network, a high signal-to-noise ratio (SNR) still image of a fake eye is used as physical information. Combined with an unlabeled pair learning method, high SNR denoising under different noise distributions is achieved.
It improves the denoising performance and robustness of the OCT system under different noise distributions, reduces computational costs, and enables image denoising capabilities across devices and platforms.
Smart Images

Figure CN122023177B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of image denoising technology, and particularly relates to an image denoising method and system based on physical information guidance. Background Technology
[0002] The statements in this section are merely background information related to the present invention and do not necessarily constitute prior art.
[0003] Optical coherence tomography (OCT) is a technique that uses weak interference signals to perform three-dimensional measurements of samples. In recent years, to achieve high resolution and large imaging depth, OCT systems have incorporated visible and near-infrared light to create VNOCT (Visible Noise-Free Computation). Currently, VNOCT has become a powerful tool for the early diagnosis of highly blinding fundus diseases. However, in VNOCT systems, the intensity and distribution of speckle noise vary due to differences in the safe power of the incident light and the phase modulation effects of optical path devices on different wavelengths. The presence of speckle noise makes it difficult to observe the layer structure in fundus images and results in inaccurate layer segmentation, thus affecting clinical diagnosis.
[0004] Currently, speckle noise suppression methods for OCT retinal images are mainly divided into low-rank estimation, supervised, unsupervised, and label-pair denoising methods. Low-rank estimation methods require the assumption of structural invariance between adjacent frames, performing registration of multiple B-scans and iterative estimation to minimize low-rank errors to achieve denoising. This method has weak generalization ability and long computation time. Supervised methods obtain high-resolution images averaged from multiple B-scans as ground values for label pair training; their performance upper limit is limited by image registration averaging methods. Unsupervised methods denoise by internally mining statistical information from noisy images and simulating Gaussian noise to form soft label pairs, achieving denoising performance close to supervised methods. Label-pair denoising methods map the noise distribution of strongly noisy images to high-quality images, forming label pairs between the mapped images and high-quality images for supervised denoising; however, their generalization ability is weak.
[0005] In summary, current methods for suppressing speckle noise in retinal images generally suffer from limited generalization ability, high computational cost, and denoising performance limited by registration or label generation methods. Summary of the Invention
[0006] To overcome the shortcomings of the prior art, this invention provides a physical information-guided image denoising method and system, which uses a high signal-to-noise ratio (SNR) still image of a fake eye as a physical information-guided speckle noise estimation network to achieve high SNR denoising under different noise distributions.
[0007] To achieve the above objectives, the present invention adopts the following technical solution:
[0008] In a first aspect, the present invention provides an image denoising method based on physical information guidance, comprising:
[0009] The live fundus image is input into the trained speckle noise estimation network to obtain the speckle noise prediction result of the live fundus image;
[0010] Among them, a dataset image is constructed based on live fundus images under different noise modes, and static average images of fake eyes under different noise modes are included in the dataset image.
[0011] The registered average live fundus image and the pseudo-eye static average image are randomly fused into a feature map to train the speckle noise estimation network. During the training process, the estimated loss is calculated based on the prediction results of the speckle noise estimation network and the mask region features of the pseudo-eye static average image with a stable signal-to-noise ratio, and the noise distribution perception loss is calculated based on the feature maps corresponding to the prediction results of the speckle noise estimation network and the pseudo-eye static average image with a stable signal-to-noise ratio, thereby guiding the training of the speckle noise estimation network.
[0012] In a second aspect, the present invention provides an image denoising system guided by physical information, comprising:
[0013] The noise estimation module is configured to input a live fundus image into a trained speckle noise estimation network to obtain the speckle noise prediction result of the live fundus image.
[0014] The training module is configured to: construct dataset images based on live fundus images under different noise modes, and incorporate static average images of fake eyes under different noise modes into the dataset images;
[0015] The registered average live fundus image and the pseudo-eye static average image are randomly fused into a feature map to train the speckle noise estimation network. During the training process, the estimated loss is calculated based on the prediction results of the speckle noise estimation network and the mask region features of the pseudo-eye static average image with a stable signal-to-noise ratio, and the noise distribution perception loss is calculated based on the feature maps corresponding to the prediction results of the speckle noise estimation network and the pseudo-eye static average image with a stable signal-to-noise ratio, thereby guiding the training of the speckle noise estimation network.
[0016] Thirdly, the present invention provides an electronic device including a memory and a processor, and computer instructions stored in the memory and running on the processor, wherein the computer instructions, when executed by the processor, perform the method described in the first aspect.
[0017] Fourthly, the present invention provides a computer-readable storage medium for storing computer instructions, which, when executed by a processor, perform the method described in the first aspect.
[0018] The above one or more technical solutions have the following beneficial effects:
[0019] In this invention, live fundus images under different noise modes are acquired, and the average image of the pseudo-eye still image is acquired. The average image of the pseudo-eye still image with a stable signal-to-noise ratio is used as physical information to guide the speckle noise estimation network to achieve high signal-to-noise ratio denoising under different noise distributions.
[0020] In this invention, the estimation loss is calculated based on the prediction results of the speckle noise estimation network and the mask region features of the static average image of the fake eye with a stable signal-to-noise ratio. The feature extraction capability of the speckle noise estimation network is enhanced by estimating images with different noise patterns. The prediction results of the speckle noise estimation network and the average image of the static fake eye image are trained using a non-labeled pair learning method to achieve high signal-to-noise ratio denoising of fundus images under different noise distributions during acquisition.
[0021] Advantages of additional aspects of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description
[0022] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.
[0023] Figure 1 This is a schematic diagram of image acquisition averaging in Embodiment 1 of the present invention;
[0024] Figure 2 This is a schematic diagram of the speckle noise estimation network in Embodiment 1 of the present invention;
[0025] Figure 3 This is a schematic diagram of the training process of the speckle noise estimation network in Embodiment 1 of the present invention;
[0026] Figure 4 This is a noise reduction effect diagram in Embodiment 1 of the present invention. Detailed Implementation
[0027] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0028] It should be noted that the terminology used herein is for the purpose of describing particular implementations only and is not intended to limit the exemplary implementations of the present invention.
[0029] Where there is no conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.
[0030] Example 1
[0031] This embodiment discloses a physically-guided image denoising method, including:
[0032] The live fundus image is input into the trained speckle noise estimation network to obtain the speckle noise prediction result of the live fundus image;
[0033] Among them, a dataset image is constructed based on live fundus images under different noise modes, and static average images of fake eyes under different noise modes are included in the dataset image.
[0034] The speckle noise estimation network is trained by randomly fusing the registered average live fundus image and the average static image of the fake eye into a feature map. During the training process, the prediction results of the speckle noise estimation network are used to estimate the loss by combining the mask region features of the average static image of the fake eye with a stable signal-to-noise ratio. The noise distribution perception loss is also calculated by combining the prediction results of the speckle noise estimation network with the feature maps of different scales corresponding to the average static image of the fake eye with a stable signal-to-noise ratio. This process guides the training of the speckle noise estimation network.
[0035] This embodiment addresses the denoising problem under different noise distributions by proposing an asymmetric label denoising method guided by physical information. By constructing a speckle noise estimation network (SNE), a structure extraction network, and a noise distribution estimation loss, noise independence mapping is achieved. The noisy structure is transferred and mapped to the structure of a high signal-to-noise ratio static fake eye image, ultimately achieving high signal-to-noise ratio denoising under different noise distributions.
[0036] The image denoising method based on physical information guidance proposed in this embodiment will be described in detail below:
[0037] 1. Dataset image preparation.
[0038] In OCT live imaging, to obtain high signal-to-noise ratio (SNR) images, multiple B-scan scans are performed on the same location, and the scanned images are registered and aligned. The aligned images are then superimposed along the channels and averaged as the final output. However, image registration inevitably leads to pixel movement, which disrupts the spatiotemporal characteristics of speckle noise, thus limiting the image's SNR capability. Figure 1 As shown.
[0039] Therefore, this embodiment obtains multiple average images of still images as high signal-to-noise ratio (SNR) images, and uses these high SNR still images as physical information to guide the speckle noise estimation network model to achieve high SNR denoising. To improve the model's ability to perceive B-Scan images with different noise patterns, image registration for six noise patterns is performed on fundus images at a total number of 1, 2, 4, 8, 16, and 40 frames. The registered images are then superimposed along the channels, and the average value is taken as the dataset image. Superposition is stopped after 40 frames because when the averaging number reaches 40, the SNR of the averaged image basically stabilizes and no longer increases.
[0040] To construct the SNE mapping relationship under different noise modes, the images acquired by the fake eye still image were averaged across six noise modes for a total of 1, 2, 4, 8, 16, and 40 frames. The average image of the fake eye still images was then included in the dataset. The physical structure of the fake eye still image acquisition ensures the invariance of pixel spatial position in the B-Scan image during repeated scanning, thus eliminating the need for registration.
[0041] 2. Construction of speckle noise estimation model.
[0042] OCT images are single-channel images, with each pixel value representing the tissue echo intensity. Therefore, when estimating the speckle noise estimation network (SNE) of an OCT image, the number of input and output channels must be set to 1. The SNE network, i.e., the speckle noise estimation network, is as follows: Figure 2 As shown.
[0043] Figure 2 In These represent the number of input batches, the number of feature map channels, the feature map height, and the feature map width, respectively. The SNE network input image is a random fusion of the average liveness registration images from frames 1, 2, 4, 8, 16, and 40 with the average static pseudo-eye images from frames 1, 2, 4, 8, and 16. The output image is the denoised prediction value of the randomly fused input image.
[0044] To enhance the estimation capability of the SNE network, the feature map channel dimension is expanded to 32 times its original size at the first convolution (Conv3x3), followed by feature mapping using the ReLU activation function. Then, two dilated convolutions (DConv3x3) with a dilation rate of 2 are applied to weight the contributions of surrounding pixels, without changing the number of channels. Since speckle noise in OCT images has global randomness, a larger receptive field is not required for feature extraction when estimating SNE. Next, a second convolution (Conv3x3) with ReLU activation is applied for further feature extraction from the dilated convolutional feature map. Finally, a Conv3x3 with ReLU activation is applied to compress the feature map channels. To ensure that the feature map size remains unchanged before and after convolution, all convolution kernel sizes are set to 3, Conv3x3 padding is 1, dilation rate is 1, and DConv3x3 padding is 2 with a dilation rate of 2. The final compressed dimension is... Feature maps and label images Loss calculation, here The value is 1. Furthermore, to enhance the SNE estimation capability at different feature layer scales, a multi-scale feature extraction network, ResNet50, is introduced during training, and the feature map outputs of ResNet50 downsampled by 2, 4, 8, and 16 times are used in… Loss calculation.
[0045] 3. Training strategy design.
[0046] In the dataset image preparation, registered fundus images and still images have been acquired. To achieve high signal-to-noise ratio denoising of unlabeled fundus images, the following steps are required: Figure 3 The training process is shown below.
[0047] During training, a binary mask is randomly generated. The average registered fundus images with total frames of 1, 2, 4, 8, 16, and 40 are randomly fused with the average still images with total frames of 1, 2, 4, 8, and 16 to form a single feature map, which is then input into the SNE network for training. The average still images of the fake eye (40 frames) serve as the ground truth labels for the speckle noise estimation model. Loss calculation.
[0048] During SNE training, firstly, the predicted values from the fused feature image are compared with the features from the 1-mask region selected from the average of 40 still frames of the fake eye. By estimating images with different noise patterns, the feature extraction capability of the SNE network is enhanced.
[0049] Secondly, to improve the denoising capabilities of fundus images with different noise patterns, an unlabeled pair learning method was used to train the prediction results of the SNE fused feature image and the average of 40 frames of still pseudo-eye images. Specifically, the prediction results and the average of 40 frames were used together to extract feature maps downsampled by 2, 4, 8, and 16 times through a ResNet50 network, and the features were then processed at the same scale. Figure 1 One-to-one correspondence, and calculation Loss is used to enhance the speckle noise estimation model's ability to perceive noise under different receptive fields. To reduce the interference of structural information during unlabeled training on the training results, [the following is used]. The loss function utilizes the semantic information of the image to constrain the network. Through these steps, high signal-to-noise ratio denoising is ultimately achieved for fundus images acquired by the VNOCT system under different noise distributions. The results are as follows: Figure 4 As shown.
[0050] 4. Loss function construction.
[0051] To achieve multi-scale learning of image structure and SNE estimation, the loss function is constructed as follows:
[0052] The first part is the SNE estimation loss function:
[0053] (1)
[0054] In the formula, , and These are weighting coefficients, which are 1.0, 0.5, and 0.1 in this embodiment.
[0055] The model training is constrained by three parts, the first being... The separation loss function is used to achieve pixel-level mapping and semantic information perception, enabling accurate separation of structural and noise features. Its calculation formula is as follows:
[0056] (2)
[0057] In the formula, This represents the model prediction value of the fused feature image. This represents the average image label value of the fake eye at rest for 40 frames. Indicates the first The first batch line, number Column pixel prediction values Indicates the first The first batch line, number Column pixel label values, For batch size, and Indicates the height and width of the image. Indicates the first One batch, Indicates the first row pixels, Indicates the first Column pixels.
[0058] The second part involves calculating the gradient loss using the model's predicted values from the fused feature images and the average image label values from 40 still frames of the fake eye. To ensure that edge pixels in the network prediction value structure of the fused feature image are not lost, the calculation formula is as follows:
[0059] (3)
[0060] In the formula, For the image in the first The directional gradient value in the width direction at each pixel. For the image in the first The directional gradient value in the height direction at each pixel. and These represent the gradient values in the width and height directions of the model's predicted values, respectively, representing the fused feature images. and These represent the width and height gradient values of the average image of the pseudo-eye over 40 static frames, respectively. Edge preservation is achieved by imposing pixel-level constraints on the image gradient values.
[0061] Finally, structural similarity is used to globally constrain the predictions of the SNE model, and the calculation formula is shown below:
[0062] (4)
[0063] In the formula, and These represent the mean of the model's predicted values and the mean of the average image label values for 40 frames of still images from the fake eye, respectively. and These represent the variance of the model's predicted values and the variance of the average image label values for 40 frames of still images from the fake eye, respectively. The covariance of the model's predicted values and the covariance of the average image label values for 40 frames of still images from the fake eye are given. This indicates that the constant is 0.001.
[0064] Building upon this, in order to incorporate multi-scale information into the loss calculation and thus improve the model's SNE estimation capability under different receptive fields, a noise distribution-aware loss is introduced, calculated as follows:
[0065] (5)
[0066] In the formula, and Let represent the mean of the i-th channel of the model-predicted feature map and the mean of the i-th channel of the average image label feature map of the fake eye over 40 still frames, respectively. and Let represent the standard deviation of the i-th channel of the model's predicted feature map and the standard deviation of the i-th channel of the average image of the fake eye at rest for 40 frames, respectively. This represents the total number of channels.
[0067] Total model loss As shown in the following formula:
[0068] (6)
[0069] In dual-band OCT systems, the incident power differs between the visible and near-infrared bands, and the spectral absorption characteristics of biological tissues differ in the two bands, resulting in significant inconsistencies in the noise distribution of the acquired dual-band OCT images. To address this issue, this embodiment employs a label-free pair denoising method to achieve simultaneous denoising of dual-band OCT images. The label-free pair denoising method proposed in this embodiment has no requirements for live-body imaging OCT images, fundamentally avoiding the inherent limitations of image registration on the system's signal-to-noise ratio. Compared to existing denoising methods, the method in this embodiment has stronger robustness and generalization ability, and is more easily implemented for OCT image denoising across devices and platforms.
[0070] Example 2
[0071] The purpose of this embodiment is to provide a physically guided image denoising system, including:
[0072] The noise estimation module is configured to input a live fundus image into a trained speckle noise estimation network to obtain the speckle noise prediction result of the live fundus image.
[0073] The training module is configured to: construct dataset images based on live fundus images under different noise modes, and incorporate the average image of the pseudo-eye still images into the dataset images;
[0074] The registered average live fundus image and the pseudo-eye still image are randomly fused into a feature map to train the speckle noise estimation network. During the training process, the estimated loss is calculated based on the prediction results of the speckle noise estimation network and the mask region features of the pseudo-eye still average image with a stable signal-to-noise ratio, and the noise distribution perception loss is calculated based on the feature maps corresponding to the prediction results of the speckle noise estimation network and the pseudo-eye still average image with a stable signal-to-noise ratio, thereby guiding the training of the speckle noise estimation network.
[0075] In further embodiments, the following is also provided:
[0076] An electronic device includes a memory and a processor, as well as computer instructions stored in the memory and running on the processor. When executed by the processor, the computer instructions perform the method described in Embodiment 1. For brevity, further details are omitted here.
[0077] It should be understood that in this embodiment, the processor can be a central processing unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc.
[0078] Memory may include read-only memory and random access memory, and provides instructions and data to the processor. A portion of memory may also include non-volatile random access memory. For example, memory may also store information about the device type.
[0079] A computer-readable storage medium for storing computer instructions, which, when executed by a processor, perform the method described in Embodiment 1.
[0080] The method in Embodiment 1 can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules within the processor. The software modules can reside in readily available storage media in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory; the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method. To avoid repetition, a detailed description is not provided here.
[0081] A computer program product includes a computer program that, when executed by a processor, implements the method described in Embodiment 1.
[0082] The present invention also provides at least one computer program product tangibly stored on a non-transitory computer-readable storage medium. The computer program product includes computer-executable instructions, such as instructions included in program modules, which execute in a device on a target real or virtual processor to perform the processes / methods described above. Typically, program modules include routines, programs, libraries, objects, classes, components, data structures, etc., that perform specific tasks or implement specific abstract data types. In various embodiments, the functionality of program modules can be combined or divided among program modules as needed. The machine-executable instructions for the program modules can execute within a local or distributed device. In a distributed device, the program modules can reside in both local and remote storage media.
[0083] The computer program code used to implement the methods of the present invention may be written in one or more programming languages. This computer program code may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the computer or other programmable data processing device, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a computer, partially on a computer, as a stand-alone software package, partially on a computer and partially on a remote computer, or entirely on a remote computer or server.
[0084] In the context of this invention, computer program code or related data may be carried by any suitable carrier to enable a device, apparatus, or processor to perform the various processes and operations described above. Examples of carriers include signals, computer-readable media, and the like. Examples of signals may include electrical, optical, radio, sound, or other forms of propagation signals, such as carrier waves, infrared signals, etc.
[0085] Those skilled in the art will recognize that the units and algorithm steps described in conjunction with the embodiments herein can be implemented in electronic hardware or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0086] While the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of the present invention are still within the scope of protection of the present invention.
Claims
1. A physically-guided image denoising method, characterized in that, include: The live fundus image is input into the trained speckle noise estimation network to obtain the speckle noise prediction result of the live fundus image; Among them, a dataset image is constructed based on live fundus images under different noise modes, and static average images of fake eyes under different noise modes are included in the dataset image. The registered average live fundus image and the pseudo-eye static average image are randomly fused into a feature map to train the speckle noise estimation network. During training, the estimated loss is calculated based on the prediction results of the speckle noise estimation network and the mask region features of the pseudo-eye static average image with a stable signal-to-noise ratio, and the noise distribution perception loss is calculated based on the feature maps corresponding to the prediction results of the speckle noise estimation network and the pseudo-eye static average image with a stable signal-to-noise ratio, thereby guiding the training of the speckle noise estimation network. The estimated loss includes image gradient loss, structural similarity loss, and separation loss.
2. The image denoising method based on physical information guidance as described in claim 1, characterized in that, A dataset was constructed based on live fundus images under different noise patterns. Static average images of fake eyes under different noise patterns were then included in the dataset. Specifically: Live fundus images were acquired at different total frame rates. Image registration was performed on live fundus images with different total frame counts. The average value of the image-registered live fundus images was then taken as the dataset image after being superimposed along the channels. Obtain still images of the fake eye with the same total number of frames as the images acquired from the live fundus, and include the average image of the still images of the fake eye into the dataset.
3. The image denoising method based on physical information guidance as described in claim 1, characterized in that, The speckle noise estimation network processes the feature map obtained by randomly fusing the registered average live fundus image and the pseudo-eye static average image as follows: The channel dimension of the random fused feature map is expanded using the first convolution and then feature mapping is performed. The contributions of surrounding pixels in the feature map after the first convolution are weighted by dilated convolution. The second convolution is used to extract features from the feature map after dilated convolution, and then channel compression is performed on the extracted feature map to obtain the speckle noise prediction result.
4. The image denoising method based on physical information guidance as described in claim 1, characterized in that, The in vivo fundus images or the average static images of the fake eye under different noise modes are specifically: in vivo fundus images or the average static images of the fake eye under different total frames corresponding to the noise modes.
5. The image denoising method based on physical information guidance as described in claim 1, characterized in that, The registered average live fundus image and the pseudo-eye static average image are randomly fused into a feature map. The fused feature map is input into the speckle noise estimation network for training. During the training process, the 1-mask region of the predicted value of the speckle noise estimation network and the pseudo-eye static average image with a stable signal-to-noise ratio are selected to calculate the estimated loss. The training of the speckle noise estimation network is guided by the estimated loss.
6. The image denoising method based on physical information guidance as described in claim 1, characterized in that, The predictions of the speckle noise estimation network were trained on the pseudo-eye still average image using an unlabeled pair learning method.
7. The image denoising method based on physical information as described in claim 1 or 6, characterized in that, The registered average live fundus image and the pseudo-eye static average image are randomly fused. The fused feature map is input into the speckle noise estimation network for training. During the training process, the predicted values of the speckle noise estimation network and the pseudo-eye static average image with a stable signal-to-noise ratio are extracted at different magnifications. The noise distribution perception loss is calculated one-to-one with the sampling feature maps of the same scale. The training of the speckle noise estimation network is guided by the noise distribution perception loss.
8. The image denoising method based on physical information as described in claim 1 or 5, characterized in that, The separation loss is used to achieve pixel-level mapping relationships and semantic information perception, and to separate structural and noise features.
9. The image denoising method based on physical information guidance as described in claim 8, characterized in that, The image gradient loss for: in, For the image in the first The directional gradient value in the width direction at each pixel. For the image in the first The directional gradient value in the height direction at each pixel. and These represent the gradient values in the width and height directions of the model's predicted values, respectively, representing the fused feature images. and These represent the width and height gradient values of the static average image of the fake eye where the signal-to-noise ratio tends to stabilize; This is the separation loss function.
10. The image denoising method based on physical information as described in claim 8, characterized in that, The structural similarity loss Specifically: in, and Let represent the mean of the model's predicted values and the mean of the average label values of the pseudo-eye at rest when the signal-to-noise ratio tends to stabilize, respectively. and Let V represent the variance of the model predictions and the variance of the average label values of the pseudo-eye at rest when the signal-to-noise ratio tends to stabilize, respectively. The covariance of the model predictions and the covariance of the pseudo-eye static average image label values with a stable signal-to-noise ratio; This represents a constant of 0.001; For structural similarity.
11. The image denoising method based on physical information as described in claim 8, characterized in that, The separation Specifically: in, This represents the model prediction value of the fused feature image. The average label value of a static image of a fake eye where the signal-to-noise ratio tends to stabilize. Indicates the first The first batch line, number Column pixel prediction values Indicates the first The first batch line, number Column pixel label values, For batch size, and Indicates the height and width of the image. Indicates the first One batch, Indicates the first row pixels, Indicates the first Column pixels.
12. The image denoising method based on physical information as described in claim 1, characterized in that, The noise distribution perception loss Specifically: in, and Let represent the mean of the i-th channel of the feature map predicted by the model and the mean of the i-th channel of the static average image of the fake eye with a stable signal-to-noise ratio, respectively. and Let $\begin{pmatrix}$ represent the standard deviation of the i-th channel of the model-predicted feature map and the standard deviation of the i-th channel of the pseudo-eye static average image with a stable signal-to-noise ratio, respectively. This represents the total number of channels.
13. A physically-guided image denoising system, characterized in that, include: The noise estimation module is configured to input a live fundus image into a trained speckle noise estimation network to obtain the speckle noise prediction result of the live fundus image. The training module is configured to: construct dataset images based on live fundus images under different noise modes, and incorporate static average images of fake eyes under different noise modes into the dataset images; The registered average live fundus image and the pseudo-eye static average image are randomly fused into a feature map to train the speckle noise estimation network. During training, the estimated loss is calculated based on the prediction results of the speckle noise estimation network and the mask region features of the pseudo-eye static average image with a stable signal-to-noise ratio, and the noise distribution perception loss is calculated based on the feature maps corresponding to the prediction results of the speckle noise estimation network and the pseudo-eye static average image with a stable signal-to-noise ratio, thereby guiding the training of the speckle noise estimation network. The estimated loss includes image gradient loss, structural similarity loss, and separation loss.
14. An electronic device, characterized in that, It includes a memory and a processor, as well as computer instructions stored in the memory and running on the processor, which, when executed by the processor, perform the method according to any one of claims 1-12.
15. A computer-readable storage medium, characterized in that, Used to store computer instructions, which, when executed by a processor, perform the method described in any one of claims 1-12.
16. A computer program product, characterized in that, Includes a computer program, which, when executed by a processor, implements the method according to any one of claims 1-12.
Citation Information
Patent Citations
Neural network based enhancement of intensity images
US20200286208A1
KR20230153091A