Speckle noise image processing apparatus

By employing an anisotropic gradient regularized self-supervised learning algorithm, and utilizing speckle noise sequence images generated by a laser source, the problems of equipment complexity and robustness in laser speckle noise removal are solved, achieving efficient and low-cost laser image denoising.

CN115063595BActive Publication Date: 2025-11-04HUST SUZHOU INST FOR BRAINMATICS +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210770815.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-30
Publication Date
2025-11-04
Estimated Expiration
2042-06-30

AI Technical Summary

Technical Problem

Existing technologies for removing laser speckle noise suffer from problems such as high equipment complexity, high cost, large transmission loss, long response time, and poor noise removal robustness. In particular, traditional methods cannot effectively remove speckle noise in applications that require laser coherence for measurement.

Method used

An anisotropic gradient regularization self-supervised learning algorithm is adopted. By using the self-supervised learning algorithm and the anisotropic gradient regularization loss function, speckle noise generated by laser illumination is removed from the speckle sequence image which is partially independent in time, thus achieving efficient image denoising.

Benefits of technology

It improves the practicality and feasibility of laser image speckle removal, reduces reliance on optical components, lowers equipment complexity and cost, and improves noise reduction performance and robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115063595B_ABST
    Figure CN115063595B_ABST
Patent Text Reader

Abstract

The application discloses a speckle noise image processing device, and belongs to the technical field of image processing. The image processing device comprises a coherent signal source, a signal sensing element and an image processing unit. The image processing unit comprises a network structure subunit. The network structure subunit is used for removing speckle noise in a noisy image by using a self-supervised learning algorithm. Network parameters of the network structure subunit are obtained by training based on a total loss function. The total loss function comprises an anisotropic gradient regularization loss function. The anisotropic gradient regularization loss function is determined based on image values of noisy images at adjacent time points. The algorithm only needs light source illumination, and does not need to additionally add hardware such as an SLM, an LED and an oLSR. With the aid of speckle sequence images which are partially independent in time, the smoothness of a homogeneous region of an image is accurately constrained by using the anisotropic gradient regularization loss function, the residual speckle noise after processing the noisy image by using a traditional network structure is removed, and speckle removal of a laser image is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and in particular to a speckle noise image processing device. Background Technology

[0002] Lasers, with their high power density, good coherence, good directionality, long lifespan, and excellent modulation characteristics, are widely used in industrial and biomedical fields such as manufacturing, measurement, imaging, display, diagnosis, and treatment. Examples include laser 3D topography measurement, laser projection display, laser holographic measurement, and coherent photobiological imaging. When a laser beam illuminates an optically rough surface or biological tissue, its high coherence causes the scattered light to form a spatially alternating, randomly distributed speckle pattern on the far-field detection or imaging plane, known as laser speckle. As the shape and position of the irradiated object change over time, the deformation, displacement, and motion of the object cause dynamic changes in the speckle pattern. Therefore, in applications such as laser display, holographic measurement, and bioimaging, laser speckle degrades image quality or measurement accuracy and is often a noise that needs to be removed. Furthermore, imaging and measurement using signal sources such as ultrasound and microwaves also face similar speckle noise as laser speckle, and their spatiotemporal statistical characteristics are similar to those of laser speckle.

[0003] To eliminate the impact of laser speckle on coherent light imaging and measurement, optical devices such as optical laser speckle reducers (oLSRs) and spatial light modulators (SLMs) can be introduced to remove speckle noise. However, these speckle-reducing optical devices generally have limitations in terms of lifespan, transmission loss, and response time. Furthermore, some speckle pattern remains after speckle removal, which is detrimental to the miniaturization and convenience of the overall equipment. In particular, in applications requiring measurements based on the coherence of laser light, where imaging is achieved directly through laser illumination, speckle reduction devices cannot be used; instead, image processing is necessary to remove the speckle effect.

[0004] In recent years, deep learning-based image processing methods have been widely used in image denoising, offering significant improvements in denoising performance and processing speed compared to traditional methods. Conventional deep learning methods require input to the network, consisting of image pairs—noisy and noiseless images—for end-to-end network learning to denoise. This necessitates using laser illumination to obtain the noisy image and incoherent LED illumination to obtain the noiseless image for network training, increasing optical path complexity and cost, and exhibiting poor robustness against strong speckle noise. Summary of the Invention

[0005] This application provides a speckle noise image processing device that uses anisotropic gradient regularization self-supervised learning to remove speckle noise from noisy images.

[0006] This application provides a speckle noise image processing apparatus, the image processing apparatus comprising:

[0007] A coherent signal source that generates a noisy image with speckle noise during imaging;

[0008] Signal sensing elements are used to acquire noisy images at different times;

[0009] The image processing unit includes a network structure subunit. The network structure subunit employs a self-supervised learning algorithm to remove speckle noise from the noisy image. The network parameters of the network structure subunit are obtained based on a total loss function, which includes an anisotropic gradient regularization loss function. The anisotropic gradient regularization loss function is determined based on the image values ​​of the noisy image at adjacent time points.

[0010] Furthermore, the anisotropic gradient regularization loss function is determined in the following manner:

[0011]

[0012] Where x represents the network output image containing residual noise that has not been fully repaired, vec represents expanding the matrix into a one-dimensional vector, B represents the number of image frames in a training batch, H and W represent the height and width of the noisy image, respectively, and D is the anisotropic search radius. The operator representing the computation of anisotropic gradients, specifically,

[0013]

[0014] Where j is the loop variable; h and w are the pixel coordinates of the image, and x... b,h,w R represents the image value at pixel (h,w) of the b-th frame within the image batch. h,w It is the pixel coordinate of the distance D from pixel (h,w). A set that consists of.

[0015] Furthermore, the total loss function also includes a training loss, which is determined based on a self-supervised learning algorithm, including Noise2Noise, Noise2Self, or Noise2Void.

[0016] Furthermore, the network structure includes a U-Net network structure, a DnCNN network structure, a DRSNet network structure, or a U-Former network structure.

[0017] Furthermore, the U-Net network structure employs global residual learning and batch normalization. The batch normalization includes three steps: normalization, scaling, and translation, to standardize the input of the convolutional kernel.

[0018] Furthermore, the image processing unit also includes a network training set generation subunit. During training, a noisy image at a certain time interval t is selected from the noisy image sequence acquired by the signal sensing element for training, where 10ms < t < 100ms.

[0019] Furthermore, the image processing unit also includes a fast computation subunit, which is used to calculate anisotropic gradients during training through convolution operations.

[0020] Furthermore, the fast calculation subunit is also used to fill the edges of the noisy image with constant values ​​of row D or column D, where D is the anisotropic search radius.

[0021] Furthermore, the coherent signal source is laser, ultrasound, or microwave.

[0022] Furthermore, the image processing unit also includes a data preprocessing subunit and a data postprocessing subunit. The data preprocessing subunit is used to perform a logarithmic transformation on the noisy image, and the transformed noisy image is used as the input image of the network structure. The data postprocessing subunit is used to perform an inverse exponential transformation on the output image of the network structure to obtain a denoised image.

[0023] One or more technical solutions provided in the embodiments of this application have at least the following technical effects or advantages:

[0024] This application employs anisotropic gradient regularization self-supervised learning for speckle noise removal. This algorithm requires only laser illumination and does not require additional hardware such as SLM, LED, or oLSR. By utilizing speckle sequence images where the speckle noise is temporally independent, the anisotropic gradient regularization loss function accurately constrains the smoothness of homogeneous regions in the image, removing the speckle noise remaining after processing noisy images with traditional network structures, thus achieving speckle removal from laser images. Since noisy images contain multiplicative noise components, logarithmic transformation makes the noise level more spatially uniform, improving the network's denoising performance. This algorithm enhances the practicality and feasibility of speckle removal in laser images. Attached Figure Description

[0025] To more clearly illustrate the technical solutions in the embodiments of this utility model, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this utility model. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0026] Figure 1 This is a schematic diagram of the structure of a speckle noise image processing device according to the present invention;

[0027] Figure 2 This is a block diagram of the image processing unit of the present invention;

[0028] Figure 3 This is a schematic diagram of the AGR-N2N algorithm based on the U-NET network structure of the present invention. The Training part is the training process, and the Inference part is the inference process.

[0029] Figure 4 The network structure diagrams of DRSNet and DnCNN of the present invention are shown in (a) the network structure of DRSNet and (b) the network structure of DnCNN containing 17 convolutional layers.

[0030] Figure 5 The network structure diagram of the U-Former of the present invention is shown in the figure. (a) Overall structure of U-Former, (b) LeWinTransformer module, (c) W-MSAs module in LeWin Transformer module.

[0031] Figure 6 The denoising results of different denoising methods of the present invention are shown in Figures (a) to (h), which are respectively a 405nm laser image (single frame exposure time 20ms, average of 2 frames), BM3D, N2N(ori), N2N, TV-N2N, AGR-N2N, reference image, and white light image. The dashed area in (h) is... Figure 5 The details to be shown in the text;

[0032] Figure 7 for Figure 6 Enlarged views of the two ROIs, Figures (a) to (h) are Figure 6 (h) The area marked by the dashed line on the left, as shown in Figures (i) to (p). Figure 6 (h) The area marked by the dashed line on the right. Figures (a) to (h) are respectively the 405nm laser image (single frame exposure time 20ms, average of 2 frames), ADF, BM3D, N2N(ori), N2N, TV-N2N, AGR-N2N, and reference image. The image order of Figures (i) to (p) is the same as that of Figures (a) to (h). Detailed Implementation

[0033] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other.

[0034] Figure 1 This is a schematic diagram of the image processing apparatus of the present invention. Figure 1 As shown in the figure, this application provides an image processing device, which includes a coherent signal source 1, a signal sensing element 3, and an image processing unit 4. The coherent signal source 1 images a sample 5, generating a noisy image with speckle noise during the imaging process. The signal sensing element 3 is used to acquire noisy images at different times.

[0035] like Figure 2 As shown, the image processing unit 4 includes a network structure subunit 100. The network structure subunit 100 uses a self-supervised learning algorithm to remove speckle noise from the noisy image. The network parameters of the network structure subunit 100 are trained based on a total loss function, which includes an anisotropic gradient regularization loss function. The anisotropic gradient regularization loss function is determined based on the image values ​​of the noisy image at adjacent time points.

[0036] In some embodiments, the coherent signal source 1 of the image processing apparatus provided by the present invention can be a 473nm single-mode blue laser (MSL-FN-473, 100mW), or a 632.8nm HeNe laser (Melles Griot, USA, 15mW), or a 405nm blue-violet laser (OM-12A405-30-G, Oeabt, China, 30mW). In other embodiments, the coherent signal source can also be ultrasound or microwave.

[0037] Furthermore, the signal sensing element 3 can be a CMOS area array camera (acA2040-120um, Basler, Germany, 2048×1536 pixels, 12-bit) connected to the stereo microscope 2 (Z16 APO, Leica, Germany).

[0038] In some embodiments, the image processing unit 4 uses an NVIDIA RTX 3090 GPU as its computing platform and employs anisotropic gradient regularization self-supervised learning for speckle removal in biological images. This algorithm requires only laser illumination and does not require additional hardware such as SLM, LED, and oLSR. By leveraging the cell movement of biological tissues, speckle sequence images with temporally partially independent speckle noise can be obtained. The anisotropic gradient regularization loss function accurately constrains the smoothness of homogeneous regions in the image, removing residual speckle noise after processing noisy images using traditional network structures, thus achieving speckle removal in laser images. This algorithm improves the practicality and feasibility of speckle removal in laser images.

[0039] Furthermore, self-supervised learning algorithms include Noise2Noise(N2N), Noise2Self, or Noise2Void. Figure 3 This is a schematic diagram of the overall AGR-N2N algorithm based on the U-NET network structure of the present invention. The Training section describes the training process, and the Inference section describes the inference process. The network structure subunit 100 is introduced below using the AGR-N2N algorithm based on the U-NET network structure as an example.

[0040] Furthermore, the total loss function also includes the training loss of the network structure. For example... Figure 3 As shown, the training loss L N2N The L1 loss is used to reduce the influence of outliers between the input and target images, and its training results are better than L2 loss. In deep learning training, gradient backpropagation is used to optimize the network parameters and reduce the overall loss L(P). The calculation method of L(P) is shown in the following formula:

[0041]

[0042] Where, λ AGR Here, is the anisotropic gradient regularization coefficient, x represents the noisy image captured by the camera, y represents the noisy image at adjacent time points in the sampling sequence, H and W represent the number of pixels corresponding to the height and width of the image, respectively, P represents the trainable parameters in the network, B represents the number of images in a mini-batch, and f(x, P) represents the output image of the network with parameter P after processing x.

[0043] Furthermore, the anisotropic gradient regularization loss function L AGR Determined in the following ways:

[0044]

[0045] Where x represents the network output image containing residual noise that has not been fully repaired, vec represents expanding the matrix into a one-dimensional vector, B represents the number of image frames in a training batch, H and W represent the height and width of the noisy image, respectively, and D is the anisotropic search radius, which is optimally set to 5 and is also the default value for AGR. Other settings, such as 3 or 4, can also achieve similar results and are all better than TV loss.

[0046] The operator representing the computation of anisotropic gradients, specifically,

[0047]

[0048] Where j is the loop variable; h and w are the pixel coordinates of the image, and x... b,h,w R represents the image value at pixel (h,w) of the b-th frame after processing within the image batch. h,w It is the pixel coordinate of the distance D from pixel (h,w). The set of components. AGR selects the minimum gradient by min and searches for the direction of the minimum gradient. It can calculate the gradient along homogeneous regions of the image, avoiding the introduction of gradients containing the true structural information of the image into the loss function. This allows for accurate constraint on the smoothness of homogeneous regions of the image, avoiding ambiguity of image structural information. After calculating the anisotropic gradient, the gradient values ​​at each pixel of the entire image are averaged to obtain the AGR loss term.

[0049] The network structure adopts the U-Net architecture, which reduces the computational cost through multi-layer downsampling and quickly achieves a large receptive field. The convolutional kernel size is set to 3×3. Unlike the traditional U-Net structure used for segmentation tasks, this invention uses global residual learning (…). Figure 3 The network employs Residual Connections (RUD) and BNorm (Batch Normalization) to improve learning efficiency. BNorm involves three steps: normalization, scaling, and translation, to standardize the input to the convolutional kernels, thus improving network learning efficiency. The network uses ReLU activation functions to enhance its non-linear mapping capabilities. Zeros are padded before convolutions to ensure that the output image of the convolutional layer has the same two-dimensional dimensions as the input image.

[0050] The network structure may also include a DnCNN network structure, a DRSNet network structure, or a U-Former network structure. Figure 4 The diagram illustrates the network architectures of DnCNN and DRSNet, where Conv(A, B, C) represents the number of convolutional kernel groups in that layer. DRSNet achieves a large receptive field through dilated convolutions. Figure 5The paper demonstrates the network structure of U-Former, which improves network performance through the attention mechanism of transformer.

[0051] Furthermore, the image processing unit also includes a network training set generation subunit 200. During training, noisy images with a certain time interval t are selected from the noisy image sequence acquired by the signal sensing element for training, where 10ms < t < 100ms, in order to improve the independence of the noisy images.

[0052] Furthermore, the image processing unit also includes a fast computation subunit 300, which is used to calculate anisotropic gradients through convolution operations during training, thereby achieving efficient computation.

[0053] Furthermore, the fast computation subunit 300 is also used to fill the edges of the noisy image with D row constant values ​​or D column constant values, where D is the anisotropic search radius, to avoid the edge effect caused by calculating anisotropic gradient regularization through convolution.

[0054] Furthermore, the image processing unit also includes a data preprocessing subunit 101 and a data postprocessing subunit 102. The data preprocessing subunit 101 is used to perform a logarithmic transformation on the noisy image, and the transformed noisy image is used as the input image of the network structure. The data postprocessing subunit 102 is used to perform an inverse exponential transformation on the output image of the network structure to obtain a denoised image. Since the noisy image contains multiplicative noise components, the logarithmic transformation makes the noise level more uniformly distributed in space, improves the denoising performance of the network, and enhances the robustness to multiplicative noise denoising.

[0055] The training and usage of the above-mentioned device are illustrated below with specific embodiments.

[0056] In practice, during data acquisition, the illumination source was switched to illuminate the mouse cerebral cortex separately, and 400 consecutive noisy images were captured for each scene. When acquiring laser images at wavelengths of 473nm and 632.8nm, the camera exposure time was 10ms. However, when acquiring laser images at a wavelength of 405nm, a 20ms exposure time was used due to the weaker image intensity. During the acquisition process, different microscope magnifications were adjusted to obtain multi-scale image information, covering optical magnifications of 0.44×, 1.00×, 1.25×, 1.58×, and 2.02×.

[0057] A total of 271 noisy image sequences of mouse cerebral cortex were captured, with each scene consisting of 400 frames. Three different time points were taken from the 400 frames of each scene at equal intervals, and two images were selected from each adjacent time point to form an image pair. The two images in a pair were separated by 30ms, resulting in a total of 813 image pairs forming the training set.

[0058] To verify the robustness of the speckle removal algorithm under different noise levels, a single 2D noisy input image was averaged using 1, 3, 5, and 10 frames of noisy images. The more frames averaged, the weaker the speckle noise becomes. Images averaged using the four methods were grouped into four different training datasets, each containing 813 image pairs.

[0059] All networks are trained for 100 generations. The learning rate decays from 10. -4 The exponential decay to 10 -5 The network was trained using the PyTorch 1.10.0 deep learning framework on an NVIDIA RTX 3090 GPU. The Adam optimization algorithm was used to adaptively update the network parameters, with momentum parameters β1 and β2 set to 0.9 and 0.999, respectively. Each mini-batch contained 8 images, randomly cropped to 256×256 pixels. To augment the training data, the training images were randomly rotated 0° or 90° before being input to the network, and randomly flipped horizontally and vertically. Once trained, the network can be directly used for speckle denoising of laser images.

[0060] During the testing phase, four test sets, corresponding to the four training sets, each contained 108 pairs of images. Each pair included an averaged noisy image and a reference image from the same scene. The test sets included noisy images of the cerebral cortex tissues of different mice. The noisy images were acquired in the same way as the training sets, but the reference images were acquired differently. After averaging 400 consecutive frames of noisy images acquired in the scene, a 5×5 spatial window mid-range filter was applied to remove residual speckle noise particles, resulting in a high-quality reference image.

[0061] In denoising experiments, AGR-N2N was compared with TV-N2N, N2N, N2N(ori), BM3D, and ADF to study the improvement effect of AGR-N2N. BM3D was used to process the original noisy image as a control method. N2N(ori) and N2N denoising networks were trained in the original domain and the logarithmic transform domain, respectively, to verify the improvement effect in the logarithmic transform domain. In subsequent descriptions, when referring to the N2N series networks, unless otherwise specified, it indicates that the network was trained in the logarithmic transform domain. Furthermore, TV-N2N based on TV regularization and AGR-N2N based on AGR were trained to verify the superiority of AGR.

[0062] The original noisy images collected in the experiment were denoised using the ADF, BM3D, N2N(ori), N2N, TV-N2N, and AGR-N2N algorithms respectively. Except for the network names with the suffix "(ori)" indicating that the network is trained in the original domain, other N2N series networks are trained in the logarithmic transformation domain. Figure 5 shows the comparison of the denoising results of different denoising algorithms for the 405nm laser noisy images, and the exposure time of the camera is 20ms.

[0063] Figure 6 are the denoising results of different denoising methods. Among them, Figures (a) to (h) are the 405nm laser images (single-frame exposure time 20ms, averaged over 2 frames), BM3D, N2N(ori), N2N, TV-N2N, AGR-N2N, reference image, and white light image respectively. MS represents the MS-SSIM evaluation index.

[0064] As Figure 6 shown, the noisy image in Figure (a) is obtained by averaging two frames of laser noisy images. The two selected regions of interest are shown in Figure 6 , and the positions of the regions of interest are marked by dotted lines in (f). As shown in Figures (b) and the local details Figure 7 of (b) and (j), the output of ADF retains obvious speckle noise, and the local weighted filtering of ADF also causes blurring of the edges of the image structure. As can be seen in Figures (c) and the local details Figure 7 of (c) and (k), BM3D misidentifies some speckle noise as real structures, resulting in the retention of noise and some bright spots remaining in the denoised image. In addition, there is also a phenomenon that small blood vessels are filtered out as noise. However, it has been significantly improved compared with the noisy input image, and the MS-SSIM index of the image denoised by BM3D is 0.9582.

[0065] Figure 7 are Figure 6 the enlarged views of the two ROIs in Figure 7 Figures (a) to (h) of Figure 6 are the regions marked by the left dotted line in (h) of Figure 7 Figures (i) to (p) of Figure 6 are the regions marked by the right dotted line in (h) of Figure 7 The image order of Figures (i) to (p) is the same as that of Figure 7 Figures (a) to (h), and the images are all 150×150 pixels.

[0066] Because speckle noise in biological images is not completely independent and contains some static speckles, self-supervised learning of adjacent noisy images can retain static speckle noise. This results in residual noise in both the N2N(ori) and N2N denoising results trained in the original domain and the logarithmic transform domain. Figure 6 As shown in (d) and (e), the images denoised by N2N(ori) and N2N respectively achieve MS-SSIM scores of 0.9431 and 0.9507, although these scores do not surpass the traditional image processing algorithm BM3D. However, N2N's denoising performance exceeds that of N2N(ori), with the logarithmic domain-trained N2N(ori) achieving a 0.0076 improvement in MS-SSIM performance compared to the original domain-trained N2N(ori). By transforming the original noisy laser image into the logarithmic transform domain, the noise characteristics of different brightness regions in the image become similar, reducing the learning difficulty of the network. The network can process different brightness regions using similar denoising methods, thus better handling speckle noise of varying intensities. This improves network accuracy while enhancing the network's generalization ability and denoising robustness during inference. Figure 6 Examples (e) through (g) show the denoising results of networks trained in the logarithmic transform domain. These networks are more stable in denoising strong noise in high-brightness areas of the image, so the overall denoising results in the logarithmic transform domain are better than those in the original domain. In datasets with averaged sampling of 1, 3, 5, and 10 frames, the networks trained in the logarithmic domain also outperformed those trained in the original domain in all three evaluation metrics (N²N(ori)).

[0067] The TV regularization method includes gradients caused by the true structural edges and textures of the image in the loss function, resulting in some details of the denoised image being smoothed and blurred. Larger TV regularization coefficients lead to poorer test set evaluation results. Therefore, the optimal TV-N2N network strikes a trade-off between blurred image boundaries and residual speckle noise, resulting in the retention of some speckle noise in the denoising result. Figure 7 As shown in (e) and (m), the TV-N2N denoising results exhibit weakened speckle noise retention. This trade-off results in a higher MS-SSIM evaluation for TV-N2N.

[0068] Learning using AGR constraints, by selecting the direction of minimum gradient for loss calculation, avoids including the gradient of the real structure in the loss function, thus better preserving the structural information of the image. Sharpness and clarity are better at structural edges in the image, and fine structures are better preserved. Simultaneously, the removal of static speckle is also more effective. Figure 6 of (f), Figure 7 (f) and Figure 7In (n), the denoising result of AGR-N2N can remove static speckle noise better than TV-N2N, while preserving the image structure information without blurring it, and the image structure edges are still clear.

Claims

1. A speckle noise image processing device, characterized in that, The image processing device includes: A coherent signal source that generates a noisy image with speckle noise during imaging; Signal sensing elements are used to acquire noisy images at different times; The image processing unit includes a network structure subunit. The network structure subunit employs a self-supervised learning algorithm to remove speckle noise from the noisy image. The network parameters of the network structure subunit are trained based on a total loss function, which includes an anisotropic gradient regularization loss function and also the training loss of the network structure. The calculation method is shown in the following formula: (1), (2), in, For the total loss function, For training loss, Here are the anisotropic gradient regularization coefficients. This represents a noisy image acquired by a signal sensing element. This represents the noisy image at adjacent time points in the sampling sequence; H and W represent the number of pixels corresponding to the height and width of the image, respectively. This represents the trainable parameters in the network. Indicates the number of images in a mini-batch; The parameter is Network processing The output image after that; The total loss function based on anisotropic gradient regularization is determined based on the image values ​​of the noisy images at adjacent time points; The anisotropic gradient regularization loss function is determined in the following way: in, This indicates that the network output image contains residual noise that has not been fully repaired. This means expanding the matrix into a one-dimensional vector, where D is the anisotropic search radius; Operators that calculate anisotropic gradients in, `h` is the loop variable; `h` and `w` are the pixel coordinates of the image. This represents the image value at pixel (h, w) of the b-th frame image within the image batch after network processing. It is the pixel coordinate of the distance D from pixel (h, w). A set consisting of ).

2. The image processing apparatus according to claim 1, characterized in that, The training loss is determined based on a self-supervised learning algorithm, which includes Noise2Noise, Noise2Self, or Noise2Void.

3. The image processing apparatus according to claim 1, characterized in that, The network structure includes U-Net, DnCNN, DRSNet, or U-Former network structures.

4. The image processing apparatus according to claim 3, characterized in that, The U-Net network structure employs global residual learning and batch normalization. Batch normalization includes three steps: normalization, scaling, and translation, to standardize the input of the convolutional kernel.

5. The image processing apparatus according to any one of claims 1-3, characterized in that, The image processing unit further includes a network training set generation subunit, which selects noisy images at a certain time interval t from the noisy image sequence obtained from the signal sensing element during training for training, where 10ms < t < 100ms.

6. The image processing apparatus according to any one of claims 2-3, characterized in that, The image processing unit further includes a fast computation subunit, which is used to calculate anisotropic gradients during training through convolution operations.

7. The image processing apparatus according to claim 6, characterized in that, The fast calculation subunit is also used to fill the edges of the noisy image with D row constant values ​​or D column constant values, where D is the anisotropic search radius.

8. The image processing apparatus according to any one of claims 1-3, characterized in that, The coherent signal source is laser, ultrasound, or microwave.

9. The image processing apparatus according to any one of claims 1-3, characterized in that, The image processing unit further includes a data preprocessing subunit and a data postprocessing subunit. The data preprocessing subunit is used to perform a logarithmic transformation on the noisy image, and the transformed noisy image is used as the input image of the network structure. The data postprocessing subunit is used to perform an inverse exponential transformation on the output image of the network structure to obtain a denoised image.