A deep learning-based acoustic image enhancement method and device, and electronic equipment

By combining a deep learning-based conditional diffusion model and a dual-path GAN model with Rayleigh distribution and Sobel filter, the problem of noise interference in acoustic images is solved, achieving high-quality image enhancement and edge preservation, and improving image recognition performance.

CN120823112BActive Publication Date: 2025-12-05HANGZHOU ZHAOHUA ELECTRONICS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511257240.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-04
Publication Date
2025-12-05
Estimated Expiration
2045-09-04

AI Technical Summary

Technical Problem

Existing acoustic images are susceptible to speckle noise, environmental reverberation, and mechanical interference in applications such as sound source imaging, industrial gas leak imaging, partial discharge imaging, and UAV detection imaging, resulting in poor image quality, blurred edges, and loss of details. Traditional algorithms such as median filtering and histogram equalization can lead to blurred edges or over-enhancement, making it difficult to clearly identify anomalies.

Method used

A deep learning-based acoustic image enhancement method is adopted. By embedding a conditional diffusion model with a noise prior model and a dual-path GAN model, a Rayleigh-distributed noise scheduler and a SwingTransformer architecture are used, combined with a Sobel filter to achieve noise removal and edge enhancement. Multi-scale constraints and shared encoder connections between the generator and discriminator improve image quality.

Benefits of technology

It effectively removes noise, preserves image edge details, improves image quality, and makes anomalies easier to identify.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120823112B_ABST
    Figure CN120823112B_ABST
Patent Text Reader

Abstract

The embodiment of the specification discloses a kind of based on deep learning's acoustic image enhancement method, device and electronic equipment.The method includes obtaining original image, inputting original image to conditional diffusion model, obtain denoising image, noise prior model is embedded in conditional diffusion model, noise prior model is used to embed Rayleigh distribution as noise prior into conditional diffusion model;Denoising image is input to double-path GAN model, and enhanced acoustic image is obtained, double-path GAN model includes generator path for processing texture recovery and discriminator path for strengthening edge constraint.In the present embodiment, the enhanced recovery of acoustic image in high-noise environment can be realized by the cooperation processing of conditional diffusion model embedded with noise prior model and double-path GAN model, the edge details of image can be effectively removed while the noise influence is effectively removed, the quality of enhanced acoustic image is improved, and the anomaly in image is more easily identified.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present specification belong to the field of acoustic wave image processing, and particularly relate to an acoustic wave image enhancement method and device based on deep learning and electronic equipment. BACKGROUND

[0002] For acoustic imaging application scenarios such as sound source point imaging, industrial gas leakage imaging, partial discharge imaging, and unmanned aerial vehicle detection imaging, the obtained acoustic wave images are easily affected by speckle noise, environmental reverberation, mechanical interference, and other noise interference, and have image quality defects such as edge blurring and detail loss. In addition, traditional algorithms (such as median filtering) can cause edge blurring, and histogram equalization can easily cause over-enhancement. Therefore, the quality of existing acoustic wave images is poor, and it is difficult to clearly identify abnormalities in the images. SUMMARY

[0003] Embodiments of the present disclosure provide an acoustic wave image enhancement method and device based on deep learning and electronic equipment, aiming to solve one or more of the above problems and other potential problems.

[0004] According to a first aspect of the present disclosure, an acoustic wave image enhancement method based on deep learning is provided. The method includes obtaining an original image, inputting the original image into a conditional diffusion model to obtain a denoised image, the conditional diffusion model being embedded with a noise prior model, the noise prior model being used to embed a Rayleigh distribution as a noise prior into the conditional diffusion model, so that in a forward diffusion process of the conditional diffusion model, a noise scheduler based on the Rayleigh distribution replaces standard Gaussian noise, and in a reverse denoising process of the conditional diffusion model, image texture is preserved based on tail features of the Rayleigh distribution; inputting the denoised image into a dual-path GAN model to obtain an enhanced acoustic wave image, the dual-path GAN model including a generator path for processing texture recovery and a discriminator path for strengthening edge constraints, the generator being a SwinTransformer architecture based on image complexity adjustment window size, and the discriminator being embedded with a Sobel filter as an edge-specific layer on a multi-scale constraint.

[0005] According to a second aspect of the present disclosure, a deep learning-based acoustic image enhancement device is provided, the device comprising a first image processing module configured to obtain an original image, input the original image into a conditional diffusion model to obtain a denoised image, the conditional diffusion model embedded with a noise prior model, the noise prior model used to embed a Rayleigh distribution as a noise prior into the conditional diffusion model, so that in a forward diffusion process of the conditional diffusion model, a noise scheduler based on the Rayleigh distribution replaces standard Gaussian noise, and in a reverse denoising process of the conditional diffusion model, image texture is preserved based on tail features of the Rayleigh distribution; a second image processing module configured to input the denoised image into a dual-path GAN model to obtain an enhanced acoustic image, the dual-path GAN model comprising a generator path for processing texture recovery and a discriminator path for strengthening edge constraints, the generator being a Swin Transformer architecture with window size adjusted based on image complexity, and the discriminator embedding a Sobel filter as an edge-specific layer under multi-scale constraints.

[0006] According to a third aspect of the present disclosure, an electronic device is provided, comprising one or more processors, and a memory associated with the one or more processors, the memory being configured to store program instructions, the program instructions being configured to perform the method provided by the first aspect when read and executed by the one or more processors.

[0007] According to a fourth aspect of the present disclosure, a computer program product is provided, comprising a computer program configured to implement the method provided by the first aspect when executed by a processor.

[0008] The method provided by the embodiments of the present disclosure can realize enhancement and recovery of acoustic images in a high-noise environment through the cooperation of the conditional diffusion model embedded with the noise prior model and the dual-path GAN model, can effectively remove the influence of noise while better restoring the edge details of the image, improve the quality of the enhanced acoustic image, and make the abnormalities in the image easier to be identified. BRIEF DESCRIPTION OF DRAWINGS

[0009] The above and other features, advantages, and aspects of the present disclosure will become more apparent by describing in detail some embodiments thereof with reference to the attached drawings in which:

[0010] Figure 1 A flowchart of a deep learning-based acoustic image enhancement method of some embodiments of the present disclosure is shown;

[0011] Figure 2 A schematic diagram of an original image of some embodiments of the present disclosure is shown;

[0012] Figure 3A schematic diagram showing a denoised image of some embodiments of the present disclosure;

[0013] Figure 4 A schematic diagram showing a super-resolution reconstructed image of some embodiments of the present disclosure;

[0014] Figure 5 A schematic diagram showing a multi-scale feature fusion processed image of some embodiments of the present disclosure;

[0015] Figure 6 A schematic diagram showing a structure of a deep learning based acoustic image enhancement device of some embodiments of the present disclosure;

[0016] Figure 7 A schematic block diagram of an electronic device of some embodiments of the present disclosure is shown. DETAILED DESCRIPTION

[0017] For the purpose of making the objects, technical solutions and advantages of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the embodiments of the present application and corresponding drawings. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.

[0018] The terms “comprise” and “have” and any variations thereof in the specification and claims and above drawings are intended to cover not exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed or optionally further includes other steps or units inherent to these processes, methods, products or devices. Depending on the context, the word “if” as used herein can be interpreted as “when” or “upon” or “in response to determining” or “in response to detecting”.

[0019] Figure 1 A flowchart of a deep learning based acoustic image enhancement method 100 of some embodiments of the present disclosure is shown. The method 100 may, for example, be performed by a terminal, which can include but is not limited to a mobile phone, a tablet computer, a desktop computer, a server, etc. As shown in FIG. 1, the method 100 can include the following steps. Figure 1As shown in FIG. 1, at block 102, the method 100 can acquire an original image, input the original image into a conditional diffusion model, and obtain a denoised image, the conditional diffusion model being embedded with a noise prior model, the noise prior model being used to embed a Rayleigh distribution as a noise prior into the conditional diffusion model, so as to replace a standard Gaussian noise with a noise scheduler based on the Rayleigh distribution in a forward diffusion process of the conditional diffusion model, and to retain image texture based on tail features of the Rayleigh distribution in an inverse denoising process of the conditional diffusion model.

[0020] In this embodiment, the noise prior model may, for example, be a Gaussian noise model, a mixed Gaussian model, etc. Assuming that an input image x is contaminated by speckle noise, the noise prior model can be expressed as y = x n, where n is multiplicative noise, and n obeys a Rayleigh distribution. The probability density function thereof is:

[0021]

[0022] wherein, is a scale parameter representing noise intensity. By logarithmic transformation, the multiplicative noise is converted into additive noise, that is, logy = logx + logn, wherein logn approximately obeys a Gaussian distribution, but retains spatial correlation of Rayleigh characteristics.

[0023] After embedding the noise prior model, in the forward diffusion process (i.e., the process of adding noise) of the conditional diffusion model, a noise scheduler based on the Rayleigh distribution is used to replace the standard Gaussian noise, so as to first estimate the scale parameter of the noise (e.g., which can be calculated from local statistics of the image such as local mean and variance, etc. by maximum likelihood estimation), and then at each time step t of diffusion, a noise increment is sampled from the Rayleigh prior. In the inverse denoising process, noise residuals can be predicted, and tail features of the Rayleigh distribution are used to retain image texture and avoid excessive smoothing. The tail features of the Rayleigh distribution are manifested as thick-tailedness (i.e., the Rayleigh distribution decays slowly at a larger x value, the tail is more "thick", and the probability mass is higher at extreme values (such as high gradients and high frequencies)), and asymmetry (i.e., right skewness, which means that the tail of the Rayleigh distribution extends to the right). In this way, when predicting the noise residuals, higher probability is given to high gradient areas, avoiding excessive smoothing due to misjudgment of texture as noise, and by suppressing diffusion of noise at extreme values (i.e., reducing noise addition to high gradient areas), the denoising process is more inclined to retain details of the original data. As an example, assuming that the original image to be enhanced this time is as shown in FIG. 2A, Figure 2 After processing by the conditional diffusion model of the present application, a denoised image as shown in FIG. 2B can be obtained. Figure 3

[0024] Existing diffusion models (such as DDPM) usually assume a Gaussian noise prior, which is suitable for additive white noise but has limited effect on multiplicative speckle noise, which can lead to texture loss. The model of the present application can introduce a Rayleigh distribution as a dedicated prior, emphasizing the non-Gaussian and spatial correlation of the noise, which is different from the traditional Gaussian or Poisson distribution prior, and can better adapt to the statistical characteristics of the image.

[0025] At block 104, the method 100 can input the denoised image to a dual-path GAN model to obtain an enhanced acoustic image, the dual-path GAN model including a generator path for processing texture restoration and a discriminator path for strengthening edge constraints, the generator being a Swin Transformer architecture with window size adjusted based on image complexity, and the discriminator embedding a Sobel filter as an edge-specific layer under multi-scale constraints.

[0026] In the present embodiment, the denoised image is also processed using a dual-path GAN model, the generator of which is based on a Swin Transformer architecture, and the architecture is adjusted so that it can adaptively adjust the size of the shift window according to the complexity of the input image, enhancing the efficiency of long-distance dependency modeling. The Swin Transformer originally uses a fixed-size window, and in the present application, after extracting local features (such as edges, gradients, texture density, etc.) of the image, a gradient map of the image can be calculated by a Sobel operator, and regions with gradient amplitude higher than a threshold value are regarded as complex regions, so that the window size can be dynamically adjusted according to the complexity, for example, a smaller window (such as 4x4) can be used for complex regions to capture local details, and a larger window (such as 8x8 or 16x16) can be used for simple regions to improve computational efficiency.

[0027] The discriminator of the dual-path GAN model is a multi-scale structure, i.e., multiple branches process different resolution inputs (e.g., original image, 2x down-sampling, 4x down-sampling), and each branch can use a Patch GAN variant. In addition, the discriminator adds an edge-specific layer by embedding a Sobel filter under multi-scale constraints, so as to embed a Sobel edge map (i.e., a gradient map of the original image) in the input of the generator, so as to guide the generator to pay attention to the edge information of the image, and improve the edge gradient consistency between the generated image and the real image.

[0028] Through the dual-path setting, one path of the model processes texture restoration, and the other path strengthens edge constraints, separating texture and edge processing and avoiding the trade-off problem of a single path. In addition, compared with existing technologies (such as ESRGAN), through the setting of the generator and the discriminator, combined with the global attention of the Transformer and the adversarial nature of the GAN, the texture fidelity is improved. As an example, the reconstructed image obtained after the denoised image is processed by the generator path can be as followsFigure 4 As shown, after the reconstructed image is constrained by the discriminator, the generator performs multi-scale feature fusion, and the final enhanced acoustic image can be obtained as follows: Figure 5 As shown.

[0029] In one possible implementation, the diffusion process of the conditional diffusion model is divided into multiple layers, each layer corresponding to a different level of abstraction. In the forward diffusion process, the conditional diffusion model adds noise increments stepwise from the higher to the lower layers based on Rayleigh prior sampling at each time step. In the reverse denoising process, the conditional diffusion model predicts the noise residuals of each layer and reconstructs the global structure layer by layer from the higher layers based on the noise residuals.

[0030] In this embodiment, the conditional diffusion model can progressively reconstruct the image through a hierarchical structure to avoid detail loss caused by a single time step. Specifically, the diffusion process can be divided into multiple layers, each corresponding to a different level of abstraction. Higher layers process global structures (such as contours), while lower layers process local details (such as textures). The specific number of layers and the starting point for higher / lower layers can be set according to the actual situation. After layering, the input noisy image y is mapped to a hierarchical latent representation. , , ..., ,in For each layer, an encoder is used to generate initial latent variables. During forward diffusion, noise is gradually added starting from higher layers. Noise increment. At time step t (from 0 to T), sampling is based on Rayleigh priors and is subject to upper-level constraints:

[0031]

[0032] in, It is a noise scheduling parameter (which will gradually decrease). At time step t , At time step t-1 .

[0033] In the reverse denoising process, the noise prediction network of the model can be used to predict the noise residual at each level. ,in These are the upper-level conditions. The reconstruction process begins at the upper levels: ,in , It is the posterior variance. Reconstruction at different time steps can be achieved through inter-layer fusion, and after the global structure is reconstructed at a high level, it can be passed down to lower levels to refine details, ensuring that it can be gradually restored.

[0034] In an implementation, in the forward diffusion process of the conditional diffusion model, the time step length corresponding to the high layer is greater than the time step length corresponding to the low layer.

[0035] In this embodiment, in the forward diffusion process, the high layer uses a large-step coarse time step, for example , and the low layer uses a small-step fine time step, for example , so as to capture multi-scale information.

[0036] In an implementation, after the dual-path GAN model performs super-resolution reconstruction on the denoised image based on the generator path, the structure boundary of the reconstructed image is constrained based on the discriminator path to obtain an enhanced acoustic image; the generator path and the discriminator path are connected through a shared encoder.

[0037] In this embodiment, the dual-path can be connected through a shared encoder, so that the generator path is responsible for outputting the reconstructed image, and the discriminator path provides adversarial feedback. Specifically, after the denoised image is input into the dual-path GAN model, it will be first passed through the generator path as a low-resolution image to extract multi-scale features (which can capture local / global patterns by Swin layers) and upsample to generate a high-resolution reconstructed image. Then, the reconstructed image is passed into the discriminator path, which is downsampled by multiple scales to strengthen the image contour and structure boundary constraint, so as to ensure that the edges of the generated image are clear, and finally an enhanced acoustic image is obtained. In addition, the discriminator path can also calculate a discrimination score according to the similarity value (which can be determined by cosine similarity calculation, for example) between the real image and the generated image, and feed back the gradient to the generator to continuously optimize the model. The mapping relationship between different discrimination scores and similarity values can be pre-set according to experience.

[0038] In an implementation, the generator is obtained by stacking Swin Transformer blocks, and the generator includes a convolutional layer for shallow feature extraction, a Swin Transformer layer, and an upsampling module.

[0039] In this embodiment, a CNN pre-layer can be added before the Transformer to extract shallow local features such as edges and colors of the image through the convolutional layer of the CNN, provide more accurate local clues for the Transformer, reduce the local bias of the Transformer, make the texture recovery more natural, and avoid overfitting. The Swin Transformer layer can capture long-distance dependencies (such as the overall structure of an object, semantic association) in the image through the ShiftedWindow Attention (SW-MSA) and Window Attention (W-MSA) mechanisms. The upsampling module is used to enlarge the low-resolution feature map to the target image size.

[0040] In an implementation, the total constraint of the discriminator is a weighted sum of the multi-scale constraint and the boundary constraint, the multi-scale constraint is a sum of independent losses corresponding to each scale, and the boundary constraint is a minimum difference between the generated image and the real image on the edge gradient.

[0041] In this embodiment, the discriminator can operate on multiple scales (1x, 2x, 4x, etc.), and the independent loss of each scale is calculated as follows:

[0042]

[0043] where s is the scale, is the corresponding branch, real represents the real image, fake represents the generated image, and E is an edge extraction function.

[0044] In addition, after adding the edge consistency term, the generator is forced to match the boundary gradient of the real image by calculating the gradient difference, and the boundary constraint can be obtained as follows:

[0045]

[0046] where, represents the gradient information of the edge extraction of the image.

[0047] Finally, the total constraint of the adversarial mechanism can be a weighted sum of the multi-scale constraint and the boundary constraint:

[0048]

[0049] where, , is a weight, which can be set in advance according to actual conditions. Through this constraint, the edge can be optimized specially, and the blurring phenomenon can be reduced.

[0050] In an implementation, the loss function of the dual-path GAN model is determined based on a weighted sum of gradient loss of a gradient map, edge detection loss of an edge map, and perception loss of a feature map.

[0051] In this embodiment, the loss function of the dual-path GAN model can be represented as:

[0052]

[0053] where, , , is a weight hyperparameter, which can be set in advance according to actual conditions. is the gradient loss, is the edge detection loss, is the perception loss.

[0054] Gradient loss is used to calculate the gradient difference of the generated image and the real image , which emphasizes the consistency of edge intensity, and can be expressed as:

[0055]

[0056] Edge detection loss emphasizes the accuracy at the boundary position, and can be expressed as:

[0057]

[0058] where E is an edge extraction function.

[0059] Perception loss can be expressed as:

[0060]

[0061] where, is the i-th layer feature map, which captures high-level semantics.

[0062] After being combined, the gradient loss provides low-level edge supervision, the edge detection loss strengthens the position accuracy, and the perception loss ensures the overall visual consistency.

[0063] Figure 6 A structural schematic diagram of a deep learning-based acoustic image enhancement apparatus 600 of some embodiments of the present disclosure is shown. Each of the embodiments in the present specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other. Each embodiment focuses on the difference from other embodiments. In particular, for the apparatus embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the part of the method embodiment. As shown in Figure 6 The apparatus 600 includes a first image processing module 601 configured to obtain an original image, input the original image into a conditional diffusion model to obtain a denoised image, the conditional diffusion model being embedded with a noise prior model, the noise prior model being used to embed Rayleigh distribution as a noise prior into the conditional diffusion model, so that in a forward diffusion process of the conditional diffusion model, a noise scheduler based on Rayleigh distribution replaces standard Gaussian noise, and in a reverse denoising process of the conditional diffusion model, image texture is preserved based on the tail features of Rayleigh distribution; a second image processing module 602 configured to input the denoised image into a double-path GAN model to obtain an enhanced acoustic image, the double-path GAN model including a generator path for processing texture recovery and a discriminator path for strengthening edge constraints, the generator being a Swin Transformer architecture based on image complexity adjustment window size, and the discriminator being embedded with a Sobel filter as an edge-specific layer under multi-scale constraints.

[0064] In an implementation, the diffusion process of the conditional diffusion model is divided into multiple layers, each layer corresponding to a different level of abstraction; in the forward diffusion process, the conditional diffusion model adds noise increments based on the Rayleigh prior sampling step by step from high layer to low layer; in the reverse denoising process, the conditional diffusion model predicts the noise residual of each layer and reconstructs the global structure layer by layer from high layer based on the noise residual.

[0065] In an implementation, in the forward diffusion process of the conditional diffusion model, the time step length corresponding to the high layer is greater than the time step length corresponding to the low layer.

[0066] In an implementation, after the dual-path GAN model performs super-resolution reconstruction on the denoised image based on the generator path, the structure boundary of the reconstructed image is constrained based on the discriminator path to obtain an enhanced acoustic image; the generator path and the discriminator path are connected through a shared encoder.

[0067] In an implementation, the generator is obtained by stacking Swin Transformer blocks, and the generator includes a convolution layer for shallow feature extraction, a Swin Transformer layer, and an upsampling module.

[0068] In an implementation, the total constraint of the discriminator is a weighted sum of a multi-scale constraint and a boundary constraint, the multi-scale constraint is a sum of independent losses corresponding to each scale, and the boundary constraint is a minimum difference between the generated image and the real image in the edge gradient.

[0069] In an implementation, the loss function of the dual-path GAN model is determined based on a weighted sum of a gradient loss of a gradient map, an edge detection loss of an edge map, and a perception loss of a feature map.

[0070] In the above embodiments, all or part of the embodiments can be implemented by software, hardware, firmware or any combination thereof. When implemented by software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions according to the embodiments of the present specification are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in or transmitted by a computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center through a wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) manner. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server, data center, etc. integrated with one or more available media sets. The available media can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a digital versatile disc (DVD)), or a semiconductor medium (for example, a solid state disk (SSD)) and the like.

[0071] Figure 7 A block diagram of an electronic device 700 that can implement various embodiments of the present disclosure is shown. As shown, the electronic device 700 includes a processor 710, a disk drive 720, an input / output interface 730, a network interface 740, and a memory 750. The processor 710, the disk drive 720, the input / output interface 730, the network interface 740, and the memory 750 can be communicatively connected through a communication bus 760. Figure 7

[0072] The processor 710 can be implemented in the form of a general-purpose CPU, a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is configured to execute related programs to implement the technical solutions provided by the present application.

[0073] ​The memory 750 can be implemented in the form of a ROM (Read Only Memory), a RAM (Read Access Memory), a static memory, a dynamic memory device, etc. The memory 750 can store an operating system 751 for controlling the operation of the electronic device 700, a basic input / output system (BIOS) 752 for controlling the low-level operation of the electronic device 700. In addition, a web browser 753, a data storage management system 754, etc. can also be stored. In summary, when the technical solutions provided in the present application are implemented by software or firmware, the relevant program codes are stored in the memory 750 and executed by the processor 710.

[0074] The input / output interface 730 is configured to connect an input / output module to realize information input and output. The input / output module can be configured in the device as a component (not shown in the figure) or externally connected to the device to provide corresponding functions. The input device can include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device can include a display, a speaker, a vibrator, a prompt light, etc.

[0075] The network interface 740 is configured to connect a communication module (not shown in the figure) to realize the communication interaction between the device and other devices. The communication module can realize communication through a wired manner (such as USB, network cable, etc.) or through a wireless manner (such as mobile network, WIFI, Bluetooth, etc.).

[0076] The bus 760 includes a channel for transmitting information between various components (such as the processor 710, the disk drive 720, the input / output interface 730, the network interface 740, and the memory 750) of the device.

[0077] It should be noted that although the above device only shows the processor 710, the disk drive 720, the input / output interface 730, the network interface 740, the memory 750, the bus 760, etc., in the specific implementation process, the device can also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device can only contain the components necessary to implement the method of the present application, and does not necessarily contain all the components shown in the figure.

[0078] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0079] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing. Furthermore, although operations are depicted in a specific order, this should be understood as requiring that such operations be performed in the specific order shown or in sequential order, or requiring that all illustrated operations be performed to achieve the desired result. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the foregoing discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single implementation. Conversely, various features described in the context of a single implementation may also be implemented individually or in any suitable sub-combination in multiple implementations.

[0080] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.

Claims

1. A deep learning-based acoustic image enhancement method, characterized by, The method comprises: obtaining an original image, inputting the original image into a conditional diffusion model to obtain a denoised image, the conditional diffusion model being embedded with a noise prior model, the noise prior model being used to embed Rayleigh distribution as noise prior into the conditional diffusion model, so that in a forward diffusion process of the conditional diffusion model, a noise scheduler based on Rayleigh distribution replaces standard Gaussian noise, and in a reverse denoising process of the conditional diffusion model, image texture is retained based on tail features of Rayleigh distribution; inputting the denoised image into a dual-path GAN model to obtain an enhanced acoustic wave image, the dual-path GAN model comprising a generator path for processing texture recovery and a discriminator path for strengthening edge constraint, the generator being a Swin Transformer architecture based on image complexity adjustment window size, and the discriminator being embedded with a Sobel filter as an edge dedicated layer under multi-scale constraint.

2. The method of claim 1, wherein, The diffusion process of the conditional diffusion model is divided into multiple layers, each layer corresponding to a different level of abstraction; in the forward diffusion process of the conditional diffusion model, noise increments are gradually added based on Rayleigh prior sampling at time steps from high layers to low layers; in the reverse denoising process of the conditional diffusion model, noise residuals of each layer are predicted, and global structures are reconstructed layer by layer starting from high layers based on the noise residuals. 3.The method of claim 2, wherein, In the forward diffusion process of the conditional diffusion model, the time step length corresponding to the high layer is greater than the time step length corresponding to the low layer. 4.The method of claim 1, wherein, After the dual-path GAN model performs super-resolution reconstruction on the denoised image based on the generator path, the structure boundary of the reconstructed image is constrained based on the discriminator path to obtain the enhanced acoustic wave image; the generator path and the discriminator path are connected through a shared encoder.

5. The method of claim 1, wherein the method is based on deep learning. The generator is obtained by stacking Swin Transformer blocks, and the generator comprises a convolution layer, a Swin Transformer layer and an upsampling module for shallow feature extraction. 6.The method of claim 1, wherein, The total constraint of the discriminator is a weighted sum of multi-scale constraint and boundary constraint, the multi-scale constraint is a sum of independent losses corresponding to each scale, and the boundary constraint is a minimum difference between the edge gradient of the generated image and the real image.

7. The method of claim 1, wherein the method is based on deep learning. The loss function of the dual-path GAN model is determined based on a weighted sum of gradient loss of a gradient map, edge detection loss of an edge map and perception loss of a feature map. 8.A deep learning-based acoustic image enhancement device, characterized by The device comprises: a first image processing module configured to obtain an original image, input the original image into a conditional diffusion model to obtain a denoised image, the conditional diffusion model being embedded with a noise prior model, the noise prior model being used to embed Rayleigh distribution as noise prior into the conditional diffusion model, so that in a forward diffusion process of the conditional diffusion model, a noise scheduler based on Rayleigh distribution replaces standard Gaussian noise, and in a reverse denoising process of the conditional diffusion model, image texture is retained based on tail features of Rayleigh distribution; The second image processing module is configured to input the de-noised image into a dual-path GAN model to obtain an enhanced acoustic image, the dual-path GAN model comprising a generator path for processing texture restoration and a discriminator path for strengthening edge constraint, the generator being a Swin Transformer architecture with window size adjusted based on image complexity, and the discriminator embedding a Sobel filter as an edge-specific layer under multi-scale constraint. 9.An electronic device comprising: one or more processors, and a memory associated with the one or more processors, the memory for storing program instructions that, when read and executed by the one or more processors, perform the steps of the deep learning based acoustic image enhancement method of any one of claims 1-7. 10.A computer program product comprising a computer program that, when executed by a processor, implements the deep learning based acoustic image enhancement method of any one of claims 1-7.

Citation Information

Patent Citations

  • Ambient noise source adaptive positioning method and device based on beam forming

    CN114863943A

  • X-ray image noise reduction method based on conditional diffusion model

    CN117522723A