Kidney ultrasonic image denoising method and device
By adopting a dual-branch deep neural network architecture in ultrasonic image denoising, processing high-frequency and low-frequency signals of the image respectively, and combining the attention mechanism, the problems of poor denoising effect and insufficient real-time performance in the existing technology are solved, and more efficient image denoising and real-time processing are achieved.
Patent Information
- Application Number
- CN202510095957.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-06-06
AI Technical Summary
In the prior art, when denoising ultrasonic images, it is difficult to effectively retain image structure information, resulting in poor denoising effect. The neural network-based methods lack lightweight design, making it difficult to meet real-time requirements.
A kidney ultrasound image denoising method based on a dual branch deep neural network is designed. By dividing the ultrasound image into high-frequency signals and low-frequency signals, it is used to process information fusion, combining the detail channel attention module and the structural non-local attention module, the information interaction and parameter quantity of the neural network are optimized.
It significantly improves the denoising effect of ultrasonic images, enhances the retention of image structure information, and improves the real-time performance of denoising processing, reducing the complexity and computing cost of neural networks.
Smart Images

Figure CN120107097A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing and analysis, and more specifically to a method and device for renal ultrasound image denoising based on deep learning. Background Art
[0002] As a non-invasive, real-time, and low-cost auxiliary examination tool, ultrasound images have an irreplaceable position in clinical diagnosis. The real-time dynamic image display function of ultrasound is very popular among clinical workers, so it is used as a guiding tool for various clinical operations, making clinical organ and tumor biopsies, as well as vascular punctures, etc. visualized. Among them, pathological examination of renal tissue through puncture biopsy is currently an important means of diagnosing glomerular and tubulointerstitial diseases, and it is also a powerful measure to guide treatment and judge prognosis. Because the kidney is a retroperitoneal organ, patients are often placed in the prone position for puncture. Doctors need to use an ultrasound probe to display the kidney from the back of the patient. Ultrasound passes through multiple layers of tissue such as the skin, subcutaneous tissue, and psoas muscle to reach the kidney. By displaying the morphology and clear contours of the kidney, the renal cortex and renal medulla are distinguished, and the path of puncturing the renal cortex at the lower pole of the kidney is accurately located, guiding the doctor to complete the kidney puncture. However, currently, ultrasound images usually have a lot of noise and lack clarity, which may increase the risk of kidney puncture and lead to unsatisfactory sampling, which is mainly reflected in the following two situations: First, for patients with obese body and thick muscle fibers, the ultrasound image is prone to unclear kidney structure and contour; Second, in order to meet the tissue sufficiency requirements of pathological diagnosis, we often need to perform a second puncture on the kidney to ensure sufficient tissue for pathological diagnosis. However, once a hematoma appears around the lower pole of the kidney after the first puncture, it often leads to unclear contour of the lower pole of the kidney. In this case, a second puncture will significantly increase the risk of bleeding, unsatisfactory sampling, and damage to surrounding organs. Therefore, developing an effective denoising algorithm to improve the clarity of ultrasound images can significantly improve the success rate of puncture and reduce puncture-related risks.
[0003] The causes of noise in ultrasound medical images are complex and of many types, including speckle noise, thermal noise, electronic noise, motion noise, etc. These noises are intertwined, causing the image quality to be affected to varying degrees. Speckle noise is a very important type of noise. Its cause is that when ultrasound propagates in biological tissues, the reflection and scattering characteristics of different tissues are different, resulting in the coherent superposition of echo signals, forming randomly distributed granular or speckled noise. This type of noise has certain locality and coherence, which seriously affects the clarity and detail visualization of the image. Locality means that speckle noise is usually particularly obvious in local areas of the image. Coherence means that speckle noise of adjacent pixels in a local area often has a strong correlation, that is, the noise values in a region tend to be similar.
[0004] The development of ultrasound medical image denoising algorithms has gone from simple spatial domain filtering to frequency domain filtering, and then to advanced algorithms based on statistical models and machine learning. Early denoising techniques mainly focused on spatial domain filtering, such as average filtering, median filtering, etc. Although these methods are simple to implement, they often sacrifice image detail information and lead to blurred edges. With the advancement of technology, frequency domain filtering methods have begun to be widely studied, such as using low-pass filters to reduce high-frequency noise. However, frequency domain filtering also has some problems, such as improper filter design may cause image distortion.
[0005] With the development of signal processing technology, model-based denoising methods have gradually attracted attention. These methods usually assume that the image or noise follows a certain statistical model, such as Gaussian mixture model (GMM), non-local means (NLM), etc. By estimating the model parameters, noise can be effectively removed while maintaining image details. For example, the non-local means algorithm achieves denoising by finding similar pixel blocks in the image. This method is better than simple spatial filtering and frequency domain filtering in preserving image texture and details.
[0006] In recent years, with the rise of machine learning and deep learning technologies, learning-based methods have made significant progress in the field of ultrasound image denoising. These methods learn the characteristics of noise and prior knowledge of images from a large amount of training data to achieve more intelligent denoising. Convolutional neural network (CNN) is one of the most popular deep learning architectures, which automatically extracts image features through multi-layer convolution and pooling operations and learns denoising mapping through training. In addition to CNN, other deep learning architectures, such as recurrent neural network (RNN) and generative adversarial network (GAN), have also been used for ultrasound image denoising. In addition to deep learning, other machine learning methods such as support vector machine (SVM) and random forest have also been used for ultrasound image denoising. These methods usually require less parameter adjustment, but their performance is not as good as deep learning methods when dealing with complex noise.
[0007] Although the existing denoising algorithms have made significant progress in some aspects, there are still some challenges and limitations. The existing deep learning-based ultrasound medical image denoising algorithms have not been carefully designed for the characteristics of ultrasound noise images, resulting in poor retention of structural information in the image and poor denoising effect. On the other hand, the current neural network-based ultrasound image denoising methods rarely consider lightweight design, which makes it difficult to meet the real-time requirements of renal puncture for ultrasound image denoising algorithms. Summary of the invention
[0008] The present invention aims to solve at least one of the technical problems existing in the prior art. To this end, the present invention provides a method for denoising renal ultrasound images based on a branch deep neural network, designs an innovative dual-branch neural network architecture, divides the ultrasound renal image into high-frequency signals and low-frequency signals for separate processing, and then fuses the processing results, which can better remove noise in the ultrasound image and improve the real-time performance of renal ultrasound image denoising.
[0009] To achieve the above object, a kidney ultrasound image denoising method of the present invention comprises the following steps:
[0010] Step 1, prepare a kidney ultrasound image dataset;
[0011] Step 2, constructing a kidney ultrasound denoising model, the denoising model includes a high-frequency branch, a low-frequency branch and a fusion module, the high-frequency branch is used to process the high-frequency part of the noisy ultrasound image, that is, the detail information of the image, the low-frequency branch is used to process the low-frequency part of the noisy ultrasound image, that is, the structural information of the image, and the fusion module is used to fuse the information of the high-frequency branch and the low-frequency branch to obtain a denoised ultrasound image;
[0012] Step 3, training the denoising model using a small batch stochastic gradient descent method;
[0013] Step 4, using the trained denoising model to perform denoising on the renal ultrasound image. The preparation of the renal ultrasound comprises the following steps:
[0014] Step 101, scanning and imaging the kidney ultrasound phantom by high-quality ultrasound equipment, collecting n original ultrasound images {x1, x2, ... xn}, then generating speckle noise, and superimposing the speckle noise on the original ultrasound image to form n noisy ultrasound images {y1, y2, ... yn};
[0015] Step 102: Combine n original ultrasound images with n noisy ultrasound images, that is, pair the original ultrasound image x1 with the image y1 after superimposing noise into (x1, y1), and pair the original ultrasound image x2 with the image y2 after superimposing noise into (x2, y2), and so on, to form n pairs of ultrasound images {(x1, y1), (x2, y2) ... (xn, yn)}, which are recorded as data set D;
[0016] Step 103: Divide the data set D into a training set D1 and a test set D2.
[0017] Specifically, in the high-frequency branch, the high-frequency component of the image is first converted into a feature map with m channels through the MCONV module, where m=g×k, g is the number of groups, and k is the number of channels in each group. Then, a multi-level detail channel attention module is used for cascade processing. The detail channel attention module includes a separation convolution module S1 and a channel attention module CA. The input and output of the separation convolution module S1 are both feature maps with a size of h×w×m, and the input and output of the channel attention module CA are both feature maps with a size of h×w×m, where h is the height of the feature map and w is the width of the feature map. In the high-frequency branch, the structure of the separation convolution module S1 includes a 3×3 depth separation convolution, a Relu activation function, a 1×1 point convolution, and a residual connection in parallel with the above three. In addition, there is a point addition operator O1 for performing point-by-point addition operation on the output of the 1×1 point convolution and the original input connected by the residual connection.
[0018] The channel attention module CA adopts a lightweight attention mechanism to adjust the weights of each channel of the feature map. The channel attention module CA includes a global average pooling GAP, a grouped cross-domain fully connected layer GFC, and a point multiplication operator O2. The global average pooling GAP calculates the average value of each channel in the feature map to obtain a mean. After calculating the m channels, a vector of length m is obtained. The vector is divided into g groups, each group has a length of k. The grouped cross-domain fully connected layer GFC performs full connection in a grouped manner, that is, constructs a k×k fully connected matrix. There are g such matrices in total. In order to ensure information interaction between groups, each group will introduce an element from the adjacent group as a cross-domain element. Therefore, g (k+1)×k fully connected matrices are constructed;
[0019] Each group introduces a cross-domain element from an adjacent group. First, an adjacent group needs to be selected. If the current group has only one adjacent group, the adjacent group is directly selected. If the current group has two adjacent groups, the correlation between the two groups and the current group is first compared. The correlation calculation process is as follows:
[0020]
[0021] Among them, a i Represents the i-th element of the current group, Indicates the mean of all elements in the current group, b i represents the i-th element of the adjacent group, Represents the mean of all elements in the current group;
[0022] Then, let the correlation degree of the first adjacent group be corr1, and the correlation degree of the second adjacent group be corr2, and the probability of selecting the first adjacent group is: The probability of selecting the second adjacent group is:
[0023] After the adjacent groups are selected according to the probabilities p1 and p2, an element is randomly selected from the k elements of the selected adjacent groups as the cross-domain element.
[0024] The step of randomly selecting an element from the k elements of the selected adjacent groups as the cross-domain element comprises the following steps:
[0025] The probability of randomly selecting an element from the k elements of the selected adjacent group follows the following formula: Among them, z i represents the i-th element in the adjacent group, S i The distance between the i-th element in the adjacent group and the current group is measured, and the calculation method is as follows: v i Represents the i-th element in the current group.
[0026] Preferably, the step of randomly selecting an element from the k elements of the selected adjacent groups as the cross-domain element comprises the following steps:
[0027] The probability of randomly selecting an element from the k elements of the selected adjacent group follows the following formula: Among them, z i represents the i-th element in the adjacent group, H i The calculation method is as follows: i =S i ×N i ,
[0028] Among them, S i The distance between the i-th element in the adjacent group and the current group is measured, and the calculation method is as follows: v i Represents the i-th element in the current group, N i Represents element z i Specificity in adjacent groups, specificity N iThe calculation method is as follows:
[0029]
[0030] N i represents the specificity between element z i and other elements in adjacent groups. The larger N i is, the greater the specificity of z i within its own group. Multiplying S i and H i not only increases the information interaction between groups but also improves the neural network's attention and processing ability to noisy information.
[0031] Si measures the distance between the i-th element in the adjacent group and the current group. Selecting elements with probability p(z i ) makes it easier to select elements with a greater distance from the current group. A greater distance represents a greater difference, which further promotes the information interaction between groups.
[0032] Furthermore, in the low-frequency branch, first, the LCONV module converts the low-frequency component into a feature map with c channels, where c < m, and then cascaded processing is performed through a multi-level structured non-local attention module. The structured non-local attention module includes two parts: a separable convolution module S2 and a non-local module NL. The input and output of the separable convolution module S2 are both feature maps of size h × w × c, and the input and output of the non-local module NL are also feature maps of size h × w × c;
[0033] The structure of the separable convolution module S2 is the same as that of the separable convolution module S1, except for the difference in the number of channels of the output feature map. The number of channels of the output feature map of the separable convolution module S2 is c, and the number of channels of the output feature map of the separable convolution module S1 is m;
[0034] The non-local module NL adjusts the weight of the image on the two-dimensional plane by paying attention, so that the structural features of the image are more prominent. The input of the non-local module NL is a feature map of size h×w×c, denoted as a three-dimensional matrix F. First, the three-dimensional matrix F is transformed into three two-dimensional matrices B1, B2 and B3, respectively. Among them, the size of matrix B1 is hw×c, the size of matrix B2 is c×hw, and the size of matrix B3 is hw×c. B1 and B2 are matrix multiplied to obtain a two-dimensional matrix B4 with a size of hw×hw. Then, a softmax operation is performed on all elements of B4 to make the sum of all elements equal to 1. The matrix after softmax calculation is denoted as B5. B5 and B3 are matrix multiplied to obtain a matrix B6 of size hw×c. B6 is operated once with 1×1 convolution to obtain a three-dimensional matrix B7 of size h×w×c. B7 is the output feature map of the non-local module NL.
[0035] The fusion module includes a splicing module CAT and a convolution RCONV connected in sequence, wherein the splicing module CAT splices the feature maps output by the high-frequency branch and the low-frequency branch on the channel, and the convolution RCONV processes and transforms the spliced feature maps to obtain a denoised ultrasound image.
[0036] Specifically, the training of the denoising model comprises the following steps:
[0037] Step 301: Set the hyperparameters for denoising model training, including sample batch size b, learning rate lr, and number of training rounds e.
[0038] Step 302: Take any b sample pairs from the training set D1, and use the noisy ultrasound images in the sample pairs as inputs of the denoising model.
[0039] Step 303: The denoising model performs forward inference on the noisy ultrasound image to obtain b predicted denoised images.
[0040] Step 304: Calculate the mean square error between the predicted denoised image and the original ultrasound image, use the mean square error as a loss function, and use the back propagation algorithm to adjust the parameters of the neural network in the noise model.
[0041] Step 305: Repeat steps 302 to 304 for a total of e times.
[0042] A kidney ultrasound image denoising device, comprising:
[0043] processor;
[0044] and, a memory for storing executable instructions of the processor;
[0045] The processor is configured to implement a method for denoising a renal ultrasound image by executing the aforementioned executable instructions.
[0046] The beneficial effects of the method of the present invention are as follows: First, according to the characteristics that ultrasonic kidney images have complex and diverse noises and texture structures are very important for diagnosis, an innovative dual-branch neural network architecture is designed to divide ultrasonic kidney images into high-frequency signals and low-frequency signals for separate processing, and then the processing results are fused. Second, according to the characteristics that high-frequency signals are dominated by noise and detail information and less texture structure information, the high-frequency branch designs a detail channel attention module DCAM module, which uses channel attention to filter out noise and retain detail information; on the other hand, in order to lightweight the neural network model, a grouped cross-domain fully connected layer GFC is designed, which not only ensures sufficient information interaction, but also reduces the number of parameters of the neural network. Third, according to the characteristics that low-frequency signals are dominated by texture structure information and less noise and detail information, the low-frequency branch designs a structural non-local attention module SFAM module, which enhances contextual texture information through a global attention mechanism; in addition, the non-local module NL in the SFAM module utilizes the global information of the entire image, which is conducive to suppressing speckle noise with local coherence. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 A schematic diagram of a process flow of an embodiment of the present invention is shown;
[0048] Figure 2 The overall architecture diagram of the denoising model according to the embodiment of the present invention is shown;
[0049] Figure 3 A schematic diagram of a DCAM module according to an embodiment of the present invention is shown;
[0050] Figure 4 A schematic diagram of a separation convolution module according to an embodiment of the present invention is shown;
[0051] Figure 5 A schematic diagram of a channel attention module CA according to an embodiment of the present invention is shown;
[0052] Figure 6 A schematic diagram of a non-local module NL according to an embodiment of the present invention is shown. DETAILED DESCRIPTION
[0053] In order to make the purpose, technical solution and advantages of the present invention clearer, the present invention will be further described in detail below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0054] It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.
[0055] Embodiment 1
[0056] To achieve the above purpose, Figure 1 As shown, a kidney ultrasound image denoising method comprises the following steps:
[0057] Step 1: Prepare ultrasound image dataset
[0058] Step 1.1: Scan and image the renal ultrasound phantom using high-quality ultrasound equipment to collect n original ultrasound medical images {x1, x2, ... xn}, then use MATLAB software to generate speckle noise, and superimpose the speckle noise on the original ultrasound medical images to form n noisy ultrasound images {y1, y2, ... yn}.
[0059] Step 1.2: Combine n original ultrasound images with n noisy ultrasound images, that is, pair the original ultrasound image x1 with the image y1 after superimposing the noise into (x1, y1), and pair the original ultrasound image x2 with the image y2 after superimposing the noise into (x2, y2), and so on, to form n pairs of ultrasound images {(x1, y1), (x2, y2)...(xn, yn)} (referred to as data set D).
[0060] Step 1.3: Use 70% of the data set D as the training set (denoted as D1) and 30% as the test set (denoted as D2).
[0061] Step 2: Construction of denoising model
[0062] like Figure 2 As shown in the figure, the denoising model uses a deep neural network and is divided into two branches (high-frequency branch and low-frequency branch). The high-frequency branch is used to process the high-frequency part of the noisy ultrasound image, that is, the image detail information, and the low-frequency branch is used to process the low-frequency part of the noisy image, that is, the image structure information. The high-frequency branch and the low-frequency branch are fused through the fusion module to obtain the final denoised ultrasound image.
[0063] In the high-frequency branch, the high-frequency component is first converted into a feature map with m channels through the MCONV module, where m = g × k (g is the number of groups, k is the number of channels in each group), and then a multi-level DCAM module (detail channel attention module) is used for cascade processing. Figure 3 As shown, the DCAM module includes two parts: a separation convolution module S1 and a channel attention module CA. The input and output of the separation convolution module S1 are both feature maps with a size of h×w×m, and the input and output of the channel attention module CA are both feature maps with a size of h×w×m.
[0064] like Figure 4 As shown in the figure, the structure of the separation convolution module S1 includes a 3×3 depth separation convolution, a Relu activation function, a 1×1 point convolution, and a residual connection in parallel with the above three. In addition, there is a point addition operator O1, which is used to perform point-by-point addition operations on the output of the 1×1 point convolution and the original input connected by the residual connection.
[0065] like Figure 5 As shown, the channel attention module CA adopts a lightweight attention mechanism to adjust the weights of each channel of the feature map so that important channel information can get more attention from the network model. The channel attention module CA includes global average pooling GAP, group cross-domain fully connected layer GFC, and point multiplication operator O2. GAP calculates the average value of each channel in the feature map to obtain a mean value. After calculating all m channels, a vector of length m is obtained. The vector is divided into g groups, each with a length of k. GFC performs full connection in a grouped manner, that is, constructs a k×k fully connected matrix, and there are g such matrices in total. However, in order to ensure information interaction between groups, each group will introduce an element from the adjacent group as a cross-domain element. Therefore, g (k+1)×k fully connected matrices will actually be constructed. In this embodiment: m=8, g=3, k=4. Group 1 and group 3 have only one adjacent group, and group 2 has two adjacent groups (group 1 and group 3, respectively).
[0066] In order to introduce an element from an adjacent group as a cross-domain element, each group first needs to select an adjacent group. If the current group has only one adjacent group, the adjacent group is directly selected. Figure 5 As shown in , group 1 has only one adjacent group (i.e., group 2), and group 3 also has only one adjacent group (i.e., group 2), so group 1 and group 3 directly select group 2 when selecting adjacent groups. Figure 5 Group 2 in has group 1 and group 3 as its adjacent groups), then first compare the correlation between the two groups and the current group. The correlation calculation process is as follows:
[0067]
[0068] Among them, a i Represents the i-th element of the current group, Indicates the mean of all elements in the current group, b i represents the i-th element of the adjacent group, Represents the mean of all elements in the current group;
[0069] Then, let the correlation degree of the first adjacent group be corr1, and the correlation degree of the second adjacent group be corr2, then the probability of selecting the first adjacent group is: The probability of selecting the second adjacent group is:
[0070] After the adjacent groups are selected according to the probabilities p1 and p2, an element is randomly selected from the k elements of the selected adjacent groups as the cross-domain element.
[0071] Randomly select an element from the k elements of the selected adjacent group, using one of the following two methods. The first method:
[0072] The probability of randomly selecting an element from the k elements of the selected adjacent group follows the following formula: Among them, z i represents the i-th element in the adjacent group, S i The distance between the i-th element in the adjacent group and the current group is measured, and the calculation method is as follows: v i Represents the i-th element in the current group;
[0073] Second method:
[0074] The probability of randomly selecting an element from the k elements of the selected adjacent group follows the following formula:
[0075] Among them, z i represents the i-th element in the adjacent group, H i The calculation method is as follows: i =S i ×N i , where (do I need to explain Si again here? Or we can use the expression: the meaning and calculation method of Si are the same as those in the first method), N i Represents element z i Specificity in adjacent groups, specificity N i The calculation method is as follows:
[0076]
[0077] N i Represents the element z i The specificity between other elements in the adjacent group, N i The larger the z i Has greater specificity within its own group, v i Represents the i-th element in the current group, and S i and H iMultiplication not only increases the information interaction between groups but also improves the neural network's attention and processing ability for noisy information.
[0078] Si measures the distance between the i-th element in adjacent groups and the current group. Selecting elements with probability p(z i ) makes it easier to select elements with a greater distance from the current group. A greater distance represents a greater difference, which further promotes information interaction between groups.
[0079] In the low-frequency branch, the low-frequency components are first converted into a feature map with c channels by the LCONV module, where c < m, and then cascaded processing is performed through multiple levels of the SFAM module (structured non-local attention module). The SFAM module consists of two parts: the separable convolution module S2 and the non-local module NL. The input and output of the separable convolution module S2 are both feature maps of size h×w×c, and the input and output of the non-local module NL are also feature maps of size h×w×c.
[0080] The structure of the separable convolution module S2 is basically the same as that of the separable convolution module S1, except for the difference in the number of channels of the output feature map. The number of channels of the output feature map of S2 is c, while the number of channels of the output feature map of S1 is m.
[0081] As Figure 6 shown, the non-local module NL performs a series of attention calculations, adjusting the weights of the image in the two-dimensional plane through attention, making the structural features of the image more prominent. The input of the non-local module NL is a feature map of size h×w×c, which is actually a three-dimensional matrix (denoted as F). First, the three-dimensional matrix F is dimensionally transformed into 3 two-dimensional matrices. Among them, the size of matrix B1 is hw×c, the size of matrix B2 is c×hw, and the size of matrix B3 is hw×c. B1 and B2 are multiplied matrix-wise to obtain a two-dimensional matrix B4, and the size of B4 is hw×hw. Then, a softmax operation is performed on all elements of B4 to make the sum of all elements equal to 1. The matrix after softmax calculation is denoted as B5. B5 and B3 are multiplied matrix-wise to obtain a matrix B6 of size hw×c. A 1×1 convolution is performed on B6 once to obtain a three-dimensional matrix B7 of size h×w×c, and B7 is the output feature map of the non-local module NL.
[0082] The fusion module includes a concatenation module CAT and a convolution RCONV connected in sequence. The concatenation module CAT concatenates the feature maps output by the high-frequency branch and the low-frequency branch on the channels, and the convolution RCONV processes and transforms the concatenated feature maps to obtain the denoised ultrasonic image.
[0083] Step 3: Train the denoising model using mini-batch stochastic gradient descent (SGD)
[0084] Step 3.1: Set the hyperparameters for denoising model training, including sample batch size b, learning rate lr, and number of training rounds e.
[0085] Step 3.2: Take any b sample pairs from the training set D1 (denoted as {(x i1 ,y i1 ),(x i2 ,y i2 ),...,(x ib ,y ib ),}), the noisy ultrasound image {y i1 ,y i2 ,...,y ib} as the input of the denoising model.
[0086] Step 3.3: The denoising model performs forward inference on the noisy ultrasound image to obtain b predicted images.
[0087] Step 3.4: Calculation With {x i1 ,x i2 ,...,x ib The mean square error between} is used as the loss function, and the back propagation algorithm is used to adjust the parameters of the neural network in the noise model.
[0088] Step 3.5: Repeat steps 3.2 to 3.4 for a total of e times.
[0089] Embodiment 2
[0090] A kidney ultrasound image denoising device, comprising:
[0091] processor;
[0092] and, a memory for storing executable instructions of the processor;
[0093] The processor is configured to execute the kidney ultrasound image denoising method in the first embodiment by executing the executable instructions.
[0094] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
Claims
1. A method for denoising a kidney ultrasound image, characterized in that: The following steps are involved: Step 1, prepare a kidney ultrasound image dataset; Step 2, constructing a kidney ultrasound denoising model, the denoising model includes a high-frequency branch, a low-frequency branch and a fusion module, the high-frequency branch is used to process the high-frequency part of the noisy ultrasound image, that is, the detail information of the image, the low-frequency branch is used to process the low-frequency part of the noisy ultrasound image, that is, the structural information of the image, and the fusion module is used to fuse the information of the high-frequency branch and the low-frequency branch to obtain a denoised ultrasound image; Step 3, training the denoising model using a small batch stochastic gradient descent method; Step 4: Use the trained denoising model to denoise the kidney ultrasound image.
2. A kidney ultrasound image denoising method according to claim 1, characterized in that: The preparation for renal ultrasound includes the following steps: Step 101, scanning and imaging the kidney ultrasound phantom by high-quality ultrasound equipment, collecting n original ultrasound images {x1, x2, ... xn}, then generating speckle noise, and superimposing the speckle noise on the original ultrasound image to form n noisy ultrasound images {y1, y2, ... yn}; Step 102: Combine n original ultrasound images with n noisy ultrasound images, that is, pair the original ultrasound image x1 with the image y1 after superimposing noise into (x1, y1), and pair the original ultrasound image x2 with the image y2 after superimposing noise into (x2, y2), and so on, to form n pairs of ultrasound images {(x1, y1), (x2, y2) ... (xn, yn)}, which are recorded as data set D; Step 103: Divide the data set D into a training set D1 and a test set D2.
3. A kidney ultrasound image denoising method according to claim 1, characterized in that: In the high-frequency branch, the high-frequency component of the image is first converted into a feature map with m channels through the MCONV module, where m=g×k, g is the number of groups, and k is the number of channels in each group. Then, a multi-level detail channel attention module is used for cascade processing. The detail channel attention module includes two parts: a separation convolution module S1 and a channel attention module CA. The input and output of the separation convolution module S1 are both feature maps with a size of h×w×m, and the input and output of the channel attention module CA are both feature maps with a size of h×w×m, where h is the height of the feature map and w is the width of the feature map.
4. A kidney ultrasound image denoising method according to claim 3, characterized in that: In the high-frequency branch, the structure of the separation convolution module S1 includes a 3×3 depth separation convolution, a Relu activation function, a 1×1 point convolution, and a residual connection connected in parallel with the above three. In addition, there is a point addition operator O1 for performing point-by-point addition operation on the output of the 1×1 point convolution and the original input connected by the residual connection; The described channel attention module CA adopts a lightweight attention mechanism to adjust the weights of each channel of the feature map. The channel attention module CA includes global average pooling GAP, grouped cross-domain fully connected layer GFC, and dot product operator O2. The global average pooling GAP calculates the average value for each channel in the feature map to obtain a mean value. After calculating for all m channels, a vector of length m is obtained. The vector is divided into g groups, each group with a length of k. The grouped cross-domain fully connected layer GFC performs full connection in a grouped manner, that is, constructs a k×k fully connected matrix, and there are g such matrices in total. To ensure information interaction between groups, each group introduces an element from an adjacent group as a cross-domain element. Therefore, g (k + 1)×k fully connected matrices are constructed; For each of the described groups to introduce a cross-domain element from an adjacent group, first, an adjacent group needs to be selected. If the current group has only one adjacent group, then directly select that adjacent group. If the current group has two adjacent groups, first compare the correlation degrees of the two adjacent groups with the current group. The calculation process of the correlation degree is as follows: Among them, a i Represents the i-th element of the current group, Indicates the mean of all elements in the current group, b i represents the i-th element of the adjacent group, Represents the mean of all elements in the current group; Then, let the correlation degree of the first adjacent group be corr1, and the correlation degree of the second adjacent group be corr2, and the probability of selecting the first adjacent group is: The probability of selecting the second adjacent group is: After selecting the adjacent group according to probabilities p1 and p2, then randomly select an element from the k elements of the selected adjacent group as the cross-domain element.
5. A kidney ultrasound image denoising method according to claim 4, characterized in that: The step of randomly selecting an element from the k elements of the selected adjacent group as the cross-domain element includes the following steps: The probability of randomly selecting an element from the k elements of the selected adjacent group follows the following formula: Among them, z i represents the i-th element in the adjacent group, S i The distance between the i-th element in the adjacent group and the current group is measured, and the calculation method is as follows: v i Represents the i-th element in the current group.
6. A method for denoising a renal ultrasound image according to claim 4, characterized in that: The step of randomly selecting an element from the k elements of the selected adjacent group as the cross-domain element includes the following steps: The probability of randomly selecting an element from the k elements of the selected adjacent group follows the following formula: Among them, z i represents the i-th element in the adjacent group, H i The calculation method is as follows: i =S i ×N i , where S i The distance between the i-th element in the adjacent group and the current group is measured, and the calculation method is as follows: v i Represents the i-th element in the current group, N i Represents element z i Specificity in adjacent groups, specificity N i The calculation method is as follows: N i Represents the element z i The specificity between other elements in the adjacent group, N i The larger the z i Has greater specificity within its own grouping.
7. A method for denoising a renal ultrasound image according to claim 4, characterized in that: In the described low-frequency branch, first, the LCONV module is used to convert the low-frequency component into a feature map with c channels, where c < m. Then, cascade processing is performed through a multi-level structured non-local attention module. The structured non-local attention module includes two parts: a separable convolution module S2 and a non-local module NL. The input and output of the separable convolution module S2 are both feature maps of size h×w×c, and the input and output of the non-local module NL are both feature maps of size h×w×c; The structure of the separable convolution module S2 is the same as that of the separable convolution module S1, except for the difference in the number of channels of the output feature map. The number of channels of the output feature map of the separable convolution module S2 is c, and the number of channels of the output feature map of the separable convolution module S1 is m; The non-local module NL adjusts the weight of the image on the two-dimensional plane by paying attention, so that the structural features of the image are more prominent. The input of the non-local module NL is a feature map of size h×w×c, denoted as a three-dimensional matrix F. First, the three-dimensional matrix F is transformed into three two-dimensional matrices B1, B2 and B3, respectively. Among them, the size of matrix B1 is hw×c, the size of matrix B2 is c×hw, and the size of matrix B3 is hw×c. B1 and B2 are matrix multiplied to obtain a two-dimensional matrix B4 with a size of hw×hw. Then, a softmax operation is performed on all elements of B4 to make the sum of all elements equal to 1. The matrix after softmax calculation is denoted as B5. B5 and B3 are matrix multiplied to obtain a matrix B6 of size hw×c. B6 is operated once with 1×1 convolution to obtain a three-dimensional matrix B7 of size h×w×c. B7 is the output feature map of the non-local module NL.
8. The method for denoising a renal ultrasound image according to claim 1, characterized in that: The fusion module includes a splicing module CAT and a convolution RCONV connected in sequence, wherein the splicing module CAT splices the feature maps output by the high-frequency branch and the low-frequency branch on the channel, and the convolution RCONV processes and transforms the spliced feature maps to obtain a denoised ultrasound image.
9. The method for denoising a renal ultrasound image according to claim 1, characterized in that: The training of the denoising model comprises the following steps: Step 301: Set the hyperparameters for denoising model training, including sample batch size b, learning rate lr, and number of training rounds e. Step 302: Take any b sample pairs from the training set D1, and use the noisy ultrasound images in the sample pairs as inputs of the denoising model. Step 303: The denoising model performs forward inference on the noisy ultrasound image to obtain b predicted denoised images. Step 304: Calculate the mean square error between the predicted denoised image and the original ultrasound image, use the mean square error as a loss function, and use the back propagation algorithm to adjust the parameters of the neural network in the noise model. Step 305: Repeat steps 302 to 304 for a total of e times, where e is a preset number of times.
10. A kidney ultrasound image denoising device, comprising: processor; and, a memory for storing executable instructions of the processor; Wherein, the processor is configured to execute a method for denoising a renal ultrasound image by executing the executable instructions of any one of claims 1 to 9.