A method for small object anomaly detection in medical images based on multi-step iterative optimization
By adopting a robust principal component analysis network based on multi-step iterative optimization in the detection of small target abnormalities in medical images, the problems of individual differences, subjectivity and resource limitation in the existing methods are solved, and higher detection accuracy and efficiency are achieved.
Patent Information
- Application Number
- CN202410480509.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-22
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-04-22
AI Technical Summary
The existing medical image abnormality detection methods have problems such as individual differences, subjectivity, limited resources, time-consuming and labor-consuming, and easy to miss detection of small-sized targets, which are difficult to effectively improve the accuracy and efficiency of detection.
Using a medical image small target abnormality detection method based on multi-step iterative optimization, a robust principal component analysis network is established, and the sparse constraint function in the target extraction module is used to approximate the target and extract the module, and a background estimation module and image reconstruction module are combined to realize the detection of small targets in medical images.
Improve the accuracy and interpretability of abnormal abnormality detection of small targets in medical images, reduce the work burden of doctors, and provide an effective solution to cope with the growth in the demand for medical image examination.
Smart Images

Figure CN118396943B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular to a method for detecting small object anomalies in medical images based on multi-step iterative optimization. Background Art
[0002] In order to improve the accuracy and efficiency of anomaly detection in medical images, small target anomaly detection technology in medical images has been widely used in the screening of various diseases. This technology analyzes small changes in images by carefully observing them. However, the quality of medical images is affected by many factors such as exposure, contrast, noise, and artifacts, making small target anomaly detection difficult and complex.
[0003] Traditional methods for detecting abnormalities in medical images rely mainly on the manual vision and experience of doctors, but this approach has some problems. First, individual differences and subjective factors may lead to inconsistent or inaccurate diagnostic results. Second, the limited resources of professional doctors make it difficult to cope with the growing demand for medical image examinations. Furthermore, the manual analysis process is time-consuming and labor-intensive, which can easily lead to doctor fatigue and misdiagnosis. Finally, small-sized objects have large position variations in the image and are easily missed. These objects are usually only a few pixels in size, and detecting them is a challenge.
[0004] In response to these problems, the medical image small object anomaly detection method based on deep unfolding network (DUN) has attracted widespread attention. DUN, also known as algorithm unfolding, is an emerging technology that effectively connects iterative algorithms and neural networks. By unfolding the iterative solution algorithm of the existing model at the iteration level to build a network and updating the hyperparameters in a network-based manner, DUN builds a precise connection between iterative algorithms and neural networks, showing great potential, which can effectively improve the accuracy and efficiency of medical image anomaly detection and reduce the workload of doctors.
[0005] Nevertheless, complex matrix operations remain an obstacle to studying DUN-based anomaly detection. Some studies have tried to model images using robust principal component analysis (RPCA), but directly learning parameters in singular value thresholding (SVT) and soft thresholding (ST) may ignore the inherent correlation of images and complicate the selection of regularization parameters.
[0006] To overcome these challenges, this paper proposes a novel robust principal component analysis network anomaly detection method. The core idea of this method is to avoid directly applying soft thresholds to deep features, and instead use neural layers to approximate the sparse constraint function within the target extraction module. By converting the traditional iterative optimization algorithm into a deep learning framework, this method aims to improve the accuracy and interpretability of detection, providing an effective solution for small target anomaly detection in medical images. Summary of the invention
[0007] Purpose of the invention: Aiming at the problem of small target anomaly detection in medical images, the present invention proposes a method for small target anomaly detection in medical images based on multi-step iterative optimization, establishes an interpretable framework based on a multi-step iterative optimization network, models the task of medical image anomaly detection as a robust principal component analysis (RPCA) problem, and solves the optimization steps through network simulation, including proximal networks and sparsely constrained neural layers. This enables the detection of anomalies from medical images.
[0008] The technical solution adopted by the present invention is: a method for detecting anomalies of small targets in medical images based on multi-step iterative optimization. The anomaly detection method is constructed using three modules, including a background estimation module, which uses a convolutional layer to imitate a proximal function, thereby approximating the background and eliminating steps such as singular value thresholds; a target extraction module, which is used to extract the required small targets under the guidance of an image reconstruction module and a background estimation module; an image reconstruction module, which combines the target and background outputs in an iterative manner; and specifically comprises the following steps:
[0009] Step 1: Input a medical image, convert it into a grayscale image, and perform preprocessing, including: image scaling, Gamma transformation, random horizontal flipping, random vertical flipping, center cropping, random angle rotation, and standardization.
[0010] Step 2: Separate the preprocessed image D into low-rank background B and sparse target T through a physical model.
[0011] Step 3: Separate the target and original image D obtained in step 2, and update the parameters to initialize D 0 =D, T 0 =0.
[0012] Step 4: Initialize the updated parameters in step 3 to D 0 , T 0 Send to the background estimation module, target extraction module and image reconstruction module.
[0013] Step 5: Pass the parameters in step 4 through N synthesis stages, each of which corresponds to an iterative matrix low-rank sparse decomposition process to simulate the multiple iterative update operations in the model-driven method. That is, the updated parameters D n-1 and T n-1 is fed into the nth decomposition stage, where n∈{1,...,N}.
[0014] Step 6: The characteristic low-rank background B and sparse target T obtained in step 5 are constrained to restore the low-rank background B that can be paired with the given D.
[0015] Step 7: The output of step 6 is the low-rank background Bn and the sparse target T n Calculate the target segmentation loss function L s .
[0016] Step 8: The output D of the image reconstruction module in step 6 n , the least squares error with the original image is used to measure the image reconstruction performance, and the image reconstruction loss function L b .
[0017] Step 9: Based on the above two losses L s , L b Calculate the loss of the entire network model and calculate L total .
[0018] Step 10: According to the loss of the entire network model and L total , and update the model parameters.
[0019] Step 11: Repeat steps 2 to 10 until the number of training times reaches the expected value.
[0020] Furthermore, the image D in step 2 is represented as:
[0021] D=B+T
[0022] where D, B, T∈R i×j , where R represents the image matrix, i and j are the height and width of the image respectively.
[0023] Furthermore, the image D∈R in step 3 H×W , where R represents the image matrix, H and W are the height and width of the image respectively.
[0024] Furthermore, the expression of the background estimation module in step 4 is:
[0025] B n =D n-1 -T n-1 +F n (D n-1 -T n-1 )
[0026] Among them: F n (·) is a 3×3 convolution group, which includes a l B The intermediate layer [Conv+BN+ReLU] composed of a convolutional layer (Conv), a normalization layer (BN layer) and a ReLU activation function layer, and two convolutional layers, namely the feature extraction layer [Conv+BN+ReLU] and the image reconstruction layer [Conv].
[0027] The target extraction module uses the updated background B n Target Tn-1 And the reconstruction result D n-1 As input, this method assigns the parameter learning task to ε, denoted as ε n , which is a learnable scalar that is independent of each reconstruction stage and does not share parameters.
[0028] Here the expression of the target extraction module is:
[0029] T n =T n-1 +D n-1 -B n -ε n G n (T n-1 +D n-1 -B n )
[0030] Among them, G n (·) includes an initial convolutional layer [Conv], l T An intermediate layer, which consists of a convolutional layer (Conv) and a ReLU activation function layer [Conv+ReLU] and a reconstruction layer [Conv].
[0031] The image reconstruction module converts the decomposition task into an image reconstruction task. The module adopts a simple and suitable CNN architecture to learn image features and effectively map the decomposed background and target.
[0032] Here the expression of the image reconstruction module is:
[0033] D n =M n (B n +T n )
[0034] Among them, M n (·) and F n (·) is similar to the convolutional layer, which has three types of convolutional layers, namely, feature extraction layer [Conv+BN+ReLU], intermediate layer l D [Conv+BN+ReLU], where l D =3. Image reconstruction layer [Conv].
[0035] Furthermore, in step 6, the detection problem is transformed into:
[0036]
[0037] Among them, μ is a penalty coefficient, α represents a positive trade-off parameter, and ||·|| F represents the F-norm, which is defined as Xij is the element in the i-th row and j-th column of the matrix X. R(B) and S(T) are the prior knowledge of the constrained background and target image, respectively. During the optimization process, the estimates of the background B and the target T are updated alternately through an iterative method. Each iteration aims to minimize L(B, T) while satisfying the constraints of the image.
[0038] Furthermore, the target segmentation loss function L in step 7 is s The calculation formula is as follows:
[0039]
[0040] Among them, N t is the total number of pixels in the training sample, TP (True Positives) is the number of correctly detected target pixels, FP (False Positives) is the number of background pixels incorrectly marked as targets, and FN (False Negatives) is the number of undetected target pixels. The target segmentation loss is measured using SoftloU (Soft Intersection over Union), which is a pixel-level performance evaluation indicator used to evaluate the accuracy of target segmentation.
[0041] Furthermore, the loss L of the image reconstruction in step 8 is b The calculation formula is as follows:
[0042]
[0043] Among them, N t is the total number of pixels in the training sample, and N is the total number of pixels in each image. ||·|| F represents the F-norm.
[0044] Further, in step 9, L total The calculation formula is:
[0045] L total =L s +γL b
[0046] Wherein, γ is a regularization parameter, which is set to 0.01 in the present invention.
[0047] The present invention discloses a method for detecting anomalies of small targets in medical images. The method is a learnable deep network architecture, which is derived from the RPCA model and combines the accuracy of data-driven networks with the interpretability of model-driven networks. The target extraction module and background estimation module of the method for detecting anomalies of small targets in medical images approximately segment the target and the background. The nonlinear proximal mapping problem of the target is effectively handled, and the complex matrix calculation in the background estimation is replaced by the neural layer, so that the traditional iterative optimization algorithm is transformed into a deep learning framework to improve the accuracy and interpretability of the detection. The present invention proposes an image reconstruction module, which combines the background and the target to reconstruct the image.
[0048] Beneficial effects of the present invention:
[0049] 1. The method proposed in this paper proposes an interpretable framework based on a multi-step iterative optimization network for small object anomaly detection in medical images. The detection task is modeled as a robust principal component analysis (RPCA) problem, and the optimization steps are solved by network simulation, including proximal networks and sparsely constrained neural layers.
[0050] 2. The interpretable architecture of the method of the present invention based on a multi-step iterative optimization network effectively guides the neural layer to learn low-rank background and sparse targets, facilitating the detection task in a nearly "white box" manner. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 It is a flow chart of a method for detecting small object anomalies in medical images based on multi-step iterative optimization according to the present invention. DETAILED DESCRIPTION
[0052] The present invention is further described below in conjunction with the accompanying drawings and specific implementation methods.
[0053] The present invention relates to a method for detecting small object anomalies in medical images based on multi-step iterative optimization. The method utilizes the spatial correlation and structural consistency in radiological images, as well as the semantic information and contextual information in the deep feature space, to extract and reconstruct common anatomical structures in the image and identify patterns that have not been seen or modified in the image. The specific steps are as follows:
[0054] Reference Figure 1 , a method for detecting small object anomalies in medical images based on multi-step iterative optimization, comprising the following steps:
[0055] Step 1: Input a medical image, convert it into a grayscale image, and perform preprocessing, including: image scaling, Gamma transformation, random horizontal flipping, random vertical flipping, center cropping, random angle rotation, and standardization.
[0056] Step 2: Separate the preprocessed image D into low-rank background B and sparse target T through the physical model. Here, the image D is represented as:
[0057] D=B+T
[0058] where D, B, T∈R i×j , where R represents the image matrix, i and j are the height and width of the image respectively.
[0059] Step 3: Separate the target and original image D obtained in step 2, and update the parameters to initialize D 0 =D, T 0 =0.
[0060] Among them, the image D∈R H×W , where R represents the image matrix, H and W are the height and width of the image respectively.
[0061] Step 4: Initialize the updated parameters in step 3 to D 0 , T 0 Send to the background estimation module, target extraction module and image reconstruction module.
[0062] Here the expression of the background estimation module is:
[0063] B n =D n-1 -T n-1 +F n (D n-1 -T n-1 )
[0064] Among them: F n (·) is a 3×3 convolution group, which includes a l B The intermediate layer [Conv+BN+ReLU] composed of a convolutional layer (Conv), a normalization layer (BN layer) and a ReLU activation function layer, and two convolutional layers, namely the feature extraction layer [Conv+BN+ReLU] and the image reconstruction layer [Conv].
[0065] The target extraction module uses the updated background B n Target T n-1 And the reconstruction result D n-1 As input, this method assigns the parameter learning task to ε, denoted as ε n , which is a learnable scalar that is independent of each reconstruction stage and does not share parameters.
[0066] Here the expression of the target extraction module is:
[0067] T n =T n-1 +D n-1 -Bn -ε n G n (T n-1 +D n-1 -B n )
[0068] Among them, G n (·) includes an initial convolutional layer [Conv], l T An intermediate layer, which consists of a convolutional layer (Conv) and a ReLU activation function layer [Conv+ReLU] and a reconstruction layer [Conv].
[0069] The image reconstruction module converts the decomposition task into an image reconstruction task. The module adopts a simple and suitable CNN architecture to learn image features and effectively map the decomposed background and target.
[0070] Here the expression of the image reconstruction module is:
[0071] D n =M n (B n +T n )
[0072] Among them, M n (·) and F n (·) is similar to the convolutional layer, which has three types of convolutional layers, namely, feature extraction layer [Conv+BN+ReLU], intermediate layer l D [Conv+BN+ReLU], where l D =3. Image reconstruction layer [Conv].
[0073] Step 5: Pass the parameters in step 4 through N synthesis stages, each of which corresponds to an iterative matrix low-rank sparse decomposition process to simulate the multiple iterative update operations in the model-driven method. That is, the updated parameters D n-1 and T n-1 is fed into the nth decomposition stage, where n∈{1,...,N}.
[0074] Step 6: The characteristic low-rank background B and sparse target T obtained in step 5 are constrained to restore the low-rank background B that can be paired with the given D.
[0075] Here the detection problem is transformed into:
[0076]
[0077] Among them, μ is a penalty coefficient, α represents a positive trade-off parameter, and ||·|| F represents the F-norm, which is defined as X ij is the element in the i-th row and j-th column of the matrix X. R(B) and S(T) are the prior knowledge of the constrained background and target image, respectively. During the optimization process, the estimates of the background B and the target T are updated alternately through an iterative method. Each iteration aims to minimize L(B, T) while satisfying the constraints of the image.
[0078] Step 7: The output of step 6 is the low-rank background B n and the sparse target T n Calculate the target segmentation loss function L s .
[0079] Target segmentation loss function L s The calculation formula is as follows:
[0080]
[0081] Among them, N t is the total number of pixels in the training sample, TP (True Positives) is the number of correctly detected target pixels, FP (False Positives) is the number of background pixels incorrectly marked as targets, and FN (False Negatives) is the number of undetected target pixels. The target segmentation loss is measured using SoftloU (Soft Intersection over Union), which is a pixel-level performance evaluation indicator used to evaluate the accuracy of target segmentation.
[0082] Step 8: The output Dn of the image reconstruction module in step 6 is measured by the least squares error with the original image to measure the image reconstruction performance. The image reconstruction loss function L b .
[0083] The loss of image reconstruction is L b The calculation formula is as follows:
[0084]
[0085] Among them, N t is the total number of pixels in the training sample, and N is the total number of pixels in each image. ||·|| F represents the F-norm.
[0086] Step 9: Based on the above two losses L s , L b Calculate the loss of the entire network model and calculate L total .
[0087] L total The calculation formula is:
[0088] L total =L s +γL b
[0089] Wherein, γ is a regularization parameter, which is set to 0.01 in the present invention.
[0090] Step 10: According to the loss of the entire network model and L total , and update the model parameters.
[0091] Step 11: Repeat steps 2 to 10 until the number of training sessions reaches the desired target.
[0092] The above are only preferred embodiments of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions under the concept of the present invention belong to the protection scope of the present invention. It should be pointed out that for ordinary technicians in this technical field, some improvements and modifications without departing from the principle of the present invention should be regarded as the protection scope of the present invention.
Claims
1. A method for detecting small object anomalies in medical images based on multi-step iterative optimization, characterized in that: The following steps are involved: Step 1: Input a medical image, convert it into a grayscale image, and perform preprocessing, including: image scaling, Gamma transformation, random horizontal flip, random vertical flip, center cropping, random angle rotation, and standardization; Step 2: Separate the preprocessed image D into low-rank background B and sparse target T through a physical model; Step 3: Update the parameters initialized to D 0 =D, T 0 =0; In step 3, the image D∈R H×W , where R represents the image matrix, H and W are the height and width of the image respectively; Step 4: Initialize the updated parameters in step 3 to D 0 , T 0 , sent to the background estimation module, target extraction module and image reconstruction module; The expression of the background estimation module in step 4 is: B n =D n-1 -T n-1 +F n (D n-1 -T n-1 ) Among them: F n (·) is a 3×3 convolution group, which includes a l B It consists of an intermediate layer consisting of a convolutional layer, a normalization layer, and a ReLU activation function layer; the two convolutional layers are the feature extraction layer and the image reconstruction layer; The target extraction module uses the updated background B n Target T n-1 And the reconstruction result D n-1 As input, the parameter learning task is assigned to ε, denoted as ε n , which is a learnable scalar that is independent of each reconstruction stage and does not share parameters; Here the expression of the target extraction module is: T n =T n-1 +D n-1 -B n -ε n G n (T n-1 +D n-1 -B n ) Among them, G n (·) includes an initial convolutional layer, l T An intermediate layer and a reconstruction layer. The intermediate layer consists of a convolutional layer and a ReLU activation function layer. The image reconstruction module converts the decomposition task into an image reconstruction task. This module adopts a simple and suitable CNN architecture to learn image features and effectively map the decomposed background and target; Here the expression of the image reconstruction module is: D n =M n (B n +T n ) Among them, M n (·) and F n (·) is similar to the convolutional layer, which has three types of convolutional layers: feature extraction layer, intermediate layer l D , image reconstruction layer, where l D =3; Step 5: The parameters in step 4 are passed through N synthesis stages, and each stage is iteratively updated according to the background estimation module, target extraction module and image reconstruction module to generate a new reconstructed image D n , where n∈{1,...,N}; Step 6: After all decomposition stages in step 5 are completed, optimize B by constraining the background R(B) n To better match the given image D, and further optimize the target T through the prior knowledge S(T) of the target image n ; In step 6, the detection problem is transformed into: Among them, μ is a penalty coefficient, α represents a positive trade-off parameter, ‖·‖ F represents the F-norm, which is defined as X ij is the element in the i-th row and j-th column of the matrix X; R(B) and S(T) are the prior knowledge of the constrained background and target image, respectively; during the optimization process, the background B is updated alternately by an iterative method n and target T n Each iteration aims to minimize L(B,T) while satisfying the image constraints. Step 7: Output the low-rank background B n and the sparse target T n Calculate the target segmentation loss function L s ; Step 8: The output D of the image reconstruction module n , the image reconstruction performance is measured using the least squares error with the original image D, and the image reconstruction loss function L b ; Step 9: Based on the above two losses L s , L b Calculate the loss of the entire network model and calculate L total ; Step 10: According to the loss of the entire network model and L total , update the model parameters; Step 11: Repeat steps 2 to 10 until the number of training times reaches the expected value.
2. The method for detecting small object anomalies in medical images based on multi-step iterative optimization according to claim 1, characterized in that: The image D in step 2 is represented as: D=B+T where D,B,T∈R i×j , where R represents the image matrix, i and j are the height and width of the image respectively.
3. The method for detecting small object anomalies in medical images based on multi-step iterative optimization according to claim 2, characterized in that: The target segmentation loss function L in step 7 s The calculation formula is as follows: Among them, N t is the total number of pixels in the training sample, TP is the number of correctly detected target pixels, FP is the number of background pixels incorrectly marked as targets, and FN is the number of undetected target pixels. The target segmentation loss is measured using SoftIoU, which is a pixel-level performance evaluation indicator used to evaluate the accuracy of target segmentation.
4. The method for detecting small object anomalies in medical images based on multi-step iterative optimization according to claim 3, characterized in that: The image reconstruction loss L in step 8 is b The calculation formula is as follows: Among them, N t is the total number of pixels in the training sample, and N is the total number of pixels in each image; ‖·‖ F represents the F-norm.
5. The method for detecting small object anomalies in medical images based on multi-step iterative optimization according to claim 4, characterized in that: In step 9, L total The calculation formula is: THE total =L s +γL b Here, γ is the regularization parameter and is set to 0.01.
Citation Information
Patent Citations
Low-rank expression and learning dictionary-based hyperspectral image abnormity detection algorithm
CN105427300A
Real-time magnetic resonance image reconstruction method and system based on deep expansion neural network
CN115439383A