Transform-based low-quality medical image super-resolution reconstruction method
Through the Transformer-based super-resolution reconstruction method of low-quality medical images, combined with pixel value correction and shadow correction processing, the defects of image detail recovery and global structural information capture in the prior art are solved, and the completeness and accuracy of super-segment processing are improved.
Patent Information
- Application Number
- CN202510408786.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-06-03
AI Technical Summary
The existing medical imaging super-resolution reconstruction technology has defects in details recovery, global dependency modeling and computing efficiency, and is particularly difficult to capture the global structural information of images. The deep learning methods have high computational complexity and high data demands.
The super-resolution reconstruction method of low-quality medical images based on Transformer is adopted to super-segment the image to be processed through the Transformer model, and combined with pixel value correction and shadow correction processing to improve the resolution and quality of the image.
It solves the problem of color distortion caused by inaccurate color relationship modeling during the super-score process, and improves the completeness and accuracy of super-score processing, especially in the recovery of shadow details.
Smart Images

Figure CN120088137A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision, and particularly to a super-resolution reconstruction method for low-quality medical images based on Transformer. Background Art
[0002] In the medical field, especially in the diagnosis of ophthalmic diseases, high-quality medical images are the key to accurate diagnosis and treatment. Retinopathy of Prematurity (ROP) is an ophthalmic disease that affects the visual development of premature infants. Its severity is usually divided into five stages according to the international ROP classification standard, from mild vascular abnormalities to severe fibrovascular proliferation, which may ultimately lead to retinal detachment, seriously affecting the visual function and even the quality of life of children. However, in clinical practice, due to equipment limitations, poor imaging conditions or patient factors, the quality of the obtained medical images is often unsatisfactory, which not only increases the difficulty of doctor diagnosis but also may lead to misdiagnosis or missed diagnosis, thus affecting treatment decisions. Therefore, inventing a method that can effectively improve the resolution of low-quality medical images is of great significance for improving the accuracy of early ROP diagnosis and promoting precision medicine.
[0003] Existing medical image super-resolution reconstruction technologies still have obvious defects in terms of detail restoration, global dependence modeling, computational efficiency, etc. In particular, traditional methods are difficult to capture the global structural information of images, while deep learning methods, although having superior performance, have high computational complexity and large data requirements. Summary of the Invention
[0004] Aiming at the deficiencies in the prior art, the present invention provides a super-resolution reconstruction method for low-quality medical images based on Transformer.
[0005] The present invention achieves the above technical objectives through the following technical means.
[0006] A super-resolution reconstruction method for low-quality medical images based on Transformer, comprising:
[0007] The server obtains a medical image that needs to be super-resolved, obtaining the image to be processed and the set of pixel values corresponding to each pixel point in the image to be processed; uses a Transformer model to perform super-resolution processing on the image to be processed, obtaining the target image to be processed; obtains the pixel values corresponding to each pixel point in the target image to be processed, obtaining the first set of pixel values; performs pixel value correction processing on the image to be processed according to a preset accuracy threshold and the first pixel values in the first set of pixel values, obtaining the corrected processed image; obtains the target atmospheric light intensity value corresponding to the corrected processed image through an atmospheric light intensity value acquisition method; based on the target atmospheric light intensity value, combines the dark channel prior theory and atmospheric scattering to perform shadow correction processing on the corrected processed image, obtaining the target image.
[0008] Further, the process of the pixel value correction processing is as follows:
[0009] By calculating the accuracy corresponding to the first pixel values in the first set of pixel values, obtaining the first set of accuracies; determining the first accuracies in the first set of accuracies that are lower than the preset accuracy threshold, obtaining the target set of accuracies; obtaining the first pixel values corresponding to each target accuracy in the target set of accuracies, obtaining the target first set of pixel values; obtaining the position information of each target first pixel value in the first set of pixel values in the first set of pixel values, obtaining the first set of position information; calculating the occurrence probability of the pixel value corresponding to each pixel point in the image to be processed at each first position in the first set of position information and the occurrence probability of each first pixel value in the target image to be processed at each first position in the first set of position information, determining the pixel value with the highest occurrence probability at each first position as the second pixel value, obtaining the second set of pixel values. If different pixel values have the same occurrence probability at the same position, use priority selection, that is, select the first pixel value for probability calculation as the second pixel value; use the second pixel values in the second set of pixel values to replace the target first pixel values in the target first set of pixel values, obtaining the target set of pixel values; use the target pixel values in the target set of pixel values to replace the pixel values of each pixel point in the image to be processed, obtaining the corrected processed image.
[0010] Even further, the first set of accuracies π d The calculation formula is as follows:
[0011]
[0012] In the formula, d i represents taking the accuracy with the maximum probability as the probability distribution of the accuracy of the i-th pixel value in the target image to be processed; Y represents the matrix representation corresponding to the first set of pixel values; T represents the accuracy parameter; softmax represents the logistic regression function; W dis the vector projection matrix corresponding to the first set of pixel values; h i represents the i-th first pixel value in the first set of pixel values.
[0013] Furthermore, the occurrence probability π of each first pixel value of the target image to be processed at each first position in the first set of position information s The calculation formula is as follows:
[0014]
[0015] In the formula, s represents the number of target first pixel values in the target first pixel value set; y i represents the i-th first position information in the first set of position information; Z represents the matrix representation corresponding to the target image to be processed; e represents the natural exponential term; W s represents the pre-input projection matrix; W s T represents the pre-input projection matrix W s is the transpose matrix of; h i and h i+1 represent the first pixel value corresponding to the i-th pixel point and the first pixel value corresponding to the (i + 1)-th pixel point in the target image to be processed;
[0016] Furthermore, the calculation of the occurrence probability of the pixel value corresponding to each pixel point in the image to be processed at each first position in the first set of position information is the same as the calculation of the occurrence probability of each first pixel value of the target image to be processed at each first position in the first set of position information.
[0017] Further, the process of the shadow correction process is as follows:
[0018] Roughly estimate the transmittance of the corrected image to obtain the first transmittance image; obtain the grayscale image corresponding to the corrected image through the grayscale image acquisition method to obtain the first corrected image; according to the first corrected image, the first transmittance in the first transmittance image, and the preset filtering window information, calculate the first linear coefficient and the second linear coefficient when filtering the first transmittance image through the minimum cost function; obtain the refined transmittance image corresponding to the corrected image according to the first corrected image, the first linear coefficient, and the second linear coefficient to obtain the second transmittance image; perform shadow correction processing on the corrected image according to the target atmospheric light intensity value and the second transmittance image to obtain the target image.
[0019] Furthermore, the calculation formula of the first transmittance image t(x, y) is as follows:
[0020]
[0021] In the formula, ω is the correction coefficient; Ω(x, y) represents the minimum filter window in the atmospheric scattering model with the center region of (x, y), c represents the three color channels of r, g, and b, A represents the target atmospheric light intensity value, and A c represents the value of the target atmospheric strong light intensity in the color channel c; (x', y') represents traversing the neighboring pixels within Ω(x, y); I c (x, y) represents the pixel value at the coordinate (x, y) in the color channel c of the dark channel prior image I after the correction process.
[0022] Furthermore, the first linear coefficient a x',y' and the second linear coefficient b x',y' are calculated as follows:
[0023]
[0024] In the formula, G(x, y) represents the first corrected image; represents the preset filter window information; E(a x',y' , b x',y' ) represents the minimum cost function; represents the filter kernel corresponding to the preset filter window information; γ x',y' represents the edge-aware smoothing parameter; ε represents the regularization parameter; represents the product of the arithmetic square root of the local variance within the filter window and the arithmetic square root of the local variance within a window with a radius of 1; represents the average variance of all pixels in the first corrected image.
[0025] Furthermore, the second transmittance image is calculated as follows:
[0026]
[0027] In the formula, represents the mean value of the first linear coefficient; represents the mean value of the second linear coefficient.
[0028] Furthermore, the formula for the target image J(x, y) is as follows:
[0029]
[0030] In the formula, I(x, y) represents the dark channel prior image after the correction process; A represents the target atmospheric light intensity value; max{} represents the maximum value operation; t 0 represents a constant, and t 0 ∈(0, 1).
[0031] The beneficial effects of the present invention are as follows:
[0032] (1) By calculating the accuracy rate of the first pixel value set, screening out the target accuracy rate set lower than the threshold, obtaining its position information and calculating the occurrence probability of each pixel in the target image, and replacing the target first pixel value with the pixel value with the highest probability, a corrected processed image is obtained, which solves the problem of color distortion caused by inaccurate color relationship modeling in the super-resolution process of the self-attention mechanism of the Transformer, and improves the integrity of the super-resolution process.
[0033] (2) By obtaining the rough transmittance and grayscale image of the corrected processed image, calculating the refined transmittance, obtaining the second transmittance image, and performing shadow correction in combination with the target atmospheric light intensity value, a target image is obtained. The transmittance reflects the light transmittance of the object and the shadow shape. By performing shadow correction through the transmittance image, the shadow details of the image after super-resolution processing are more accurate and delicate, the shadow gradient effect is restored, and the accuracy of the super-resolution process is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 It is a flowchart of the method for extracting low-quality medical image super-resolution reconstruction based on Transformer according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0035] The present invention will be further described below in conjunction with the drawings and specific embodiments, but the protection scope of the present invention is not limited thereto.
[0036] A method for extracting low-quality medical image super-resolution reconstruction based on Transformer according to the present invention specifically includes the following steps:
[0037] Step 1: The server obtains a medical image that needs to be super-resolved, and obtains the image to be processed and the pixel value set corresponding to each pixel point in the image to be processed.
[0038] Step 2: Use the Transformer model to perform super-resolution processing on the image to be processed, and obtain a target image to be processed.
[0039] Step 3: Obtain the pixel value corresponding to each pixel point in the target image to be processed, and obtain a first pixel value set.
[0040] Step 4: Perform pixel value correction processing on the image to be processed according to a preset accuracy rate threshold and the first pixel value in the first pixel value set, and obtain a corrected processed image.
[0041] Although the self-attention mechanism of Transformer is good at capturing long-range dependencies, it has certain deficiencies in processing color information. Color information in images often exhibits local correlations and certain patterns, while the self-attention mechanism may overemphasize global information during the calculation process, ignoring the subtle changes and consistencies of colors in local regions. For example, when processing an image containing rich textures and color gradients, the self-attention mechanism may not be able to accurately model the mutual relationships of colors within local regions, resulting in color deviations during the super-resolution process, and thus color distortion, leading to relatively low accuracy in the super-resolution processing of the image to be processed.
[0042] To reduce the color distortion of the image during super-resolution processing and enhance the image effect of super-resolution processing, in this embodiment, the correct rate corresponding to the first pixel value in the first pixel value set is calculated to obtain the first correct rate set; the first correct rate in the first correct rate set that is lower than the preset correct rate threshold is determined to obtain the target correct rate set; the first pixel value corresponding to each target correct rate in the target correct rate set is obtained to obtain the target first pixel value set; the position information of each target first pixel value in the first pixel value set in the target first pixel value set is obtained to obtain the first position information set; the occurrence probability of the pixel value corresponding to each pixel point in the image to be processed at each first position in the first position information set and the occurrence probability of each first pixel value in the target image to be processed at each first position in the first position information set are calculated, and the pixel value with the highest occurrence probability at each first position is determined as the second pixel value to obtain the second pixel value set. If different pixel values have the same occurrence probability at the same position, preferential selection is adopted, that is, the pixel value that is the first to perform probability calculation is selected as the second pixel value; the target first pixel value in the target first pixel value set is replaced with the second pixel value in the second pixel value set to obtain the target pixel value set; the pixel value of each pixel point in the image to be processed is replaced with the target pixel value in the target pixel value set to obtain the corrected processed image.
[0043] Among them, the correct rate corresponding to each first pixel value in the first pixel value set can be calculated by the method shown in the following formula to obtain the first correct rate set π d :
[0044]
[0045] In the formula, d i represents the probability distribution that the correct rate of taking the maximum probability is the correct rate of the i-th pixel value in the target image to be processed; Y represents the matrix representation corresponding to the first pixel value set; T represents the correct rate parameter, which can be determined by user input or system default; softmax represents the logistic regression function; Wd is the vector projection matrix corresponding to the first pixel value set, which can be obtained by performing vector projection on the first pixel values in the first pixel value set through a conventional vector projection method; h i represents the i-th first pixel value in the first pixel value set.
[0046] Among them, the occurrence probability π of each first pixel value of the target image to be processed at each first position in the first position information set can be calculated by the method shown in the following formula s :
[0047]
[0048] In the formula, s represents the number of target first pixel values in the target first pixel value set; y i represents the i-th first position information in the first position information set; Z represents the matrix representation corresponding to the target image to be processed; e represents the natural exponential term; W s represents the pre-input projection matrix; W s T represents the pre-input projection matrix W s 's transpose matrix; h i and h i+1 represent the first pixel value corresponding to the i-th pixel point and the first pixel value corresponding to the (i + 1)-th pixel point in the target image to be processed. The occurrence probability of the pixel value corresponding to each pixel point in the image to be processed at each first position in the first position information set can be obtained in the same way.
[0049] Step 5: Obtain the atmospheric light intensity value corresponding to the corrected image through a general method for obtaining atmospheric light intensity values (such as the maximum brightness method or the random forest regression method) to obtain the target atmospheric light intensity value.
[0050] Step 6: Perform shadow correction processing on the corrected image in Step 4 according to the target atmospheric light intensity value in Step 5 to obtain the target image.
[0051] Although Transformer performs well in global information processing, for some fine local details in the shadow area (such as the subtle gradient at the shadow edge, the small reflection of objects in the shadow, etc.), its restoration ability is limited. Because the self-attention calculation of Transformer mainly focuses on the correlation between different positions, capturing local details is not its strong point, and the restored shadow is not accurate and delicate enough in details, so the gradient effect of the shadow part in the image cannot be restored, resulting in low accuracy when performing super-resolution processing on the image to be processed. Therefore, in this embodiment, based on the target atmospheric light intensity value, combined with the dark channel prior theory and atmospheric scattering, the corrected image is subjected to shadow correction processing to output the target image, and the process is as follows:
[0052] Roughly estimate the transmittance of the corrected processed image to obtain a first transmittance image; obtain the grayscale image corresponding to the corrected processed image through a general grayscale image acquisition method (such as the weighted average method) to obtain a first corrected processed image; according to the first corrected processed image, the first transmittance in the first transmittance image, and the preset filtering window information, calculate the first linear coefficient and the second linear coefficient when filtering the first transmittance image through a minimum cost function; obtain the refined transmittance image corresponding to the corrected processed image according to the first corrected processed image, the first linear coefficient, and the second linear coefficient to obtain a second transmittance image; perform shadow correction processing on the corrected processed image according to the target atmospheric light intensity value and the second transmittance image to obtain a target image. Among them, the first transmittance image can be understood as being composed of first transmittance pixels corresponding to each pixel point of the corrected processed image; the first transmittance is the rough transmittance corresponding to each pixel point of the corrected processed image; the second transmittance image can be understood as being composed of second transmittance pixels corresponding to each pixel point of the corrected processed image; the second transmittance is the refined transmittance corresponding to each pixel point of the corrected processed image.
[0053] The calculation formula of the first transmittance image is as follows:
[0054]
[0055] In the formula, t(x, y) represents the rough transmittance corresponding to the pixel point with coordinates (x, y) in the first transmittance image; ω in the formula is a correction coefficient, which is taken as 0.9 in this embodiment and can be determined by user input or system default; Ω(x, y) represents the minimum filter window in the atmospheric scattering model with (x, y) as the central region, c represents the r, g, b three color channels, A represents the target atmospheric light intensity value, and A c represents the value of the target atmospheric strong light intensity in the color channel c; (x', y') represents traversing the neighboring pixels within Ω(x, y); I c (x, y) represents the pixel value of the coordinate (x, y) in the color channel c of the dark channel prior image I after correction processing.
[0056] The calculation formulas of the first linear coefficient and the second linear coefficient are as follows:
[0057]
[0058] G(x, y) in the formula represents the first corrected processed image; a x',y' represents the first linear coefficient; b x',y' represents the second linear coefficient; represents the preset filtering window information, which can be determined by user input or system default; E(ax',y' , b x',y' ) represents the minimum cost function; represents the filter kernel corresponding to the preset filter window information, which is obtained by smoothing the weights using two-dimensional Gaussian filtering; γ x',y' represents the edge-aware smoothing parameter, which can be determined by user input or by system default; ε represents the regularization parameter used to limit x',y' the value of a from becoming too large, which can be determined by user input or by system default; represents the product of the arithmetic square root of the local variance within the filter window and the arithmetic square root of the local variance within a window with a radius of 1; represents the average variance of all pixels in the first corrected processed image.
[0059] The calculation formula for the second transmittance image is as follows:
[0060]
[0061] In the formula, represents the second transmittance image; represents the mean value of the first linear coefficient; represents the mean value of the second linear coefficient.
[0062] The calculation formula for the target image is as follows:
[0063]
[0064] In the formula, J(x, y) represents the target image; I(x, y) represents the dark channel prior image after the correction process; A represents the target atmospheric light intensity value; max{} represents the maximum value operation; t 0 represents a constant, and t 0 ∈ (0, 1), which can be determined by user input or by system default.
[0065] The above embodiments are the preferred embodiments of the present invention, but the present invention is not limited to the above embodiments. Without departing from the essence of the present invention, any obvious improvements, substitutions, or modifications that those skilled in the art can make all fall within the protection scope of the present invention.
Claims
1. A Transformer-based low-quality medical image super-resolution reconstruction method, characterized by: The server obtains a medical image that needs to be super-resolution processed, and obtains the image to be processed and a set of pixel values corresponding to each pixel point in the image to be processed; Perform super-resolution processing on the image to be processed using the Transformer model to obtain a target image to be processed; obtain a pixel value corresponding to each pixel point in the target image to be processed to obtain a first pixel value set; Performing pixel value correction processing on the image to be processed according to a preset accuracy threshold and a first pixel value in the first pixel value set to obtain a corrected image; The target atmospheric light intensity value corresponding to the corrected image is obtained by using an atmospheric light intensity value obtaining method; Based on the target atmospheric light intensity value, combined with the dark channel prior theory and atmospheric scattering, the shadow correction processing is performed on the corrected image to obtain the target image.
2. The Transformer-based low-quality medical image super-resolution reconstruction method according to claim 1, characterized in that: The process of pixel value correction is as follows: A first accuracy set is obtained by calculating the accuracy corresponding to the first pixel value in the first pixel value set; a first accuracy in the first accuracy set that is lower than the preset accuracy threshold is determined to obtain a target accuracy set; and a first pixel value corresponding to each target accuracy in the target accuracy set is obtained to obtain a target first pixel value set; Acquire position information of each target first pixel value in the target first pixel value set in the first pixel value set to obtain a first position information set; Calculate the probability of occurrence of the pixel value corresponding to each pixel point in the to-be-processed image at each first position in the first position information set and the probability of occurrence of each first pixel value in the target to-be-processed image at each first position in the first position information set, determine the pixel value with the highest probability of occurrence of each first position information as the second pixel value, and obtain a second pixel value set. If different pixel values have the same probability of occurrence at the same position, a priority selection is adopted, that is, the first pixel value for which probability calculation is performed is selected as the second pixel value; Using the second pixel value in the second pixel value set to replace the target first pixel value in the target first pixel value set, to obtain a target pixel value set; The pixel value of each pixel in the image to be processed is replaced by the target pixel value in the target pixel value set to obtain a corrected image.
3. The Transformer-based low-quality medical image super-resolution reconstruction method according to claim 2, characterized in that: The first accuracy rate set π d The calculation formula is as follows: In the formula, d i represents the probability distribution of the accuracy of the i-th pixel value in the target image to be processed with the maximum probability as the accuracy; Y represents the matrix representation corresponding to the first pixel value set; T represents the accuracy parameter; softmax represents the logistic regression function; W d is the vector projection matrix corresponding to the first pixel value set; h i Represents the i-th first pixel value in the first pixel value set.
4. The Transformer-based low-quality medical image super-resolution reconstruction method according to claim 3, characterized in that: The probability of occurrence of each first pixel value of the target image to be processed at each first position in the first position information set is π s The calculation formula is as follows: In the formula, s represents the number of target first pixel values in the target first pixel value set; y i represents the i-th first position information in the first position information set; Z represents the matrix representation corresponding to the target image to be processed; e represents the natural exponential term; W s Represents the pre-input projection matrix; W s T Represents the pre-entered projection matrix W s The transposed matrix of i and h i+1 Represents the first pixel value corresponding to the i-th pixel point and the first pixel value corresponding to the i+1-th pixel point in the target image to be processed.
5. The Transformer-based low-quality medical image super-resolution reconstruction method according to claim 4, characterized in that: The calculation of the occurrence probability of the pixel value corresponding to each pixel point in the image to be processed at each first position in the first position information set is similar to the calculation of the occurrence probability of each first pixel value of the target image to be processed at each first position in the first position information set.
6. The Transformer-based low-quality medical image super-resolution reconstruction method according to claim 1, characterized in that: The process of the shadow correction processing is as follows: Roughly estimating the transmittance of the corrected image to obtain a first transmittance image; acquiring a grayscale image corresponding to the corrected image by a grayscale image acquisition method to obtain a first corrected image; According to the first corrected processed image, the first transmittance in the first transmittance image and the preset filtering window information, the first linear coefficient and the second linear coefficient when filtering the first transmittance image are calculated through the minimum cost function; according to the first corrected processed image, the first linear coefficient and the second linear coefficient, the refined transmittance image corresponding to the corrected processed image is obtained to obtain the second transmittance image; according to the target atmospheric light intensity value and the second transmittance image, the shadow correction processing is performed on the corrected processed image to obtain the target image.
7. The Transformer-based low-quality medical image super-resolution reconstruction method according to claim 6, characterized in that: The calculation formula of the first transmittance image t(x, y) is as follows: In the formula, ω is the correction coefficient; Ω(x,y) represents the minimum filter window with (x,y) as the center in the atmospheric scattering model, c represents the three color channels of r, g, and b, A represents the target atmospheric light intensity value, and A c represents the value of the target atmospheric intensity in color channel c; (x', y') represents the neighborhood pixels traversed within Ω(x, y); I c (x, y) represents the pixel value of coordinate (x, y) in the color channel c of the dark channel prior image I after correction.
8. The Transformer-based low-quality medical image super-resolution reconstruction method according to claim 7, characterized in that: The first linear coefficient a x',y' and the second linear coefficient b x',y' The calculation formula is as follows: In the formula, G(x,y) represents the first corrected image; ω ξ1 (x', y') represents the preset filter window information; E(a x',y' ,b x',y' ) represents the minimum cost function; Indicates the filter kernel corresponding to the preset filter window information; γ x',y' represents the edge-aware smoothing parameter; ε represents the regularization parameter; represents the product of the square root of the local variance within the filter window and the square root of the local variance within the window with a radius of 1; Represents the average square error of all pixels in the first rectified image.
9. The Transformer-based low-quality medical image super-resolution reconstruction method according to claim 8, characterized in that: The second transmittance image The calculation formula is as follows: In the formula, represents the mean of the first linear coefficient; Represents the mean of the second linear coefficient.
10. The Transformer-based low-quality medical image super-resolution reconstruction method according to claim 9, characterized in that: The target image J(x,y) is calculated as follows: In the formula, I(x,y) represents the dark channel prior image after correction; A represents the target atmospheric light intensity value; max{} represents the maximum value operation; t0 represents a constant, and t0∈(0,1).
Citation Information
Cited By
Unmanned aerial vehicle charging intelligent take-off and landing control system based on image recognition
CN121187195A