CT image training data construction method and system based on multi-scale similarity matching
By using a multi-scale similarity matching method, image segmentation points are automatically adjusted and a loss function is constructed, which solves the problem of similarity measurement between low-dose CT images and normal-dose CT images, achieving efficient and accurate image pairing and denoising effects, and improving diagnostic accuracy.
Patent Information
- Application Number
- CN202511146742.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-15
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2045-08-15
AI Technical Summary
Existing technologies struggle to accurately measure the similarity between low-dose CT images and normal-dose CT images, impacting image quality and diagnostic accuracy. Furthermore, manual matching methods are time-consuming, labor-intensive, and prone to errors.
A multi-scale similarity matching method is adopted, which dynamically adjusts the image segmentation points through automatic discrete point learning and unsupervised learning, and constructs similarity loss and difference loss by combining contrastive learning to optimize the image matching process and select high-quality image pairs.
It improves the accuracy and efficiency of image similarity calculation, enhances the diagnostic accuracy of low-dose CT image denoising models, reduces costs, and minimizes manual intervention.
Smart Images

Figure CN121095694A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, and in particular to a method and system for constructing CT image training data based on multi-scale similarity matching. Background Technology
[0002] In real-world clinical settings, minimizing patient radiation dose is crucial because radiation exposure can impact patient health, particularly the risk of hematologic malignancies, which is directly proportional to the cumulative dose, with an additional relative risk (ERR) of 1.96 / 100mGv. Based on this, low-dose CT imaging (LDCT) has been widely adopted; however, this technique introduces significant noise and artifacts, negatively impacting image quality and diagnostic accuracy. Currently, commonly used supervised learning methods simulate low-dose CT images by adding Poisson noise to the sinusoidal domain of normal-dose CT (NDCT) images. This simulated low-dose CT image is then projected to obtain a low-dose CT image, which is then denoised by pairing LDCT and NDCT images. However, the distribution of noise in this simulation differs from real-world clinical scenarios, making it difficult to mimic the complexity of real-world conditions. While manual pairing can obtain more realistic image pairs, it is time-consuming, labor-intensive, and prone to errors, significantly hindering the development of large-scale real-world datasets. Furthermore, due to noise and spatial shift between LDCT and NDCT images, traditional Euclidean distance cannot accurately measure their similarity.
[0003] Therefore, there is an urgent need for an automated, robust, and efficient image similarity matching method to automatically pair LDCT and NDCT images and provide a high-quality dataset for training. Summary of the Invention
[0004] To address the aforementioned issues, this invention proposes a method and system for constructing CT image training data based on multi-scale similarity matching. By introducing automatic discrete point learning based on the contrastive learning approach, the optimal discrete point partitioning rules are learned from small samples in an unsupervised manner. This adapts to the actual distribution characteristics of the images, improves the accuracy of similarity calculation, and more accurately selects matching image pairs, providing a high-quality dataset for training denoising models.
[0005] To achieve the above objectives, the present invention adopts the following technical solution: In a first aspect, the present invention provides a method for constructing CT image training data based on multi-scale similarity matching, comprising: Acquire unpaired normal-dose CT images and low-dose CT images; Define the number of scales and segmentation points, and perform multi-scale discretization on normal-dose CT images and low-dose CT images respectively; wherein, a trainable mapping function is used to determine the set of segmentation points in an unsupervised manner. For each scale, the discretized normal-dose CT image and low-dose CT image are subtracted pixel by pixel to generate a normalized difference map; the mean of the difference map is calculated to obtain the single-scale similarity; among them, similarity loss and difference loss are constructed based on the idea of contrastive learning, and the parameters of the trainable mapping function are optimized. Weighted fusion of multi-scale similarity scores is used to select image pairs with similarity greater than a threshold as training data.
[0006] Secondly, the present invention provides a CT image training data construction system based on multi-scale similarity matching, comprising: The acquisition module is configured to acquire unpaired normal-dose CT images and low-dose CT images; The segmentation module is configured to define the number of scales and segmentation points, and to perform multi-scale discretization processing on normal dose CT images and low dose CT images respectively; wherein, the set of segmentation points is determined by a trainable mapping function in an unsupervised manner. The similarity calculation module is configured to subtract the discretized normal-dose CT image and low-dose CT image at each scale pixel by pixel to generate a normalized difference map; calculate the mean of the difference map to obtain the single-scale similarity; and construct similarity loss and difference loss based on the idea of contrastive learning to optimize the parameters of the trainable mapping function. The training data construction module is configured to weight and fuse multi-scale similarity scores, and select image pairs with similarity greater than a threshold as training data.
[0007] Thirdly, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in the method for constructing CT image training data based on multi-scale similarity matching as described in the first aspect.
[0008] Fourthly, the present invention provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps in the method for constructing CT image training data based on multi-scale similarity matching as described in the first aspect.
[0009] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention introduces automatic discrete point learning, leveraging unsupervised learning to explore the most suitable discrete data distribution partitioning method from samples without human intervention. By automatically optimizing discrete points, the multi-scale discrete similarity measure (MSP) can more accurately evaluate image similarity. Experiments show that the MSP algorithm using automatic discrete point learning, when combined with various models, significantly improves both visual effects and numerical indicators, powerfully promoting the development of low-dose CT image denoising technology and improving the accuracy of clinical diagnosis.
[0010] Advantages of additional aspects of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description
[0011] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute a limitation thereof.
[0012] Figure 1 LDCT images under different methods provided in embodiments of the present invention; Figure 1 (a) is an NDCT image; Figure 1 (b) is an LDCT image simulated by adding sinusoidal noise to (a); Figure 1 (c) are manually matched LDCT images; Figure 1 (d) LDCT images matched using Euclidean distance; Figure 1 (e) Images matched using multi-scale discrete similarity measure as a similarity metric; Figure 2 The main flowchart of a method for constructing CT image training data based on multi-scale similarity matching provided in an embodiment of the present invention; Figure 3 A flowchart of a whole-image matching method based on multi-scale discrete similarity measurement provided in an embodiment of the present invention; Figure 4 The flowchart illustrates the image patch selection method based on multi-scale discrete similarity measurement provided in this embodiment of the invention. Detailed Implementation
[0013] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0014] In the field of medical imaging, acquiring real LDCT-NDCT paired images faces significant challenges: due to radiation exposure restrictions and ethical requirements, it is not feasible to acquire two CT images of the same patient at the same time point.
[0015] Figure 1 The existing methods and the method proposed in this invention were compared for NDCT images ( Figure 1 (a) Results of matching highly similar LDCT images.
[0016] To address the challenge of acquiring paired images in LDCT-NDCT, existing techniques typically employ simulated images to provide large-scale, high-quality paired training data for denoising models—a crucial supplementary method for model training in medical settings where real paired data is extremely scarce. However, the synthetic noise generated by these methods cannot fully simulate the complex noise characteristics of clinical CT scans, such as… Figure 1 As shown in (b).
[0017] Another type of conventional technique obtains more realistic image pairs through manual pairing, but this process is time-consuming, labor-intensive, and prone to introducing errors, such as... Figure 1 As shown in (c), this severely restricts the construction of large-scale real datasets.
[0018] Meanwhile, due to noise differences and spatial shifts between LDCT and NDCT images, traditional Euclidean distance is insufficient to accurately measure their similarity. Figure 1 As shown in (d).
[0019] Therefore, this invention proposes a method for constructing CT image training data based on multi-scale similarity matching, aiming to replace manual operation and achieve low-cost, rapid pairing of highly similar real-world LDCT-NDCT images, such as... Figure 1 As shown in (e).
[0020] Example 1 like Figure 2 As shown in the figure, this embodiment discloses a method for constructing CT image training data based on multi-scale similarity matching, including the following steps: S1: Acquire unpaired normal-dose CT images and low-dose CT images; S2: Define the number of scales and segmentation points, and perform multi-scale discretization processing on normal dose CT images and low dose CT images respectively; wherein, a trainable mapping function is used to determine the set of segmentation points in an unsupervised manner. S3: Subtract the discretized normal-dose CT image and low-dose CT image at each scale pixel by pixel to generate a normalized difference map; calculate the mean of the difference map to obtain the single-scale similarity; among them, similarity loss and difference loss are constructed based on the idea of contrastive learning, and the parameters of the trainable mapping function are optimized. S4: Weighted fusion of multi-scale similarity scores, selecting image pairs with similarity greater than a threshold as training data.
[0021] Next, combined Figure 3 This embodiment provides a detailed description of a method for constructing CT image training data based on multi-scale similarity matching.
[0022] In S1, the normal dose CT (NDCT) dataset is defined as Low-dose CT (LDCT) datasets are defined as follows: Among them, the normal dose CT (NDCT) dataset Compared with low-dose CT (LDCT) datasets There is no matching relationship between them.
[0023] From normal dose CT dataset and low-dose CT datasets m normal dose CT images were randomly sampled from each sample. and m low-dose CT images .
[0024] In S2, the number of scales and segmentation points are set for normal dose CT images. and low-dose CT images Multi-scale segmentation is performed based on the number of scales and segmentation points.
[0025] Define the number of scales , This indicates the number of different scales to use; for example, setting two scales.
[0026] Each scale corresponds to a set of discrete segmentation points. .in, n is used to adjust the degree of discretization refinement. When the number of segment points n is small, the focus is on capturing the overall structural information of the image, such as large-sized tissue contours. When n is large, more fine details can be captured, such as the texture of small blood vessels.
[0027] For determining the segmentation points, this embodiment introduces an automatic discrete point learning algorithm. First, a trainable mapping function is used to learn the segmentation rules that adapt to the data distribution, and then the pixel values are discretized based on these rules.
[0028] In the process of multi-scale discretization, since CT images from different patients and different parts of the body have different pixel value ranges and structural features, fixing the segmentation points, such as manually setting {0, 100, 200, 255}, cannot adapt to the differences in tissue distribution of all CT images. Therefore, this embodiment introduces a mapping function to dynamically learn the optimal segmentation points, making the discretization results more closely match the structural features of the real image.
[0029] First, define a trainable mapping function. This function determines the function mapping method. This represents the set of trainable parameters in the trainable mapping function. This example uses the Sigmoid function to define the trainable mapping function: ; in, Input the pixel values of the CT image (values range from 0 to 255). This is the set of trainable parameters.
[0030] This nonlinear function automatically adjusts the center position and stretching amplitude of the mapping through a learnable nonlinear transformation, making the output more sensitive in key pixel areas. This adaptively generates discrete segmentation points, replacing the traditional fixed-interval division method, and better reflects the actual pixel distribution characteristics of the image.
[0031] Apply a trainable mapping function to each image Map the input data from the set of integers 0-255 to The set of real numbers.
[0032] To extract more stable and robust structural features, this embodiment will input the mapped function into a discrete function for discretization, discretizing it to... The set of integers. The discretization process is implemented using a rounding function based on the STE (Straight-Through Estimator) method. Since conventional discretization operations are mathematically non-differentiable, gradients cannot propagate effectively backward, hindering model training. The STE method, however, uses the original rounding operation in forward propagation and an approximate derivative (e.g., constant to 1) to replace the true gradient in backward propagation, enabling gradient backpropagation and thus optimizing the parameters of the mapping function. Through this mechanism, a dynamic set of segmented points adapted to the current data distribution can be obtained, i.e. This completes the self-supervised feature extraction training of the mapping function.
[0033] Any image represented by the automatic discrete point learning algorithm is: ; At this point, each pixel value is calculated using the following formula. Mapping to discrete values: ; in, From the trained mapping function Dynamically determined, The output value is a discretized value used for subsequent multi-scale image discrete similarity measurement.
[0034] In this embodiment, a multi-scale quantity and a trainable mapping function are first defined to dynamically adapt to the distribution of CT image data, solving the problem that fixed segmentation points cannot match the image features of different patients and body parts. The mapping function is optimized through self-supervised training, and discretization using the STE method allows the segmentation points to dynamically adjust according to data features. When n is small, the focus is on capturing the overall structure and resisting large shifts; when n is large, details are refined to enhance robustness to minute structures. By constructing similarity and difference loss optimization parameters, the discretization results are made to closely resemble real images, allowing similarity calculations to focus more on stable structures, effectively reducing noise and minor shift interference, and laying the foundation for the subsequent construction of training data for low-dose CT denoising models.
[0035] In S3, after discrete processing, such as Figure 4 As shown, the calculation is performed for each low-dose CT image. Compared with input normal dose CT image set Similarity vectors between = The specific steps are as follows: Each scale has a unique set of parameters, represented as: { .in, Used to control the mapping function at scale s. Used to control the number of pixel value discretization regions at scale s. -1 equals the number of discrete intervals in the multi-scale mask division. The normal-dose CT image and the low-dose CT image discretized at the s-scale are subtracted pixel-by-pixel, and the low-dose CT image is calculated sequentially. With each normal dose CT image The normalized difference plot between them is represented as: ; Then, a similarity score is calculated for each difference map, and the similarity is summed using trainable weight parameters to obtain low-dose CT images. Compared with input normal dose CT images Final similarity between ); Controlling the contribution ratio of each scale to the "overall similarity score" is key to improving matching accuracy and robustness. ; in, For multi-scale weights, defined as S represents the number of scales. Multi-scale weights are used to fuse image similarity information at different scales, enabling image matching to consider both global structure and local details.
[0036] At the same time, construct similarity loss and difference loss Optimize parameters This maximizes the similarity between matched pairs and the difference between matched and unmatched pairs. ; ; ; in, This represents hyperparameters.
[0037] In S4, weighted fusion across multiple scales is used to derive the final similarity score. Image pairs with a final similarity greater than a threshold are selected as training data.
[0038] To verify the effectiveness of this embodiment, the Multi-scale Similarity Purification (MSP) and Division Points Estimation (DPE) algorithms provided in this embodiment were used to automatically filter and match low-dose CT (LDCT) and normal-dose CT (NDCT) image patches. After using the selected image pairs to train the denoising model, significant improvements were achieved compared to traditional methods in both visual effects (such as noise suppression and structural fidelity) and numerical metrics (such as PSNR and SSIM) of manually paired data.
[0039] Visually, CT reconstructed images using data sanitization strategies exhibit clearer details, including blood vessels and tissues. Numerically, the number of patches is reduced by approximately 93.1%. The method in this embodiment, combined with the RED-CNN model, improves the fid index by 47.8%, the kID index by 77.3%, and the CLIP-FID index by 26.4%. The method in this embodiment, combined with the Lit-former model, improves the fid index by 18.2%, the kID index by 30.5%, and the CLIP-FID index by 22.2%. The method in this embodiment, combined with the NAFNet model, improves the fid index by 37.0%, the kID index by 49.4%, and the CLIP-FID index by 60.8%.
[0040] Using the MSP and DPE algorithms in this embodiment for whole-image matching and patch selection, there was no significant difference in visual effect and numerical indicators compared to manually matched and patch-selected CT restoration images. On the RedCNN model, compared to manual matching, the whole-image matching method based on MSP and DPE algorithms improved FID by 1.51%, Kid by 0.00%, and Clip-FID by 6.02%. On the Liformer model, compared to manual matching, the whole-image matching method based on MSP and DPE algorithms improved FID by 0.07%, Kid by 0.59%, and Clip-FID by 1.28%. On the NAFNet model, compared to manual matching, the whole-image matching method based on MSP and DPE algorithms only decreased FID by 0.76%, Kid improved by 1.20%, and Clip-FID improved by 16.67%. Furthermore, in the restoration of low-dose CT images that are equivalent to normal-dose CT images, the data filtering strategy of this embodiment, combined with RED-CNN, litformer, and NAFNet, achieves good results in both image quality and numerical indicators compared to other algorithms and CT denoising and image denoising methods.
[0041] This specific embodiment addresses the challenge of acquiring true LDCT-NDCT paired images in medical CT imaging. It achieves automated pairing through multi-scale similarity matching, reducing costs and improving efficiency. This solves the problems of significant differences between simulated and clinically real noise, and the time-consuming and error-prone nature of manual pairing. Unsupervised automatic discrete point learning is introduced, combined with contrastive learning to construct a loss function and optimize parameters, making similarity calculation more accurate and overcoming the inaccuracy of traditional Euclidean distance similarity measurement due to noise and spatial offset. Through multi-scale discretization and weighted fusion, both global structure and local details are considered, selecting high-quality image pairs to provide excellent training data for low-dose CT denoising models, thus promoting the development of related technologies and improving the accuracy of clinical diagnosis.
[0042] Example 2 This embodiment provides a CT image training data construction system based on multi-scale similarity matching, including: The acquisition module is configured to acquire unpaired normal-dose CT images and low-dose CT images; The segmentation module is configured to define the number of scales and segmentation points, and to perform multi-scale discretization processing on normal dose CT images and low dose CT images respectively; wherein, the set of segmentation points is determined by a trainable mapping function in an unsupervised manner. The similarity calculation module is configured to subtract the discretized normal-dose CT image and low-dose CT image at each scale pixel by pixel to generate a normalized difference map; calculate the mean of the difference map to obtain the single-scale similarity; and construct similarity loss and difference loss based on the idea of contrastive learning to optimize the parameters of the trainable mapping function. The training data construction module is configured to weight and fuse multi-scale similarity scores, and select image pairs with similarity greater than a threshold as training data.
[0043] Example 3 This embodiment provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps in the CT image training data construction method based on multi-scale similarity matching as described in Embodiment 1 above.
[0044] Example 4 This embodiment provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the steps in the CT image training data construction method based on multi-scale similarity matching as described in Embodiment 1 above.
[0045] The steps or modules involved in Embodiments 2 to 4 above correspond to those in Embodiment 1. For specific implementation details, please refer to the relevant description section of Embodiment 1. The term "computer-readable storage medium" should be understood as a single medium or multiple media including one or more instruction sets; it should also be understood as including any medium capable of storing, encoding, or carrying an instruction set for execution by a processor and enabling the processor to perform any of the methods in this invention.
[0046] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for constructing CT image training data based on multi-scale similarity matching, characterized in that, include: Acquire unpaired normal-dose CT images and low-dose CT images; Define the number of scales and segmentation points, and perform multi-scale discretization on normal-dose CT images and low-dose CT images respectively; wherein, a trainable mapping function is used to determine the set of segmentation points in an unsupervised manner. For each scale, the discretized normal-dose CT image and low-dose CT image are subtracted pixel by pixel to generate a normalized difference map; the mean of the difference map is calculated to obtain the single-scale similarity; among them, similarity loss and difference loss are constructed based on the idea of contrastive learning, and the parameters of the trainable mapping function are optimized. Weighted fusion of multi-scale similarity scores is used to select image pairs with similarity greater than a threshold as training data.
2. The method for constructing CT image training data based on multi-scale similarity matching as described in claim 1, characterized in that, The method of determining the set of segmentation points using a trainable mapping function in an unsupervised manner specifically includes: ; in, Input the pixel values of the CT image. The set of trainable parameters; the trainable mapping function dynamically learns and adapts the segmentation point generation logic to the data features through nonlinear transformation, mapping the input data from the integer set of 0-255 to... The set of real numbers.
3. The method for constructing CT image training data based on multi-scale similarity matching as described in claim 2, characterized in that, The defined number of scales and segmentation points are used to perform multi-scale discretization processing on normal-dose CT images and low-dose CT images, respectively, specifically including: Multiple scales are defined. For each scale, normal-dose CT images and low-dose CT images are divided into multiple scales based on segmentation points, and the value of each pixel is then... Mapping to discrete values: ; in, From the trained mapping function Dynamically determined, This is the discretized output value.
4. The method for constructing CT image training data based on multi-scale similarity matching as described in claim 1, characterized in that, The step of subtracting the discretized normal-dose CT image and low-dose CT image at each scale pixel by pixel to generate a normalized difference map specifically includes: ; in, For discretized low-dose CT images, Discretized normal-dose CT images; Low-dose CT images, For normal dose CT images, { } represents the parameter.
5. The method for constructing CT image training data based on multi-scale similarity matching as described in claim 1, characterized in that, The construction of similarity loss and difference loss based on the contrastive learning approach specifically includes: ; ; ; in, For similarity loss, For difference loss, For single-scale similarity, The similarity after weighted fusion This is a hyperparameter.
6. A CT image training data construction system based on multi-scale similarity matching, characterized in that, include: The acquisition module is configured to acquire unpaired normal-dose CT images and low-dose CT images; The segmentation module is configured to define the number of scales and segmentation points, and to perform multi-scale discretization processing on normal dose CT images and low dose CT images respectively; wherein, the set of segmentation points is determined by a trainable mapping function in an unsupervised manner. The similarity calculation module is configured to subtract the discretized normal-dose CT image and low-dose CT image at each scale pixel by pixel to generate a normalized difference map; calculate the mean of the difference map to obtain the single-scale similarity; and construct similarity loss and difference loss based on the idea of contrastive learning to optimize the parameters of the trainable mapping function. The training data construction module is configured to weight and fuse multi-scale similarity scores, and select image pairs with similarity greater than a threshold as training data.
7. The CT image training data construction system based on multi-scale similarity matching as described in claim 6, characterized in that, The method of determining the set of segmentation points using a trainable mapping function in an unsupervised manner specifically includes: ; in, Input the pixel values of the CT image. The set of trainable parameters; the trainable mapping function dynamically learns and adapts the segmentation point generation logic to the data features through nonlinear transformation, mapping the input data from the integer set of 0-255 to... The set of real numbers.
8. The CT image training data construction system based on multi-scale similarity matching as described in claim 7, characterized in that, The defined number of scales and segmentation points are used to perform multi-scale discretization processing on normal-dose CT images and low-dose CT images, respectively, specifically including: Multiple scales are defined. For each scale, normal-dose CT images and low-dose CT images are divided into multiple scales based on segmentation points, and the value of each pixel is then... Mapping to discrete values: ; in, From the trained mapping function Dynamically determined, This is the discretized output value.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps in the method for constructing CT image training data based on multi-scale similarity matching as described in any one of claims 1-5.
10. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in the method for constructing CT image training data based on multi-scale similarity matching as described in any one of claims 1-5.
Citation Information
Patent Citations
Unsupervised low-dose CT (Computed Tomography) denoising model training method, denoising method and device
CN117094902A
Real scene low-dose CT image denoising method and system
CN118172278A