Early lung cancer screening method based on transfer learning

Through the deep three-dimensional residual network and multi-scale feature enhancement mechanism, combined with the layer thickness-noise collaborative correction loss function and the adversarial domain adaptation module, the cross-domain adaptation problems of existing lung cancer screening methods in micronodules detection are solved, and more efficient early-stage lung cancer screening is achieved.

CN120544852AInactive Publication Date: 2025-08-26HUNAN UNIV OF SCI & ENG
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510657145.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-08-26
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

When existing lung cancer screening methods deal with micronodules with small size, blurred boundaries or strong heterogeneity, it is difficult to effectively characterize three-dimensional morphological characteristics, and the differences in cross-domain distribution affect the generalization performance of the model, resulting in unstable detection results.

Method used

The deep three-dimensional residual network is adopted to combine multi-scale feature enhancement mechanism and transfer learning strategy, and the loss function optimization model is optimized by layer thickness-noise collaborative correction, and an adversarial domain adaptive module is embedded to achieve cross-domain robustness and accuracy improvement.

Benefits of technology

The robustness and accuracy of the model for micronodules detection is significantly improved, the ability to discriminate weak boundary areas is enhanced, and the adaptability between different data domains is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120544852A_ABST
    Figure CN120544852A_ABST
Patent Text Reader

Abstract

The invention discloses an early lung cancer screening method based on transfer learning. The method comprises the following steps: S1, obtaining preprocessed three-dimensional volume data; s2, generating a multi-scale enhanced feature map; s3, obtaining a preliminary micro-nodule detection result; s4, obtaining a first-stage optimization model; s5, performing second-stage training on the first-stage optimization model by adopting a double-normal-form transfer learning strategy to obtain a transfer optimization model; s6, embedding an adversarial domain adaptive module in the migration optimization model to obtain a cross-domain robust model; and S7, in a clinical reasoning stage, outputting a micro-nodule detection result for the new chest low-dose spiral CT original image data by using the cross-domain robust model. According to the invention, the model can adaptively emphasize the local gradient difference between the lung micronodule and the background texture, thereby improving the discrimination capability of the weak boundary region.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of lung cancer screening, and in particular to an early lung cancer screening method based on transfer learning. Background Art

[0002] With the widespread application of medical imaging technology, low-dose spiral CT has become an important means of early screening for lung cancer, especially small lung nodules. CT scanning can non-invasively obtain high-resolution three-dimensional lung images, thereby preliminarily locating and evaluating suspicious lesions. At present, most medical institutions rely on manual film reading and traditional computer-aided detection systems based on convolutional neural networks to identify and distinguish lung nodules. Although this has improved the detection rate of early lung cancer to a certain extent, it still has significant shortcomings in dealing with small, fuzzy-bounded or highly heterogeneous micronodules.

[0003] The current mainstream lung nodule detection methods mainly rely on two-dimensional slice-level feature extraction, ignoring the importance of volume information in spatial structure modeling, and it is difficult to effectively characterize the three-dimensional morphological characteristics of micronodules. At the same time, traditional deep models are easily disturbed by voxel resolution differences when dealing with multi-scale structural changes and inconsistent image layer thickness problems, resulting in distorted feature expression. In addition, in actual clinical applications, there are differences in the CT scanning protocols used by different medical institutions, resulting in obvious cross-domain distribution differences between imaging data, affecting the generalization performance of the model.

[0004] To alleviate the above problems, some studies have introduced residual network structures or attention mechanisms to significantly enhance the nodule area, but most have failed to combine layer thickness changes and noise distribution for joint optimization, which can easily lead to unstable detection results. On the other hand, existing transfer learning methods mostly focus on fine-tuning model parameters, failing to fully explore the information of gradient differences between the source domain and the target domain, and lack adaptability to small sample data scenarios.

[0005] Therefore, there is an urgent need for an improved method that integrates deep three-dimensional residual networks, multi-scale feature enhancement mechanisms and transfer learning strategies to enhance the robustness and accuracy of early lung cancer micronodule detection systems across different data domains. Summary of the Invention

[0006] One purpose of the present invention is to propose an early lung cancer screening method based on transfer learning. The present invention enables the model to adaptively emphasize the local gradient difference between lung micronodules and background texture, thereby improving the discrimination ability of weak boundary areas.

[0007] According to an embodiment of the present invention, a method for early lung cancer screening based on transfer learning includes the following steps:

[0008] S1. Acquire and preprocess raw chest low-dose spiral CT image data to obtain preprocessed three-dimensional volume data;

[0009] S2. Input the preprocessed 3D volume data into a deep residual network backbone structure, which is used to extract 3D volume features and output encoded feature maps, and perform significance weighting to generate multi-scale enhanced feature maps;

[0010] S3. Input the multi-scale enhanced feature map into the candidate generation branch and the authenticity discrimination branch respectively to generate nodule candidate boxes and nodule authenticity scores, and obtain preliminary micro-nodule detection results;

[0011] S4. Based on the preliminary micronodule detection results, a thickness-noise collaborative correction loss function is constructed. This function is then used to jointly optimize the deep residual network to obtain the first-stage optimization model.

[0012] S5. A dual-paradigm transfer learning strategy was used to train the first-stage optimized model in the second stage. Gradient complementary transfer fine-tuning was performed on the public lung nodule dataset and the target hospital small sample dataset to obtain the transfer optimized model.

[0013] S6. An adversarial domain adaptation module is embedded in the transfer optimization model. Through synchronous training, the differences in the distribution of low-dose spiral thoracic CT images from different centers are narrowed, resulting in a cross-domain robust model.

[0014] S7. In the clinical reasoning stage, the cross-domain robust model is used to output micronodule detection results for new chest low-dose spiral CT raw image data.

[0015] Optionally, the S1 includes the following steps:

[0016] S11. Collect chest low-dose spiral CT original image data D raw ;

[0017] S12. Perform dose normalization on the original chest low-dose spiral CT image data and calculate the voxel value I of each CT slice. k The mean and standard deviation of the pixel intensity of (x, y, z) are used to adjust the intensity of each CT slice voxel value according to the uniformly set target intensity range to obtain the normalized CT slice voxel value. All normalized CT slice voxel values ​​constitute the dose normalization result;

[0018] S13. Perform slice thickness consistency correction on the dose normalization results. Slice thickness consistency correction is used to eliminate spatial resolution deviation caused by inconsistent slice thickness under different CT scanning protocols, thereby generating slice thickness consistency reconstruction data.

[0019] S14. Perform pyramid voxel resampling on the thickness consistency reconstruction data, and generate multiple volume data with different resolutions by downsampling layer by layer. Each downsampling level corresponds to a sampling ratio, and all downsampled volume data sets constitute a multi-scale resolution volume P = {D (s)}, D (s) Represents preprocessed 3D volume data at a fixed downsampling rate.

[0020] Optionally, the S2 includes the following steps:

[0021] S21. Multi-scale resolution volume P = D (s) The input is fed into the deep residual network backbone structure, which uses an adaptive residual skip connection unit with lung micronodule scale sensitivity. The adaptive residual skip connection unit is gated by the scale sensitivity gating function G scale (s) Control the fusion of residual information at different scales;

[0022] S22. The multi-scale resolution volume is subjected to the first layer of 3D convolution operation of the deep residual network backbone structure to obtain the initial input feature map F (s) (x, y, z, c0), c0 is the initial channel index, using the scale-sensitive gating function G scale (s) Input feature map F for different scale levels s (s) (x, y, z, c) is adaptively weighted fused to generate a scale-enhanced fusion feature map F scale-fused (x,y,z,c);

[0023] S23. Scale-enhanced fusion feature map F scale-fused (x, y, z, c) is input to the 3D dilated convolution residual module, and feature extraction is performed at the spatial voxel position (x, y, z) and channel index c respectively. Each group of 3D dilated convolution outputs a feature map F atrous,i (x, y, z, c), each group of output feature maps are concatenated and fused in the channel dimension to form a dilated convolutional fusion feature map F atrous (x,y,z,c);

[0024] S24. The dilated convolution fusion feature map F output by the 3D dilated convolution residual module atrous (x, y, z, c) introduces the self-correction attention mechanism of lung micronodule features. The self-correction attention mechanism includes self-correction spatial attention weight and self-correction channel attention weight. The calculation of self-correction spatial attention weight and self-correction channel attention weight are based on the local gradient difference between the typical texture of lung micronodules and the background texture, so as to automatically highlight the significance of lung micronodule areas. The self-correction attention mechanism of lung micronodule features is used to adjust the feature map F atrous(x, y, z, c) is weighted by element-by-element saliency enhancement to obtain the final multi-scale enhanced feature map F enhanced (x,y,z,c):

[0025] F enhanced (x,y,z,c)=F atrous (x,y,z,c)·A SC-channel (c)·A SC-spatial (x,y,z);

[0026] Among them, F enhanced (x, y, z, c) represents the multi-scale enhanced feature map after the lung micronodule feature self-correction attention mechanism significantly enhanced, A SC-channel (c) is the self-correction channel attention weight, A SC-spatial (x,y,z) is the self-corrected spatial attention weight.

[0027] Optionally, S3 includes the following steps:

[0028] S31. Calculate the feature response at each spatial voxel position and channel index in the multi-scale enhanced feature map to obtain a complete feature representation of the multi-scale enhanced feature map, and input the complete feature representation of the multi-scale enhanced feature map into a joint nodule candidate generation and evolution discrimination module, which includes a stage-aware candidate evolution module and a dynamic authenticity estimator;

[0029] S32. The computation phase perceives the feature response at each spatial position in the candidate evolution module, generates an initial nodule candidate frame set, and selects the nodule candidate frame B in the initial nodule candidate frame set. k The corresponding three-dimensional receptive field area is extracted from the position in the multi-scale enhanced feature map as the feature block R k , the feature block R k The similarity is compared with the lung micronodule stage feature vector, and the similarity comparison result is used as the candidate evolution confidence η k ;

[0030] S33. Calculate the spatial overlap between each pair of nodule candidate frames and the stage weighting factor λ of each nodule candidate frame k , taking spatial overlap and stage weighting factor as joint criteria, performing medical risk-oriented non-maximum suppression operation, eliminating nodule candidate frames whose spatial overlap exceeds the set threshold and whose stage weighting factor is lower than the threshold, and the retained nodule candidate frames constitute the refined nodule candidate frame set B refined ;

[0031] S34. Extract each nodule candidate box B in the refined nodule candidate box set jThe three-dimensional RoI area in the multi-scale enhanced feature map constitutes the corresponding RoI feature block R j , the RoI feature block R j The residual information of the upper and lower levels corresponding to the region in the deep residual network backbone structure is spliced ​​and fused to form a dynamic feature representation, which is input into the attention-guided mapping network to obtain the nodule candidate box B. j Authenticity score of lung micronodules j ;

[0032] S35. Match and combine the refined nodule candidate frame set with the corresponding lung micronodule authenticity score set to generate the preliminary micronodule detection result R initial .

[0033] Optionally, the S4 includes the following steps:

[0034] S41. The preliminary micronodule detection results R initial Serves as the basic input for constructing the thickness-noise collaborative correction loss function;

[0035] S42. For each nodule candidate box B j , extract the three-dimensional RoI feature block R in the multi-scale enhanced feature map j Obtain the corresponding dynamic feature representation, combine the layer thickness parameters and noise intensity estimation of the original CT image data, and jointly participate in the layer thickness-noise collaborative correction loss function The construction of

[0036] S43. Loss function for thickness-noise collaborative correction , set the layer thickness sensitivity Layer thickness sensitivity According to the nodule candidate box B j Calculation of the normalized error between the predicted center coordinates and the true center coordinates in the tangential direction;

[0037] S44. Loss function for thickness-noise collaborative correction In the noise suppression settings, set the noise suppression Noise suppression term According to each nodule candidate box B j Corresponding dynamic feature representation Calculation of structural similarity differences between the feature representations reconstructed with the denoised ones;

[0038] S45. The final thickness-noise collaborative correction loss function is obtained by linearly combining the layer thickness sensitivity term and the noise suppression term according to the weighted coefficient

[0039] S46. The layer thickness-noise collaborative correction loss function is used to perform end-to-end parameter update on the deep residual network, and finally the first-stage optimization model is obtained.

[0040] Optionally, the S5 includes the following steps:

[0041] S51. Use the first-stage optimized model as the initialization model for the second-stage training, and construct a dual-paradigm transfer learning training process. The dual-paradigm transfer learning training process includes a general feature pre-training stage and a target domain gradient complementary transfer stage.

[0042] S52. In the general feature pre-training stage, the natural image training set is used to migrate the deep residual network backbone structure in the first stage optimization model, and the image classification loss function is minimized. Optimize low-level convolution parameters to form a migration optimization model M pre ;

[0043] S53. In the target domain gradient complementary migration stage, a joint training set D is constructed, which includes the public lung nodule dataset and the target hospital small sample dataset. joint =D public ∪D target , where D public To provide a fully annotated public lung nodule dataset, D target This is a small sample dataset of the target hospital with a sample size of less than 100 cases;

[0044] S54. Calculate the public lung nodule dataset and the target hospital small sample dataset in the migration optimization model M respectively. pre The loss gradient on and Constructing gradient complementary weight ω comp Used to guide gradient merging strategy:

[0045]

[0046] S55. According to the gradient complementary weight ω comp Constructing weighted transfer optimization objective function

[0047]

[0048] S56. During the target domain training process, the structures of the nodule candidate generation branch and the authenticity discrimination branch remain unchanged, and only the authenticity score of the lung micronodules in the candidate frame is scored. j Introducing a lightweight target domain normalization adjustment factor δ norm The adjustment factor is dynamically adjusted according to the overall score distribution of the target hospital samples to define the updated lung micronodule authenticity score. where δ normUsed to suppress cross-domain rating deviation;

[0049] S57. The transfer optimization model obtained after the general feature pre-training stage and the target domain gradient complementary transfer stage is defined as the final transfer optimization model M trans .

[0050] Optionally, the S6 includes the following steps:

[0051] S61. In the migration optimization model M trans The adversarial domain adaptation module is embedded in the model, which includes a feature encoder, a gradient reversal layer, and a domain classifier.

[0052] S62. Input CT images from the public lung nodule dataset and the target hospital small sample dataset into the forward structure of the transfer optimization model to extract mid- and high-level shared feature representations.

[0053] S63. Input the mid- and high-level shared feature representations into the feature encoder in the adversarial domain adaptation module to obtain domain discriminant feature representations. Input the domain discriminant feature representations into the gradient reversal layer, outputting the reversal features. Input the reversal features into the domain classifier for discrimination to obtain domain discriminant probability.

[0054] S64. Define adversarial domain classification loss function It is used to measure the accuracy of the domain classifier in distinguishing the domain origin of the sample and drive the feature extraction process to produce cross-domain inseparable features. The domain classification loss function It is defined in the form of binary cross entropy loss;

[0055] S65. Domain Classification Loss Function and migration optimization objective function Together they form the cross-domain joint optimization objective function

[0056]

[0057] Among them, λ adv is the weighting coefficient of the domain adversarial loss in the joint optimization objective;

[0058] S66. Adopting cross-domain joint optimization objective function The deep residual network backbone structure, multi-level residual dense fusion module, multi-scale attention enhancement module, feature encoder and domain classifier are trained end-to-end to obtain the final cross-domain robust model M robust .

[0059] Optionally, the S7 includes the following steps:

[0060] S71. During the clinical reasoning phase, the trained cross-domain robust model is fed with the original low-dose spiral CT chest image data of the patient to be tested. The model then performs a full-process reasoning operation to obtain candidate micronodule detection results, including spatial coordinates, structural dimensions, classification scores, and significant regional thermal distribution.

[0061] S72. Locate the candidate nodule area output by the cross-domain robust model and obtain the spatial position coordinates corresponding to each micronodule

[0062] S73. Extract structural size information for each micronodule candidate area, including the long diameter size With short diameter size

[0063] S74. Assign a malignancy probability score to each candidate nodule region based on the updated lung micronodule authenticity score of the domain classifier, wherein the malignancy probability score corresponds to the lung micronodule authenticity score in a one-to-one manner;

[0064] S75. Generate interpretable heatmaps based on back-propagation gradients of high-order convolutional channels in a cross-domain robust model.

[0065] The beneficial effects of the present invention are:

[0066] (1) The present invention introduces a layer thickness-noise collaborative correction loss function, which embeds the scanning layer thickness and image noise intensity parameters in the original CT image into the loss function design, thereby achieving adaptive optimization of image resolution inconsistency and quality differences during the training process. The loss function contains two parts: a layer thickness sensitivity term and a noise suppression term. The former performs normalized error measurement based on the prediction deviation of the candidate nodule center point, and the latter measures the feature denoising effect through the structural similarity index, which effectively improves the robustness and accuracy of the detection model under low-dose CT data.

[0067] (2) The present invention proposes a joint modeling strategy of multi-scale enhanced residual backbone structure and self-correction attention mechanism of lung micronodule features to improve the model's recognition ability of the fuzzy edge areas of small nodules. Multi-scale volume input is generated by multi-level voxel downsampling, and the fusion strength of features at different levels is controlled by scale-sensitive gating function, which significantly enhances the model's perception of structures of different granularities. At the same time, the self-correction space and channel attention mechanism are combined to perform element-by-element significance weighting on the feature map, so that the model can adaptively emphasize the local gradient difference between lung micronodules and background textures, thereby improving the discrimination ability of weak boundary areas.

[0068] (3) The present invention introduces a cross-domain optimization strategy that combines dual-paradigm transfer learning with adversarial domain adaptation, which significantly improves the generalization ability and robustness of the model on small sample data sets from different hospitals. In the first stage of migration, natural image tasks are used for general feature pre-training. In the second stage, small samples of the target domain and public large data sets are introduced to form a joint training set. The learning signals of the two are dynamically fused through the gradient complementation strategy. On this basis, an adversarial domain classifier is embedded, and the gradient reversal mechanism is used to force the distribution of mid- and high-level features to approach the domain-indistinguishable direction, thereby achieving feature consistency of cross-center CT images. The score offset problem is significantly alleviated through the joint optimization of domain classification loss and task loss. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:

[0070] Figure 1 This is a flowchart of an early lung cancer screening method based on transfer learning proposed in the present invention. DETAILED DESCRIPTION

[0071] The present invention will now be described in further detail with reference to the accompanying drawings, which are simplified schematic diagrams that illustrate the basic structure of the present invention in a schematic manner.

[0072] refer to Figure 1 , an early lung cancer screening method based on transfer learning, comprising the following steps:

[0073] S1. Acquire and preprocess raw chest low-dose spiral CT image data to obtain preprocessed three-dimensional volume data;

[0074] S2. Input the preprocessed 3D volume data into a deep residual network backbone structure, which is used to extract 3D volume features and output encoded feature maps, and perform significance weighting to generate multi-scale enhanced feature maps;

[0075] S3. Input the multi-scale enhanced feature map into the candidate generation branch and the authenticity discrimination branch respectively to generate nodule candidate boxes and nodule authenticity scores, and obtain preliminary micro-nodule detection results;

[0076] S4. Based on the preliminary micronodule detection results, a thickness-noise collaborative correction loss function is constructed. This function is then used to jointly optimize the deep residual network to obtain the first-stage optimization model.

[0077] S5. A dual-paradigm transfer learning strategy was used to train the first-stage optimized model in the second stage. Gradient complementary transfer fine-tuning was performed on the public lung nodule dataset and the target hospital small sample dataset to obtain the transfer optimized model.

[0078] S6. An adversarial domain adaptation module is embedded in the transfer optimization model. Through synchronous training, the differences in the distribution of low-dose spiral thoracic CT images from different centers are narrowed, resulting in a cross-domain robust model.

[0079] S7. In the clinical reasoning stage, the cross-domain robust model is used to output micronodule detection results for new chest low-dose spiral CT raw image data.

[0080] In this embodiment, S1 includes the following steps:

[0081] S11. Collect chest low-dose spiral CT original image data D raw The original image data of low-dose spiral CT of the chest is represented as a set of CT slice voxels in three-dimensional format. The original image data of low-dose spiral CT of the chest consists of multi-layer CT slice voxel values ​​I k (x, y, z), where k represents the k-th CT slice and (x, y, z) represents the spatial voxel position of the CT slice voxel in three-dimensional space;

[0082] S12. Perform dose normalization on the original chest low-dose spiral CT image data. Dose normalization is used to unify the pixel intensity scale of all CT slices under different scanning conditions, enhance the model's adaptability to different image grayscale distributions, and calculate the voxel value I of each CT slice. k The pixel intensity mean and standard deviation of (x, y, z) are used to adjust the intensity of each CT slice voxel value according to the uniformly set target intensity range to obtain the normalized CT slice voxel value All normalized CT slice voxel values ​​constitute the dose normalization results;

[0083] S13. Perform slice thickness consistency correction on the dose normalization results. Slice thickness consistency correction is used to eliminate spatial resolution deviations caused by inconsistent slice thicknesses under different CT scanning protocols. Linear interpolation is used to generate the voxel values ​​of the intermediate CT slices between any two consecutive slices. The interpolation weight is calculated based on the relative ratio of the current target position to the original slice position, ultimately generating slice thickness consistency reconstruction data.

[0084] S14. Perform pyramid voxel resampling on the thickness consistency reconstruction data. Pyramid voxel resampling is used to construct a three-dimensional volume representation with multiple spatial resolutions, thereby enhancing the network's ability to represent lung structures at different scales. Multiple volume data with different resolutions are generated by downsampling layer by layer. Each downsampling level corresponds to a sampling ratio. The set of all downsampled volume data constitutes a multi-scale resolution volume P = {D (s)}, D (s) Represents preprocessed 3D volume data at a fixed downsampling rate.

[0085] In this embodiment, S2 includes the following steps:

[0086] S21. Multi-scale resolution volume P = D (s) The input is fed into the deep residual network backbone structure, which uses an adaptive residual skip connection unit with lung micronodule scale sensitivity to achieve adaptive selective fusion for lung micronodule features of different scales. The adaptive residual skip connection unit is gated by the scale sensitivity gating function G scale (s) controls the fusion of residual information at different scales, and the scale-sensitive gating function is dynamically adjusted according to the typical diameter range of lung micronodules corresponding to the scale level s;

[0087] S22. The multi-scale resolution volume is subjected to the first layer of 3D convolution operation of the deep residual network backbone structure to obtain the initial input feature map F (s) (x, y, z, c0), c0 is the initial channel index, using the scale-sensitive gating function G scale (s) Input feature map F for different scale levels s (s) (x, y, z, c) is adaptively weighted fused to generate a scale-enhanced fusion feature map F scale-fused (x,y,z,c):

[0088]

[0089] Where S represents the total number of scale levels in the multi-scale resolution volume, and c is the channel index;

[0090] The scale-sensitive gating function is based on the typical diameter d of lung micronodules at scale level s. s Dynamic determination, when the typical diameter of lung micronodules d s When the size is in the range of 3-5 mm, the weight of the scale-sensitive gating function is adaptively increased to optimize the network's ability to express the scale characteristics of early micronodules.

[0091] The scale-enhanced fusion feature map formula in S22 performs weighted integration of feature information from multi-scale resolution volume data. Specifically, for input volumes at different scale levels, the system assigns adaptive weights to the initial input feature maps at each scale through a scale-sensitive gating function. The scale-sensitive gating function is dynamically adjusted according to the typical diameter of lung micronodules at each scale level, thereby achieving scale response enhancement of tiny structures such as 3-5mm lung micronodules in the feature space. After all scales are weighted, the generated fusion feature map comprehensively expresses multi-scale information and improves the network's ability to resolve small-sized structures.

[0092] S23. Scale-enhanced fusion feature map F scale-fused (x, y, z, c) is input to the 3D dilated convolution residual module. The 3D dilated convolution residual module includes N groups of parallel 3D dilated convolution units. Each group of dilated convolution units uses a different dilation rate and performs feature extraction at the spatial voxel position (x, y, z) and channel index c. Each group of 3D dilated convolution outputs a feature map F atrous,i (x, y, z, c), each group of output feature maps are concatenated and fused in the channel dimension to form a dilated convolutional fusion feature map F atrous (x, y, z, c), the hollow sampling interval with lung micronodule feature extraction specificity captures contextual information in a larger spatial range, while avoiding the weakening of the edge texture of small-sized lung micronodules caused by the excessive receptive field of CT slice voxel features;

[0093] S24. The dilated convolution fusion feature map F output by the 3D dilated convolution residual module atrous (x, y, z, c) introduces the self-correction attention mechanism of lung micronodule features. The self-correction attention mechanism includes self-correction spatial attention weight and self-correction channel attention weight. The calculation of self-correction spatial attention weight and self-correction channel attention weight are based on the local gradient difference between the typical texture of lung micronodules and the background texture, so as to automatically highlight the significance of lung micronodule areas. The self-correction attention mechanism of lung micronodule features is used to adjust the feature map F atrous (x, y, z, c) is weighted by element-by-element saliency enhancement to obtain the final multi-scale enhanced feature map F enhanced (x,y,z,c):

[0094] F enhanced (x,y,z,c)=F atrous (x,y,z,c)·A SC-channel (c)·A SC-spatial (x,y,z);

[0095] Among them, F enhanced(x, y, z, c) represents the multi-scale enhanced feature map after the lung micronodule feature self-correction attention mechanism significantly enhanced, A SC-channel (c) is the self-corrected channel attention weight, which is dynamically determined by the channel domain gradient difference between the typical texture of lung micronodules and the background texture. SC-spatial (x, y, z) is the self-corrected spatial attention weight, which is dynamically determined by the local spatial gradient significance of lung micronodules.

[0096] S24's multi-scale attention saliency enhancement formula applies saliency weight adjustment to each element in the dilated convolution fusion feature map through self-correcting spatial attention weights and self-correcting channel attention weights. The spatial attention weights reflect the saliency of local gradients and textures at each spatial position, while the channel attention weights reflect the discriminative contribution of each feature channel to distinguishing lung micronodules from background tissues. The combined effect of the two enables the network to automatically focus on spatial positions and semantic channels with micronodule structural characteristics, achieving element-by-element weighted enhancement of the feature map.

[0097] In this embodiment, S3 includes the following steps:

[0098] S31. Calculate the feature response at each spatial voxel position and channel index in the multi-scale enhanced feature map to obtain a complete feature representation of the multi-scale enhanced feature map, and input the complete feature representation of the multi-scale enhanced feature map into a joint nodule candidate generation and evolution discrimination module. The joint nodule candidate generation and evolution discrimination module includes a stage-aware candidate evolution module and a dynamic authenticity estimator. The stage-aware candidate evolution module is used to complete the generation of micro-nodule spatial candidate positions, and the dynamic authenticity estimator is used to estimate the authenticity score of the micro-nodule candidate area.

[0099] S32. The computation phase perceives the feature response at each spatial position in the candidate evolution module and generates an initial nodule candidate frame set. Each nodule candidate frame B in the initial nodule candidate frame set k Including the horizontal center offset, longitudinal center offset, tangential center offset, major diameter size and minor diameter size, the nodule candidate box represents the rough position and size of the micronodule in three-dimensional space. k In the multi-scale enhanced feature map, the corresponding three-dimensional receptive field area is extracted as the feature block R k , the feature block R k The similarity is compared with the lung micronodule stage feature vector, and the similarity value is used to measure the nodule candidate box B k Whether it has the typical morphological structure of early lung micronodules, the similarity calculation result is used as the candidate evolution confidence η k , candidate evolution confidence η kThe higher the value, the more consistent the structure of the nodule candidate frame is with the eigenvector of the lung micronodule stage, and the earlier the lesion stage.

[0100] S33. Calculate the spatial overlap between each pair of nodule candidate frames. The spatial overlap is obtained by calculating the three-dimensional intersection-union ratio between the nodule candidate frames. Calculate the stage weighting factor λ for each nodule candidate frame. k , stage weighting factor λ k According to the candidate evolution confidence η k Dynamically assign a value to indicate the medical priority of each nodule candidate frame when it is retained. Taking spatial overlap and stage weighting factor as joint criteria, perform medical risk-oriented non-maximum suppression operation to eliminate nodule candidate frames whose spatial overlap exceeds the set threshold and whose stage weighting factor is lower than the threshold. The retained nodule candidate frames constitute the refined nodule candidate frame set B. refined ;

[0101] S34. Extract each nodule candidate box B in the refined nodule candidate box set j The three-dimensional RoI area in the multi-scale enhanced feature map constitutes the corresponding RoI feature block R j , and then the RoI feature block R j The residual information of the upper and lower levels corresponding to the region in the deep residual network backbone structure is spliced ​​and fused to form a dynamic feature representation, which is input into the attention-guided mapping network to obtain the nodule candidate box B. j Authenticity score of lung micronodules j , lung micronodule authenticity score j The value range of is [0,1], which is used to indicate the possibility that the nodule candidate box is an actual early lung micronodule. The larger the value, the higher the credibility.

[0102] S35. Combine the refined nodule candidate frame set with the corresponding lung micronodule authenticity score set Perform matching and combination to generate preliminary micronodule detection results R initial , each item in the preliminary micro-nodule detection result is represented by a nodule candidate box B j and authenticity scores j Composition, nodule candidate box B j Contains the three-dimensional space center position offset and major and minor diameter dimensions, authenticity score s j Indicates the probability that the area corresponding to the nodule candidate box is a lung micronodule.

[0103] In this embodiment, S4 includes the following steps:

[0104] S41. The preliminary micronodule detection results R initial Serves as the basic input for constructing the thickness-noise collaborative correction loss function;

[0105] S42. For each nodule candidate box B j , extract the three-dimensional RoI feature block R in the multi-scale enhanced feature map j , and then obtain the corresponding dynamic feature representation, combined with the layer thickness parameters and noise intensity estimation value of the CT original image data, and jointly participate in the layer thickness-noise collaborative correction loss function The construction of

[0106] S43. Loss function for thickness-noise collaborative correction , set the layer thickness sensitivity Layer thickness sensitivity According to the nodule candidate box B j The normalized error between the predicted center coordinates and the true center coordinates in the tangential direction is calculated. The normalization factor is the current CT slice thickness parameter. The expression is the sum of the square error of the tangential distance between the predicted and true micronodule centers and the current CT slice thickness parameter. It is used to reflect the robustness of the detection frame positioning under different slice thickness conditions.

[0107] S44. Loss function for thickness-noise collaborative correction In the noise suppression settings, set the noise suppression Noise suppression term According to each nodule candidate box B j Corresponding dynamic feature representation The structural similarity difference between the denoised and reconstructed feature representations is calculated, and the noise intensity estimation value of the corresponding area of ​​the CT image is combined to assign a higher penalty weight to the high-noise area, suppress the false feature response caused by high-frequency artifacts, and improve the authenticity score of micronodules. j Reliability under high noise conditions;

[0108] S45. The final thickness-noise collaborative correction loss function is obtained by linearly combining the layer thickness sensitivity term and the noise suppression term according to the weighted coefficient

[0109] S46. The layer thickness-noise collaborative correction loss function is applied to the entire training optimization process. The layer thickness-noise collaborative correction loss function is used to perform end-to-end parameter updates on the deep residual network backbone structure, multi-level residual dense fusion module, multi-scale attention enhancement module, nodule candidate generation branch and authenticity discrimination branch, and finally the first-stage optimization model is obtained.

[0110] The deep residual network backbone structure is used to perform preliminary feature encoding on the preprocessed 3D volume data. The backbone structure is constructed based on the scale-sensitive residual skip connection unit, and the gating function G is introduced for different scale levels s. scale(s), dynamically controls the degree of fusion of residual information between different scale paths, and the output of the deep residual network backbone structure is the scale-enhanced fusion feature map.

[0111] The multi-level residual dense fusion module is used to enhance the multi-receptive field structure expression capability based on the scale-enhanced fusion feature map. The multi-level residual dense fusion module contains N groups of parallel three-dimensional dilated convolution units with different dilation rates. Each group of output feature maps is spliced ​​in the channel dimension to obtain a fused feature map, which can effectively retain the micronodule texture and spatial context information, and has an enhanced extraction effect on the heterogeneous areas of 3-5mm micronodules.

[0112] The multi-scale attention enhancement module is used to apply saliency enhancement processing to the fused feature map, including spatial attention weight and channel attention weight. The multi-scale attention enhancement module dynamically adjusts the response degree of the lung micronodule area in the spatial dimension and semantic channel through the self-correcting attention mechanism, and outputs the enhanced multi-scale feature map.

[0113] The nodule candidate generation branch is used to generate preliminary candidate areas for lung micronodule detection in the enhanced multi-scale feature map. The nodule candidate generation branch obtains a set of nodule candidate frames through the stage-aware candidate evolution module, and then calculates the candidate evolution confidence based on the lung micronodule stage feature vector. Finally, the refined nodule candidate frame set is obtained through screening using the medical risk-guided non-maximum suppression strategy.

[0114] The authenticity discrimination branch is used to evaluate the malignancy probability of each nodule area in the refined candidate frame and ultimately form a preliminary detection result.

[0115] In this embodiment, S5 includes the following steps:

[0116] S51. Use the first-stage optimized model as the initialization model for the second-stage training, and construct a dual-paradigm transfer learning training process. The dual-paradigm transfer learning training process includes a general feature pre-training stage and a target domain gradient complementary transfer stage.

[0117] S52. In the general feature pre-training stage, the natural image training set is used to migrate the deep residual network backbone structure in the first stage optimization model, and the image classification loss function is minimized. Optimize the low-level convolution parameters to enable the deep residual network to obtain universal visual structure recognition capabilities across categories, forming a migration optimization model M pre ;

[0118] S53. In the target domain gradient complementary migration stage, a joint training set D is constructed, which includes the public lung nodule dataset and the target hospital small sample dataset. joint =D public ∪D target , where Dpublic To provide a fully annotated public lung nodule dataset, D target This is a small sample dataset of the target hospital with a sample size of less than 100 cases;

[0119] S54. Calculate the public lung nodule dataset and the target hospital small sample dataset in the migration optimization model M respectively. pre The loss gradient on and Constructing gradient complementary weight ω comp Used to guide gradient merging strategy:

[0120]

[0121] The larger the gradient complementarity weight value is, the stronger the consistency of the gradient directions of the two domains is, and the higher the stability of the model migration is;

[0122] The S54 gradient complementary weight formula is used to measure the consistency of the loss gradient direction between the public dataset and the target hospital small sample dataset during transfer training. The principle is to calculate the difference between the gradient vectors of the two domain data and give the complementary weight in the form of a normalized ratio. The higher the weight value, the more consistent the gradient direction of the two domains, and the stronger the stability of transfer learning during network parameter optimization. Conversely, a low weight indicates a large transfer conflict, and it is necessary to reduce the influence of the target domain gradient and balance the learning process.

[0123] S55. According to the gradient complementary weight ω comp Constructing weighted transfer optimization objective function Used to balance the impact of public data and target data in training:

[0124]

[0125] The transfer optimization objective function As a training supervision signal, the parameters of the deep residual network backbone structure, multi-level residual dense fusion module, and multi-scale attention enhancement module are updated to improve the model's transfer robustness in the small sample scenario of the target hospital;

[0126] The weighted transfer optimization objective function formula of S55 linearly weights the loss functions of the public dataset and the small sample dataset of the target domain on the model according to the gradient complementary weights. During training, the influence of the public data and target data in the loss feedback is dynamically adjusted, and the weights change adaptively with the gradient direction differences. The weighted objectives ensure that parameter optimization can maintain the representation ability on the public large dataset while enhancing the adaptation and optimization of the small sample data in the target domain.

[0127] S56. During the target domain training process, the structures of the nodule candidate generation branch and the authenticity discrimination branch remain unchanged, and only the authenticity score of the lung micronodules in the candidate frame is scored.j Introducing a lightweight target domain normalization adjustment factor δ norm The adjustment factor is dynamically adjusted according to the overall score distribution of the target hospital samples to define the updated lung micronodule authenticity score. where δ norm Used to suppress cross-domain scoring deviation and enhance the consistency of the final detection results;

[0128] S57. The transfer optimization model obtained after the general feature pre-training stage and the target domain gradient complementary transfer stage is defined as the final transfer optimization model M trans .

[0129] In this embodiment, S6 includes the following steps:

[0130] S61. In the migration optimization model M trans The adversarial domain adaptation module is embedded in the model. The adversarial domain adaptation module includes a feature encoder, a gradient reversal layer, and a domain classifier. The feature encoder is used to extract mid- and high-level feature representations of the input low-dose spiral CT chest image. The gradient reversal layer is used to reverse the back gradient propagation of the features to the domain classifier. The domain classifier is used to discriminate the domain origin of the input sample features.

[0131] S62. CT images from the public lung nodule dataset and the target hospital small sample dataset are fed into the forward structure of the transfer optimization model to extract mid- and high-level shared feature representations. The mid- and high-level shared feature representations are generated by the deep residual network backbone and the multi-level residual dense fusion module.

[0132] S63. Input the mid- and high-level shared feature representations into the feature encoder in the adversarial domain adaptation module to obtain domain discriminant feature representations. Input the domain discriminant feature representations into the gradient reversal layer, outputting the reversal features. Input the reversal features into the domain classifier for discrimination to obtain domain discriminant probability.

[0133] S64. Define adversarial domain classification loss function It is used to measure the accuracy of the domain classifier in distinguishing the domain origin of the sample and drive the feature extraction process to produce cross-domain inseparable features. The domain classification loss function It is defined in the form of binary cross entropy loss:

[0134]

[0135] in, is the true domain label of the j-th sample, is the predicted value of its domain classification probability, and N is the total number of samples participating in domain adversarial training;

[0136] S65. Domain Classification Loss Function and migration optimization objective function Together they form the cross-domain joint optimization objective function

[0137]

[0138] Among them, λ adv is the weighting coefficient of the domain adversarial loss in the joint optimization objective;

[0139] The cross-domain joint optimization objective function formula of S65 combines the transfer optimization objective function and the adversarial domain classification loss function according to weighted coefficients. During end-to-end training, the formula optimizes the main task of lung micronodule detection through transfer loss on the one hand, and drives the network to actively narrow the feature distribution differences of images from different centers and different devices through adversarial loss on the other hand. The domain classifier and feature encoder work together adversarially to force the sample distribution in the feature space to be consistent, thereby improving the cross-domain generalization robustness of the model.

[0140] S66. Adopting cross-domain joint optimization objective function The deep residual network backbone structure, multi-level residual dense fusion module, multi-scale attention enhancement module, feature encoder and domain classifier are trained end-to-end, so that the transfer optimization model can retain the ability to identify lung micronodules while minimizing the feature distribution differences between chest low-dose spiral CT images from different centers and different equipment, and the final cross-domain robust model M is obtained. robust .

[0141] In this embodiment, S7 includes the following steps:

[0142] S71. During the clinical reasoning phase, the trained cross-domain robust model is fed with the original low-dose spiral CT chest image data of the patient to be tested. The model then performs a full-process reasoning operation to obtain candidate micronodule detection results, including spatial coordinates, structural dimensions, classification scores, and significant regional thermal distribution.

[0143] S72. Locate the candidate nodule area output by the cross-domain robust model and obtain the spatial position coordinates corresponding to each micronodule spatial position coordinates is the center position of the nodule candidate area;

[0144] S73. Extract structural size information for each micronodule candidate area, including the long diameter size With short diameter size Long diameter Indicates the maximum axial length of the nodule in the cross section, the short diameter Indicates its minimum axial length;

[0145] S74. Based on the updated lung micronodule authenticity score of the domain classifier, assign a malignancy probability score to each candidate nodule region. The malignancy probability score corresponds one-to-one with the lung micronodule authenticity score and represents the probability level of the currently detected nodule being malignant. A higher value indicates a greater likelihood of malignancy.

[0146] S75. Based on the back-propagation gradient of the high-order convolution channel of the cross-domain robust model, an interpretable heat map is generated. The interpretable heat map is used to visualize the saliency response intensity of the nodule candidate region in the spatial dimension.

[0147] Example 1:

[0148] The Radiology Department of a hospital in City A admitted a 42-year-old male patient surnamed Li, who was participating in this year's high-risk lung cancer screening program organized by his unit. The hospital arranged a low-dose spiral CT scan for him using a GE Revolution device with parameters set at 120 kV, 60 mA-second, and a slice thickness of 1.25 mm. The entire scan lasted 22 seconds and ultimately output 426 DICOM-format slice files with a total size of 362 megabytes.

[0149] After the scan is completed, the radiologist transmits this set of images through the hospital intranet to the newly deployed artificial intelligence-assisted lung cancer screening system. The system is deployed on the central server of the imaging department and is equipped with two high-performance graphics processing units to support the reasoning process of the deep learning model. After the image is successfully uploaded, the system automatically starts the preprocessing process.

[0150] The system first performs image intensity normalization, calculates the pixel average value of each CT slice to be 312, and the standard deviation to be 47.6, judging that the image brightness is too high. Then, it performs grayscale compression and stretching adjustments according to the preset target intensity range to make the overall image contrast clearer. Then, the system detects that the layer thickness of the data is 1.25 mm, which exceeds the system reference standard of 0.75 mm. Therefore, it automatically starts the layer thickness consistency correction program to construct multi-scale resampled volume data to eliminate the characteristic errors caused by inconsistent spatial resolution.

[0151] After image normalization, the system enters the feature extraction phase, loading its proprietary deep residual network structure and sequentially analyzing image data at different scales. Unlike traditional methods that only process two-dimensional slices, this system uses three-dimensional voxel analysis to accurately capture tissue boundary changes within the three-dimensional structure. In the right upper lobe, near the oblique fissure, the system noticed a small, low-grayscale region with distinct boundary features. It automatically labeled it as a suspected micronodule candidate and assigned the number ROI_008734.

[0152] The system further conducts a detailed analysis of the area. First, the system uses the attention mechanism module to compare the texture of the area with the surrounding background and finds that the local gradient variation of the suspected area is significantly higher than the average level. Therefore, the area is given a higher channel and spatial significance score. The comprehensive score exceeds the risk threshold preset by the system. The system automatically enters it into the candidate nodule list and calls the authenticity discriminant sub-model for dynamic analysis.

[0153] During the authenticity scoring stage, the system analyzed the upper and lower level feature expressions of the candidate nodule, combined with its spatial position, morphological stability, and boundary sharpness indicators, and output a 56% probability that it was a benign micronodule, and suggested follow-up observation. The entire analysis process was completed in less than two minutes, and the system generated a complete nodule detection report at 9:42.

[0154] As a control, the hospital also input the patient's image into the traditional two-dimensional U-Net model for analysis. The model failed to identify the nodule area and judged it as negative. The hospital also tried to use a standard V-Net structure model, which took a long time to run and output multiple false positive areas, requiring manual investigation for confirmation. In comparison, the system proposed in the present invention not only accurately identified the suspicious area, but also reduced a large number of false positives in non-target areas, thereby improving work efficiency.

[0155] To evaluate the overall effectiveness of the system, the hospital collected data from 168 patients as clinical evaluation samples, including 120 first-time physical examination patients and 48 patients with a history of chronic lung disease. The data came from three CT devices from different manufacturers, with layer thickness settings ranging from 0.625 mm to 1.5 mm, and there were significant differences in image quality.

[0156] When using the traditional two-dimensional model to analyze this batch of data, the micronodule detection rate was 68% and the false positive rate was about 22%. After adopting the system of the present invention, the micronodule detection rate was increased to 91% and the false positive rate was reduced to 6%. Especially in scanning images with severe noise and thick layers, the system can still stably output detection results with clear structure and accurate position, which is significantly better than traditional methods.

[0157] To verify the system's generalization capability across hospitals, the research team introduced 200 open-source lung nodule annotation data as pre-training data and conducted secondary migration training on a local small sample dataset. While maintaining the stability of the backbone parameters, the system effectively eliminated cross-device and cross-regional discrimination bias by dynamically adjusting the classification score distribution of target hospital samples. The final model's evaluation indicators on the full test set showed an AUC value improvement of nearly 10%, demonstrating good domain adaptability.

[0158] More importantly, the system generates a visual heat map after each inference, showing doctors the areas the model focuses on and their significance levels. In actual use, the average reading time for doctors has been shortened from three minutes to one minute and fifteen seconds, greatly improving diagnostic efficiency. At the same time, the system provides five types of auxiliary information for each suspected area, including structural size, spatial position, malignancy probability, and significance score, to help doctors make clinical judgments.

[0159] Patient Li was re-examined three months later, and no obvious changes were found in the nodule, which was later confirmed to be a benign fibrous nodule. This event verified the practical effectiveness of the present invention in early lung cancer screening, while also avoiding unnecessary excessive medical intervention.

[0160] The present invention introduces a layer thickness-noise collaborative correction loss function, embeds the scanning layer thickness and image noise intensity parameters in the original CT image into the loss function design, thereby achieving adaptive optimization of image resolution inconsistency and quality differences during the training process. The loss function contains two parts: a layer thickness sensitivity term and a noise suppression term. The former performs normalized error measurement based on the prediction deviation of the candidate nodule center point, and the latter measures the feature denoising effect through the structural similarity index, which effectively improves the robustness and accuracy of the detection model under low-dose CT data.

[0161] This paper proposes a joint modeling strategy of multi-scale enhanced residual backbone structure and self-correction attention mechanism of lung micronodule features to improve the model's recognition ability of the blurred edge areas of small nodules. Multi-scale volume input is generated through multi-level voxel downsampling, and the fusion strength of features at different levels is controlled with the help of scale-sensitive gating function, which significantly enhances the model's perception of structures of different granularities. At the same time, the self-correction space and channel attention mechanism are combined to perform element-by-element significance weighting on the feature map, so that the model can adaptively emphasize the local gradient difference between lung micronodules and background textures, thereby improving the discrimination ability of weak boundary areas.

[0162] The present invention introduces a cross-domain optimization strategy that combines dual-paradigm transfer learning with adversarial domain adaptation, which significantly improves the generalization ability and robustness of the model on small sample data sets from different hospitals. In the first stage of migration, natural image tasks are used for general feature pre-training. In the second stage, small samples of the target domain and public large data sets are introduced to form a joint training set. The learning signals of the two are dynamically fused through the gradient complementarity strategy. On this basis, an adversarial domain classifier is embedded, and the gradient reversal mechanism is used to force the distribution of mid- and high-level features to approach the domain-indistinguishable direction, thereby achieving feature consistency of cross-center CT images. The score offset problem is significantly alleviated through the joint optimization of domain classification loss and task loss.

[0163] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.

Claims

1. A method for early lung cancer screening based on transfer learning, characterized in that: The steps include: S1. Acquire and preprocess raw chest low-dose spiral CT image data to obtain preprocessed three-dimensional volume data; S2. Input the preprocessed 3D volume data into a deep residual network backbone structure, which is used to extract 3D volume features and output encoded feature maps, and perform significance weighting to generate multi-scale enhanced feature maps; S3. Input the multi-scale enhanced feature map into the candidate generation branch and the authenticity discrimination branch respectively to generate nodule candidate boxes and nodule authenticity scores, and obtain preliminary micro-nodule detection results; S4. Based on the preliminary micronodule detection results, a thickness-noise collaborative correction loss function is constructed. This function is then used to jointly optimize the deep residual network to obtain the first-stage optimization model. S5. A dual-paradigm transfer learning strategy was used to train the first-stage optimized model in the second stage. Gradient complementary transfer fine-tuning was performed on the public lung nodule dataset and the target hospital small sample dataset to obtain the transfer optimized model. S6. An adversarial domain adaptation module is embedded in the transfer optimization model. Through synchronous training, the differences in the distribution of low-dose spiral thoracic CT images from different centers are narrowed, resulting in a cross-domain robust model. S7. In the clinical reasoning stage, the cross-domain robust model is used to output micronodule detection results for new chest low-dose spiral CT raw image data.

2. The method for early lung cancer screening based on transfer learning according to claim 1, characterized in that: Said S1 comprises the following steps: S11. Collect chest low-dose spiral CT original image data D raw ; S12. Perform dose normalization on the original chest low-dose spiral CT image data and calculate the voxel value I of each CT slice. k The mean and standard deviation of the pixel intensity of (x, y, z) are used to adjust the intensity of each CT slice voxel value according to the uniformly set target intensity range to obtain the normalized CT slice voxel value. All normalized CT slice voxel values ​​constitute the dose normalization result; S13. Perform slice thickness consistency correction on the dose normalization results. Slice thickness consistency correction is used to eliminate spatial resolution deviation caused by inconsistent slice thickness under different CT scanning protocols, thereby generating slice thickness consistency reconstruction data. S14. Perform pyramid voxel resampling on the thickness consistency reconstruction data, and generate multiple volume data with different resolutions by downsampling layer by layer. Each downsampling level corresponds to a sampling ratio, and all downsampled volume data sets constitute a multi-scale resolution volume P = {D (s) }, D (s) Represents preprocessed 3D volume data at a fixed downsampling rate.

3. The method for early lung cancer screening based on transfer learning according to claim 2, characterized in that: The S2 comprises the following steps: S21. Multi-scale resolution volume P = D (s) The input is fed into the deep residual network backbone structure, which uses an adaptive residual skip connection unit with lung micronodule scale sensitivity. The adaptive residual skip connection unit is gated by the scale sensitivity gating function G scale (s) Control the fusion of residual information at different scales; S22. The multi-scale resolution volume is subjected to the first layer of 3D convolution operation of the deep residual network backbone structure to obtain the initial input feature map F (s) (x, y, z, c0), c0 is the initial channel index, using the scale-sensitive gating function G scale (s) Input feature map F for different scale levels s (s) (x, y, z, c) is adaptively weighted fused to generate a scale-enhanced fusion feature map F scale-fused (x,y,z,c); S23. Scale-enhanced fusion feature map F scale-fused (x, y, z, c) is input to the 3D dilated convolution residual module, and feature extraction is performed at the spatial voxel position (x, y, z) and channel index c respectively. Each group of 3D dilated convolution outputs a feature map F atrous,i (x, y, z, c), each group of output feature maps are concatenated and fused in the channel dimension to form a dilated convolutional fusion feature map F atrous (x,y,z,c); S24. The dilated convolution fusion feature map F output by the 3D dilated convolution residual module atrous (x, y, z, c) introduces the self-correction attention mechanism of lung micronodule features. The self-correction attention mechanism includes self-correction spatial attention weight and self-correction channel attention weight. The calculation of self-correction spatial attention weight and self-correction channel attention weight are based on the local gradient difference between the typical texture of lung micronodules and the background texture, so as to automatically highlight the significance of lung micronodule areas. The self-correction attention mechanism of lung micronodule features is used to adjust the feature map F atrous (x, y, z, c) is weighted by element-by-element saliency enhancement to obtain the final multi-scale enhanced feature map F enhanced (x,y,z,c): F enhanced (x,y,z,c)=F atrous (x,y,z,c)·A SC-channel (c)·A SC-spatial (x,y,z); Among them, F enhanced (x, y, z, c) represents the multi-scale enhanced feature map after the lung micronodule feature self-correction attention mechanism significantly enhanced, A SC-channel (c) is the self-correction channel attention weight, A SC-spatial (x,y,z) is the self-corrected spatial attention weight.

4. The method for early lung cancer screening based on transfer learning according to claim 3, characterized in that: The S3 includes the following steps: S31. Calculate the feature response at each spatial voxel position and channel index in the multi-scale enhanced feature map to obtain a complete feature representation of the multi-scale enhanced feature map, and input the complete feature representation of the multi-scale enhanced feature map into a joint nodule candidate generation and evolution discrimination module, which includes a stage-aware candidate evolution module and a dynamic authenticity estimator; S32. The computation phase perceives the feature response at each spatial position in the candidate evolution module, generates an initial nodule candidate frame set, and selects the nodule candidate frame B in the initial nodule candidate frame set. k The corresponding three-dimensional receptive field area is extracted from the position in the multi-scale enhanced feature map as the feature block R k , the feature block R k The similarity is compared with the lung micronodule stage feature vector, and the similarity comparison result is used as the candidate evolution confidence η k ; S33. Calculate the spatial overlap between each pair of nodule candidate frames and the stage weighting factor λ of each nodule candidate frame k , taking spatial overlap and stage weighting factor as joint criteria, performing medical risk-oriented non-maximum suppression operation, eliminating nodule candidate frames whose spatial overlap exceeds the set threshold and whose stage weighting factor is lower than the threshold, and the retained nodule candidate frames constitute the refined nodule candidate frame set B refined ; S34. Extract each nodule candidate box B in the refined nodule candidate box set j The three-dimensional RoI area in the multi-scale enhanced feature map constitutes the corresponding RoI feature block R j , the RoI feature block R j The residual information of the upper and lower levels corresponding to the region in the deep residual network backbone structure is spliced ​​and fused to form a dynamic feature representation, which is input into the attention-guided mapping network to obtain the nodule candidate box B. j Authenticity score of lung micronodules j ; S35. Match and combine the refined nodule candidate frame set with the corresponding lung micronodule authenticity score set to generate the preliminary micronodule detection result R initial .

5. The method for early lung cancer screening based on transfer learning according to claim 4, characterized in that: The S4 comprises the following steps: S41. The preliminary micronodule detection results R initial Serves as the basic input for constructing the thickness-noise collaborative correction loss function; S42. For each nodule candidate box B j , extract the three-dimensional RoI feature block R in the multi-scale enhanced feature map j Obtain the corresponding dynamic feature representation, combine the layer thickness parameters and noise intensity estimation of the original CT image data, and jointly participate in the layer thickness-noise collaborative correction loss function The construction of S43. Loss function for thickness-noise collaborative correction , set the layer thickness sensitivity Layer thickness sensitivity According to the nodule candidate box B j Calculation of the normalized error between the predicted center coordinates and the true center coordinates in the tangential direction; S44. Loss function for thickness-noise collaborative correction In the noise suppression settings, set the noise suppression Noise suppression term According to each nodule candidate box B j Corresponding dynamic feature representation Calculation of structural similarity differences between the feature representations reconstructed with the denoised ones; S45. The final thickness-noise collaborative correction loss function is obtained by linearly combining the layer thickness sensitivity term and the noise suppression term according to the weighted coefficient S46. The layer thickness-noise collaborative correction loss function is used to perform end-to-end parameter update on the deep residual network, and finally the first-stage optimization model is obtained.

6. The method for early lung cancer screening based on transfer learning according to claim 5, characterized in that: The S5 comprises the following steps: S51. Use the first-stage optimized model as the initialization model for the second-stage training, and construct a dual-paradigm transfer learning training process. The dual-paradigm transfer learning training process includes a general feature pre-training stage and a target domain gradient complementary transfer stage. S52. In the general feature pre-training stage, the natural image training set is used to migrate the deep residual network backbone structure in the first stage optimization model, and the image classification loss function is minimized. Optimize low-level convolution parameters to form a migration optimization model M pre ; S53. In the target domain gradient complementary migration stage, a joint training set D is constructed, which includes the public lung nodule dataset and the target hospital small sample dataset. joint =D public ∪D target , where D public To provide a well-annotated public lung nodule dataset, D target This is a small sample dataset of the target hospital with a sample size of less than 100 cases; S54. Calculate the public lung nodule dataset and the target hospital small sample dataset in the migration optimization model M respectively. pre The loss gradient on and Constructing gradient complementary weight ω comp Used to guide gradient merging strategy: S55. According to the gradient complementary weight ω comp Constructing weighted transfer optimization objective function S56. During the target domain training process, the structures of the nodule candidate generation branch and the authenticity discrimination branch remain unchanged, and only the authenticity score of the lung micronodules in the candidate frame is scored. j Introducing a lightweight target domain normalization adjustment factor δ norm The adjustment factor is dynamically adjusted according to the overall score distribution of the target hospital samples to define the updated lung micronodule authenticity score. where δ norm Used to suppress cross-domain rating deviation; S57. The transfer optimization model obtained after the general feature pre-training stage and the target domain gradient complementary transfer stage is defined as the final transfer optimization model M trans .

7. The method for early lung cancer screening based on transfer learning according to claim 6, characterized in that: The S6 comprises the following steps: S61. In the migration optimization model M trans The adversarial domain adaptation module is embedded in the model, which includes a feature encoder, a gradient reversal layer, and a domain classifier. S62. Input CT images from the public lung nodule dataset and the target hospital small sample dataset into the forward structure of the transfer optimization model to extract mid- and high-level shared feature representations. S63. Input the mid- and high-level shared feature representations into the feature encoder in the adversarial domain adaptation module to obtain domain discriminant feature representations. Input the domain discriminant feature representations into the gradient reversal layer, outputting the reversal features. Input the reversal features into the domain classifier for discrimination to obtain domain discriminant probability. S64. Define adversarial domain classification loss function It is used to measure the accuracy of the domain classifier in distinguishing the domain origin of the sample and drive the feature extraction process to produce cross-domain inseparable features. The domain classification loss function It is defined in the form of binary cross entropy loss; S65. Domain Classification Loss Function and migration optimization objective function Together they form the cross-domain joint optimization objective function S66. Adopting cross-domain joint optimization objective function The deep residual network backbone structure, multi-level residual dense fusion module, multi-scale attention enhancement module, feature encoder and domain classifier are trained end-to-end to obtain the final cross-domain robust model M robust .

8. The method for early lung cancer screening based on transfer learning according to claim 7, characterized in that: The S7 comprises the following steps: S71. During the clinical reasoning phase, the trained cross-domain robust model is fed with the original low-dose spiral CT chest image data of the patient to be tested. The model then performs a full-process reasoning operation to obtain candidate micronodule detection results, including spatial coordinates, structural dimensions, classification scores, and significant regional thermal distribution. S72. Locate the candidate nodule area output by the cross-domain robust model and obtain the spatial position coordinates corresponding to each micronodule S73. Extracting structural dimension information for each micronodule candidate region, including major diameter and minor diameter; S74. Assign a malignancy probability score to each candidate nodule region based on the updated lung micronodule authenticity score of the domain classifier, wherein the malignancy probability score corresponds to the lung micronodule authenticity score in a one-to-one manner; S75. Generate interpretable heatmaps based on back-propagation gradients of high-order convolutional channels in a cross-domain robust model.

Citation Information

Cited By

  • Bone age assessment method and system combining deep learning and logic correction segmentation

    CN120977590A

  • Processing method of segmentation model training data set and related device

    CN121330421A

  • Low-dose CT image quality optimization method and system based on generative adversarial network

    CN121544485A

  • A lung nodule detection method and system based on adaptive multi-scale deformable attention

    CN122436200A