Large model-based pulmonary nodule analysis method, system, equipment and medium
By adopting a large-model-based method in the analysis of lung nodules, combined with YOLOv8, U-Net and Vision Transformer models, the problems of low detection sensitivity of small lesions and low accuracy of segmentation of complex lesions are solved, and higher detection accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510429587.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-04-08
AI Technical Summary
The prior art has problems with low detection sensitivity of small lesions and low accuracy of segmentation of complex lesions when analyzing lung nodules based on DR chest radiographs.
A large-model-based lung nodule analysis method is adopted, including the object detection sub-model, the target segmentation sub-model and the large-model-based classification sub-model. The object detection sub-model adopts the YOLOv8 network, and the C2f-nmODE module is introduced into the backbone network for deep feature fusion; the target segmentation sub-model adopts the U-Net network, and nmODE ordinary differential equations are introduced in the upsampling layer; the classification sub-model adopts the Vision Transformer model, combining the output of the object detection sub-model and the target segmentation sub-model as input.
It improves the detection sensitivity of small lesions and the segmentation accuracy of complex lesions, enhances the performance and robustness of the model, is suitable for real-time inference of mobile devices, and has stable performance in low-quality DR images.
Smart Images

Figure CN119941731B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence medical technology, and relates to the detection and classification of lung nodules, and in particular to the design of a lung nodule analysis method, system, equipment and medium based on a large model. Background Art
[0002] Medical image segmentation is an important research field at the intersection of computer vision and medicine. Its main purpose is to divide different structures, tissues or pathologies in medical images, making quantitative analysis and diagnosis of specific areas possible. Medical image segmentation has been widely used in many medical disciplines, especially in lung cancer detection. In lung cancer detection based on artificial intelligence medical technology, it mainly uses algorithms or models such as machine learning, deep learning, neural networks, fuzzy logic, genetic algorithms, etc. to accurately identify suspicious lesions in lung CT images, automatically perform segmentation and feature extraction, and analyze and synthesize the shape, density, texture and other characteristics of the lesions by comparing a large amount of clinical data, predict the location of lung nodules, and use the results as a reference for subsequent doctor diagnosis and treatment.
[0003] For the detection and analysis of lung nodules, the invention patent application with application number 202311602306.X discloses a lung CT image physiological detection system and method based on an attention mechanism, which includes: a sequentially connected input unit, a lung lesion classification unit, a lung nodule segmentation unit and a lung nodule benign and malignant classification unit; wherein the lung lesion classification unit is used to detect which images contain lung nodules; the lung nodule segmentation unit is used to segment the images that have been determined as lung nodules in the lung lesion classification unit, so as to determine and mark the location of the lesion area of the lung nodules in the image; the lung nodule benign and malignant classification unit: is used to output the probability of benign and malignant lung nodules in the lung nodule image according to the analysis results of the segmented lung nodules.
[0004] In the classification task of lung nodules, target detection of lung nodules is involved, and the YOLOv8 network is a common target detection network. The invention patent application with application number 202411533920.X discloses a method for detecting small targets in drone images based on improved YOLOv8, which constructs and trains a YOLOv8 benchmark network model for detecting small targets in drone captured images. YOLOv8 is divided into the following three parts: backbone network, the backbone network of YOLOv8 is a key component for extracting features from input images, consisting of a CBS module and a C2f module, where the CBS module is used for downsampling operations and the C2f module is used for feature extraction; neck network, the neck network of the YOLOv8 benchmark network model adopts the structure of path aggregation network + feature pyramid network, allowing low-level features to be secondary fused with high-level features, which helps to capture targets of different scales; detection head, the detection head is responsible for generating the final prediction output, including the location and size of the bounding box, and the category probability of each box.
[0005] In the classification task of lung nodules, in addition to target detection of lung nodules, it also involves segmentation of lung nodules, and the U-Net network is a common target segmentation network. The invention patent with application number 202410195116.9 discloses a lung nodule analysis method, system, equipment and medium based on a large model. The lung nodule segmentation network model constructed by it adopts the U-Net network. The lung nodule segmentation network model includes an encoder and a decoder. The encoder includes a U-shaped structure composed of 9 modules, and the decoder includes a U-shaped structure composed of 9 modules, each of which is composed of a recursive residual convolution layer.
[0006] When using the existing YOLO network for target detection of lung nodules, there will be problems with missed detection of small lesions, especially when dealing with tiny lesions (such as nodules with a diameter of <5mm), the vanishing gradient of the detection task will significantly reduce the convergence stability of the model; and the perception field of its standard detection head is mainly for medium and large targets, and the recall rate of early signs such as diffuse tiny nodules and ground-glass shadows common in DR images is low. In addition, the traditional U-Net segmentation network is prone to topological structure distortion and edge fragmentation when dealing with nodule areas with multiple lesions overlapping and blurred boundaries.
[0007] Therefore, when the existing technology detects, segmentes and classifies lung nodules based on DR images, there are problems such as insufficient sensitivity in detecting small lesions and low accuracy in segmenting complex lesions, resulting in low accuracy in the analysis results of lung nodules. It is necessary to provide a method with higher precision and accuracy for detecting, segmenting and classifying lung nodules. Summary of the invention
[0008] The purpose of the present invention is to provide a lung nodule analysis method, system, device and medium based on a large model in order to solve the problems of low sensitivity in detecting small lesions and low accuracy in segmenting complex lesions in the prior art when analyzing lung nodules based on DR chest radiographs.
[0009] In order to achieve the above-mentioned purpose, the present invention specifically adopts the following technical solutions:
[0010] A method for analyzing lung nodules based on a large model comprises the following steps:
[0011] Step 1: Obtain sample data and labels;
[0012] Obtain lung DR sample images and annotate them to obtain label data; the labels include lung nodule location labels, lung nodule segmentation mask labels, and lung nodule category labels;
[0013] Step 2, constructing a lung nodule analysis model;
[0014] Construct a lung nodule analysis model, which includes a target detection sub-model, a target segmentation sub-model, and a classification sub-model based on a large model;
[0015] The target detection sub-model adopts the YOLOv8 network, including the backbone network, the neck network and the detection head. The four feature maps of different sizes output by the backbone network are input into the neck network. The neck network fuses the feature maps of different sizes and then inputs them into the target detection head of the corresponding size in the detection head. The C2f-nmODE module in the backbone network includes the first CBS layer, the Split layer and the second CBS layer. The output of the first CBS layer is used as the input of the Split layer. The output of the Split layer passes through three nmODE layers and the outputs of the three nmODE layers are fused and then input into the second CBS layer together with the output of the Split layer. Among them, the nmODE layer changes the discretized neural memory ordinary differential equation into a form in which the number of neural network layers is used as the number of derivation steps.
[0016] Step 3, training a lung nodule analysis model;
[0017] Using the sample data and labels obtained in step 1, the lung nodule analysis model constructed in step 2 is trained;
[0018] Step 4, real-time classification;
[0019] The DR image of the lung to be tested is obtained and input into the pulmonary nodule analysis model trained in step 3. The pulmonary nodule analysis model outputs pulmonary nodule detection and classification results, which include whether there are pulmonary nodules, the location of the pulmonary nodules, and whether the pulmonary nodules are benign or malignant.
[0020] Further, in step 2, the backbone network includes a first convolution module, a second convolution module, a first C2f-nmODE module, a third convolution module, a second C2f-nmODE module, a fourth convolution module, a third C2f-nmODE module, a fifth convolution module, a Swin Transformer module and an SPPF-LSKA module, which are connected in sequence; the first C2f-nmODE module, the second C2f-nmODE module, the third C2f-nmODE module and the SPPF-LSKA module output a 1 / 4 feature map, a 1 / 8 feature map, a 1 / 16 feature map and a 1 / 32 feature map, respectively.
[0021] Furthermore, in step 2, the target segmentation sub-model adopts a U-Net network, including a first convolutional layer, an encoder, a decoder, and a second convolutional layer;
[0022] The encoder includes a first downsampling layer, a second downsampling layer, a third downsampling layer, a fourth downsampling layer, and a fifth downsampling layer;
[0023] The decoder includes a channel adjustment layer, a second ODE upsampling layer, a third ODE upsampling layer, a fourth ODE upsampling layer, and a fifth ODE upsampling layer.
[0024] Furthermore, the channel adjustment layer includes a 1*1 convolution layer and the first upsampling layer;
[0025] The second ODE upsampling layer, the third ODE upsampling layer, the fourth ODE upsampling layer, and the fifth ODE upsampling layer all include a 1*1 convolution layer, five ordinary differential equation solvers and a second upsampling layer. The feature map output by the downsampling layer is input into the 1*1 convolution layer. The 1*1 convolution layer serves as the input to the five ordinary differential equation solvers. The first feature map output by the channel adjustment layer or the fused feature map output by the remaining ODE upsampling layers serves as the input to the lowest ordinary differential equation solver. The output of the 1*1 convolution layer and the output of the previous ordinary differential equation solver serve as the input to the next ordinary differential equation solver. The output of the last ordinary differential equation solver serves as the input to the second upsampling layer.
[0026] Furthermore, in step 2, the classification sub-model based on the large model adopts the Vision Transformer large model; in the Vision Transformer large model, the probability masks output by the DR image and the target segmentation sub-model are divided into the same patch. When generating the patch vector, the two types of patch vectors are spliced and input into the multi-head attention function. When performing the multi-head attention function processing, the positioning box output by the target detection sub-model is input as the attention prior parameter.
[0027] Furthermore, in step 3, when training the lung nodule analysis model, the specific training method is:
[0028] First, the target segmentation sub-model is trained using lung DR sample images and lung nodule segmentation mask labels to obtain a probability mask;
[0029] The target detection sub-model is then trained using the lung DR sample images, lung nodule location labels, and the probability mask generated during the target segmentation sub-model training to obtain the lung nodule positioning frame.
[0030] Finally, the lung nodule localization box detected by the target detection sub-model, the probability mask output by the target segmentation sub-model, and the lung DR sample image are input into the classification sub-model based on the large model. The probability mask and the lung DR sample image are encoded as input data into the Patch image block in the classification sub-model based on the large model, and the lung nodule localization box is provided as prior knowledge to the multi-head attention function of the classification sub-model based on the large model as the spatial alignment offset of the attention.
[0031] Furthermore, when training the target detection sub-model, the loss function for:
[0032] ;
[0033] When training the target segmentation sub-model, the loss function It is expressed as:
[0034] ;
[0035] When training the classification sub-model based on the large model, the loss function It is expressed as:
[0036] ;
[0037] in, Represents the total number of samples in a batch, represents the true label, represents the model prediction probability, represents the weight of positive samples, represents the modulation factor, Represents the weight of negative samples; represents the weight of the constraint, Indicates the preset threshold value, represents the average lesion probability; represents the KL divergence function, represents the detection quantity distribution vector, represents the classification quantity distribution vector, represents the cross entropy weight, represents the true classification label, Represents the classification prediction result.
[0038] A pulmonary nodule analysis system based on a large model, comprising:
[0039] The sample data and label acquisition module is used to acquire lung DR sample images and annotate them to obtain label data; the labels include lung nodule location labels, lung nodule segmentation mask labels, and lung nodule category labels;
[0040] A pulmonary nodule analysis model construction module is used to construct a pulmonary nodule analysis model, which includes a target detection sub-model, a target segmentation sub-model, and a classification sub-model based on a large model;
[0041] The target detection sub-model adopts the YOLOv8 network, including the backbone network, the neck network and the detection head. The four feature maps of different sizes output by the backbone network are input into the neck network. The neck network fuses the feature maps of different sizes and then inputs them into the target detection head of the corresponding size in the detection head. The C2f-nmODE module in the backbone network includes the first CBS layer, the Split layer and the second CBS layer. The output of the first CBS layer is used as the input of the Split layer. The output of the Split layer passes through three nmODE layers and the outputs of the three nmODE layers are fused and then input into the second CBS layer together with the output of the Split layer. Among them, the nmODE layer changes the discretized neural memory ordinary differential equation into a form in which the number of neural network layers is used as the number of derivation steps.
[0042] A pulmonary nodule analysis model training module is used to train the pulmonary nodule analysis model constructed by the pulmonary nodule analysis model construction module using the sample data and labels acquired by the sample data and label acquisition module;
[0043] The real-time classification module is used to obtain the lung DR image to be tested and input the lung nodule analysis model trained by the lung nodule analysis model training module. The lung nodule analysis model outputs the lung nodule detection and classification results, which include whether there are lung nodules, the location of the lung nodules, and whether the lung nodules are benign or malignant.
[0044] A computer device comprises a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the above method.
[0045] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor is caused to execute the steps of the above method.
[0046] The beneficial effects of the present invention are as follows:
[0047] 1. In the present invention, the C2f-nmODE module in the backbone network of the YOLOv8 network target detection submodel adopts 3 nmODE layers, and each nmODE layer changes the discretized neural memory ordinary differential equation into a form with the number of neural network layers as the number of derivation steps, which can achieve deeper feature fusion under the same amount of calculation, is suitable for the detection of small lesions with high detection sensitivity, is conducive to the segmentation of complex lesions, improves the detection accuracy, and effectively solves the problems of low sensitivity for small lesion detection and low accuracy for complex lesion segmentation when analyzing lung nodules based on DR chest radiographs; at the same time, it provides reliable technical support for the subsequent intelligent diagnosis of lung nodules, and has significant clinical application value and market promotion potential.
[0048] 2. In the present invention, the backbone network of the YOLOv8 network is innovated, and the Swin Transformer module is used to replace the C2f module of the ninth layer of the original backbone network, and a window-based multi-head self-attention mechanism is introduced. At the same time, the SPPF-LSKA module is used to replace the original SPPF of the tenth layer of the original backbone network; this innovation strengthens the information exchange between features at different levels, reduces the interference of complex background on lesion detection, enables the model to more accurately detect and identify small objects in complex environments, and greatly improves the performance and robustness of the model.
[0049] 3. In the present invention, the nmODE ordinary differential equation is introduced in the upsampling layer of the encoder part on the right side of the U-shaped network, which can effectively improve the feature reconstruction quality and model performance; traditional interpolation or deconvolution methods are often prone to introduce artifacts or information loss, while the present application uses nmODE ordinary differential to achieve continuous dynamic modeling, making the upsampling process smoother and more physically meaningful, thereby reducing feature distortion and retaining more spatial detail information; in addition, nmODE can adaptively learn the sampling path, making the feature mapping more accurate, which helps to improve the segmentation accuracy, especially in tasks with complex edge details.
[0050] 4. Based on the above innovations, the present invention greatly improves the detection performance of small targets and minimizes the impact of noise artifacts; the neural memory ODE segmentation is used to significantly improve the segmentation effect of overlapping areas of multiple lesions, and the number of parameters is reduced by 40%; the lightweight classification module can realize real-time reasoning on mobile devices and adapt to the hardware conditions of primary medical institutions; the preprocessing module effectively suppresses noise and metal artifacts, and has stable performance in low-quality DR images. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 It is a schematic diagram of the process of the present invention;
[0052] Figure 2 It is a schematic diagram of the structure of the target detection sub-model in the present invention;
[0053] Figure 3 It is a schematic diagram of the structure of the C2f-nmODE module in the present invention;
[0054] Figure 4 It is a schematic diagram of the structure of the target segmentation sub-model in the present invention;
[0055] Figure 5 It is a schematic diagram of the structure of the ODE upsampling layer in the present invention. DETAILED DESCRIPTION
[0056] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments.
[0057] Therefore, based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative work shall fall within the scope of protection of the present invention.
[0058] Example 1
[0059] This embodiment provides a lung nodule analysis method based on a large model, which is used to detect, segment and classify lung nodules in lung DR images. Figure 1 As shown, the following steps are included:
[0060] Step 1: Obtain sample data and labels;
[0061] Obtain lung DR sample images and annotate them to obtain label data; the labels include lung nodule location labels, lung nodule segmentation mask labels, and lung nodule category labels.
[0062] In this embodiment, a pathology-preserving enhancement strategy is adopted to ensure that the original morphology of the lesion is not destroyed when the data is expanded. That is, the lung DR image is expanded by contrast adjustment, local brightness perturbation and other methods. These technologies are in line with the characteristics of X-ray imaging, and can enhance image details while maintaining the pathological structure, providing real and rich feature information for subsequent model training and learning.
[0063] In view of the differences in image clarity and resolution in the sample data set, the sample data was divided into different groups. When training the model later, a phased training method was adopted, namely: first, low-resolution images were used for coarse lesion positioning training, so that the model could quickly grasp the approximate area of the lesion; then the image resolution was gradually improved, so that the model could conduct in-depth learning of fine granularity, and further improve the detection accuracy of small or overlapping lesions.
[0064] Step 2, constructing a lung nodule analysis model;
[0065] A pulmonary nodule analysis model is constructed, which includes a target detection sub-model, a target segmentation sub-model, and a classification sub-model based on a large model.
[0066] like Figure 2 As shown, the target detection sub-model adopts the YOLOv8 network, including a backbone network, a neck network and a detection head. The backbone network includes a first convolution module, a second convolution module, a first C2f-nmODE module, a third convolution module, a second C2f-nmODE module, a fourth convolution module, a third C2f-nmODE module, a fifth convolution module, a SwinTransformer module and an SPPF-LSKA module connected in sequence, and the output of the previous module is used as the input of the next module (for example, the output of the first convolution module is used as the input of the second convolution module); the first C2f-nmODE module, the second C2f-nmODE module, the third C2f-nmODE module and the SPPF-LSKA module output 1 / 4 feature map, 1 / 8 feature map, 1 / 16 feature map and 1 / 32 feature map respectively.
[0067] like Figure 3 As shown, the C2f-nmODE modules in the backbone network (i.e., the first C2f-nmODE module, the second C2f-nmODE module, and the third C2f-nmODE module) include the first CBS layer, the Split layer, the nmODE layer, and the second CBS layer, and the nmODE layer adopts a three-layer structure. The output of the first CBS layer is used as the input of the Split layer. The output of the Split layer passes through three nmODE layers, and the outputs of the three nmODE layers are fused and then input into the second CBS layer together with the output of the Split layer. Among them, the nmODE layer changes the discretized neural memory ordinary differential equation into a form in which the number of neural network layers is used as the number of derivation steps.
[0068] In this embodiment, four feature maps of different sizes output by the backbone network are input into the neck network, and the neck network fuses the feature maps of different sizes and then inputs the target detection head of the corresponding size in the detection head. The SwinTransformer module is used to replace the C2F module of the ninth layer in the backbone network, and the SPPF-LSKA module is used to replace the original SPPF of the tenth layer. The above modifications strengthen the information exchange between features at different levels, reduce the interference of complex background on lesion detection, enable the target detection sub-model to more accurately detect and identify small objects in complex environments, and greatly improve the performance and robustness of the target detection sub-model.
[0069] like Figure 4 As shown, the target segmentation sub-model adopts the U-Net network architecture, which includes a first convolutional layer, an encoder, a decoder, and a second convolutional layer.
[0070] The encoder includes a first downsampling layer, a second downsampling layer, a third downsampling layer, a fourth downsampling layer, and a fifth downsampling layer.
[0071] The decoder includes a channel adjustment layer, a second ODE upsampling layer, a third ODE upsampling layer, a fourth ODE upsampling layer, and a fifth ODE upsampling layer.
[0072] The channel adjustment layer includes a 1*1 convolution layer and the first upsampling layer;
[0073] The decoder uses a fixed channel size, the fused feature map channel map is the same as the number of final predicted categories, and each ODE upsampling layer uses 5 solvers with a step size of 0.2. Each solver shares the solution parameters to improve the computational efficiency during the network decoding process. Figure 5 As shown, specifically: the second ODE upsampling layer, the third ODE upsampling layer, the fourth ODE upsampling layer, and the fifth ODE upsampling layer all include a 1*1 convolution layer, five ordinary differential equation solvers and a second upsampling layer. The feature map output by the downsampling layer is input into the 1*1 convolution layer, and the 1*1 convolution layer is used as the input of the five ordinary differential equation solvers. The first feature map output by the channel adjustment layer or the fused feature map output by the remaining ODE upsampling layers is used as the input of the lowest ordinary differential equation solver. The output of the 1*1 convolution layer and the output of the previous ordinary differential equation solver are used as the input of the next ordinary differential equation solver, and the output of the last ordinary differential equation solver is used as the input of the second upsampling layer.
[0074] The process of discretization solution of ordinary differential equation solver is:
[0075] ;
[0076] in, represents the discretization step size, which is 0.2; represents the learnable memory weight matrix, represents the input of the skip connection, represents the learning memory bias matrix, , Represent the current hidden state and the next hidden state respectively.
[0077] The classification sub-model based on the large model adopts the Vision Transformer large model. When the Vision Transformer large model is applied to medical images, there will be problems such as small samples and difficult convergence. Therefore, this embodiment integrates the target detection sub-model (providing a lung nodule positioning frame), the target segmentation sub-model (providing a probability mask) and the original DR sample image as the input of the classification sub-model based on the large model; in the Vision Transformer large model, the DR image and the probability mask output by the target segmentation sub-model are both divided into the same patch. When generating the patch vector, the two types of patch vectors are spliced and input into the multi-head attention function. When performing the multi-head attention function processing, the positioning frame output by the target detection sub-model is input as the attention prior parameter.
[0078] Specifically, the Vision Transformer large model generally uses a fixed-size patch. This network divides the original image and the probability mask output by the target segmentation sub-model into the same patch, specifically:
[0079] ;
[0080] ;
[0081] Then, when generating the Patch vector, the two types of Patch vectors are concatenated and input into the multi-head attention function:
[0082] ;
[0083] When performing multi-head attention function processing, the lung nodule positioning box output by the target detection sub-model is As the attention prior, it is used as a parameter input, that is, setting the weight ; If the center point of the positioning box falls in a certain Patch, the degree of attention the network pays to this area is affected by setting the bias weight:
[0084] ;
[0085] ;
[0086] in, represents the original image, represents the original image block, represents the segmentation probability mask, represents the probability mask block, Indicates a channel connection operation. represents a full connection operation, represents the final feature image block, , , Represent the query vector, key vector, and transposition operation respectively. represents the scaling factor, represents the attention weight matrix, represents the attention weight value, represents the i-th image block.
[0087] Step 3, training a lung nodule analysis model;
[0088] The sample data and labels obtained in step 1 are used to train the lung nodule analysis model constructed in step 2.
[0089] When training the lung nodule analysis model, the specific training method is:
[0090] First, the target segmentation sub-model is trained using lung DR sample images and lung nodule segmentation mask labels to obtain a probability mask;
[0091] The target detection sub-model is then trained using the lung DR sample images, lung nodule location labels, and the probability mask generated during the target segmentation sub-model training to obtain the lung nodule positioning frame.
[0092] Finally, the lung nodule localization box detected by the target detection sub-model, the probability mask output by the target segmentation sub-model, and the lung DR sample image are input into the classification sub-model based on the large model. The probability mask and the lung DR sample image are encoded as input data into the Patch image block in the classification sub-model based on the large model, and the lung nodule localization box is provided as prior knowledge to the multi-head attention function of the classification sub-model based on the large model as the spatial alignment offset of the attention.
[0093] When training the target detection sub-model, the loss function distinguishes easy samples from hard samples according to the IoU size between the predicted box and the true box, and uses the loss function To give higher weights to difficult samples. The loss function for:
[0094] ;
[0095] Represents the total number of samples in a batch, represents the true label, represents the model prediction probability, represents the weight of positive samples, represents the modulation factor, represents the weight of negative samples.
[0096] When training the target segmentation sub-model, a new pathology consistency constraint is added. After the target segmentation sub-model is trained, pseudo labels are generated for the training data, and the detection sub-model is optimized and trained together to constrain the consistency between the features in the positioning box and the semantics of the lesion. Specifically, the image is first segmented at the pixel level using the Unet network to generate a lesion probability map; then the yolo detection network is used to detect the candidate lesion bounding box and calculate the average lesion probability inside it. , and based on the preset threshold Construct a consistency constraint and incorporate it into the loss function. It is expressed as:
[0097] ;
[0098] represents the weight of the constraint, Indicates the preset threshold value, represents the average lesion probability.
[0099] When training the classification sub-model based on the large model, the LoRA adapter is used to inject low-rank decomposition parameters into the projection matrix of each multi-head attention layer. By adding a low-rank matrix adaptation layer on the original weight matrix, efficient task adaptation can be performed without significantly changing the original model. The updated hidden state is updated through learnable parameters Adjust the fusion ratio of the adapter output and the original features, that is:
[0100] ;
[0101] In order to achieve the consistency constraint of the number of lesions, a dual-path supervision mechanism is adopted. The number of lung nodules output by the statistical detection network , generate the quantity distribution vector through the Poisson distribution encoder (Corresponding to clinical t-level classification). In addition, an auxiliary output layer is added at the end of the classification subnetwork to generate the predicted distribution of the number of lesions. The two distribution vectors are used as the detection quantity encoding path and the classification confidence path respectively, and then the loss function based on the KL divergence constraint term is constructed. It is expressed as:
[0102] ;
[0103] in, represents the original hidden state, represents the learning parameters, represents the low-rank matrix adaptation function, represents the updated hidden state, represents the KL divergence function, represents the detection quantity distribution vector, represents the classification quantity distribution vector, represents the cross entropy weight, represents the true classification label, Represents the classification prediction result.
[0104] Step 4, real-time classification;
[0105] The DR image of the lung to be tested is obtained and input into the pulmonary nodule analysis model trained in step 3. The pulmonary nodule analysis model outputs pulmonary nodule detection and classification results, which include whether there are pulmonary nodules, the location of the pulmonary nodules, and whether the pulmonary nodules are benign or malignant.
[0106] Example 2
[0107] This embodiment provides a large model-based pulmonary nodule analysis system for detecting, segmenting, and classifying pulmonary nodules in lung DR images, including:
[0108] The sample data and label acquisition module is used to obtain lung DR sample images and annotate them to obtain label data; the labels include lung nodule location labels, lung nodule segmentation mask labels, and lung nodule category labels.
[0109] In this embodiment, a pathology-preserving enhancement strategy is adopted to ensure that the original morphology of the lesion is not destroyed when the data is expanded. That is, the lung DR image is expanded by contrast adjustment, local brightness perturbation and other methods. These technologies are in line with the characteristics of X-ray imaging, and can enhance image details while maintaining the pathological structure, providing real and rich feature information for subsequent model training and learning.
[0110] In view of the differences in image clarity and resolution in the sample data set, the sample data was divided into different groups. When training the model later, a phased training method was adopted, namely: first, low-resolution images were used for coarse lesion positioning training, so that the model could quickly grasp the approximate area of the lesion; then the image resolution was gradually improved, so that the model could conduct in-depth learning of fine granularity, and further improve the detection accuracy of small or overlapping lesions.
[0111] The pulmonary nodule analysis model construction module is used to construct a pulmonary nodule analysis model. The pulmonary nodule analysis model includes a target detection sub-model, a target segmentation sub-model, and a classification sub-model based on a large model.
[0112] like Figure 2As shown, the target detection sub-model adopts the YOLOv8 network, including a backbone network, a neck network and a detection head. The backbone network includes a first convolution module, a second convolution module, a first C2f-nmODE module, a third convolution module, a second C2f-nmODE module, a fourth convolution module, a third C2f-nmODE module, a fifth convolution module, a SwinTransformer module and an SPPF-LSKA module connected in sequence, and the output of the previous module is used as the input of the next module (for example, the output of the first convolution module is used as the input of the second convolution module); the first C2f-nmODE module, the second C2f-nmODE module, the third C2f-nmODE module and the SPPF-LSKA module output 1 / 4 feature map, 1 / 8 feature map, 1 / 16 feature map and 1 / 32 feature map respectively.
[0113] like Figure 3 As shown, the C2f-nmODE modules in the backbone network (i.e., the first C2f-nmODE module, the second C2f-nmODE module, and the third C2f-nmODE module) include the first CBS layer, the Split layer, the nmODE layer, and the second CBS layer, and the nmODE layer adopts a three-layer structure. The output of the first CBS layer is used as the input of the Split layer. The output of the Split layer passes through three nmODE layers, and the outputs of the three nmODE layers are fused and then input into the second CBS layer together with the output of the Split layer. Among them, the nmODE layer changes the discretized neural memory ordinary differential equation into a form in which the number of neural network layers is used as the number of derivation steps.
[0114] In this embodiment, four feature maps of different sizes output by the backbone network are input into the neck network, and the neck network fuses the feature maps of different sizes and then inputs the target detection head of the corresponding size in the detection head. The SwinTransformer module is used to replace the C2F module of the ninth layer in the backbone network, and the SPPF-LSKA module is used to replace the original SPPF of the tenth layer. The above modifications strengthen the information exchange between features at different levels, reduce the interference of complex background on lesion detection, enable the target detection sub-model to more accurately detect and identify small objects in complex environments, and greatly improve the performance and robustness of the target detection sub-model.
[0115] like Figure 4 As shown, the target segmentation sub-model adopts the U-Net network architecture, which includes a first convolutional layer, an encoder, a decoder, and a second convolutional layer.
[0116] The encoder includes a first downsampling layer, a second downsampling layer, a third downsampling layer, a fourth downsampling layer, and a fifth downsampling layer.
[0117] The decoder includes a channel adjustment layer, a second ODE upsampling layer, a third ODE upsampling layer, a fourth ODE upsampling layer, and a fifth ODE upsampling layer.
[0118] The channel adjustment layer includes a 1*1 convolution layer and the first upsampling layer;
[0119] The decoder uses a fixed channel size, the fused feature map channel map is the same as the number of final predicted categories, and each ODE upsampling layer uses 5 solvers with a step size of 0.2. Each solver shares the solution parameters to improve the computational efficiency during the network decoding process. Figure 5 As shown, specifically: the second ODE upsampling layer, the third ODE upsampling layer, the fourth ODE upsampling layer, and the fifth ODE upsampling layer all include a 1*1 convolution layer, five ordinary differential equation solvers and a second upsampling layer. The feature map output by the downsampling layer is input into the 1*1 convolution layer, and the 1*1 convolution layer is used as the input of the five ordinary differential equation solvers. The first feature map output by the channel adjustment layer or the fused feature map output by the remaining ODE upsampling layers is used as the input of the lowest ordinary differential equation solver. The output of the 1*1 convolution layer and the output of the previous ordinary differential equation solver are used as the input of the next ordinary differential equation solver, and the output of the last ordinary differential equation solver is used as the input of the second upsampling layer.
[0120] The process of discretization solution of ordinary differential equation solver is:
[0121] ;
[0122] in, represents the discretization step size, which is 0.2; represents the learnable memory weight matrix, represents the input of the skip connection, represents the learning memory bias matrix, , Represent the current hidden state and the next hidden state respectively.
[0123] The classification sub-model based on the large model adopts the Vision Transformer large model. When the Vision Transformer large model is applied to medical images, there will be problems such as small samples and difficult convergence. Therefore, this embodiment integrates the target detection sub-model (providing a lung nodule positioning frame), the target segmentation sub-model (providing a probability mask) and the original DR sample image as the input of the classification sub-model based on the large model; in the Vision Transformer large model, the DR image and the probability mask output by the target segmentation sub-model are both divided into the same patch. When generating the patch vector, the two types of patch vectors are spliced and input into the multi-head attention function. When performing the multi-head attention function processing, the positioning frame output by the target detection sub-model is input as the attention prior parameter.
[0124] Specifically, the Vision Transformer large model generally uses a fixed-size patch. This network divides the original image and the probability mask output by the target segmentation sub-model into the same patch, specifically:
[0125] ;
[0126] ;
[0127] Then, when generating the Patch vector, the two types of Patch vectors are concatenated and input into the multi-head attention function:
[0128] ;
[0129] When performing multi-head attention function processing, the lung nodule positioning box output by the target detection sub-model is As the attention prior, it is used as a parameter input, that is, setting the weight ; If the center point of the positioning box falls in a certain Patch, the degree of attention the network pays to this area is affected by setting the bias weight:
[0130] ;
[0131] ;
[0132] in, represents the original image, represents the original image block, represents the segmentation probability mask, represents the probability mask block, Indicates a channel connection operation. represents a full connection operation, represents the final feature image block, , , Represent the query vector, key vector, and transposition operation respectively. represents the scaling factor, represents the attention weight matrix, represents the attention weight value, represents the i-th image block.
[0133] The pulmonary nodule analysis model training module is used to train the pulmonary nodule analysis model constructed by the pulmonary nodule analysis model construction module using the sample data and labels acquired by the sample data and label acquisition module.
[0134] When training the lung nodule analysis model, the specific training method is:
[0135] First, the target segmentation sub-model is trained using lung DR sample images and lung nodule segmentation mask labels to obtain a probability mask;
[0136] The target detection sub-model is then trained using the lung DR sample images, lung nodule location labels, and the probability mask generated during the target segmentation sub-model training to obtain the lung nodule positioning frame.
[0137] Finally, the lung nodule localization box detected by the target detection sub-model, the probability mask output by the target segmentation sub-model, and the lung DR sample image are input into the classification sub-model based on the large model. The probability mask and the lung DR sample image are encoded as input data into the Patch image block in the classification sub-model based on the large model, and the lung nodule localization box is provided as prior knowledge to the multi-head attention function of the classification sub-model based on the large model as the spatial alignment offset of the attention.
[0138] When training the target detection sub-model, the loss function distinguishes easy samples from hard samples according to the IoU size between the predicted box and the true box, and uses the loss function To give higher weights to difficult samples. The loss function for:
[0139] ;
[0140] Represents the total number of samples in a batch, represents the true label, represents the model prediction probability, represents the weight of positive samples, represents the modulation factor, represents the weight of negative samples.
[0141] When training the target segmentation sub-model, a new pathology consistency constraint is added. After the target segmentation sub-model is trained, pseudo labels are generated for the training data, and the detection sub-model is optimized and trained together to constrain the consistency between the features in the positioning box and the semantics of the lesion. Specifically, the image is first segmented at the pixel level using the Unet network to generate a lesion probability map; then the yolo detection network is used to detect the candidate lesion bounding box and calculate the average lesion probability inside it. , and based on the preset threshold Construct a consistency constraint and incorporate it into the loss function. It is expressed as:
[0142] ;
[0143] represents the weight of the constraint, Indicates the preset threshold value, represents the average lesion probability.
[0144] When training the classification sub-model based on the large model, the LoRA adapter is used to inject low-rank decomposition parameters into the projection matrix of each multi-head attention layer. By adding a low-rank matrix adaptation layer on the original weight matrix, efficient task adaptation can be performed without significantly changing the original model. The updated hidden state is updated through learnable parameters Adjust the fusion ratio of the adapter output and the original features, that is:
[0145] ;
[0146] In order to achieve the consistency constraint of the number of lesions, a dual-path supervision mechanism is adopted. The number of lung nodules output by the statistical detection network , generate the quantity distribution vector through the Poisson distribution encoder (Corresponding to clinical t-level classification). In addition, an auxiliary output layer is added at the end of the classification subnetwork to generate the predicted distribution of the number of lesions. The two distribution vectors are used as the detection quantity encoding path and the classification confidence path respectively, and then the loss function based on the KL divergence constraint term is constructed. It is expressed as:
[0147] ;
[0148] in, represents the original hidden state, represents the learning parameters, represents the low-rank matrix adaptation function, represents the updated hidden state, represents the KL divergence function, represents the detection quantity distribution vector, represents the classification quantity distribution vector, represents the cross entropy weight, represents the true classification label, Represents the classification prediction result.
[0149] The real-time classification module is used to obtain the lung DR image to be tested and input the lung nodule analysis model trained by the lung nodule analysis model training module. The lung nodule analysis model outputs the lung nodule detection and classification results, which include whether there are lung nodules, the location of the lung nodules, and whether the lung nodules are benign or malignant.
[0150] It should be noted that the final output is the risk coefficient of benign or malignant lung nodules. Doctors will subsequently further determine whether the corresponding lung nodules are benign, malignant, and the degree of malignancy based on the analysis results and clinical diagnosis.
[0151] In addition, the system superimposes the YOLO target positioning frame (i.e., detection frame) on the original DR chest radiograph to indicate possible lung nodule areas in the image. For each positioning frame, an additional cropped image will be generated and the segmentation results and the risk factor predicted by the classification sub-model will be superimposed to assist further doctor judgment. In addition, the system will also generate a detection heat map to show the attention paid to different positions of the image during model operation. Specifically, for the complete input image, it is divided into multiple grid cells (x, y), and the highest confidence among all candidate boxes in the cell is taken as the thermal value of the position, that is:
[0152] ;
[0153] in, represents the confidence of the kth candidate box in the grid cell (x, y), Represents the number of candidate boxes in the grid cell, represents the probability that the kth candidate box in the grid cell (x, y) contains the target, represents the predicted bounding box of the kth candidate box in the grid cell (x, y), represents the ground-truth bounding box.
[0154] Example 3
[0155] A computer device comprises a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of a lung nodule analysis method based on a large model.
[0156] The computer device may be a desktop computer, a notebook, a PDA, a cloud server, etc. The computer device may interact with the user through a keyboard, a mouse, a remote control, a touch pad, or a voice control device.
[0157] The memory includes at least one type of readable storage medium, and the readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (for example, SD or D interface display memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, disk, optical disk, etc. In some embodiments, the memory can be an internal storage unit of the computer device, such as a hard disk or memory of the computer device. In other embodiments, the memory can also be an external storage device of the computer device, such as a plug-in hard disk equipped on the computer device, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. Of course, the memory can also include both the internal storage unit of the computer device and its external storage device. In this embodiment, the memory is often used to store the operating system and various application software installed on the computer device, such as the program code of the large model-based pulmonary nodule analysis method, etc. In addition, the memory can also be used to temporarily store various types of data that have been output or are to be output.
[0158] The processor may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip in some embodiments. The processor is generally used to control the overall operation of the computer device. In this embodiment, the processor is used to run the program code stored in the memory or process data, such as running the program code of the large model-based pulmonary nodule analysis method.
[0159] Example 4
[0160] A computer-readable storage medium stores a computer program, which, when executed by a processor, causes the processor to perform the steps of a large model-based lung nodule analysis method.
[0161] The computer-readable storage medium stores an interface display program, and the interface display program can be executed by at least one processor so that the at least one processor performs the steps of the large model-based lung nodule analysis method as described above.
[0162] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment method can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, disk, CD), and includes several instructions for a terminal device (which can be a mobile phone, computer, server or network device, etc.) to execute the large model-based lung nodule analysis method described in the embodiment of the present application.
Claims
1. A lung nodule analysis method based on a large model, characterized in that: The following steps are involved: Step 1: Obtain sample data and labels; Obtain lung DR sample images and annotate them to obtain label data; the labels include lung nodule location labels, lung nodule segmentation mask labels, and lung nodule category labels; Step 2, constructing a lung nodule analysis model; Construct a lung nodule analysis model, which includes a target detection sub-model, a target segmentation sub-model, and a classification sub-model based on a large model; The target detection sub-model adopts the YOLOv8 network, including the backbone network, the neck network and the detection head. The four feature maps of different sizes output by the backbone network are input into the neck network. The neck network fuses the feature maps of different sizes and then inputs them into the target detection head of the corresponding size in the detection head. The C2f-nmODE module in the backbone network includes the first CBS layer, the Split layer and the second CBS layer. The output of the first CBS layer is used as the input of the Split layer. The output of the Split layer passes through three nmODE layers and the outputs of the three nmODE layers are fused and then input into the second CBS layer together with the output of the Split layer. Among them, the nmODE layer changes the discretized neural memory ordinary differential equation into a form in which the number of neural network layers is used as the number of derivation steps. Step 3, training a lung nodule analysis model; Using the sample data and labels obtained in step 1, the lung nodule analysis model constructed in step 2 is trained; Step 4, real-time classification; Obtain the lung DR image to be tested and input it into the lung nodule analysis model trained in step 3. The lung nodule analysis model outputs the lung nodule detection and classification results, which include whether there are lung nodules, the location of the lung nodules, and whether the lung nodules are benign or malignant. In step 2, the backbone network includes a first convolution module, a second convolution module, a first C2f-nmODE module, a third convolution module, a second C2f-nmODE module, a fourth convolution module, a third C2f-nmODE module, a fifth convolution module, a Swin Transformer module and an SPPF-LSKA module, which are connected in sequence; the first C2f-nmODE module, the second C2f-nmODE module, the third C2f-nmODE module and the SPPF-LSKA module output a 1 / 4 feature map, a 1 / 8 feature map, a 1 / 16 feature map and a 1 / 32 feature map, respectively.
2. A lung nodule analysis method based on a large model as claimed in claim 1, characterized in that: In step 2, the target segmentation sub-model adopts a U-Net network, including the first convolutional layer, encoder, decoder and the second convolutional layer; The encoder includes a first downsampling layer, a second downsampling layer, a third downsampling layer, a fourth downsampling layer, and a fifth downsampling layer; The decoder includes a channel adjustment layer, a second ODE upsampling layer, a third ODE upsampling layer, a fourth ODE upsampling layer, and a fifth ODE upsampling layer.
3. A lung nodule analysis method based on a large model as claimed in claim 2, characterized in that: The channel adjustment layer includes a 1*1 convolution layer and the first upsampling layer; The second ODE upsampling layer, the third ODE upsampling layer, the fourth ODE upsampling layer, and the fifth ODE upsampling layer all include a 1*1 convolution layer, five ordinary differential equation solvers and a second upsampling layer. The feature map output by the downsampling layer is input into the 1*1 convolution layer. The 1*1 convolution layer serves as the input to the five ordinary differential equation solvers. The first feature map output by the channel adjustment layer or the fused feature map output by the remaining ODE upsampling layers serves as the input to the lowest ordinary differential equation solver. The output of the 1*1 convolution layer and the output of the previous ordinary differential equation solver serve as the input to the next ordinary differential equation solver. The output of the last ordinary differential equation solver serves as the input to the second upsampling layer.
4. A lung nodule analysis method based on a large model as claimed in claim 1, characterized in that: In step 2, the classification sub-model based on the large model adopts the Vision Transformer large model; in the Vision Transformer large model, the probability masks output by the DR image and the target segmentation sub-model are divided into the same patch. When generating the patch vector, the two types of patch vectors are spliced and input into the multi-head attention function. When performing the multi-head attention function processing, the positioning box output by the target detection sub-model is input as the attention prior parameter.
5. A lung nodule analysis method based on a large model as claimed in claim 1, characterized in that: In step 3, when training the lung nodule analysis model, the specific training method is: First, the target segmentation sub-model is trained using lung DR sample images and lung nodule segmentation mask labels to obtain a probability mask; The target detection sub-model is then trained using the lung DR sample images, lung nodule location labels, and the probability mask generated during the target segmentation sub-model training to obtain the lung nodule positioning frame. Finally, the lung nodule localization box detected by the target detection sub-model, the probability mask output by the target segmentation sub-model, and the lung DR sample image are input into the classification sub-model based on the large model. The probability mask and the lung DR sample image are encoded as input data into the Patch image block in the classification sub-model based on the large model, and the lung nodule localization box is provided as prior knowledge to the multi-head attention function of the classification sub-model based on the large model as the spatial alignment offset of the attention.
6. A method for analyzing lung nodules based on a large model as claimed in claim 5, characterized in that: When training the target detection sub-model, the loss function for: ; When training the target segmentation sub-model, the loss function It is expressed as: ; When training the classification sub-model based on the large model, the loss function It is expressed as: ; in, Represents the total number of samples in a batch, represents the true label, represents the model prediction probability, represents the weight of positive samples, represents the modulation factor, Represents the weight of negative samples; represents the weight of the constraint, Indicates the preset threshold value, represents the average lesion probability; represents the KL divergence function, represents the detection quantity distribution vector, represents the classification quantity distribution vector, represents the cross entropy weight, represents the true classification label, Represents the classification prediction result.
7. A pulmonary nodule analysis system based on a large model, characterized in that: include: The sample data and label acquisition module is used to acquire lung DR sample images and annotate them to obtain label data; the labels include lung nodule location labels, lung nodule segmentation mask labels, and lung nodule category labels; A pulmonary nodule analysis model construction module is used to construct a pulmonary nodule analysis model, which includes a target detection sub-model, a target segmentation sub-model, and a classification sub-model based on a large model; The target detection sub-model adopts the YOLOv8 network, including the backbone network, the neck network and the detection head. The four feature maps of different sizes output by the backbone network are input into the neck network. The neck network fuses the feature maps of different sizes and then inputs them into the target detection head of the corresponding size in the detection head. The C2f-nmODE module in the backbone network includes the first CBS layer, the Split layer and the second CBS layer. The output of the first CBS layer is used as the input of the Split layer. The output of the Split layer passes through three nmODE layers and the outputs of the three nmODE layers are fused and then input into the second CBS layer together with the output of the Split layer. Among them, the nmODE layer changes the discretized neural memory ordinary differential equation into a form in which the number of neural network layers is used as the number of derivation steps. A pulmonary nodule analysis model training module is used to train the pulmonary nodule analysis model constructed by the pulmonary nodule analysis model construction module using the sample data and labels acquired by the sample data and label acquisition module; A real-time classification module is used to obtain the lung DR image to be tested and input the lung nodule analysis model trained by the lung nodule analysis model training module. The lung nodule analysis model outputs the lung nodule detection and classification results, which include whether there is a lung nodule, the location of the lung nodule, and whether the lung nodule is benign or malignant; In the pulmonary nodule analysis model construction module, the backbone network includes a first convolution module, a second convolution module, a first C2f-nmODE module, a third convolution module, a second C2f-nmODE module, a fourth convolution module, a third C2f-nmODE module, a fifth convolution module, a Swin Transformer module and an SPPF-LSKA module, which are connected in sequence; the first C2f-nmODE module, the second C2f-nmODE module, the third C2f-nmODE module and the SPPF-LSKA module output 1 / 4 feature maps, 1 / 8 feature maps, 1 / 16 feature maps and 1 / 32 feature maps, respectively.
8. A computer device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that: A computer program is stored, and when the computer program is executed by a processor, the processor is caused to execute the steps of the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Lung CT image physiological detection system and method based on attention mechanism
CN117764923A
Lung CXR image automatic segmentation method, system, device and medium
CN118154863A
Unmanned aerial vehicle image small target detection method based on improved YOLOv8
CN119495036A
Pulmonary nodule CT image analysis method based on YOLO-CSC model
CN116912212A
Cervical cancer cell detection method based on Yolov5l model
CN117274220A