Pulmonary nodule detection incremental learning method and device, equipment, product and storage medium
By applying incremental learning methods of elastic weight integration and feature distillation in lung nodule detection, the catastrophic forgetting problems faced by deep learning models in the incremental learning process are solved, and the continuous improvement of model performance and reduction of time and space costs are achieved.
Patent Information
- Application Number
- CN202411997014.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-06
AI Technical Summary
Deep learning models face catastrophic forgetting problems in lung nodules detection. Direct training with new data sets may lead to a significant decline in the detection performance of the model on the original data set samples.
The incremental learning method based on elastic weight integration (EWC) method and feature distillation method is used to update and strengthen the preset deep learning model parameters to obtain the trained third deep learning model.
It effectively alleviates catastrophic forgetting and reduces the time and space overhead of the incremental update process. It can incrementally update the model without saving the original data set, which has low time and space costs.
Smart Images

Figure CN119941655A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to an incremental learning method, device, electronic device, product and storage medium for lung nodule detection. Background Art
[0002] Automatic nodule detection uses deep learning technology to accurately identify potential nodules in lung computed tomography (CT) images and annotate their exact location and probability information. The deep learning-based computer-aided diagnosis (CAD) system automatically recognizes and learns nodule features by learning and training models based on rich medical imaging data, without relying on manually designed features.
[0003] The environment around different types of nodules is also diverse. Due to the diversity of lung nodule characteristics, the model needs to have the ability to perform incremental learning as new samples arrive. However, directly training with a new dataset may cause the model's detection performance for the original dataset samples to drop significantly, thereby causing catastrophic forgetting. Summary of the invention
[0004] To solve related technical problems, the embodiments of the present application provide an incremental learning method, device, electronic device, product and storage medium for lung nodule detection.
[0005] The technical solution of the embodiment of the present application is implemented as follows:
[0006] The present application embodiment provides an incremental learning method for lung nodule detection, comprising:
[0007] Preprocessing the lung CT file to obtain a first data set including image information of the lung nodules and location information of the lung nodules;
[0008] Obtain a second data set to be learned;
[0009] Based on the Elastic Weight Consolidation (EWC) method, the first data set, and the second data set, parameters of the preset first deep learning model are updated to obtain first parameters and an updated second deep learning model;
[0010] Based on the feature distillation method and the second data set, the second deep learning model is enhanced to obtain second parameters;
[0011] The first deep learning model is trained according to the first parameter and the second parameter to obtain a trained third deep learning model; the third deep learning model is used for incremental learning of lung nodule detection.
[0012] In the above scheme, the method of updating the parameters of the preset first deep learning model based on the elastic weight integration EWC method, the first data set, and the second data set to obtain the first parameters includes:
[0013] Determine a third data set in the first data set that meets a preset confidence level and a first network parameter of the first deep learning model;
[0014] Determine a parameter weight matrix of the first deep learning model according to the third data set and the first network parameter; the parameter weight matrix reflects the importance of the first network parameter to the third data set;
[0015] The first parameter is determined based on the parameter weight matrix and the second data set.
[0016] In the above solution, determining the parameter weight matrix of the first deep learning model according to the third data set and the first network parameters includes:
[0017] Determining a gradient parameter of the third data set and the first network parameter;
[0018] The parameter weight matrix is determined based on the gradient parameters.
[0019] In the above solution, determining the first parameter based on the parameter weight matrix and the second data set includes:
[0020] Determining a penalty parameter of a loss function in the first deep learning model based on the parameter weight matrix; a larger value of the penalty parameter indicates that the first network parameter is more important and less likely to be changed;
[0021] The first deep learning model is updated according to the penalty parameter and back propagation of the second data set to determine the first parameter.
[0022] In the above solution, the second deep learning model is enhanced based on the feature distillation method and the second data set to obtain the second parameter, including:
[0023] Inputting the second data set into the teacher network and the student network in the second deep learning model to obtain a first feature map output by the teacher network and a second feature map output by the student network;
[0024] Generating candidate target information of a region proposal network (RPN) based on the second feature map;
[0025] Inputting the candidate target information into a first prediction module in the teacher network and a second prediction module in the student network to obtain first prediction information of the pulmonary nodules output by the first prediction module and second prediction information of the pulmonary nodules output by the second prediction module;
[0026] The second parameter is determined based on the first feature map, the second feature map, the first prediction information, and the second prediction information.
[0027] In the above solution, the second prediction module includes a classification layer; the method further includes:
[0028] The classification layer is frozen to preserve the classification ability learned by the student network in the previous stage.
[0029] In the above solution, determining the second parameter based on the first feature map, the second feature map, the first prediction information, and the second prediction information includes:
[0030] Determine a first mean square error loss between the first feature map and the second feature map and a second mean square error loss between the first prediction information and the second prediction information;
[0031] The second parameter is determined based on the first mean square error loss and the second mean square error loss.
[0032] In the above solution, the first deep learning model is trained according to the first parameter and the second parameter to obtain a trained third deep learning model, including:
[0033] determining a first weighting factor for the first parameter and a second weighting factor for the second parameter;
[0034] Determine a first loss function based on the first parameter, the second parameter, the first weight factor, and the second weight factor;
[0035] The first deep learning model is trained using the first loss function to obtain a trained third deep learning model.
[0036] The present application also provides an incremental learning device for lung nodule detection, comprising:
[0037] A preprocessing unit, used for preprocessing the lung CT file to obtain a first data set including image information of the lung nodules and location information of the lung nodules;
[0038] An acquisition unit, used for acquiring a second data set to be learned;
[0039] An updating unit, configured to update parameters of a preset first deep learning model based on an EWC method, the first data set, and the second data set to obtain first parameters and an updated second deep learning model;
[0040] An enhancement processing unit, configured to perform enhancement processing on the second deep learning model based on a feature distillation method and the second data set to obtain a second parameter;
[0041] A training unit is used to train the first deep learning model according to the first parameter and the second parameter to obtain a trained third deep learning model; the third deep learning model is used for incremental learning of lung nodule detection.
[0042] The present application also provides an electronic device, including:
[0043] A memory for storing executable instructions;
[0044] The processor is used to implement any step of the above-mentioned method when executing the executable instructions stored in the memory.
[0045] The present application also provides a computer program product, including a computer program, which implements any step of the above method when executed by a processor.
[0046] The embodiment of the present application also provides a computer-readable storage medium storing executable instructions for implementing any step of the above-described method when executed by a processor.
[0047] The embodiments of the present application provide an incremental learning method, device, electronic device, product and storage medium for lung nodule detection, wherein the method comprises: preprocessing a lung CT file to obtain a first data set including image information of lung nodules and location information of the lung nodules; obtaining a second data set to be learned; updating parameters of a preset first deep learning model based on the EWC method, the first data set and the second data set to obtain first parameters and an updated second deep learning model; enhancing the second deep learning model based on the feature distillation method and the second data set to obtain second parameters; training the first deep learning model according to the first parameters and the second parameters to obtain a trained third deep learning model; and using the third deep learning model to perform a training on the first deep learning model. Regarding incremental learning of lung nodule detection, the solution of an embodiment of the present application is to update the parameters of a preset first deep learning model based on the EWC method, a first data set including image information of lung nodules and location information of lung nodules, and a second data set to be learned, to obtain first parameters and an updated second deep learning model; based on the feature distillation method and the second data set, the second deep learning model is enhanced to obtain second parameters; the first deep learning model is trained according to the first parameters and the second parameters to obtain a trained third deep learning model; the third deep learning model is used for incremental learning of lung nodule detection, which effectively alleviates catastrophic forgetting, reduces the time and space overhead of the incremental update process, and can incrementally update the model without saving the original data set, with low time and space costs. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 A schematic diagram of the process of an incremental learning method for lung nodule detection is provided for an embodiment of the present application;
[0049] Figure 2 A flowchart of an incremental learning method for lung nodule detection is provided for an embodiment of the present application;
[0050] Figure 3 This is a schematic diagram of the EWC-P parameter update process for this application;
[0051] Figure 4 This is a schematic diagram of the characteristic distillation process of this application;
[0052] Figure 5 This is a schematic diagram of the structure of an incremental learning device for lung nodule detection according to an embodiment of the present application;
[0053] Figure 6 A schematic diagram of a hardware entity structure of an electronic device in an embodiment of the present application. DETAILED DESCRIPTION
[0054] The present application is further described in detail below in conjunction with the accompanying drawings and embodiments.
[0055] Automatic detection of lung nodules uses deep learning technology to accurately identify potential nodules in lung CT images and annotate their exact location and probability information. The deep learning-based CAD system automatically recognizes and learns nodule features by learning and training models based on rich medical imaging data, without relying on manually designed features. This method greatly reduces the workload of doctors and effectively prevents misjudgments and misdiagnoses that may be caused by visual fatigue. By improving the accuracy and detection efficiency of lung nodule diagnosis, this technology significantly improves the quality of medical image interpretation and provides reliable auxiliary support for clinical work.
[0056] Pulmonary nodules show wide variation in size, shape, and location, and include multiple types, such as ground glass nodules, solid nodules, and semi-solid nodules. The environment around different types of nodules is also diverse. Due to the diversity of pulmonary nodule characteristics, the model needs to have the ability to perform incremental learning as new samples arrive. Direct training with a new dataset may cause the model's detection performance for the original dataset samples to drop significantly, leading to catastrophic forgetting. On the other hand, if the original dataset is retained and all the data is mixed for training the new model, the time and space costs will increase significantly. Therefore, this application aims to find a low-cost incremental update method based on the original model to improve the performance of the model in pulmonary nodule detection.
[0057] Catastrophic forgetting is a key issue in incremental learning, and a variety of methods have emerged in recent years to address this challenge. One of the methods is to adjust the model network framework, that is, to increase the number of parameters on the basis of the original model and expand the model to adapt to the new data set. At the same time, the original model parameters are kept unchanged to maintain the effect on the original data set, thereby alleviating the forgetting phenomenon. Another method is based on experience replay, which divides the memory and retains all or part of the original data set so that the two parts can be used in combination when new data arrives to alleviate the forgetting of the original data. In this method, the way to select the original data usually includes random selection and interval selection. Although saving enough old samples can mitigate catastrophic forgetting in practice, it also increases time and space costs.
[0058] In addition, regularization-based methods limit the update process by introducing different parameter regularization terms to prevent over-updating of important parameters in the original model. The advantage of this type of method is that it does not require modifying the network structure or additional memory space. However, parameter restriction methods may hinder the training process and make it difficult to learn new data. Each method has its own advantages and disadvantages, and needs to be weighed according to specific needs in practical applications.
[0059] Based on this, the embodiment of the present application provides an incremental learning method for lung nodule detection, which is applied to electronic devices. The functions implemented by the method can be implemented by calling program codes by a processor in the electronic device. Of course, the program codes can be stored in a computer storage medium. It can be seen that the electronic device at least includes a processor and a storage medium. As an example, the electronic device can be a mobile phone, a computer, a terminal, an information transceiver device, a tablet device, a personal digital assistant, etc.
[0060] Figure 1 A schematic diagram of the process of an incremental learning method for lung nodule detection is provided for an embodiment of the present application; Figure 1 As shown, the method includes:
[0061] Step S101: preprocessing the lung CT file to obtain a first data set including image information of lung nodules and location information of the lung nodules;
[0062] Step S102: Acquire a second data set to be learned;
[0063] Step S103: updating parameters of a preset first deep learning model based on the EWC method, the first data set, and the second data set to obtain first parameters and an updated second deep learning model;
[0064] Step S104: Based on the feature distillation method and the second data set, the second deep learning model is enhanced to obtain second parameters.
[0065] Step S105: Train the first deep learning model according to the first parameter and the second parameter to obtain a trained third deep learning model; the third deep learning model is used for incremental learning of lung nodule detection.
[0066] It should be noted that the incremental learning method for pulmonary nodule detection can be determined according to actual conditions and is not limited here. As an example, the incremental learning method for pulmonary nodule detection can be an incremental learning method for pulmonary nodule detection based on EWC and knowledge distillation.
[0067] In step 101, the specific preprocessing process of preprocessing the lung CT file to obtain the first data set including the image information of the lung nodules and the position information of the lung nodules can be determined according to the actual situation and is not limited here. As an example, the preprocessing process can be to extract the slices containing nodules in the lung CT file, save them as image information after lung parenchyma segmentation, find the position information of the nodules in the image by calculating the relative coordinates, and use the image information and the position information as the first data set. The first data set can be determined according to the actual situation and is not limited here. As an example, the first data set can be the original data set, which can be recorded as
[0068] In practical applications, the lung computed tomography CT file is first preprocessed, the slices containing nodules are taken out, the lung parenchyma is segmented and saved as an image format, and the specific position of the nodule in the image is found by calculating the relative coordinates. Finally, the image file position and the nodule position information are recorded, and there is a corresponding relationship between the image file position and the nodule position information.
[0069] In step 102, the second data set may be determined according to actual conditions, which is not limited here. As an example, the second data set may be a new data set, which may be recorded as
[0070] In step 103, the first deep learning model may be any deep learning model, which is not limited here. The first deep learning model may be understood as an original model.
[0071] The specific updating process of updating the parameters of the preset first deep learning model based on the EWC method, the first data set, and the second data set to obtain the first parameters and the updated second deep learning model can be determined according to actual conditions and is not limited here. As an example, the updating of the parameters of the preset first deep learning model based on the EWC method, the first data set, and the second data set to obtain the first parameters may include: determining a third data set in the first data set that meets the preset confidence and the first network parameter of the first deep learning model; determining the parameter weight matrix of the first deep learning model according to the third data set and the first network parameter; the parameter weight matrix reflects the importance of the first network parameter to the third data set; determining the first parameter based on the parameter weight matrix and the second data set. Among them, the first parameter can be determined according to actual conditions and is not limited here. As an example, the first parameter may be an EWC-P loss function, which may also be called a network loss. In practical applications, the network loss may be recorded as L(θ) or LF .
[0072] In step 104, the second deep learning model is enhanced based on the feature distillation method and the second data set, and the specific processing process in obtaining the second parameter can be determined according to actual conditions and is not limited here. As an example, the second deep learning model is enhanced based on the feature distillation method and the second data set to obtain the second parameter may include inputting the second data set into the teacher network and the student network in the second deep learning model to obtain the first feature map output by the teacher network and the second feature map output by the student network; generating candidate target information of the region proposal network RPN based on the second feature map; inputting the candidate target information into the first prediction module in the teacher network and the second prediction module in the student network to obtain the first prediction information of the lung nodules output by the first prediction module and the second prediction information of the lung nodules output by the second prediction module; determining the second parameter based on the first feature map, the second feature map, the first prediction information, and the second prediction information. Among them, the second parameter can be determined according to actual conditions and is not limited here. As an example, the second parameter may be a distillation loss function, which may be denoted as L Distill .
[0073] In step 104, the specific training process of training the first deep learning model according to the first parameter and the second parameter to obtain the trained third deep learning model can be determined according to actual conditions and is not limited here. As an example, the training of the first deep learning model according to the first parameter and the second parameter to obtain the trained third deep learning model may include determining a first weight factor of the first parameter and a second weight factor of the second parameter; determining a first loss function based on the first parameter, the second parameter, the first weight factor, and the second weight factor; and training the first deep learning model using the first loss function to obtain a trained third deep learning model.
[0074] The use of the third deep learning model for incremental learning of lung nodule detection can be understood as the third deep learning model can learn new data and incrementally update the model without saving the original data set, which has a low time and space cost, and can effectively alleviate the catastrophic forgetting phenomenon and reduce the time and space overhead of the incremental update process.
[0075] In an embodiment of the present application, the parameters of a preset first deep learning model are updated based on the EWC method, a first data set including image information of lung nodules and location information of lung nodules, and a second data set to be learned, so as to obtain first parameters and an updated second deep learning model; the second deep learning model is enhanced based on the feature distillation method and the second data set to obtain second parameters; the first deep learning model is trained according to the first parameters and the second parameters to obtain a trained third deep learning model; the third deep learning model is used for incremental learning of lung nodule detection, which effectively alleviates the catastrophic forgetting phenomenon, reduces the time and space overhead of the incremental update process, and can incrementally update the model without saving the original data set, with low time and space costs.
[0076] In one embodiment, the updating of parameters of a preset first deep learning model based on the elastic weight integration EWC method, the first data set, and the second data set to obtain the first parameter includes:
[0077] Determine a third data set in the first data set that meets a preset confidence level and a first network parameter of the first deep learning model;
[0078] Determine a parameter weight matrix of the first deep learning model according to the third data set and the first network parameter; the parameter weight matrix reflects the importance of the first network parameter to the third data set;
[0079] The first parameter is determined based on the parameter weight matrix and the second data set.
[0080] The preset confidence level can be determined according to actual conditions and is not limited here. As an example, the preset confidence level can be samples with a confidence level ranging from 75 to 95.
[0081] The specific determination process of determining the third data set in the first data set that meets the preset reliability can be determined according to actual conditions and is not limited here. As an example, the determination of the third data set in the first data set that meets the preset reliability can be to select a data set that meets the preset reliability from the first data set as the third data set; the first data set can be the original data set, which can be recorded as In practical applications, the original data set can be The samples with confidence levels ranging from 75 to 95 are selected as the third data set.
[0082] The specific determination process in determining the first network parameter of the first deep learning model can be determined according to actual conditions and is not limited here. As an example, the first network parameter of the first deep learning model can be extracted from the first deep learning model; wherein the first deep learning model can be the original model; the first network parameter can be referred to as the network parameter, which can be recorded as θ1, θ 2。。。。。。 θ n In practical applications, the network parameters θ1, θ 2。。。。。。 θ n .
[0083] The specific determination process in determining the parameter weight matrix of the first deep learning model according to the third data set and the first network parameters can be determined according to actual conditions and is not limited here. As an example, determining the parameter weight matrix of the first deep learning model according to the third data set and the first network parameters may include determining the gradient parameters of the third data set and the first network parameters; and determining the parameter weight matrix based on the gradient parameters.
[0084] The specific determination process in determining the first parameter based on the parameter weight matrix and the second data set can be determined according to actual conditions and is not limited here. As an example, determining the first parameter based on the parameter weight matrix and the second data set may include determining a penalty parameter of a loss function in the first deep learning model based on the parameter weight matrix; the larger the value of the penalty parameter, the more important the first network parameter is and the less likely it is to be changed; and the first deep learning model is back-propagated according to the penalty parameter and the second data set to update the parameters and determine the first parameter.
[0085] In one embodiment, determining a parameter weight matrix of the first deep learning model according to the third data set and the first network parameters includes:
[0086] Determining a gradient parameter of the third data set and the first network parameter;
[0087] The parameter weight matrix is determined based on the gradient parameters.
[0088] In this embodiment, the specific determination process in determining the gradient parameters of the third data set and the first network parameter can be determined according to actual conditions, which is not limited here; as an example, the determination of the gradient parameters of the third data set and the first network parameter can be understood as performing a gradient calculation on the third data set and the first network parameter to determine the gradient parameters. The determination of the gradient parameters by performing a gradient calculation on the third data set and the first network parameter can specifically be performing a gradient derivative calculation on the third data set and the first network parameter to determine the gradient parameters. Among them, the gradient parameters can be determined according to actual conditions, which is not limited here; as an example, the gradient parameters can be recorded as J n .
[0089] The specific determination process in determining the parameter weight matrix based on the gradient parameter can be determined according to actual conditions, which is not limited here; as an example, determining the parameter weight matrix based on the gradient parameter can be calculating the parameter weight matrix based on the gradient parameter. The parameter weight matrix can be determined according to actual conditions, which is not limited here; as an example, the parameter weight matrix can be recorded as P n .
[0090] For ease of understanding, here is an example. First, from the original data set Select samples with confidence levels ranging from 75 to 95 and calculate a gradient derivative with the original model network parameters, and use the gradient to calculate the parameter weight matrix P n This matrix reflects the importance of each parameter in the original model network to the original data set. The specific process is to assume that the original data set is The new dataset is The parameter θ is equal to the given data and The optimal parameter value under the condition is obtained, so the model parameter update is expressed as a conditional probability optimization process, and the model parameter θ can be expressed by referring to the following formula (1):
[0091]
[0092] Assumptions are independent of each other. We can use the conditional probability formula and Bayes' rule to decompose the above formula and calculate the conditional probability You can refer to the following formula (2):
[0093]
[0094] In the above formula (2), the first term on the right side depends only on the new data set The loss function is -L N(θ). The third term is a constant and has nothing to do with the parameters. The optimization objective can be referred to in the following formula (3):
[0095]
[0096] The current optimization goal can minimize the loss on the new data or maximize the posterior distribution of the parameters on the original data set. Can effectively avoid catastrophic forgetting.
[0097] because The calculation is difficult, and the Laplace approximation method is used to approximate a Gaussian distribution. θ represents the original model parameters, maximize Transform to Maximize The Hessian matrix is n×n. This method has a high computational complexity. The expectation of the Hessian matrix is negated and converted into the Fisher information matrix F. The Fisher information matrix is approximated as a first-order parameter weight matrix through the Jacobian matrix to simplify the calculation. The Jacobian matrix can be calculated through the current gradient, which can be referred to as the following formula (4):
[0098]
[0099] Where n represents the batch sample of the nth iteration, is the number of samples, and L is the loss function to be minimized. The parameter weight matrix calculation method can refer to the following formula (5):
[0100]
[0101] The expansion formula can refer to the following formula (6):
[0102]
[0103] In one embodiment, determining the first parameter based on the parameter weight matrix and the second data set includes:
[0104] Determining a penalty parameter of a loss function in the first deep learning model based on the parameter weight matrix; a larger value of the penalty parameter indicates that the first network parameter is more important and less likely to be changed;
[0105] The first deep learning model is updated according to the penalty parameter and back propagation of the second data set to determine the first parameter.
[0106] It should be noted that the larger the value of the penalty parameter, the more important the first network parameter is, and the less likely it is to be changed, which can be understood as the more important the parameter is, the larger the penalty term is, and the less likely it is to be changed. Among them, the penalty parameter can be determined according to actual conditions and is not limited here; as an example, the penalty parameter can be a regularization penalty term.
[0107] The specific determination process of determining the penalty parameter of the loss function in the first deep learning model based on the parameter weight matrix can be determined according to actual conditions and is not limited here; as an example, the penalty parameter of the loss function in the first deep learning model based on the parameter weight matrix can be to introduce a regularization penalty term into the loss function based on the parameter weight matrix.
[0108] The first deep learning model is updated according to the penalty parameter and the back propagation of the second data set, and the specific determination process in determining the first parameter can be determined according to actual conditions and is not limited here. Among them, the first parameter can be an EWC-P loss function.
[0109] For ease of understanding, here is an example to introduce a regularized penalty term into the loss function based on the parameter weight matrix. The more important the parameter, the larger the penalty term and the less likely it is to be changed. Finally, in the training process of new samples, the loss function with the regularized penalty term is back-propagated to update the parameters. The EWC-P loss function can refer to the following formula (7):
[0110]
[0111] In formula (7), L N (θ) is the loss for calculating new data, and λ is a parameter representing the correlation between new and old data. i is the i-th value of the parameter weight matrix, θ i To optimize the i-th parameter in the model, is the i-th parameter in the original model.
[0112] In one embodiment, the step of performing enhancement processing on the second deep learning model based on the feature distillation method and the second data set to obtain the second parameter includes:
[0113] Inputting the second data set into the teacher network and the student network in the second deep learning model to obtain a first feature map output by the teacher network and a second feature map output by the student network;
[0114] Generate candidate target information of a region proposal network RPN based on the second feature map;
[0115] Inputting the candidate target information into a first prediction module in the teacher network and a second prediction module in the student network to obtain first prediction information of the pulmonary nodules output by the first prediction module and second prediction information of the pulmonary nodules output by the second prediction module;
[0116] The second parameter is determined based on the first feature map, the second feature map, the first prediction information, and the second prediction information.
[0117] In this embodiment, the second data set can be determined according to actual conditions and is not limited here. As an example, the second data set can be a new sample.
[0118] The step of inputting the second data set into the teacher network and the student network in the second deep learning model to obtain the first feature map output by the teacher network and the second feature map output by the student network may be the step of inputting the second data set into the backbone network of the teacher network and the backbone network of the student network in the second deep learning model to obtain the first feature map output by the teacher network and the second feature map output by the student network. This process may be referred to as feature map distillation loss. The first feature map and the second feature map may be determined based on actual conditions and are not limited here. As an example, the first feature map may be a feature map of the teacher network, and the first feature map may be denoted as F T ; The second feature map can be a feature map of the student network; The second feature map can be denoted by F S .
[0119] The specific generation process of the candidate target information of the region proposal network RPN generated based on the second feature map can be determined according to actual conditions, and is not limited here. As an example, the candidate target information of the region proposal network RPN generated based on the second feature map can be the candidate target information of the region proposal network RPN generated by passing the second feature map through its own region. Among them, the candidate target information can be determined according to actual conditions, and is not limited here. As an example, the candidate target information can include candidate boxes. In actual applications, the feature map of the student network passes through its own region generation network RPN to generate candidate boxes.
[0120] The first prediction module and the second prediction module can be determined according to actual conditions, and are not limited here. As an example, the first prediction module and the second prediction module can both include a region of interest pooling layer (ROI pooling) and a prediction output layer, wherein the prediction output layer includes a classification layer and a regression layer. The classification layer outputs two categories, namely, nodule probability and non-nodule probability.
[0121] The first prediction information and the second prediction information can be determined according to actual conditions and are not limited here. As an example, the first prediction information can be recorded as t T The second prediction information can be recorded as t S .
[0122] The candidate target information is input into the first prediction module in the teacher network and the second prediction module in the student network to obtain the first prediction information of the lung nodules output by the first prediction module and the second prediction information of the lung nodules output by the second prediction module. This process can be understood as the nodule position prediction distillation loss.
[0123] The specific determination process in determining the second parameter based on the first feature map, the second feature map, the first prediction information, and the second prediction information can be determined according to actual conditions and is not limited here. As an example, determining the second parameter based on the first feature map, the second feature map, the first prediction information, and the second prediction information may include determining a first mean square error loss between the first feature map and the second feature map and a second mean square error loss between the first prediction information and the second prediction information; determining the second parameter based on the first mean square error loss and the second mean square error loss.
[0124] In one embodiment, the second prediction module includes a classification layer; the method further includes:
[0125] The classification layer is frozen to preserve the classification ability learned by the student network in the previous stage.
[0126] In this embodiment, the main consideration is that if only new samples are used for distillation, the classification ability learned by the student network in the original dataset may be forgotten, resulting in an increase in the false positive rate. To solve this problem, this proposal adopts a strategy of freezing the classification layer of the student network to maintain the classification ability learned by the student network in the previous stage, thereby avoiding the occurrence of forgetting.
[0127] In one embodiment, determining the second parameter based on the first feature map, the second feature map, the first prediction information, and the second prediction information includes:
[0128] Determine a first mean square error loss between the first feature map and the second feature map and a second mean square error loss between the first prediction information and the second prediction information;
[0129] The second parameter is determined based on the first mean square error loss and the second mean square error loss.
[0130] In this embodiment, the specific determination process of determining the first mean square error loss between the first feature map and the second feature map and the second mean square error loss between the first prediction information and the second prediction information can be determined according to actual conditions and is not limited here. As an example, the determination of the first mean square error loss between the first feature map and the second feature map may be to calculate the first mean square error loss between the first feature map and the second feature map; the determination of the second mean square error loss between the first prediction information and the second prediction information may be to calculate the second mean square error loss between the first prediction information and the second prediction information; wherein, the first mean square error loss and the second mean square error loss may both be referred to as mean square error loss; the mean square error loss may also be referred to as L2 regression loss, and the L2 regression loss may be recorded as L R In practical applications, the first feature map can be denoted by F T ; The second feature map can be denoted as F S ; The first mean square error loss can be recorded as L R (F T , F S ). The first prediction information can be recorded as t T The second prediction information can be recorded as t S ; The second mean square error loss can be recorded as L R (F T , F S ).
[0131] The specific determination process of determining the second parameter based on the first mean square error loss and the second mean square error loss can be determined according to actual conditions and is not limited here. As an example, the determination of the second parameter based on the first mean square error loss and the second mean square error loss can be the determination of the second parameter based on the first mean square error loss and the second mean square error loss according to a preset algorithm; wherein the preset algorithm can be determined according to actual conditions and is not limited here. As an example, the preset algorithm can be an addition algorithm. The second parameter can be a distillation loss, and the distillation loss can be denoted as L Distill .
[0132] In practical applications, the preprocessed CT images are simultaneously input into the backbone networks of the teacher network and the student network to obtain two feature maps. Then, the mean square error is used to calculate the correlation between the two feature maps. The feature map of the student network is passed through its own region generation network (RPN) to generate candidate boxes, which are then passed to the prediction modules of the two networks. The prediction module contains a ROI pooling layer and a prediction output layer, where the prediction output layer contains a classification layer and a regression layer. The classification layer outputs two categories, namely nodule probability and non-nodule probability, and the classification layer is then optimized using cross entropy loss.
[0133] However, since only new samples are used for distillation, the classification ability learned by the student network in the original dataset may be forgotten, resulting in an increase in the false positive rate. To solve this problem, this proposal adopts a strategy of freezing the classification layer of the student network to maintain the classification ability learned by the student network in the previous stage, thereby avoiding the occurrence of forgetting. Finally, by calculating the distillation loss of the regression result, the distillation losses obtained by the two calculations are added together to obtain the complete distillation loss. This method helps to maintain the learning results of the student network at different stages and improve the robustness of the model. The distillation loss is calculated as shown in the following formula (8):
[0134] L Distill =L R (F T ,F S )+L R (t T ,t S ) (8);
[0135] Where L R is the L2 regression loss, F T , F S is the feature map obtained by image I through the model backbone network, t T , t S It is the location of the lung nodule prediction box predicted by Roi pooling features.
[0136] In one embodiment, the training of the first deep learning model according to the first parameter and the second parameter to obtain a trained third deep learning model includes:
[0137] determining a first weighting factor for the first parameter and a second weighting factor for the second parameter;
[0138] Determine a first loss function based on the first parameter, the second parameter, the first weight factor, and the second weight factor;
[0139] The first deep learning model is trained using the first loss function to obtain a trained third deep learning model.
[0140] In this embodiment, the specific determination process of determining the first weight factor of the first parameter and the second weight factor of the second parameter can be determined according to actual conditions and is not limited here. The first weight factor represents the proportion of the first parameter; the second weight factor represents the proportion of the second parameter.
[0141] The first weight factor and the second weight factor can be determined according to actual conditions and are not limited here. As an example, the second weight factor can be recorded as α; and the first weight factor can be recorded as 1-α.
[0142] The specific determination process in determining the first loss function based on the first parameter, the second parameter, the first weight factor, and the second weight factor can be determined according to actual conditions and is not limited here. As an example, the first loss function determined based on the first parameter, the second parameter, the first weight factor, and the second weight factor can be determined based on the first parameter, the second parameter, the first weight factor, and the second weight factor according to a preset algorithm; wherein the first loss function can be referred to as a loss function, which can be recorded as L total The preset algorithm can be determined according to actual conditions and is not limited here. As an example, the preset algorithm can refer to the following formula (9):
[0143] L total =αL Distill +(1-α)L F (9);
[0144] Among them, α is the balancing factor that controls the weights of each item in the loss function.
[0145] The loss function can be understood as the distillation loss and the network loss L F The weighted sum of .
[0146] The first deep learning model is trained using the first loss function to obtain a trained third deep learning model, which can be understood as using the distillation loss and the network loss L F The first deep learning model is trained by the weighted sum of to obtain a trained third deep learning model.
[0147] In this application, CT image preprocessing; EWC-P method training optimization model; feature distillation method to strengthen the optimization model, can effectively alleviate the catastrophic forgetting phenomenon, reduce the time and space overhead of the incremental update process, and can incrementally update the model without saving the original data set, with low time and space costs.
[0148] For ease of understanding, an example is given here to illustrate that the incremental learning method for lung nodule detection can be specifically an incremental learning method for lung nodule detection based on elastic weight integration EWC and knowledge distillation. In order to effectively alleviate the catastrophic forgetting phenomenon and reduce the time and space overhead of the incremental update process, this proposal proposes an incremental learning method for lung nodule detection based on elastic weight integration and feature distillation. This proposal does not need to save the original data set, and can incrementally update the model, which has a low time and space cost.
[0149] The process of solving the technical problem in this application is: CT image preprocessing; EWC-P method training optimization model; feature distillation method to strengthen the optimization model. The flowchart of this application is as follows Figure 2 As shown, Figure 2 A flowchart of an incremental learning method for lung nodule detection is provided for an embodiment of the present application.
[0150] Below Figure 1 The constructed technical processes are described separately, and the specific steps include:
[0151] Step 1: Data preprocessing.
[0152] First, the lung computed tomography CT file is preprocessed, the slices containing nodules are taken out, the lung parenchyma is segmented and saved as an image format, and the specific position of the nodule in the image is found by calculating the relative coordinates. Finally, the image file position and nodule position information are recorded.
[0153] Step 2: EWC-P method parameter update.
[0154] When using the EWC (Elastic Weight Consolidation) loss function, the more important the parameters in the original data set, the higher the loss value, and accordingly, the less likely these parameters are to change. This helps to retain the original model's ability to distinguish lung nodules in the original data set and improves the ability of the optimized model to reduce false positives. However, due to the sample imbalance problem in the lung nodule data set, it is not very reasonable to use all samples in the original data set to calculate parameter weights.
[0155] In focal loss, confidence is used to distinguish between easy and difficult samples, where high confidence represents easy samples and low confidence represents difficult samples. However, when there are many easy samples in the dataset, they may dominate the loss function, so the loss function should pay more attention to difficult samples. However, in the lung nodule dataset, overly difficult samples may be missed or mislabeled nodules, and over-focusing on these samples is also disadvantageous.
[0156] In view of these considerations, this proposal decided to exclude samples that are too simple or too difficult when calculating the parameter weight matrix. This is because these samples may have an adverse effect on the calculation of the weight matrix, especially when there are a large number of simple samples and a small number of extremely difficult samples. By eliminating these samples, the parameter weights can be calculated more accurately, improving the stability of the model during the optimization process of the loss function.
[0157] The parameter update flow chart of the EWC-P method is as follows: Figure 3 As shown, Figure 3 This is a schematic diagram of the EWC-P parameter update process for this application. First, from the original data set Select samples with confidence levels ranging from 75 to 95 and calculate a gradient derivative with the original model network parameters, and use the gradient to calculate the parameter weight matrix P n . This matrix reflects the importance of each parameter in the original model network to the original data set. Then, based on the parameter weight matrix, a regularization penalty term is introduced into the loss function. The more important the parameter, the larger the penalty term and the less likely it is to be changed. Finally, in the new sample training process, the loss function with the regularization penalty term is back-propagated to update the parameters.
[0158] The following is the proof and derivation process of the EWC-P method. Assume that the original data set is The new dataset is The parameter θ is equal to the given data and Therefore, the model parameter update is expressed as a conditional probability optimization process, and the model parameter θ can refer to the previous formula (1).
[0159] Assumptions They are independent of each other and can be decomposed using the conditional probability formula and Bayesian rule. The conditional probability can be calculated by referring to the previous formula (2).
[0160] In the above formula (2), the first term on the right side depends only on the new data set The loss function is -L N (θ). The third term is a constant and has nothing to do with the parameters, so the optimization objective becomes the previous formula (3).
[0161] The current optimization goal can minimize the loss on the new data or maximize the posterior distribution of the parameters on the original data set. Can effectively avoid catastrophic forgetting.
[0162] because The calculation is difficult, and the Laplace approximation method is used to approximate a Gaussian distribution. θ represents the original model parameters, maximize Transform to Maximize The Hessian matrix is n×n. This method has a high computational complexity. The expectation of the Hessian matrix is negated and converted into the Fisher information matrix F. The Fisher information matrix is approximated as a first-order parameter weight matrix through the Jacobian matrix to simplify the calculation. The Jacobian matrix can be calculated through the current gradient, which can refer to the previous formula (4).
[0163] Where n represents the batch sample of the nth iteration, is the number of samples, and L is the loss function to be minimized. The parameter weight matrix calculation method can refer to the previous formula (5).
[0164] The expansion formula can refer to the previous formula (6).
[0165] Finally, the EWC-P loss function can be expressed by referring to the previous formula (7).
[0166] In formula (7), L N (θ) is the loss for calculating new data, and λ is a parameter representing the correlation between new and old data. i is the i-th value of the parameter weight matrix, θ i To optimize the i-th parameter in the model, is the i-th parameter in the original model.
[0167] Step 3: Feature distillation.
[0168] Although the EWC-P method limits the changes in important parameters of the original dataset during training to preserve the detection performance of the original dataset, this method will hinder the model convergence process and cause the model to not fit the incremental dataset sufficiently. In order to enhance the learning effect of the model on the incremental dataset, this proposal proposes a feature distillation method inspired by knowledge distillation to distill the intermediate features of the model. The feature distillation process is as follows: Figure 4 As shown, Figure 4 This is a schematic diagram of the characteristic distillation process of this application.
[0169] The preprocessed CT images are simultaneously input into the backbone networks of the teacher network and the student network to obtain two feature maps. Then, the mean square error is used to calculate the correlation between the two feature maps. The feature map of the student network is passed through its own region generation network (RPN) to generate candidate boxes, which are then passed to the prediction modules of the two networks. The prediction module contains a ROI pooling layer and a prediction output layer, where the prediction output layer contains a classification layer and a regression layer. The classification layer outputs two categories, namely nodule probability and non-nodule probability, and the classification layer is then optimized using cross entropy loss.
[0170] However, since only new samples are used for distillation, the classification ability learned by the student network in the original dataset may be forgotten, resulting in an increase in the false positive rate. To solve this problem, this proposal adopts a strategy of freezing the classification layer of the student network to maintain the classification ability learned by the student network in the previous stage, thereby avoiding the occurrence of forgetting. Finally, by calculating the distillation loss of the regression result, the distillation losses obtained by the two calculations are added together to obtain the complete distillation loss. This method helps to maintain the learning results of the student network at different stages and improve the robustness of the model. The calculation of the distillation loss can refer to the previous formula (8).
[0171] Where L R is the L2 regression loss, F T , F S is the feature map obtained by image I through the model backbone network, t T , t S is the predicted box position of the lung nodule predicted by the Roi pooling feature. The final loss function is defined as the distillation loss and the network loss L F The weighted sum of can refer to the previous formula (9).
[0172] Where α is the balancing factor that controls the weights of each item in the loss function.
[0173] This application combines the Faster-R-CNN framework with the elastic weight integration method, selects original data set samples with appropriate confidence intervals to calculate the model parameter weight matrix, uses the weight matrix as a parameter during model incremental learning, and adds an L2 regularization penalty term to the network loss to ensure that the model's fit to the original data set is not destroyed and alleviate the forgetting phenomenon.
[0174] This application uses the optimized model as the student network and the original model as the teacher network, calculates the feature distillation loss, and updates the parameters by the weighted sum of the network loss and the feature distillation loss. The prediction classification layer of the student network is frozen during the distillation process to prevent forgetting.
[0175] The deep learning-based system of this application can automatically identify and learn nodule features through learning and model training of a large amount of medical imaging data, and no longer requires manual feature design. This method can significantly reduce the workload of doctors, prevent misjudgments and misdiagnoses caused by visual fatigue, and improve the accuracy and detection efficiency of pulmonary nodules.
[0176] In order to implement the method of the embodiment of the present application, the embodiment of the present application also provides a lung nodule detection incremental learning device 500, which is set on an electronic device, such as Figure 5 As shown, Figure 5 This is a schematic diagram of the structure of an incremental learning device for lung nodule detection according to an embodiment of the present application, comprising:
[0177] A preprocessing unit 501 is used to preprocess the lung computerized tomography CT file to obtain a first data set including image information of lung nodules and location information of the lung nodules;
[0178] An acquisition unit 502 is used to acquire a second data set to be learned;
[0179] An updating unit 503 is used to update parameters of a preset first deep learning model based on an elastic weight integration EWC method, the first data set, and the second data set to obtain first parameters and an updated second deep learning model;
[0180] An enhancement processing unit 504 is used to perform enhancement processing on the second deep learning model based on a feature distillation method and the second data set to obtain a second parameter;
[0181] The training unit 505 is used to train the first deep learning model according to the first parameter and the second parameter to obtain a trained third deep learning model; the third deep learning model is used for incremental learning of lung nodule detection.
[0182] Here, in one embodiment, the updating unit 503 is also used to determine a third data set in the first data set that meets a preset confidence level and a first network parameter of the first deep learning model; determine a parameter weight matrix of the first deep learning model based on the third data set and the first network parameter; the parameter weight matrix reflects the importance of the first network parameter to the third data set; and determine the first parameter based on the parameter weight matrix and the second data set.
[0183] Here, in one embodiment, the updating unit 503 is further used to determine the gradient parameters of the third data set and the first network parameters; and determine the parameter weight matrix based on the gradient parameters.
[0184] Here, in one embodiment, the updating unit 503 is also used to determine the penalty parameter of the loss function in the first deep learning model based on the parameter weight matrix; the larger the value of the penalty parameter, the more important the first network parameter is and the less likely it is to be changed; the first deep learning model is updated according to the penalty parameter and the back propagation of the second data set to determine the first parameter.
[0185] Here, in one embodiment, the enhancement processing unit 504 is also used to input the second data set into the teacher network and the student network in the second deep learning model to obtain the first feature map output by the teacher network and the second feature map output by the student network; generate candidate target information of the region proposal network RPN based on the second feature map; input the candidate target information into the first prediction module in the teacher network and the second prediction module in the student network to obtain the first prediction information of the lung nodules output by the first prediction module and the second prediction information of the lung nodules output by the second prediction module; determine the second parameter based on the first feature map, the second feature map, the first prediction information, and the second prediction information.
[0186] Here, in one embodiment, the second prediction module includes a classification layer; the apparatus 500 further includes a freezing unit for freezing the classification layer to maintain the classification capability learned by the student network in the previous stage.
[0187] Here, in one embodiment, the enhancement processing unit 504 is also used to determine a first mean square error loss between the first feature map and the second feature map and a second mean square error loss between the first prediction information and the second prediction information; and determine the second parameter based on the first mean square error loss and the second mean square error loss.
[0188] Here, in one embodiment, the training unit 505 is also used to determine a first weight factor of the first parameter and a second weight factor of the second parameter; determine a first loss function based on the first parameter, the second parameter, the first weight factor, and the second weight factor; and use the first loss function to train the first deep learning model to obtain a trained third deep learning model.
[0189] It should be noted that: the incremental learning device for pulmonary nodule detection provided in the above embodiment only uses the division of the above program modules as an example when performing incremental learning for pulmonary nodule detection. In actual applications, the above processing can be assigned to different program modules as needed, that is, the internal structure of the device can be divided into different program modules to complete all or part of the processing described above. In addition, the incremental learning device for pulmonary nodule detection provided in the above embodiment and the incremental learning method embodiment for pulmonary nodule detection belong to the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.
[0190] Based on the hardware implementation of the above-mentioned program module, an embodiment of the present application also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and when the processor executes the program, the steps in the incremental learning method for lung nodule detection provided in the above-mentioned embodiment are implemented.
[0191] An embodiment of the present application also provides a computer program product, including a computer program, which, when executed by a processor, implements the steps in the incremental learning method for lung nodule detection provided in the above embodiment.
[0192] Correspondingly, an embodiment of the present application provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps in the incremental learning method for lung nodule detection provided in the above embodiment are implemented.
[0193] It should be noted here that the description of the above storage medium and device embodiments is similar to the description of the above method embodiments, and has similar beneficial effects as the method embodiments. For technical details not disclosed in the storage medium and device embodiments of this application, please refer to the description of the method embodiments of this application for understanding.
[0194] It should be noted that Figure 6 is a schematic diagram of a hardware entity structure of an electronic device in an embodiment of the present application, such as Figure 6 As shown, the hardware entity of the electronic device 600 includes: a processor 601 and a memory 603 . Optionally, the electronic device 600 may also include a communication interface 602 .
[0195] It can be understood that the memory 603 can be a volatile memory or a non-volatile memory, and can also include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic random access memory (FRAM), a flash memory, a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM); the magnetic surface memory can be a disk memory or a tape memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), synchronous static random access memory (SSRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM, SyncLink Dynamic Random Access Memory), and direct RAM bus random access memory (DRRAM, Direct Rambus Random Access Memory).The memory 603 described in the embodiments of the present application is intended to include but is not limited to these and any other suitable types of memories.
[0196] The method disclosed in the above embodiment of the present application can be applied to the processor 601, or implemented by the processor 601. The processor 601 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the hardware integrated logic circuit or software instructions in the processor 601. The above processor 601 may be a general processor, a digital signal processor (DSP, Digital Signal Processor), or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The processor 601 can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general processor can be a microprocessor or any conventional processor, etc. In combination with the steps of the method disclosed in the embodiment of the present application, it can be directly embodied as a hardware decoding processor to execute, or it can be executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium, which is located in the memory 603, and the processor 601 reads the information in the memory 603 and completes the steps of the above method in combination with its hardware.
[0197] In an exemplary embodiment, the device may be implemented by one or more application specific integrated circuits (ASIC), DSP, programmable logic device (PLD), complex programmable logic device (CPLD), field programmable gate array (FPGA), general processor, controller, microcontroller (MCU), microprocessor, or other electronic components to execute the aforementioned method.
[0198] It should be understood that "one embodiment" or "an embodiment" mentioned throughout the specification means that specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present application, the size of the sequence number of the above-mentioned processes does not mean the order of execution, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. The above-mentioned sequence numbers of the embodiments of the present application are only for description and do not represent the advantages and disadvantages of the embodiments.
[0199] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.
[0200] The methods disclosed in several method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.
[0201] The features disclosed in several product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.
[0202] The features disclosed in several method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.
[0203] The above is only an implementation method of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
Claims
1. An incremental learning method for pulmonary nodule detection, characterized in that: include: Preprocessing a lung computed tomography (CT) file to obtain a first data set including image information of a lung nodule and location information of the lung nodule; Obtain a second data set to be learned; Based on the elastic weight integration EWC method, the first data set, and the second data set, the parameters of the preset first deep learning model are updated to obtain first parameters and an updated second deep learning model; Based on the feature distillation method and the second data set, the second deep learning model is enhanced to obtain second parameters; The first deep learning model is trained according to the first parameter and the second parameter to obtain a trained third deep learning model; the third deep learning model is used for incremental learning of lung nodule detection.
2. The method according to claim 1, characterized in that: The method of updating the parameters of the preset first deep learning model based on the elastic weight integration EWC method, the first data set, and the second data set to obtain the first parameters includes: Determine a third data set in the first data set that meets a preset confidence level and a first network parameter of the first deep learning model; Determine a parameter weight matrix of the first deep learning model according to the third data set and the first network parameter; the parameter weight matrix reflects the importance of the first network parameter to the third data set; The first parameter is determined based on the parameter weight matrix and the second data set.
3. The method according to claim 2, characterized in that The determining a parameter weight matrix of the first deep learning model according to the third data set and the first network parameters includes: Determining a gradient parameter of the third data set and the first network parameter; The parameter weight matrix is determined based on the gradient parameters.
4. The method according to claim 2, characterized in that: The determining the first parameter based on the parameter weight matrix and the second data set comprises: Determining a penalty parameter of a loss function in the first deep learning model based on the parameter weight matrix; a larger value of the penalty parameter indicates that the first network parameter is more important and less likely to be changed; The first deep learning model is updated according to the penalty parameter and back propagation of the second data set to determine the first parameter.
5. The method according to claim 1, characterized in that The step of performing enhancement processing on the second deep learning model based on the feature distillation method and the second data set to obtain the second parameter includes: Inputting the second data set into the teacher network and the student network in the second deep learning model to obtain a first feature map output by the teacher network and a second feature map output by the student network; Generate candidate target information of a region proposal network RPN based on the second feature map; Inputting the candidate target information into a first prediction module in the teacher network and a second prediction module in the student network to obtain first prediction information of the pulmonary nodules output by the first prediction module and second prediction information of the pulmonary nodules output by the second prediction module; The second parameter is determined based on the first feature map, the second feature map, the first prediction information, and the second prediction information.
6. The method according to claim 5, characterized in that The second prediction module includes a classification layer; the method further includes: The classification layer is frozen to preserve the classification ability learned by the student network in the previous stage.
7. The method according to claim 6, characterized in that The determining the second parameter based on the first feature map, the second feature map, the first prediction information, and the second prediction information includes: Determine a first mean square error loss between the first feature map and the second feature map and a second mean square error loss between the first prediction information and the second prediction information; The second parameter is determined based on the first mean square error loss and the second mean square error loss.
8. The method according to any one of claims 1 to 7, characterized in that: The step of training the first deep learning model according to the first parameter and the second parameter to obtain a trained third deep learning model includes: determining a first weighting factor for the first parameter and a second weighting factor for the second parameter; Determine a first loss function based on the first parameter, the second parameter, the first weight factor, and the second weight factor; The first deep learning model is trained using the first loss function to obtain a trained third deep learning model.
9. A pulmonary nodule detection incremental learning device, characterized in that: include: A preprocessing unit, used for preprocessing a lung computerized tomography (CT) file to obtain a first data set including image information of a lung nodule and location information of the lung nodule; An acquisition unit, used for acquiring a second data set to be learned; An updating unit, configured to update parameters of a preset first deep learning model based on an elastic weight integration (EWC) method, the first data set, and the second data set to obtain first parameters and an updated second deep learning model; An enhancement processing unit, configured to perform enhancement processing on the second deep learning model based on a feature distillation method and the second data set to obtain a second parameter; A training unit is used to train the first deep learning model according to the first parameter and the second parameter to obtain a trained third deep learning model; the third deep learning model is used for incremental learning of lung nodule detection.
10. An electronic device, characterized in that: include: A memory for storing executable instructions; A processor, used to implement the incremental learning method for pulmonary nodule detection according to any one of claims 1 to 8 when executing the executable instructions stored in the memory.
11. A computer program product, comprising a computer program, characterized in that When executed by a processor, the computer program implements the incremental learning method for pulmonary nodule detection according to any one of claims 1 to 8.
12. A computer-readable storage medium, characterized in that: Executable instructions are stored, which are used to implement the incremental learning method for lung nodule detection described in any one of claims 1 to 8 when executed by a processor.