Pulmonary nodule detection network processing method and device, electronic equipment and storage medium
Through the pseudo label generation method of teacher-student structure, the graph convolution network is used to merge the pseudo label frame and preset label frame, the false positive problem in the lung nodule detection model is solved, and the accuracy of lung nodule detection is improved.
Patent Information
- Application Number
- CN202311458893.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-03
- Publication Date
- 2025-08-05
AI Technical Summary
In the prior art, the lung nodule detection model has many false labeling errors due to the high confidence vacation positive problem, which makes the network difficult to converge, affecting the detection accuracy.
The pseudo-notation generation method of teacher-student structure is adopted to generate prediction annotation boxes through the teacher network, and the graph convolution network is used to merge the pseudo-notation boxes and preset annotation boxes to optimize false positive problems and improve detection performance.
Maximize the use of unfinished image data, optimize the false positive problem of the lung nodule detection model, and improve the annotated image performance of the detection model.
Smart Images

Figure CN120431356A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image processing technology, and in particular to a processing method and device, electronic equipment, and storage medium for a pulmonary nodule detection network. Background Art
[0002] Pulmonary nodules are focal, round, dense shadows of varying sizes, with clear or blurred margins and a diameter of 3 cm or less on lung images. Detecting and classifying pulmonary nodules can help doctors screen for lung cancer and monitor disease progression. Many established pulmonary nodule detection solutions exist, most of which rely on public datasets. The labeling process for these datasets is rigorous and requires the expertise of multiple experienced radiologists. In practice, labeling pulmonary nodule data is a time-consuming and labor-intensive task, limited by the experience and availability of physicians.
[0003] In related technologies, pseudo-labeling methods are usually used to solve the problem of lung nodule detection and labeling in natural images and pathological images. The core of the pseudo-labeling method is to use a model to generate pseudo-labels for the data and merge them with the original labels. The pseudo-labeling method is used to mine unlabeled data and guide model training. However, when merging pseudo-labels with original labels, the confidence level is usually used as a threshold to remove incorrect detection frames. For lung nodule detection, there are many high-confidence false positives. False positives refer to the model identifying content that is not a lung nodule as a lung nodule. Using only the confidence threshold can easily introduce many erroneous pseudo-labels. Using erroneous pseudo-labels to guide network learning will make it difficult for the network to converge.
[0004] Therefore, how to optimize the false positive problem and accurately identify lung nodules in lung images is an urgent problem that needs to be solved. Summary of the Invention
[0005] This disclosure provides a processing method and apparatus, electronic device, and storage medium for a pulmonary nodule detection network. The primary purpose is to address the problem of high-confidence false positives (false positives), where a model identifies something that is not a pulmonary nodule as a pulmonary nodule. Using only confidence thresholds can easily introduce numerous false positives, which can lead to network convergence difficulties.
[0006] According to a first aspect of the present disclosure, a processing method for a pulmonary nodule detection network is provided, comprising:
[0007] Inputting first training image data into a teacher network to obtain a predicted annotation box, wherein the first training image data includes a preset annotation box;
[0008] Extracting a feature vector of the predicted annotation box and inputting the feature vector into a preset graph convolutional network to obtain a pseudo annotation box;
[0009] Merging the pseudo annotation frame with the preset annotation frame to obtain a target annotation frame;
[0010] Inputting the second training image data and the target annotation box into the student network for training to obtain a trained student network, wherein the second training image data is the same as the preset annotation box in the first training image data;
[0011] The parameter weights of the teacher network are updated using the parameter weights of the trained student network.
[0012] Optionally, before inputting the first training image data into the teacher network, the method further includes:
[0013] Normalize the lung image to obtain the lung lobe segmentation image;
[0014] The lung lobe segmentation image is enhanced to obtain the first training image data and the second training image data.
[0015] Optionally, extracting the feature vector of the predicted annotation box includes:
[0016] Sorting the feature vectors of the predicted annotation box according to the confidence of the prediction result, and determining the feature vectors whose ranking order is higher than a preset ranking threshold and has a corresponding relationship with the preset annotation box as positive samples, determining the feature vectors whose ranking order is lower than the preset ranking threshold and has no corresponding relationship with the preset annotation box as negative samples, and determining the remaining feature vectors as samples to be predicted;
[0017] Generate a feature matrix based on the positive samples, negative samples and samples to be predicted;
[0018] The feature similarity between each node in the feature matrix is calculated to obtain an edge matrix.
[0019] Optionally, inputting the feature vector into a preset graph convolutional network to obtain a pseudo-annotated box includes:
[0020] Inputting the feature matrix and the edge matrix into the preset graph convolutional network, and training the preset graph convolutional network using positive samples and negative samples;
[0021] Input the feature matrix and edge matrix into the trained preset graph convolutional network to predict the category of the sample to be predicted;
[0022] The predicted annotation box corresponding to the node predicted as a positive sample in the sample to be predicted is determined as the pseudo annotation box.
[0023] Optionally, inputting the second training image data and the target annotation box into the student network for training to obtain a trained student network includes:
[0024] Obtaining the predicted annotation box output by the student network;
[0025] Calculating a target loss function based on the predicted annotation box output by the student network and the target annotation box;
[0026] The student network is trained based on the target loss function to obtain a trained student network.
[0027] In some embodiments, updating the parameter weights of the teacher network using the parameter weights of the trained student network includes:
[0028] The preset exponential moving average algorithm is called to update the parameter weights of the teacher network using the parameter weights of the trained student network.
[0029] According to a second aspect of the present disclosure, a processing device for a pulmonary nodule detection network is provided, comprising:
[0030] A first input unit is configured to input first training image data into a teacher network to obtain a predicted annotation box, wherein the first training image data includes a preset annotation box;
[0031] An extraction unit, configured to extract a feature vector of the predicted annotation box;
[0032] A second input unit is used to input the feature vector into a preset graph convolutional network to obtain a pseudo-annotated box;
[0033] a merging unit, configured to merge the pseudo annotation frame with the preset annotation frame to obtain a target annotation frame;
[0034] a training unit, configured to input second training image data and the target annotation box into a student network for training to obtain a trained student network, wherein the second training image data is identical to the preset annotation box in the first training image data;
[0035] An updating unit is used to update the parameter weights of the teacher network using the parameter weights of the trained student network.
[0036] In some embodiments, the processing device of the pulmonary nodule detection network further includes:
[0037] a processing unit, configured to perform normalization processing on the lung image to obtain a lung lobe segmentation image before the first input unit inputs the first training image data into the teacher network;
[0038] The enhancement unit is used to perform enhancement processing on the lung lobe segmentation image to obtain the first training image data and the second training image data respectively.
[0039] Optionally, the extraction unit includes:
[0040] A sorting module, configured to sort the feature vectors of the predicted annotation boxes according to the confidence of the prediction results;
[0041] a determination module, configured to determine feature vectors whose ranking order is higher than a preset ranking threshold and which have a corresponding relationship with a preset annotation box as positive samples, feature vectors whose ranking order is lower than the preset ranking threshold and which have no corresponding relationship with the preset annotation box as negative samples, and the remaining feature vectors as samples to be predicted;
[0042] A generation module, configured to generate a feature matrix from the positive samples, negative samples, and samples to be predicted;
[0043] The calculation module is used to calculate the feature similarity between each node in the feature matrix to obtain an edge matrix.
[0044] Optionally, the second input unit includes:
[0045] A training module, configured to input the feature vector into the preset graph convolutional network and train the preset graph convolutional network using positive samples and negative samples;
[0046] A prediction module is used to input the feature matrix and edge matrix into a trained preset graph convolutional network to predict the category of the sample to be predicted;
[0047] The determining module is configured to determine the predicted annotation box corresponding to the node predicted as a positive sample in the sample to be predicted as the pseudo annotation box.
[0048] Optionally, the training unit includes:
[0049] A first training module is configured to input the second training image data and the target annotation box into a student network for training to obtain a trained student network;
[0050] An output module, configured to obtain the predicted annotation box output by the student network;
[0051] A calculation module, configured to calculate a target loss function based on the predicted annotation box output by the student network and the target annotation box;
[0052] The second training module is used to train the student network using the target loss function to obtain a trained student network.
[0053] Optionally, the updating unit includes:
[0054] A first updating module is used to update the parameter weights of the teacher network using the parameter weights of the trained student network;
[0055] The second updating module is used to call a preset exponential moving average algorithm and use the parameter weights of the trained student network to update the parameter weights of the teacher network.
[0056] According to a third aspect of the present disclosure, there is provided an electronic device, including:
[0057] at least one processor; and
[0058] a memory communicatively connected to the at least one processor; wherein,
[0059] The memory stores instructions that can be executed by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to perform the method described in the first aspect.
[0060] According to a fourth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the method described in the first aspect.
[0061] According to a fifth aspect of the present disclosure, a computer program product is provided, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the method as described in the first aspect above.
[0062] The present disclosure provides a processing method and device, electronic device and storage medium for a lung nodule detection network, wherein the first training image data is input into a teacher network to obtain a predicted annotation box, wherein the first training image data contains a preset annotation box; extracting a feature vector of the predicted annotation box, and inputting the feature vector into a preset graph convolutional network to obtain a pseudo annotation box; merging the pseudo annotation box with the preset annotation box to obtain a target annotation box; inputting the second training image data and the target annotation box into a student network for training to obtain a trained student network, wherein the second training image data is the same as the preset annotation box in the first training image data; and updating the parameter weights of the teacher network using the parameter weights of the trained student network. Compared with the related art, the specific implementation process of the embodiment of the present application does not use additional annotation information, nor does it make more demanding assumptions about the data set. It maximizes the use of incompletely annotated image data, adopts a pseudo annotation generation method of a teacher-student structure, merges the pseudo annotation box generated by the teacher network and the preset annotation box, and proposes a merging strategy based on a graph convolutional network to address the false positive problem in the lung nodule detection model; optimizes the false positive problem and improves the performance of the annotated image of the lung nodule detection model.
[0063] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present application, nor is it intended to limit the scope of the present application. Other features of the present application will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.
[0065] Figure 1 A flowchart of a processing method for a pulmonary nodule detection network provided by an embodiment of the present disclosure;
[0066] Figure 2 A flowchart of a method for obtaining a pseudo-annotation frame provided in an embodiment of the present disclosure;
[0067] Figure 3 A flow chart of a method for training a student network provided in an embodiment of the present disclosure;
[0068] Figure 4 A schematic diagram of the structure of a processing device for a pulmonary nodule detection network provided in an embodiment of the present disclosure;
[0069] Figure 5 A schematic diagram of the structure of another processing device of a pulmonary nodule detection network provided in an embodiment of the present disclosure;
[0070] Figure 6 A schematic block diagram of an exemplary electronic device 500 provided in accordance with an embodiment of the present disclosure. DETAILED DESCRIPTION
[0071] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0072] The following describes the processing method and device, electronic device, and storage medium of the lung nodule detection network according to the embodiments of the present disclosure with reference to the accompanying drawings.
[0073] Figure 1 A flowchart of a processing method for a pulmonary nodule detection network provided in an embodiment of the present disclosure is provided.
[0074] like Figure 1 As shown, the method comprises the following steps:
[0075] Step 101: input first training image data into a teacher network to obtain a predicted annotation box, wherein the first training image data includes a preset annotation box.
[0076] In order to collect the predicted annotation box results output by the teacher network, the first training image data is input into the teacher network, and the teacher network predicts the first training image data to obtain the predicted annotation box. The predicted annotation box is a newly generated annotation box for the first training image data. The predicted annotation box is different from the preset annotation box contained in the first training image data. The preset annotation box is a annotation box pre-set by the staff before model training, while the annotation box in the predicted annotation box is a annotation box that has not been pre-marked.
[0077] Step 102: extract the feature vector of the predicted annotation box, and input the feature vector into a preset graph convolutional network to obtain a pseudo annotation box.
[0078] As an implementation method of the embodiment of the present application, the label prediction box result output by the teacher network can be processed using, but not limited to, a non-maximum suppression method to extract a feature vector corresponding to the predicted label box. The feature vector is then input into a preset graph convolutional network to obtain a pseudo-label box based on the preset graph convolutional network.
[0079] Step 103: Merge the pseudo annotation frame with the preset annotation frame to obtain a target annotation frame.
[0080] In order to obtain the target annotation box to train the student network, a graph convolution-based annotation merging method is used to merge the pseudo annotation box with the preset annotation box. The pseudo annotation box is obtained by inputting the feature vector into a preset graph convolution network based on the preset graph convolution network. The preset annotation box is the annotation box pre-set by the staff before model training.
[0081] Step 104 : Input the second training image data and the target annotation box into the student network for training to obtain a trained student network. The second training image data has the same preset annotation box as the first training image data.
[0082] To obtain new annotations generated by the student network, the student network needs to be trained to obtain the trained student network. The second training image data and the target annotation box are input into the student network for training to obtain the trained student network. The second training image data is identical to the preset annotation box in the first training image data. The target annotation box is obtained by merging the pseudo annotation box with the preset annotation box based on a preset graph convolutional network.
[0083] In an embodiment of the present application, the second training image data is the same as the preset annotation box in the first training image data in order to use consistency as a constraint, that is, the pixel values of the input data are different but the model is required to output the same features for the inputs of the two models.
[0084] Step 105: Use the trained parameter weights of the student network to update the parameter weights of the teacher network.
[0085] In order to update the parameter weights of the teacher network, as an implementation method of the embodiment of the present application, the parameter weights of the trained student network can be updated to the parameter weights of the teacher network using an exponential moving average method. The above update method is not intended to limit the update of the parameter weights of the teacher network to only the above formula, and other methods are not limited in the embodiments of the present application.
[0086] The present disclosure provides a processing method for a lung nodule detection network, which inputs first training image data into a teacher network to obtain a predicted annotation box, wherein the first training image data includes a preset annotation box; extracts a feature vector of the predicted annotation box, and inputs the feature vector into a preset graph convolutional network to obtain a pseudo annotation box; merges the pseudo annotation box with the preset annotation box to obtain a target annotation box; inputs second training image data and the target annotation box into a student network for training to obtain a trained student network, wherein the second training image data is the same as the preset annotation box in the first training image data; and uses the parameter weights of the trained student network to update the parameter weights of the teacher network. Compared with the related art, the specific implementation process of the embodiment of the present application does not use additional annotation information, nor does it make more demanding assumptions about the data set. It maximizes the use of incompletely annotated image data, adopts a pseudo annotation generation method of a teacher-student structure, merges the pseudo annotation box generated by the teacher network and the preset annotation box, and proposes a merging strategy based on a graph convolutional network to address the false positive problem in the lung nodule detection model; optimizes the false positive problem and improves the performance of the annotated image of the lung nodule detection model.
[0087] In order to obtain the first training image data and the second training image data, before the first training image data is input into the teacher network, the lung image is normalized to obtain a lung lobe segmentation image; the lung lobe segmentation image is enhanced to obtain the first training image data and the second training image data, respectively.
[0088] In some embodiments, the lung image is normalized, for example, by converting the original data (lung image) to a numerical range using a window width of 1500 and a window level of -500. After the conversion is complete, maximum and minimum value normalization is performed. The lung lobe segmentation region is obtained by obtaining the maximum connected domain and image morphology methods.
[0089] In some embodiments, the lung lobe segmentation image is enhanced by, but not limited to: during training, a 3D image block x0 of a fixed size at a random position is taken from the entire 3D sequence each time, and each image block is first subjected to spatial transformation enhancement, including rotation, flipping, etc. to obtain x1, and then two sets of different pixel value transformation enhancement methods are performed on x1, including saturation, brightness, etc., to obtain x A and x B , due to x A and x B With the same spatial enhancement process, the corresponding annotation box information is completely consistent. A is the first training image data, x B is the second training image data.
[0090] In deep learning, a sufficient number of samples is generally required. The more samples there are, the better the trained model and the stronger the model's generalization ability. However, in practice, insufficient samples or poor quality samples often require data augmentation to improve sample quality. Data augmentation can increase the amount of training data and improve the model's generalization ability.
[0091] As a refinement of step 102, the feature vector of the predicted annotation box is extracted and input into a preset graph convolutional network to obtain a pseudo annotation box, which can be implemented in the following ways but is not limited to: Figure 2 As shown, the method includes:
[0092] In step 201, the feature vectors of the predicted annotation box are sorted according to the confidence of the prediction result, and the feature vectors whose ranking order is higher than a preset ranking threshold and has a corresponding relationship with the preset annotation box are determined as positive samples, the feature vectors whose ranking order is lower than the preset ranking threshold and has no corresponding relationship with the preset annotation box are determined as negative samples, and the feature vectors of the remaining parts are determined as samples to be predicted.
[0093] In some embodiments, the prediction results of the teacher network can be processed using, but not limited to, a non-maximum suppression method to extract the feature vector corresponding to the prediction annotation box of the teacher network. The feature vectors are sorted according to the confidence of the prediction results, and the feature vectors with a ranking higher than a preset ranking threshold and a corresponding relationship with the preset annotation box are determined as positive samples, and the feature vectors with a ranking lower than the preset ranking threshold and no corresponding relationship with the preset annotation box are determined as negative samples, or the feature vectors of low-confidence prediction nodes that are not matched to the annotation are determined as negative samples, and the feature vectors of the remaining parts are determined as samples to be predicted.
[0094] As an implementable method, when sorting by the confidence of the prediction results, the confidence can be sorted from large to small, or from small to large. Specifically, the embodiments of the present application do not limit this.
[0095] The preset ranking threshold is an empirical value and can be flexibly set according to needs, such as 100, 200, or other values. For example, when determining positive samples and negative samples, feature vectors with a confidence ranking higher than 100 can be used as positive samples, and feature vectors with a confidence ranking lower than 100 can be used as negative samples. Alternatively, the 50 feature vectors with the lowest confidence ranking can be used as negative samples, and all feature vectors with other confidence rankings can be used as positive samples. Specifically, the division of positive samples and negative samples is not limited in the embodiments of this application.
[0096] Step 202: Generate a feature matrix based on the positive samples, negative samples, and samples to be predicted.
[0097] The positive samples include feature vectors whose ranking order is higher than the preset ranking threshold, and the negative samples include feature vectors whose ranking order is lower than the preset ranking threshold. The remaining feature vectors are samples to be predicted. A feature matrix is generated based on the feature vectors. In the embodiment of the present application, the feature matrix is an N×M feature matrix F tr .
[0098] Step 203: Calculate the feature similarity between each node in the feature matrix to obtain an edge matrix.
[0099] For the feature matrix F tr The feature similarity between each node is calculated to obtain the edge matrix F te .
[0100] Step 204: Input the feature matrix and edge matrix into the preset graph convolutional network, and train the preset graph convolutional network using positive samples and negative samples.
[0101] Step 205: Input the feature matrix and edge matrix into the trained preset graph convolutional network to predict the category of the sample to be predicted.
[0102] Step 206 : Determine the predicted annotation box corresponding to the node predicted as a positive sample in the sample to be predicted as the pseudo annotation box.
[0103] In order to train the preset graph convolutional network, the feature matrix F tr and edge matrix E te The preset graph convolutional network is input for training. The predicted annotation box information corresponding to the node whose prediction result is a positive sample in the sample to be predicted is determined as a pseudo annotation box.
[0104] In order to obtain the target annotation box To train the student network, a graph convolution-based annotation merging method is used to merge the pseudo annotation frame with the preset annotation frame to obtain the target annotation frame. As an implementation method of the embodiment of the present application, the merging process can be understood as adding the pseudo-annotation box and the preset annotation box, which is not specifically limited.
[0105] As a refinement of step 104, the second training image data and the target annotation box are input into the student network for training to obtain a trained student network, which can be achieved by but not limited to the following methods: Figure 3 As shown, the method includes:
[0106] Step 301: Input the second training image data into the student network to obtain the predicted annotation box output by the student network.
[0107] Step 302: Calculate a target loss function based on the predicted annotation box output by the student network and the target annotation box.
[0108] In one embodiment, the second training image data x is taken B The corresponding predicted annotation box and the output of the graph convolution-based annotation merging module are used to calculate the supervised learning loss L Sup :
[0109]
[0110] Among them, L sup is the supervised learning loss, L cls is the classification loss, L reg is the regression loss, y is the predicted annotation box corresponding to the second training image data, The target annotation box is obtained by merging the pseudo annotation box and the preset annotation box.
[0111] The target loss function is obtained and updated using a back-propagation algorithm.
[0112] Step 303: Continue training the student network based on the target loss function to obtain a trained student network.
[0113] As a refinement of step 105, the parameter weights of the teacher network are updated using the parameter weights of the trained student network, which can be implemented by, but not limited to, the following methods:
[0114] Call the preset exponential moving average algorithm and use the parameter weights of the trained student network to update the parameter weights of the teacher network. The formula is as follows:
[0115] θ t =α*θ t +(1-α)*θ s
[0116] where θ t is the parameter weight of the teacher network, θ s is the parameter weight of the student network, and α is a hyperparameter that controls the update speed.
[0117] The above formula is not intended to limit the updating of the parameter weights of the teacher network to only the above formula, and other methods are not limited in the embodiments of this application.
[0118] In summary, the embodiments of the present disclosure can achieve the following effects:
[0119] 1. According to the loss function, the pulmonary nodule detection model is adjusted and optimized accordingly to improve the performance of the pulmonary nodule detection model in labeling the image data to be detected.
[0120] 2. Optimized the false positive problem in the lung nodule detection model.
[0121] Corresponding to the aforementioned method for optimizing lung nodule detection, the present invention also provides an optimized device for lung nodule detection. Since the device embodiments of the present invention correspond to the aforementioned method embodiments, details not disclosed in the device embodiments can be referred to the aforementioned method embodiments and will not be further described in this invention.
[0122] Figure 4 A schematic diagram of a processing device for a pulmonary nodule detection network according to an embodiment of the present disclosure is shown in FIG. Figure 4 Shown, including:
[0123] A first input unit 41 is configured to input first training image data into a teacher network to obtain a predicted annotation box, wherein the first training image data includes a preset annotation box;
[0124] An extraction unit 42, configured to extract a feature vector of the predicted annotation box;
[0125] A second input unit 43 is used to input the feature vector into a preset graph convolutional network to obtain a pseudo-annotated box;
[0126] a merging unit 44, configured to merge the pseudo annotation frame with the preset annotation frame to obtain a target annotation frame;
[0127] A training unit 45 is configured to input second training image data and the target annotation box into a student network for training to obtain a trained student network, wherein the second training image data has the same preset annotation box as the first training image data;
[0128] The updating unit 46 is configured to update the parameter weights of the teacher network using the parameter weights of the trained student network.
[0129] The processing device of the lung nodule detection network provided by the present disclosure inputs the first training image data into the teacher network to obtain a predicted annotation box, wherein the first training image data includes a preset annotation box; extracts the feature vector of the predicted annotation box, and inputs the feature vector into the preset graph convolution network to obtain a pseudo annotation box; merges the pseudo annotation box with the preset annotation box to obtain a target annotation box; inputs the second training image data and the target annotation box into the student network for training to obtain a trained student network, wherein the second training image data is the same as the preset annotation box in the first training image data; and uses the parameter weights of the trained student network to update the parameter weights of the teacher network. Compared with the related art, the processing device of the lung nodule detection network provided by the present disclosure does not use additional annotation information, nor does it make more demanding assumptions on the data set. It maximizes the use of incompletely annotated data, adopts the pseudo annotation generation method of the teacher-student structure, and uses the annotation merging module to merge the pseudo annotation box generated by the teacher network and the preset annotation box. To address the false positive problem in the lung nodule detection model, a merging strategy based on graph convolutional networks was proposed; this optimized the false positive problem and improved the performance of the lung nodule detection model in annotating images.
[0130] Furthermore, in a possible implementation of this embodiment, as Figure 5 As shown, the extraction unit 42 includes:
[0131] A sorting module 421 is used to sort the feature vectors of the predicted annotation boxes according to the confidence of the prediction results;
[0132] Determination module 422, configured to determine feature vectors whose ranking is higher than a preset ranking threshold and which correspond to a preset annotation box as positive samples, feature vectors whose ranking is lower than the preset ranking threshold and which do not correspond to the preset annotation box as negative samples, and the remaining feature vectors as samples to be predicted;
[0133] A generating module 423 is configured to generate a feature matrix from the positive samples, negative samples and samples to be predicted;
[0134] The calculation module 424 is used to calculate the feature similarity between each node in the feature matrix to obtain an edge matrix.
[0135] Furthermore, in a possible implementation of this embodiment, as Figure 5 As shown, the second input unit 43 includes:
[0136] A training module 431 is configured to input the feature vector into a preset graph convolutional network and train the preset graph convolutional network using positive samples and negative samples;
[0137] A prediction module 432 is configured to input the feature matrix and edge matrix into a trained preset graph convolutional network to predict the category of the sample to be predicted;
[0138] The determination module 433 is configured to determine the predicted annotation box corresponding to the node predicted as a positive sample in the sample to be predicted as the pseudo annotation box.
[0139] Furthermore, in a possible implementation of this embodiment, as Figure 5 As shown, the training unit 45 includes:
[0140] A first training module 451 is configured to input the second training image data and the target annotation box into the student network for training to obtain a trained student network;
[0141] Output module 452, used to obtain the predicted annotation box output by the student network;
[0142] A calculation module 453 is used to calculate a target loss function based on the predicted annotation box output by the student network and the target annotation box;
[0143] The second training module 454 is used to train the student network using the target loss function to obtain a trained student network.
[0144] Furthermore, in a possible implementation of this embodiment, as Figure 5 As shown, the updating unit 46 includes:
[0145] A first updating module 461 is configured to update the parameter weights of the teacher network using the parameter weights of the trained student network;
[0146] The second updating module 462 is used to call a preset exponential moving average algorithm and use the parameter weights of the trained student network to update the parameter weights of the teacher network.
[0147] It should be noted that the above explanation of the method embodiment is also applicable to the device of this embodiment, and the principles are the same, which is not limited in this embodiment.
[0148] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0149] Figure 6 A schematic block diagram of an example electronic device 500 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0150] like Figure 6 As shown, the device 500 includes a computing unit 501, which can perform various appropriate actions and processes according to a computer program stored in a ROM (Read-Only Memory) 502 or a computer program loaded from a storage unit 508 into a RAM (Random Access Memory) 503. Various programs and data required for the operation of the device 500 can also be stored in the RAM 503. The computing unit 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An I / O (Input / Output) interface 505 is also connected to the bus 504.
[0151] Various components in device 500 are connected to I / O interface 505, including: an input unit 506, such as a keyboard, mouse, etc.; an output unit 507, such as various types of displays, speakers, etc.; a storage unit 508, such as a magnetic disk, optical disk, etc.; and a communication unit 509, such as a network card, modem, wireless communication transceiver, etc. The communication unit 509 allows device 500 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0152] The computing unit 501 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a CPU (Central Processing Unit), a GPU (Graphic Processing Unit), various specialized AI (Artificial Intelligence) computing chips, various computing units that run machine learning model algorithms, a DSP (Digital Signal Processor), and any suitable processor, controller, microcontroller, etc. The computing unit 501 performs the various methods and processes described above, such as the processing method for pulmonary nodule detection. For example, in some embodiments, the processing method for pulmonary nodule detection can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded into the RAM 503 and executed by the computing unit 501, one or more steps of the method described above can be performed. Alternatively, in other embodiments, the computing unit 501 may be configured to execute the aforementioned processing method for pulmonary nodule detection in any other appropriate manner (eg, by means of firmware).
[0153] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, FPGAs (Field Programmable Gate Arrays), ASICs (Application-Specific Integrated Circuits), ASSPs (Application Specific Standard Products), SOCs (System on Chips), CPLDs (Complex Programmable Logic Devices), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0154] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0155] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or apparatus. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, RAM, ROM, EPROM (Electrically Programmable Read-Only-Memory) or flash memory, optical fiber, CD-ROM (Compact Disc Read-Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0156] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (Cathode-Ray Tube) or LCD (Liquid Crystal Display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0157] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: LAN (Local Area Network), WAN (Wide Area Network), the Internet, and blockchain networks.
[0158] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact via a communication network. This client-server relationship is established by computer programs running on the respective computers, establishing a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host, a host product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosts and VPS services ("Virtual Private Servers" or simply "VPS"). The server may also be a server in a distributed system or a server integrated with blockchain.
[0159] It's important to note that artificial intelligence (AI) is the study of how computers can simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). This encompasses both hardware and software technologies. AI hardware technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, and big data processing. AI software technologies primarily encompass computer vision, speech recognition, natural language processing, machine learning / deep learning, big data processing, and knowledge graphs.
[0160] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.
[0161] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.
Claims
1. A processing method for a pulmonary nodule detection network, characterized in that: include: Inputting first training image data into a teacher network to obtain a predicted annotation box, wherein the first training image data includes a preset annotation box; Extracting a feature vector of the predicted annotation box and inputting the feature vector into a preset graph convolutional network to obtain a pseudo annotation box; Merging the pseudo annotation frame with the preset annotation frame to obtain a target annotation frame; Inputting the second training image data and the target annotation box into the student network for training to obtain a trained student network, wherein the second training image data is the same as the preset annotation box in the first training image data; The parameter weights of the teacher network are updated using the parameter weights of the trained student network.
2. The method according to claim 1, characterized in that Before inputting the first training image data into the teacher network, the method further includes: Normalize the lung image to obtain the lung lobe segmentation image; The lung lobe segmentation image is enhanced to obtain the first training image data and the second training image data.
3. The method according to claim 1, characterized in that Extracting the feature vector of the predicted annotation box includes: Sorting the feature vectors of the predicted annotation box according to the confidence of the prediction result, and determining the feature vectors whose ranking order is higher than a preset ranking threshold and has a corresponding relationship with the preset annotation box as positive samples, determining the feature vectors whose ranking order is lower than the preset ranking threshold and has no corresponding relationship with the preset annotation box as negative samples, and determining the remaining feature vectors as samples to be predicted; Generate a feature matrix based on the positive samples, negative samples and samples to be predicted; The feature similarity between each node in the feature matrix is calculated to obtain an edge matrix.
4. The method according to claim 3, characterized in that Inputting the feature vector into a preset graph convolutional network to obtain a pseudo-annotated frame includes: Inputting the feature matrix and the edge matrix into the preset graph convolutional network, and training the preset graph convolutional network using positive samples and negative samples; Input the feature matrix and edge matrix into the trained preset graph convolutional network to predict the category of the sample to be predicted; The predicted annotation box corresponding to the node predicted as a positive sample in the sample to be predicted is determined as the pseudo annotation box.
5. The method according to claim 4, characterized in that Inputting the second training image data and the target annotation frame into the student network for training to obtain a trained student network includes: Obtaining the predicted annotation box output by the student network; Calculating a target loss function based on the predicted annotation box output by the student network and the target annotation box; The student network is trained based on the target loss function to obtain a trained student network.
6. The method according to claim 1, characterized in that The updating of the parameter weights of the teacher network using the parameter weights of the trained student network includes: The preset exponential moving average algorithm is called to update the parameter weights of the teacher network using the parameter weights of the trained student network.
7. A processing device for a pulmonary nodule detection network, characterized in that: include: A first input unit is configured to input first training image data into a teacher network to obtain a predicted annotation box, wherein the first training image data includes a preset annotation box; An extraction unit, configured to extract a feature vector of the predicted annotation box; A second input unit is used to input the feature vector into a preset graph convolutional network to obtain a pseudo-annotated box; a merging unit, configured to merge the pseudo annotation frame with the preset annotation frame to obtain a target annotation frame; a training unit, configured to input second training image data and the target annotation box into a student network for training to obtain a trained student network, wherein the second training image data is identical to the preset annotation box in the first training image data; An updating unit is used to update the parameter weights of the teacher network using the parameter weights of the trained student network.
8. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 7.
9. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 7.
10. A computer program product, characterized in that The invention comprises a computer program which, when executed by a processor, implements the method according to any one of claims 1 to 7.