A target detection method and device

Through the downsampling and upsampling processing of the dual convolutional neural network model, combined with the feature fusion of the encoder and decoder, the accuracy problem of nodule segmentation in medical images is solved, and efficient nodule region segmentation is achieved.

CN115035054BActive Publication Date: 2025-06-06UNIV OF SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210588021.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-27
Publication Date
2025-06-06
Estimated Expiration
2042-05-27

AI Technical Summary

Technical Problem

The prior art is difficult to accurately segment the nodule areas in medical images using deep learning techniques.

Method used

The dual convolutional neural network model is used to obtain the category probability of the image through the first convolutional neural network model. If it is an abnormal slice, it is input to the second convolutional neural network model for downsampling and upsampling processing. The shallow and deep features are extracted in combination with the encoder and decoder, and the fusion process is performed to segment the nodule area.

Benefits of technology

The segmentation accuracy and segmentation efficiency of the nodule area are improved, the interference of non-anomalous slices on the model is reduced, and the segmentation effect is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115035054B_ABST
    Figure CN115035054B_ABST
Patent Text Reader

Abstract

The present application provides a target detection method and device, which obtains an image to be processed, inputs the image to be processed into a first convolutional neural network model, obtains the category probability of the image to be processed obtained by the first convolutional neural network model, and if the category probability characterizes that the image to be processed is an abnormal slice, the image to be processed is input into a second convolutional neural network model to reduce the interference of non-abnormal slices on the second convolutional neural network model, so that the second convolutional neural network model processes the image to be processed, improves the segmentation accuracy of the nodule area, reduces the workload of the second convolutional neural network model, and improves the segmentation efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular to a target detection method and device. Background Art

[0002] With the development of artificial intelligence technology, it is increasingly widely used in the field of medical image analysis. Among them, for the task of nodule segmentation in medical images, machine learning and deep learning have become the mainstream research methods.

[0003] Among them, deep learning technology can be used to predict the category of each pixel in the image, and the nodule area can be segmented from the image based on the category of the pixel.

[0004] However, how to use deep learning technology to accurately segment the nodule area in the image becomes a problem. Summary of the invention

[0005] This application provides the following technical solutions:

[0006] On the one hand, the present application provides a target detection method, comprising:

[0007] Get the image to be processed;

[0008] Inputting the image to be processed into a first convolutional neural network model to obtain a category probability of the image to be processed obtained by the first convolutional neural network model;

[0009] If the category probability indicates that the image to be processed is an abnormal slice, the image to be processed is input into an encoder of a second convolutional neural network model to obtain shallow features obtained by the encoder performing downsampling processing on the image to be processed;

[0010] Inputting the image to be processed and the shallow features into the decoder of the second convolutional neural network model to obtain deep features obtained by upsampling the image to be processed by the decoder, and fusing the shallow features and the deep features to obtain fused features;

[0011] The fused features are input into the output layer of the second convolutional neural network model to obtain the nodule region of the image to be processed obtained by the output layer.

[0012] Optionally, the step of inputting the image to be processed into an encoder of a second convolutional neural network model to obtain shallow features obtained by the encoder performing downsampling processing on the image to be processed includes:

[0013] Inputting the image to be processed into an encoder of a second convolutional neural network model, wherein the encoder randomly groups input channels of the image to be processed to obtain a plurality of first groups;

[0014] The encoder randomly groups the input channels in each of the first groups to obtain a plurality of second groups corresponding to the first groups;

[0015] The encoder uses a residual network model to extract features of the image to be processed based on input channels in the second group, and uses the extracted features as features corresponding to the second group;

[0016] The encoder inputs the features corresponding to the first group and the second group into the attention model to obtain the features corresponding to the first group obtained by the attention model;

[0017] The connection layer of the encoder concatenates the features corresponding to each of the first groups to obtain shallow features of the image to be processed.

[0018] Optionally, the inputting the image to be processed and the shallow features into the decoder of the second convolutional neural network model to obtain the deep features obtained by the decoder upsampling the image to be processed, and fusing the shallow features and the deep features to obtain the fused features, including:

[0019] The image to be processed and the shallow features are input into the decoder of the second convolutional neural network model, the decoder upsamples the image to be processed to obtain deep features, and uses dense skip connections to fuse the shallow features and the deep features to obtain fused features.

[0020] Optionally, the first convolutional neural network model is trained in the following manner, including:

[0021] Acquire a first training image, where the first training image is annotated with a true category probability, where the true category probability indicates whether the first training image contains a nodule or does not contain a nodule;

[0022] Performing data enhancement, label smoothing, and random erasing on the first training image to obtain a first target training image;

[0023] Inputting the first target training image into a first convolutional neural network model to obtain a category prediction probability of the first target training image predicted by the first convolutional neural network model;

[0024] Determining whether a loss function value of the first convolutional neural network model converges, the loss function value of the first convolutional neural network model representing a difference between the category prediction probability and the true category probability;

[0025] If convergence occurs, the training ends;

[0026] If convergence has not occurred, the parameters of the first convolutional neural network model are updated.

[0027] Optionally, the second convolutional neural network model is trained in the following manner:

[0028] Obtain a second training image with a true box position of a nodule marked;

[0029] Performing data enhancement on the second training image to obtain a second target training image;

[0030] Inputting the second target training image into a second convolutional neural network model to obtain a region of interest predicted by the first convolutional neural network model for the second target training image;

[0031] Determine whether a loss function value of the second convolutional neural network model converges, where the loss function value of the second convolutional neural network model represents a difference between the real frame position and the region of interest;

[0032] If convergence occurs, the training ends;

[0033] If not converged, update the parameters of the second convolutional neural network model.

[0034] Another aspect of the present application provides a target detection device, comprising:

[0035] An acquisition module, used for acquiring an image to be processed;

[0036] A classification module, used for inputting the image to be processed into a first convolutional neural network model to obtain a category probability of the image to be processed obtained by the first convolutional neural network model;

[0037] Segmentation module for:

[0038] If the category probability indicates that the image to be processed is an abnormal slice, the image to be processed is input into an encoder of a second convolutional neural network model to obtain shallow features obtained by the encoder performing downsampling processing on the image to be processed;

[0039] Inputting the image to be processed and the shallow features into the decoder of the second convolutional neural network model to obtain deep features obtained by upsampling the image to be processed by the decoder, and fusing the shallow features and the deep features to obtain fused features;

[0040] The fused features are input into the output layer of the second convolutional neural network model to obtain the nodule region of the image to be processed obtained by the output layer.

[0041] Optionally, the segmentation module is specifically used to:

[0042] Inputting the image to be processed into an encoder of a second convolutional neural network model, wherein the encoder randomly groups input channels of the image to be processed to obtain a plurality of first groups;

[0043] The encoder randomly groups the input channels in each of the first groups to obtain a plurality of second groups corresponding to the first groups;

[0044] The encoder uses a residual network model to extract features of the image to be processed based on input channels in the second group, and uses the extracted features as features corresponding to the second group;

[0045] The encoder inputs the features corresponding to the first group and the second group into the attention model to obtain the features corresponding to the first group obtained by the attention model;

[0046] The connection layer of the encoder concatenates the features corresponding to each of the first groups to obtain shallow features of the image to be processed.

[0047] Optionally, the segmentation module is specifically used to:

[0048] The image to be processed and the shallow features are input into the decoder of the second convolutional neural network model, the decoder upsamples the image to be processed to obtain deep features, and uses dense skip connections to fuse the shallow features and the deep features to obtain fused features.

[0049] Optionally, the device further comprises:

[0050] The first training module is used to:

[0051] Acquire a first training image, where the first training image is annotated with a true category probability, where the true category probability indicates whether the first training image contains a nodule or does not contain a nodule;

[0052] Performing data enhancement, label smoothing, and random erasing on the first training image to obtain a first target training image;

[0053] Inputting the first target training image into a first convolutional neural network model to obtain a category prediction probability of the first target training image predicted by the first convolutional neural network model;

[0054] Determining whether a loss function value of the first convolutional neural network model converges, the loss function value of the first convolutional neural network model representing a difference between the category prediction probability and the true category probability;

[0055] If convergence occurs, the training ends;

[0056] If convergence has not occurred, the parameters of the first convolutional neural network model are updated.

[0057] Optionally, the device further comprises:

[0058] The second training module is used to:

[0059] Obtain a second training image with a true box position of a nodule marked;

[0060] Performing data enhancement on the second training image to obtain a second target training image;

[0061] Inputting the second target training image into a second convolutional neural network model to obtain a region of interest predicted by the first convolutional neural network model for the second target training image;

[0062] Determine whether a loss function value of the second convolutional neural network model converges, where the loss function value of the second convolutional neural network model represents a difference between the real frame position and the region of interest;

[0063] If convergence occurs, the training ends;

[0064] If not converged, update the parameters of the second convolutional neural network model.

[0065] Compared with the prior art, the beneficial effects of this application are:

[0066] In the present application, by acquiring an image to be processed, the image to be processed is input into a first convolutional neural network model to obtain the category probability of the image to be processed obtained by the first convolutional neural network model. If the category probability characterizes that the image to be processed is an abnormal slice, the image to be processed is input into a second convolutional neural network model to reduce the interference of non-abnormal slices on the second convolutional neural network model, so that the second convolutional neural network model processes the image to be processed, improves the segmentation accuracy of the nodule area, reduces the workload of the second convolutional neural network model, and improves the segmentation efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0068] Figure 1 It is a flowchart of a target detection method provided in Example 1 of the present application;

[0069] Figure 2 This is a schematic diagram of the structure of a 50-layer residual network provided by this application;

[0070] Figure 3 It is a schematic diagram of the segmentation effect of a second convolutional neural network model provided by the present application;

[0071] Figure 4 It is a flowchart of a target detection method provided in Example 2 of the present application;

[0072] Figure 5 It is a flowchart of a target detection method provided in Example 3 of the present application;

[0073] Figure 6 It is a structural schematic diagram of a target detection device provided in this application. DETAILED DESCRIPTION

[0074] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0075] In order to solve the above problems, the present application provides a target detection method, which is introduced below.

[0076] Reference Figure 1 , is a flow chart of a target detection method provided in Example 1 of the present application, such as Figure 1 As shown, the method may include but is not limited to the following steps:

[0077] Step S11, obtaining an image to be processed.

[0078] In this embodiment, the images to be processed may include but are not limited to: CT images.

[0079] Step S12: input the image to be processed into a first convolutional neural network model to obtain the category probability of the image to be processed obtained by the first convolutional neural network model.

[0080] This step may include but is not limited to:

[0081] S121, preprocessing the image to be processed to obtain a target image.

[0082] Preprocessing the image to be processed may include but is not limited to:

[0083] S1211. Use a morphological method and a threshold segmentation method to clean, denoise and enhance the contrast of the image to be processed to obtain a first image.

[0084] The contrast enhancement processing of the image to be processed can be understood as: performing HU value conversion and HU value standardization, and adjusting the window width and window position to enhance the contrast between the nodule and the surrounding tissues and organs.

[0085] S1212. Resample the first image to obtain a second image with isotropic resolution.

[0086] The second image with isotropic resolution can be understood as a second image with a consistent pixel spacing.

[0087] S1213: Convert the second image into a target image of a set size.

[0088] The set size may be, but is not limited to, 512 pixels × 512 pixels.

[0089] S122. Input the target image into the first convolutional neural network model to obtain the category probability of the target image obtained by the first convolutional neural network model.

[0090] In this embodiment, the first convolutional neural network model can be but is not limited to: Figure 2 The 50-layer residual network (ResidualNetwork, ResNet) model shown in Figure 1. The 50-layer residual network can be specifically divided into 5 layers of convolution, each layer of convolution includes several residual blocks, and the convolution block in each residual block is followed by a ReLU activation function to alleviate overfitting.

[0091] In this embodiment, the first convolutional neural network model can be trained in the following ways, including:

[0092] S123: Acquire a first training image, where the first training image is annotated with a true category probability, where the true category probability represents whether the first training image contains a nodule or does not contain a nodule.

[0093] S124. Perform data enhancement, label smoothing, and random erasing on the first training image to obtain a first target training image.

[0094] Data enhancement may include, but is not limited to: any one or more of horizontal and vertical flipping, translation, scaling, and Gaussian noise removal.

[0095] Label smoothing can be understood as weakening the image category of the first training image. Weakening the image category of the first training image can reduce overfitting and improve the generalization ability and learning speed of the first convolutional neural network model.

[0096] Random erasing can be understood as randomly deleting a region in the first training image. Randomly deleting a region in the first training image can make the images for each training different, increase the diversity of the training images, and suppress overfitting.

[0097] S125. Inputting the first target training image into a first convolutional neural network model to obtain a category prediction probability of the first target training image predicted by the first convolutional neural network model;

[0098] S126. Determine whether a loss function value of the first convolutional neural network model converges, wherein the loss function value of the first convolutional neural network model represents a difference between the category prediction probability and the true category probability.

[0099] In this embodiment, the loss function of the first convolutional neural network model may be, but is not limited to, a Dice loss function.

[0100] In this embodiment, a suitable hardware and software environment can be configured to perform multiple iterative training on the first convolutional neural network model. The number of training rounds can be, but is not limited to, set to 100 rounds, and AdamW is used as the optimizer of the first convolutional neural network model training process. The initial learning rate of the optimizer can be set to 1e-3, and cosine annealing is used as the learning rate adjustment strategy to dynamically adjust the learning rate to accelerate the convergence of the first convolutional neural network model. For example, the maximum value of the learning rate is set to 0.01, the minimum value is set to 0.002, and the decay period is 5. Then in the first round of training, the learning rate is 0.01, the second round learning rate is 0.008, then 0.006, 0.004, 0.002, and then the learning rate of the sixth round returns to 0.01, so as to adjust the period.

[0101] The first convolutional neural network model can be trained using round training and batch training. The batch training method depends on the selected hardware accelerator card. For example, each round of training uses all training images for training, while each batch of training uses only a small portion of training images for training, that is, a batch. For example, if there are 100 training images in total and the batch size is set to 10, then in this round of training, only 10 training images are used for training.

[0102] If converged, the training is terminated; if not converged, step S127 is executed.

[0103] S127. Update the parameters of the first convolutional neural network model.

[0104] Step S13: If the category probability characterizes that the image to be processed is an abnormal slice, the image to be processed is input into the encoder of the second convolutional neural network model to obtain shallow features obtained by the encoder by downsampling the image to be processed.

[0105] Corresponding to the implementation of steps S121-S122, this step may include:

[0106] If the category probability characterizes that the target image is an abnormal slice, the target image is input into the encoder of the second convolutional neural network model to obtain shallow features obtained by the encoder by downsampling the image to be processed.

[0107] Shallow features may include but are not limited to: location features.

[0108] In this embodiment, the second convolutional neural network model may be, but is not limited to, a UNet model.

[0109] In this embodiment, the second convolutional neural network model can be trained in the following way:

[0110] S131, obtaining a second training image with a real frame position marked with a nodule;

[0111] S132: Perform data enhancement on the second training image to obtain a second target training image.

[0112] Performing data augmentation on the second training image may include, but is not limited to: performing any one or more of flipping, rotating, normalizing, and cutmixing on the second training image. Cutmixing is a simple and effective data augmentation method, which involves cutting off a portion of an image and pasting the portion of the image into another image to increase the diversity of training data.

[0113] S133, inputting the second target training image into a second convolutional neural network model to obtain a region of interest predicted by the first convolutional neural network model for the second target training image;

[0114] S134, judging whether the loss function value of the second convolutional neural network model converges, where the loss function value of the second convolutional neural network model represents the difference between the real frame position and the region of interest;

[0115] If converged, the training is terminated; if not converged, step S135 is executed.

[0116] In this embodiment, a suitable hardware and software environment can be configured to perform multiple iterations of training on the second convolutional neural network model. The number of training rounds can be, but is not limited to, set to 200 rounds, AdaBelief is used as the optimizer of the neural network training process, the initial learning rate is set to 1e-4, cosine annealing is used as the learning rate adjustment strategy, the learning rate is dynamically adjusted to accelerate network convergence, and the training batch size is determined according to the selected hardware acceleration card, for example, on a Tesla P40gpu with 24G video memory, it can be set to 16.

[0117] After the training is completed, the second convolutional neural network model can be tested. Specifically, the foreground (i.e., nodule) and background (e.g., part not related to the nodule) can be set, and a threshold can be set. When the pixel value in the feature map output by the second convolutional neural network model is greater than the threshold, it means that the point is the foreground, and the pixel value is set to 1; when the pixel value is less than the threshold, it means that the point is the background, and the pixel value is set to 0, thereby obtaining the nodule area.

[0118] S135. Update the parameters of the second convolutional neural network model.

[0119] Step S14: input the image to be processed and the shallow features into the decoder of the second convolutional neural network model to obtain the deep features obtained by the decoder upsampling the image to be processed, and fuse the shallow features and the deep features to obtain fused features.

[0120] Deep features may include but are not limited to: semantic features.

[0121] The fusing process of the shallow features and the deep features may include but is not limited to: splicing the shallow features and the deep features.

[0122] Step S15: input the fusion feature into the output layer of the second convolutional neural network model to obtain the nodule area of ​​the image to be processed obtained by the output layer.

[0123] In this embodiment, a nodule region may be extracted from the image to be processed, and the longest diameter of the nodule region may be calculated from three dimensions: x, y, and z.

[0124] In this embodiment, by acquiring an image to be processed, the image to be processed is input into the first convolutional neural network model to obtain the category probability of the image to be processed obtained by the first convolutional neural network model. If the category probability indicates that the image to be processed is an abnormal slice, the image to be processed is input into the second convolutional neural network model to reduce the interference of non-abnormal slices on the second convolutional neural network model, so that the second convolutional neural network model processes the image to be processed, improves the segmentation accuracy of the nodule area, reduces the workload of the second convolutional neural network model, and improves the segmentation efficiency. For example, Figure 3 As shown, the difference between the nodule area obtained by the second convolutional neural network model and the real nodule area is small, which greatly improves the segmentation accuracy of the nodule area.

[0125] As another optional embodiment of the present application, refer to Figure 4 , is a flow chart of a target detection method provided in Example 2 of the present application. This embodiment is mainly a refinement of the target detection method described in Example 1 above, such as Figure 4 As shown, the method may include but is not limited to the following steps:

[0126] Step S21, obtaining an image to be processed.

[0127] Step S22: input the image to be processed into a first convolutional neural network model to obtain the category probability of the image to be processed obtained by the first convolutional neural network model.

[0128] The detailed process of steps S21-S22 can refer to the relevant introduction of steps S11-S12 in Example 1, which will not be repeated here.

[0129] Step S23: If the category probability characterizes that the image to be processed is an abnormal slice, the image to be processed is input into the encoder of the second convolutional neural network model, and the encoder randomly groups the input channels of the image to be processed to obtain multiple first groups.

[0130] Step S24: the encoder randomly groups the input channels in each of the first groups to obtain a plurality of second groups corresponding to the first groups.

[0131] Step S25: The encoder uses the residual network model to extract features of the image to be processed based on the input channels in the second group, and uses the extracted features as features corresponding to the second group.

[0132] Step S26: The encoder inputs the features corresponding to the first group and the second group into the attention model to obtain the features corresponding to the first group obtained by the attention model.

[0133] Step S27: The connection layer of the encoder concatenates the features corresponding to each of the first groups to obtain shallow features of the image to be processed.

[0134] Steps S23-S27 are a specific implementation of step S13 in Example 1.

[0135] Step S28: input the image to be processed and the shallow features into the decoder of the second convolutional neural network model, obtain the deep features obtained by the decoder by upsampling the image to be processed, and obtain the fused features by fusing the shallow features and the deep features.

[0136] Step S29: input the fusion feature into the output layer of the second convolutional neural network model to obtain the nodule area of ​​the image to be processed obtained by the output layer.

[0137] The detailed process of steps S28-S29 can refer to the relevant introduction of steps S14-S15 in Example 1, which will not be repeated here.

[0138] In this embodiment, the encoder uses the residual network model and the attention model to extract shallow features, which can further improve the accuracy of shallow feature extraction, and further improve the accuracy of nodule area segmentation.

[0139] As another optional embodiment of the present application, refer to Figure 5 , is a flow chart of a target detection method provided in Example 3 of the present application. This embodiment is mainly a refinement of the target detection method described in Example 2 above. Figure 5 As shown, the method may include but is not limited to the following steps:

[0140] Step S31, obtaining an image to be processed.

[0141] Step S32: input the image to be processed into a first convolutional neural network model to obtain the category probability of the image to be processed obtained by the first convolutional neural network model.

[0142] Step S33: If the category probability characterizes that the image to be processed is an abnormal slice, the image to be processed is input into the encoder of the second convolutional neural network model, and the encoder randomly groups the input channels of the image to be processed to obtain multiple first groups.

[0143] Step S34: the encoder randomly groups the input channels in each of the first groups to obtain a plurality of second groups corresponding to the first groups.

[0144] Step S35: The encoder uses the residual network model to extract features of the image to be processed based on the input channels in the second group, and uses the extracted features as features corresponding to the second group.

[0145] Step S36: The encoder inputs the features corresponding to the first group and the second group into the attention model to obtain the features corresponding to the first group obtained by the attention model.

[0146] Step S37: The connection layer of the encoder concatenates the features corresponding to each of the first groups to obtain shallow features of the image to be processed.

[0147] The detailed process of steps S31-S37 can refer to the relevant introduction of steps S21-S27 in Example 2, which will not be repeated here.

[0148] Step S38: input the image to be processed and the shallow features into the decoder of the second convolutional neural network model, the decoder upsamples the image to be processed to obtain deep features, and uses dense skip connections to fuse the shallow features and the deep features to obtain fused features.

[0149] Step S38 is a specific implementation of step S28 in Example 2.

[0150] Step S39: input the fusion feature into the output layer of the second convolutional neural network model to obtain the nodule area of ​​the image to be processed obtained by the output layer.

[0151] The detailed process of step S39 can refer to the relevant introduction of step S29 in Example 2, which will not be repeated here.

[0152] In this embodiment, the encoder uses the residual network model and the attention model to extract shallow features, which can further improve the accuracy of shallow feature extraction, and further improve the accuracy of nodule area segmentation.

[0153] In addition, by using dense skip connections to fuse the shallow features and the deep features, the hierarchical and phased fusion of features can be completed, ensuring that the features at different levels can be fully utilized, further improving the accuracy of segmentation. For example, the sensitivity of feature maps at different layers to different targets is different. For example, high-level feature maps have large receptive fields, but the edge information of objects is often easily lost during repeated downsampling. In this case, low-level features are needed to supplement them. After fusing and superimposing features at different levels, the features extracted from the original encoder can be fully utilized, so it is easier to achieve a good segmentation effect.

[0154] Next, the target detection device provided by the present application is introduced. The target detection device introduced below and the target detection method introduced above can be referenced to each other.

[0155] See also Figure 6 The target detection device includes: an acquisition module 100, a classification module 200 and a segmentation module 300.

[0156] An acquisition module 100 is used to acquire an image to be processed;

[0157] A classification module 200, used for inputting the image to be processed into a first convolutional neural network model to obtain a category probability of the image to be processed obtained by the first convolutional neural network model;

[0158] The segmentation module 300 is used for:

[0159] If the category probability indicates that the image to be processed is an abnormal slice, the image to be processed is input into an encoder of a second convolutional neural network model to obtain shallow features obtained by the encoder performing downsampling processing on the image to be processed;

[0160] Inputting the image to be processed and the shallow features into the decoder of the second convolutional neural network model to obtain deep features obtained by upsampling the image to be processed by the decoder, and fusing the shallow features and the deep features to obtain fused features;

[0161] The fused features are input into the output layer of the second convolutional neural network model to obtain the nodule region of the image to be processed obtained by the output layer.

[0162] The segmentation module 300 may be specifically used for:

[0163] Inputting the image to be processed into an encoder of a second convolutional neural network model, wherein the encoder randomly groups input channels of the image to be processed to obtain a plurality of first groups;

[0164] The encoder randomly groups the input channels in each of the first groups to obtain a plurality of second groups corresponding to the first groups;

[0165] The encoder uses a residual network model to extract features of the image to be processed based on input channels in the second group, and uses the extracted features as features corresponding to the second group;

[0166] The encoder inputs the features corresponding to the first group and the second group into the attention model to obtain the features corresponding to the first group obtained by the attention model;

[0167] The connection layer of the encoder concatenates the features corresponding to each of the first groups to obtain shallow features of the image to be processed.

[0168] The segmentation module 300 may be specifically used for:

[0169] The image to be processed and the shallow features are input into the decoder of the second convolutional neural network model, the decoder upsamples the image to be processed to obtain deep features, and uses dense skip connections to fuse the shallow features and the deep features to obtain fused features.

[0170] In this embodiment, the target detection device may further include:

[0171] The first training module is used to:

[0172] Acquire a first training image, where the first training image is annotated with a true category probability, where the true category probability indicates whether the first training image contains a nodule or does not contain a nodule;

[0173] Performing data enhancement, label smoothing, and random erasing on the first training image to obtain a first target training image;

[0174] Inputting the first target training image into a first convolutional neural network model to obtain a category prediction probability of the first target training image predicted by the first convolutional neural network model;

[0175] Determining whether a loss function value of the first convolutional neural network model converges, the loss function value of the first convolutional neural network model representing a difference between the category prediction probability and the true category probability;

[0176] If convergence occurs, the training ends;

[0177] If convergence has not occurred, the parameters of the first convolutional neural network model are updated.

[0178] In this embodiment, the target detection device may further include:

[0179] The second training module is used to:

[0180] Obtain a second training image with a true box position of a nodule marked;

[0181] Performing data enhancement on the second training image to obtain a second target training image;

[0182] Inputting the second target training image into a second convolutional neural network model to obtain a region of interest predicted by the first convolutional neural network model for the second target training image;

[0183] Determine whether a loss function value of the second convolutional neural network model converges, where the loss function value of the second convolutional neural network model represents a difference between the real frame position and the region of interest;

[0184] If convergence occurs, the training ends;

[0185] If not converged, update the parameters of the second convolutional neural network model.

[0186] It should be noted that each embodiment focuses on the differences from other embodiments, and the same or similar parts between the embodiments can be referred to each other. For the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0187] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.

[0188] For the convenience of description, the above device is described in various units according to their functions. Of course, when implementing the present application, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0189] It can be known from the description of the above implementation methods that those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present application can be essentially or partly contributed to the prior art in the form of a software product, which can be stored in a storage medium such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in the various embodiments of the present application or certain parts of the embodiments.

[0190] The above is a detailed introduction to a target detection method and device provided by the present application. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea; at the same time, for general technical personnel in this field, according to the idea of ​​the present application, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.

Claims

1. A target detection method, It is characterized in that include: Get the image to be processed; Inputting the image to be processed into a first convolutional neural network model to obtain a category probability of the image to be processed obtained by the first convolutional neural network model; If the category probability represents that the image to be processed is an abnormal slice, the image to be processed is input into the encoder of the second convolutional neural network model to obtain shallow features obtained by the encoder performing downsampling processing on the image to be processed, including: inputting the image to be processed into the encoder of the second convolutional neural network model, the encoder randomly groups the input channels of the image to be processed to obtain multiple first groups; the encoder randomly groups the input channels in each of the first groups to obtain multiple second groups corresponding to the first group; the encoder uses the residual network model to extract the features of the image to be processed based on the input channels in the second group, and uses the extracted features as the features corresponding to the second group; the encoder inputs the features corresponding to the first group and the second group into the attention model to obtain the features corresponding to the first group obtained by the attention model; the connection layer of the encoder splices the features corresponding to each of the first groups to obtain shallow features of the image to be processed; Inputting the image to be processed and the shallow features into the decoder of the second convolutional neural network model to obtain deep features obtained by upsampling the image to be processed by the decoder, and fusing the shallow features and the deep features to obtain fused features; The fused features are input into the output layer of the second convolutional neural network model to obtain the nodule region of the image to be processed obtained by the output layer.

2. The method according to claim 1, It is characterized in that The step of inputting the image to be processed and the shallow features into the decoder of the second convolutional neural network model to obtain the deep features obtained by upsampling the image to be processed by the decoder, and fusing the shallow features and the deep features to obtain the fused features includes: The image to be processed and the shallow features are input into the decoder of the second convolutional neural network model, the decoder upsamples the image to be processed to obtain deep features, and uses dense skip connections to fuse the shallow features and the deep features to obtain fused features.

3. The method according to claim 1, It is characterized in that The first convolutional neural network model is trained in the following manner, including: Acquire a first training image, where the first training image is annotated with a true category probability, where the true category probability indicates whether the first training image contains a nodule or does not contain a nodule; Performing data enhancement, label smoothing, and random erasing on the first training image to obtain a first target training image; Inputting the first target training image into a first convolutional neural network model to obtain a category prediction probability of the first target training image predicted by the first convolutional neural network model; Determining whether a loss function value of the first convolutional neural network model converges, the loss function value of the first convolutional neural network model representing a difference between the category prediction probability and the true category probability; If convergence occurs, the training ends; If convergence has not occurred, the parameters of the first convolutional neural network model are updated.

4. The method according to claim 1, It is characterized in that The second convolutional neural network model is trained in the following way: Obtain a second training image with a true box position of a nodule marked; Performing data enhancement on the second training image to obtain a second target training image; Inputting the second target training image into a second convolutional neural network model to obtain a region of interest predicted by the first convolutional neural network model for the second target training image; Determine whether a loss function value of the second convolutional neural network model converges, where the loss function value of the second convolutional neural network model represents a difference between the real frame position and the region of interest; If convergence occurs, the training ends; If not converged, update the parameters of the second convolutional neural network model.

5. A target detection device, It is characterized in that include: An acquisition module, used for acquiring an image to be processed; A classification module, used for inputting the image to be processed into a first convolutional neural network model to obtain a category probability of the image to be processed obtained by the first convolutional neural network model; Segmentation module for: If the category probability indicates that the image to be processed is an abnormal slice, the image to be processed is input into an encoder of a second convolutional neural network model to obtain shallow features obtained by the encoder performing downsampling processing on the image to be processed; Inputting the image to be processed and the shallow features into the decoder of the second convolutional neural network model to obtain deep features obtained by upsampling the image to be processed by the decoder, and fusing the shallow features and the deep features to obtain fused features; Inputting the fused features into the output layer of the second convolutional neural network model to obtain a nodule region of the image to be processed obtained by the output layer; The segmentation module is specifically used for: Inputting the image to be processed into an encoder of a second convolutional neural network model, wherein the encoder randomly groups input channels of the image to be processed to obtain a plurality of first groups; The encoder randomly groups the input channels in each of the first groups to obtain a plurality of second groups corresponding to the first groups; The encoder uses a residual network model to extract features of the image to be processed based on input channels in the second group, and uses the extracted features as features corresponding to the second group; The encoder inputs the features corresponding to the first group and the second group into the attention model to obtain the features corresponding to the first group obtained by the attention model; The connection layer of the encoder concatenates the features corresponding to each of the first groups to obtain shallow features of the image to be processed.

6. The device according to claim 5, It is characterized in that The segmentation module is specifically used for: The image to be processed and the shallow features are input into the decoder of the second convolutional neural network model, the decoder upsamples the image to be processed to obtain deep features, and uses dense skip connections to fuse the shallow features and the deep features to obtain fused features.

7. The device according to claim 5, It is characterized in that The device also includes: The first training module is used to: Acquire a first training image, where the first training image is annotated with a true category probability, where the true category probability indicates whether the first training image contains a nodule or does not contain a nodule; Performing data enhancement, label smoothing, and random erasing on the first training image to obtain a first target training image; Inputting the first target training image into a first convolutional neural network model to obtain a category prediction probability of the first target training image predicted by the first convolutional neural network model; Determining whether a loss function value of the first convolutional neural network model converges, the loss function value of the first convolutional neural network model representing a difference between the category prediction probability and the true category probability; If convergence occurs, the training ends; If convergence has not occurred, the parameters of the first convolutional neural network model are updated.

8. The device according to claim 5, It is characterized in that The device also includes: The second training module is used to: Obtain a second training image with a true box position of a nodule marked; Performing data enhancement on the second training image to obtain a second target training image; Inputting the second target training image into a second convolutional neural network model to obtain a region of interest predicted by the first convolutional neural network model for the second target training image; Determine whether a loss function value of the second convolutional neural network model converges, where the loss function value of the second convolutional neural network model represents a difference between the real frame position and the region of interest; If convergence occurs, the training ends; If not converged, update the parameters of the second convolutional neural network model.

Citation Information

Patent Citations

  • Target detection method

    CN112232232A

  • Multi-pneumonia CT classification method and device based on time sequence high-dimensional feature extraction

    CN113269230A