Image-assisted diagnosis method and device based on instance segmentation and storage medium

The instance segmentation method combining Mask R-CNN and YOLOv5 algorithms solves the problem of low efficiency and accuracy in target nodule identification in image-assisted diagnosis, and achieves efficient and accurate identification of target nodules in 3D CT images.

CN117437189BActive Publication Date: 2026-05-12JINAN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
JINAN UNIVERSITY
Filing Date
2023-10-17
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

In existing technologies, image-assisted diagnostic methods are inefficient and inaccurate in identifying target nodules. In particular, the computation of non-lung areas in 3D CT images is wasteful of computing power, and the anchor frame of the 2D slicing method cannot accurately mark the lesion area.

Method used

An instance-based segmentation method was adopted, which used the Mask R-CNN model to segment the target organ region and combined it with the YOLOv5 algorithm to detect target nodules in the two-dimensional scan image. The trained neural network was used to extract and merge lesion information to determine the diagnostic results of the three-dimensional scan image.

Benefits of technology

It improves the efficiency and accuracy of target nodule identification. By removing invalid regions, it enhances the accuracy and practicality of nodule identification, and achieves efficient and accurate identification of target nodules in 3D scan images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117437189B_ABST
    Figure CN117437189B_ABST
Patent Text Reader

Abstract

The application relates to an image auxiliary diagnosis method and device based on instance segmentation and a storage medium, the method comprising the following steps: acquiring a three-dimensional scanning graph corresponding to a chest cavity of a target object, performing slice processing on the three-dimensional scanning graph to obtain a plurality of two-dimensional slice graphs; in the plurality of two-dimensional slice graphs, a target organ region is extracted by using a trained instance segmentation model Mask R-CNN, and a two-dimensional scanning graph corresponding to each two-dimensional slice graph is obtained; target nodules are detected in the target organ region of the two-dimensional scanning graph by using a trained target recognition network, and lesion information is obtained, wherein the lesion information comprises target information of the target nodules; and target nodules are merged and classified based on the target information corresponding to each two-dimensional scanning graph, so as to determine a diagnosis result corresponding to the three-dimensional scanning graph. Through the application, the problem that the efficiency and accuracy of a method for image auxiliary diagnosis in the related art are low is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of digital image processing technology, and in particular to image-assisted diagnosis methods, devices and storage media based on instance segmentation. Background Technology

[0002] In the medical field, radiologists often interpret medical images, but examining three-dimensional organ voxels (e.g., lungs, liver) layer by layer from two-dimensional CT images is a challenging task. Because CT scans contain a wealth of information about nodules in the target organ, misdiagnosis and missed diagnoses are common, leading to false negatives (FN) or false positives (FP) where non-lesion findings are misinterpreted. This makes it difficult to interpret and identify cells from CT scans, significantly limiting the detection of nodules in target organs.

[0003] In existing technologies, artificial intelligence is used for impact-assisted diagnostic testing, providing new technical means to improve the diagnosis of various diseases. However, in existing technologies, when performing three-dimensional scanning of target organs and performing target detection on the scanned three-dimensional CT images, a large amount of computing power is wasted on calculating non-lung areas, resulting in low efficiency and accuracy in identifying target nodules. At the same time, existing technologies use slicing three-dimensional CT images into two-dimensional CT images and then using a set target algorithm to identify target nodules on the two-dimensional CT images. However, the anchor frames used in these methods cannot accurately mark the lesion areas, which also easily leads to low efficiency and accuracy in identifying target nodules.

[0004] Currently, no effective solution has been proposed to address the low efficiency and accuracy of image-assisted diagnostic methods in identifying target nodules. Summary of the Invention

[0005] This application provides an image-assisted diagnosis method, apparatus, and storage medium based on instance segmentation, to at least solve the problem of low efficiency and accuracy in identifying target nodules in related technologies for image-assisted diagnosis.

[0006] In a first aspect, embodiments of this application provide an image-assisted diagnosis method based on instance segmentation, comprising: acquiring a three-dimensional scan image corresponding to the thoracic cavity of a target object; slicing the three-dimensional scan image to obtain multiple two-dimensional slice images; segmenting and extracting target organ regions in the multiple two-dimensional slice images using a trained instance segmentation model Mask R-CNN to obtain a two-dimensional scan image corresponding to each two-dimensional slice image, wherein the two-dimensional scan image includes the target organ region; detecting target nodules in the target organ region of the two-dimensional scan image using a trained target recognition network to obtain lesion information, wherein the lesion information includes target information of the target nodules, the target recognition network being a neural network based on the YOLOv5 algorithm and trained according to a preset two-dimensional sample scan image dataset and the actual diagnostic lesion information corresponding to the two-dimensional sample scan images in the two-dimensional sample scan image dataset, wherein the two-dimensional sample scan images are obtained using the Mask R-CNN model. R-CNN is generated by extracting target organ regions and labeling them with preset bounding boxes in two-dimensional sample slice images. The two-dimensional sample slice images are generated by slicing three-dimensional sample scan images. The bounding boxes are used to identify the target nodules. Based on the target information corresponding to each two-dimensional scan image, the target nodules are merged and classified to determine the diagnostic results corresponding to the three-dimensional scan images. The diagnostic results include the number, size, and location of the target nodules in the three-dimensional scan images.

[0007] In a second aspect, embodiments of this application provide an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the instance segmentation-based image-assisted diagnosis method as described in the first aspect.

[0008] Thirdly, embodiments of this application provide a storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the image-assisted diagnosis method based on instance segmentation as described in the first aspect above.

[0009] Compared to related technologies, the image-assisted diagnosis method, device, and storage medium based on instance segmentation provided in this application obtains a three-dimensional scan image corresponding to the thoracic cavity of the target object, slices the three-dimensional scan image to obtain multiple two-dimensional slice images; in the multiple two-dimensional slice images, the target organ region is segmented and extracted using a trained instance segmentation model MaskR-CNN to obtain a two-dimensional scan image corresponding to each two-dimensional slice image, the two-dimensional scan image including the target organ region; using a trained target recognition network, target nodules are detected in the target organ region of the two-dimensional scan image to obtain lesion information, the lesion information including the target information of the target nodules, the target recognition network is a neural network based on the YOLOv5 algorithm and trained according to a preset two-dimensional sample scan image dataset and the actual diagnostic lesion information corresponding to the two-dimensional sample scan images in the two-dimensional sample scan image dataset, the two-dimensional sample scan images being obtained using the MaskR-CNN model. R-CNN is generated by extracting target organ regions and labeling pre-defined bounding boxes in two-dimensional sample slice images. The two-dimensional sample slice images are generated by slicing three-dimensional sample scan images, and the bounding boxes are used to identify the target nodules. Based on the target information corresponding to each two-dimensional scan image, the target nodules are merged and classified to determine the diagnostic results corresponding to the three-dimensional scan images. The diagnostic results include the number, size, and location of the target nodules in the three-dimensional scan images. This solves the problem of low efficiency and accuracy in identifying target nodules in related technologies for image-assisted diagnosis. By combining the target recognition network with Mask R-CNN for instance segmentation, invalid regions where target nodules cannot exist are removed, avoiding the detection of invalid regions and improving the efficiency and accuracy of target nodule identification. Furthermore, by merging and classifying the nodule identification results after target nodule identification, the target nodule information in the corresponding three-dimensional scan images is statistically analyzed, thereby improving the accuracy and practicality of nodule identification.

[0010] Details of one or more embodiments of this application are set forth in the following drawings and description to make other features, objects and advantages of this application more readily apparent. Attached Figure Description

[0011] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0012] Figure 1 This is a hardware structure block diagram of the terminal of the image-assisted diagnosis method based on instance segmentation according to an embodiment of this application;

[0013] Figure 2This is a flowchart of an image-assisted diagnosis method based on instance segmentation according to an embodiment of this application;

[0014] Figure 3 This is a flowchart of an image-assisted diagnostic method based on instance segmentation according to a preferred embodiment of this application;

[0015] Figure 4 This is a structural block diagram of an image-assisted diagnostic device based on instance segmentation according to an embodiment of this application. Implementation

[0016] To make the objectives, technical solutions, and advantages of this application clearer, the application is described and illustrated below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the application. All other embodiments obtained by those skilled in the art based on the embodiments provided in this application without inventive effort are within the scope of protection of this application. Furthermore, it is understood that although the efforts made in such a development process may be complex and lengthy, for those skilled in the art related to the content disclosed in this application, modifications to design, manufacturing, or production based on the technical content disclosed in this application are merely conventional technical means and should not be construed as insufficient disclosure of the content of this application.

[0017] In this application, the reference to "embodiment" means that a specific feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment that is mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described in this application may be combined with other embodiments without conflict.

[0018] Unless otherwise defined, the technical or scientific terms used in this application shall have the ordinary meaning understood by one of ordinary skill in the art to which this application pertains. The terms “a,” “an,” “an,” “the,” and similar words used in this application do not indicate quantity limitation and may indicate singular or plural. The terms “comprising,” “including,” “having,” and any variations thereof used in this application are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or device that includes a series of steps or modules (units) is not limited to the listed steps or units, but may also include steps or units not listed, or may include other steps or units inherent to these processes, methods, products, or devices. The terms “connected,” “linked,” “coupled,” and similar words used in this application are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. “Multiple” used in this application means two or more. “And / or” describes the relationship between related objects, indicating that three relationships may exist; for example, “A and / or B” can represent: A alone, A and B simultaneously, and B alone. The terms “first,” “second,” “third,” etc., used in this application are merely to distinguish similar objects and do not represent a specific ordering of the objects.

[0019] The method embodiments provided in this example can be executed on a terminal, computer, or similar computing device. Taking running on a terminal as an example, Figure 1 This is a hardware structure block diagram of the terminal for the image-assisted diagnosis method based on instance segmentation according to an embodiment of this application. Figure 1 As shown, a terminal may include one or more ( Figure 1 Only one is shown in the diagram. A processor 102 (which may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.) and a memory 104 for storing data are also shown. Optionally, the terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the terminal described above. For example, the terminal may also include components that are more... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.

[0020] The memory 104 can be used to store computer programs, such as application software programs and modules, like the computer program corresponding to the image-assisted diagnosis method based on instance segmentation in this embodiment of the invention. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, thereby implementing the above-described method. The memory 104 may include high-speed random access memory and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to the terminal via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0021] The transmission device 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the terminal's communication provider. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a Radio Frequency (RF) module used for wireless communication with the Internet.

[0022] This embodiment provides an image-assisted diagnosis method based on instance segmentation that runs on the aforementioned terminal. Figure 2 This is a flowchart of an image-assisted diagnostic method based on instance segmentation according to an embodiment of this application, such as... Figure 2 As shown, the process includes the following steps:

[0023] Step S201: Obtain a three-dimensional scan image of the chest cavity of the target object, and slice the three-dimensional scan image to obtain multiple two-dimensional slice images.

[0024] In this embodiment, the three-dimensional scan image is obtained by using a CT scanner (e.g., an X-ray computed tomography scanner). This involves using a CT scanner with precisely collimated X-ray beams, gamma rays, ultrasound, etc., along with a highly sensitive detector, to perform a series of cross-sectional scans around a specific part of the human body, thereby obtaining a three-dimensional scan image of the chest cavity of the target object (the examinee). In this embodiment, the specific method of acquiring the three-dimensional scan image is not limited. Before executing the method of this embodiment, the corresponding three-dimensional scan image has been acquired and prepared. In this embodiment, the three-dimensional scan image can be obtained directly from a terminal or server connected to the corresponding CT scanner.

[0025] In this embodiment, after obtaining the three-dimensional scan image, the three-dimensional scan image is sliced ​​along a set slicing direction (e.g., along the Z-axis) to obtain a two-dimensional image corresponding to the thoracic cavity, which is a two-dimensional slice image. It can be understood that, due to the choice of slicing direction, after slicing the three-dimensional scan image, the generated two-dimensional slice image will contain slices of the plane where the center of the target nodule is located and slices of the plane where the center of the nodule is located. In this embodiment, both the slices of the plane where the center of the nodule is located and the slices of the plane where the center of the nodule is located are used as input images to enrich the input data and ensure the integrity of the features contained in the three-dimensional scan image.

[0026] Step S202: In multiple two-dimensional slice images, the trained instance segmentation model Mask R-CNN is used to segment and extract the dirty region of the target organ to obtain a two-dimensional scan image corresponding to each two-dimensional slice image, wherein the two-dimensional scan image includes the dirty region of the target organ.

[0027] In this embodiment, multiple two-dimensional slice images are input into a trained Mask R-CNN to obtain a two-dimensional scan image (corresponding to a two-dimensional CT image) containing only the target organ region (e.g., the lung region).

[0028] In this embodiment, the trained instance segmentation model Mask R-CNN includes convolutional layers, pooling layers, deconvolutional layers, and a loss function. The convolutional layer is a neural network layer used to extract image features. It uses a convolutional kernel to slide across the input image and calculates the dot product between the kernel and the image to obtain the output feature map. Mask R-CNN uses ResNet as its backbone network, which effectively solves the problems of gradient vanishing and overfitting. The pooling layer is a neural network layer used to reduce the dimensionality of the feature map and enhance feature invariance. It uses a fixed-size window to slide across the input feature map and performs corresponding aggregation operations (e.g., maximum value, average value) on the values ​​within the window to obtain the output feature map. Mask R-CNN uses RoI Pooling, which converts region proposals of different sizes into fixed-size feature vectors, and bilinear interpolation to avoid quantization errors and improve the accuracy of mask prediction. Align employs two pooling operations; the deconvolution layer is a neural network layer used to upsample feature maps. It uses a convolutional kernel to slide on the input feature map and calculates the dot product between the kernel and the feature map to obtain the output feature map. Mask R-CNN uses a fully convolutional network (FCN) as the mask branch to achieve pixel-level semantic segmentation; the loss function is a function used to measure the difference between the model's prediction and the true label. Mask R-CNN uses a multi-task loss function, which includes three parts: RPN loss, classification loss, and mask loss. The RPN loss and classification loss are calculated using binary cross-entropy loss and smoothing L1 loss to optimize the category of region proposal, the category of target object, and bounding box regression. The mask loss is calculated using multi-class cross-entropy loss to optimize whether each pixel belongs to the target category.

[0029] It is understood that the Mask R-CNN of this application addresses the spatial misalignment problem caused by the RoI Pool layer by introducing the RoI Align layer. The RoI Align layer introduced by Mask R-CNN maintains spatial alignment by canceling quantization and using bilinear interpolation to calculate feature values ​​at arbitrary positions. In this embodiment, the Mask R-CNN model adds a branch for predicting the target mask to the traditional Faster R-CNN model using a fully convolutional network (FCN). This branch outputs a binary mask representing the target pixels and background pixels in each region of interest (RoI). This branch runs in parallel with the bounding box recognition branch and decouples mask prediction from class prediction. That is, it generates a class-independent mask for each RoI and then selects the final mask based on the class prediction result, thereby reducing the number of parameters, improving computational efficiency, avoiding competition between classes, and improving mask quality.

[0030] Step S203: Using the trained target recognition network, target nodules are detected in the target organ region of the two-dimensional scan image to obtain lesion information. The lesion information includes the target information of the target nodule. The target recognition network is a neural network based on the YOLOv5 algorithm and trained according to a preset two-dimensional sample scan image dataset and the actual lesion information corresponding to the two-dimensional sample scan images in the two-dimensional sample scan image dataset. The two-dimensional sample scan image is generated by extracting the target organ region and adding preset bounding boxes in the two-dimensional sample slice image using Mask R-CNN. The two-dimensional sample slice image is generated by slicing the three-dimensional sample scan image. The bounding boxes are used to determine the target nodule.

[0031] In this embodiment, a target recognition network based on a neural network trained using the YOLOv5 algorithm is employed to detect target nodules in the target organ region of a two-dimensional scan image. This allows for bounding boxes to be drawn in areas containing lesions, accurately marking the lesion regions, analyzing the target nodules, and returning inference results. In this embodiment, a Mask R-CNN layer is introduced before the conventional YOLOv5 algorithm to segment and extract the target organ region from the input two-dimensional slice image. Then, the target organ region is extracted from the segmented two-dimensional scan image containing the target organ region, and YOLOv5 target detection and recognition are performed to identify the lesion. This improves recognition efficiency and accuracy, avoids unnecessary computational effort in non-target organ regions, and enhances target recognition and diagnostic efficiency.

[0032] It should be noted that in the 3D sample scan image, the target nodule is recorded in the form of sphere center coordinates and radius R. When training the target recognition network, the 2D sample image used is also generated by Mask R-CNN to extract the target organ region from the 2D sample slice image. At the same time, when extracting the target organ region, the 2D sample slice image is labeled with a corresponding bounding box. That is, based on the height h of the 2D sample slice image from the target nodule's sphere center point in the corresponding slice image on the same plane, and the sphere radius R of the target nodule, the circular radius r of the target nodule in the 2D sample image is determined. Then, based on r and the X-axis and Y-axis coordinates of the sphere center coordinates, a circumscribed square is constructed, and this circumscribed square is used as the bounding box of the target nodule in the 2D sample image. In this way, the corresponding 2D sample image is labeled with the target organ region and the bounding box, and the data format of the 2D sample image matches the data format required for training the target recognition network, thereby solving the data mismatch problem that exists when directly training the target recognition network with 3D sample scan images.

[0033] Step S204: Based on the target information corresponding to each two-dimensional scan image, the target nodules are merged and classified to determine the diagnostic results corresponding to the three-dimensional scan image. The diagnostic results include the number, size, and location of the target nodules in the three-dimensional scan image.

[0034] In this embodiment, the 3D scan image exists in a 3D data format. After slicing, it is identified in a 2D form. That is, the target nodules outlined in multiple consecutive 2D data images are actually the same. In this embodiment, by classifying and merging along the set slicing direction, consecutive target nodules with the same characteristics can be identified as the same target nodule, and detailed lesion information (maximum radius, center position) can be recorded. When the target nodules are discontinuous, that is, when a certain 2D image has no target nodule, it means that the previous target nodule has been classified and the classification of the next target nodule can be carried out. Through the classification and merging in this embodiment, target nodules in different locations in 3D form can be obtained.

[0035] Through steps S201 to S204 above, a three-dimensional scan image corresponding to the thoracic cavity of the target object is obtained, and the three-dimensional scan image is sliced ​​to obtain multiple two-dimensional slice images. In these multiple two-dimensional slice images, the target organ region is segmented and extracted using a trained instance segmentation model, Mask R-CNN, to obtain a two-dimensional scan image corresponding to each two-dimensional slice image. The two-dimensional scan image includes the target organ region. Using a trained target recognition network, target nodules are detected in the target organ region of the two-dimensional scan image to obtain lesion information. The lesion information includes the target nodule information. The target recognition network is a neural network based on the YOLOv5 algorithm, trained according to a preset two-dimensional sample scan image dataset and the actual diagnostic lesion information corresponding to the two-dimensional sample scan images in the dataset. The two-dimensional sample scan images are obtained using Mask R-CNN. R-CNN is generated by extracting target organ regions and labeling pre-defined bounding boxes in two-dimensional sample slice images. Two-dimensional sample slice images are generated by slicing three-dimensional sample scan images, and bounding boxes are used to identify target nodules. Based on the target information corresponding to each two-dimensional scan image, target nodules are merged and categorized to determine the diagnostic results corresponding to the three-dimensional scan image. The diagnostic results include the number, size, and location of target nodules in the three-dimensional scan image. This addresses the low efficiency and accuracy issues of image-assisted diagnosis methods in related technologies. By combining the target recognition network with Mask R-CNN for instance segmentation, invalid regions where target nodules are impossible are removed, avoiding the detection of invalid regions and improving the efficiency and accuracy of target nodule identification. Furthermore, by merging and categorizing the nodule identification results after target nodule identification, the target nodule information in the corresponding three-dimensional scan image is statistically analyzed, improving the accuracy and practicality of nodule identification.

[0036] It should be noted that, in this embodiment, the target recognition network enhances the recognition of small objects based on the traditional deep learning model, making it better suited for target nodule recognition scenarios. In this embodiment, the target recognition network improves the attention mechanism and feature fusion technology by introducing a Global Information Attention (GIOU) mechanism to calculate the loss function between the predicted bounding box and the ground truth bounding box. GIOU can consider the overlap and inclusion relationships between bounding boxes, improving localization accuracy. In this embodiment, the target recognition network also improves the feature extraction network and feature fusion network by using techniques such as residual connections, depthwise separable convolutions, and attention modules, increasing the network depth and width, improving feature representation capabilities, and enhancing the feature representation capabilities and semantic information of small objects.

[0037] It should be further explained that, in this embodiment, TensorRT is also introduced to accelerate the inference speed of deep learning models in detecting target nodules. Specifically, Mask R-CNN and target recognition networks are trained on 3D scan atlases. The trained models are imported and internal networks are constructed. The models are optimized using methods such as layer fusion, kernel selection, and tensor allocation. Then, based on the optimized network and specified configuration parameters, a CUDA engine for inference is constructed and serialized and saved as a file for subsequent loading and execution. Finally, through the created CUDA engine, inference tasks can be executed on the GPU. Input and output data are passed through the CUDA buffer, improving the efficiency of model inference and reducing memory usage.

[0038] In some embodiments, target nodules are merged and categorized based on the target information corresponding to each two-dimensional scan image to determine the diagnostic result corresponding to the three-dimensional scan image, which is achieved through the following steps:

[0039] Step 21: In the target information corresponding to each two-dimensional scan image, obtain the first bounding box information corresponding to the target nodule. The first bounding box information includes the corresponding first nodule bounding box and the center point coordinate parameters of the first nodule bounding box.

[0040] In this embodiment, before obtaining the first bounding box information corresponding to the target nodule, it has been determined that the target nodule is present in the corresponding two-dimensional scan image, that is, the target nodule has been detected; at the same time, for two-dimensional scan images where the target nodule has not been detected, additional process processing is performed, such as ending the merging and classification process of the target nodule.

[0041] Step 22: Based on the center point coordinate parameters, classify the first nodule bounding boxes to obtain first bounding box groups, where each first bounding box group corresponds to a target nodule.

[0042] In this embodiment, the distance between the center points of the two first nodule bounding boxes is compared to whether it is less than or equal to a set distance threshold. If it is less than or equal to the set distance threshold, the two first nodule bounding boxes are classified as belonging to the same target nodule.

[0043] Step 23: Determine the number of target nodules based on the number of groups in the first bounding box group.

[0044] In this embodiment, after traversing all the first nodule bounding boxes and completing the grouping of the bounding boxes, it means that the same target nodule appearing in the 3D scan image has been merged, and the number of target nodules determined by merging corresponds to the number of groups of the first bounding boxes.

[0045] Step 24: Based on the target node bounding box selected from all the first node bounding boxes in the first bounding box group, determine the size of the target node corresponding to the first bounding box group, and based on the center point coordinate parameters corresponding to the target node bounding box, determine the position of the target node corresponding to the first bounding box group.

[0046] In this embodiment, after merging each first bounding box group, the first bounding box group represents the corresponding target nodule. To determine the size and position of the target nodule, a target nodule bounding box is selected from each first bounding box group, and the size of the target nodule is represented by the anchor frame data of the target nodule bounding box, and the position of the target nodule is represented by the center point coordinate parameters of the target nodule bounding box.

[0047] In some optional implementations, the first aiming frame information also includes the first anchor frame size corresponding to the first nodule bounding box. The target nodule bounding box selected from all the first nodule bounding boxes in the first bounding box group is achieved by the following steps: selecting the first nodule bounding box with the largest first anchor frame size from all the first nodule bounding boxes in the first bounding box group to obtain the target nodule bounding box.

[0048] By taking the steps described above, the first bounding box information corresponding to the target nodule is obtained from the target information corresponding to each two-dimensional scan image; the first bounding box is classified according to the center point coordinate parameters to obtain the first bounding box group; the number of target nodules is determined based on the number of the first bounding box group; the size of a target nodule corresponding to the first bounding box group is determined according to the target nodule bounding box selected from all the first bounding box bounding boxes of the first bounding box group; and the position of a target nodule corresponding to the first bounding box group is determined according to the center point coordinate parameters of the target nodule bounding box. This process realizes the classification and merging of target nodules outlined in two-dimensional form, thereby obtaining target nodules at different positions in three-dimensional form.

[0049] In some embodiments, the first nodule bounding boxes are classified according to the center point coordinate parameters to obtain the first bounding box grouping, including the following steps:

[0050] Step 31: Iterate through the center point coordinate parameters of the bounding boxes of the first nodules corresponding to all two-dimensional scan images in sequence, and calculate the distance between the center points of two adjacent first nodule bounding boxes according to the center point coordinate parameters.

[0051] Step 32: Determine whether the distance is not greater than a preset distance threshold. If the distance is determined to be not greater than the preset distance threshold, classify the corresponding two first nodal bounding boxes into the corresponding first bounding box group to obtain multiple first bounding box groups.

[0052] By sequentially traversing the center point coordinates of the bounding boxes of the first nodules corresponding to all the two-dimensional scan images in the above steps, and calculating the distance between the center points of two adjacent bounding boxes of the first nodules according to the center point coordinates, it is determined whether the distance is not greater than a preset distance threshold. If the distance is not greater than the preset distance threshold, the two corresponding bounding boxes of the first nodules are classified into the corresponding bounding box group to obtain multiple bounding box groups. This realizes the merging of bounding boxes of the first nodules that belong to the same target nodule, thereby realizing the classification and merging of target nodules outlined in two-dimensional form and obtaining target nodules at different positions in three-dimensional form.

[0053] In some embodiments, the first nodule bounding boxes are classified according to the center point coordinate parameters to obtain the first bounding box grouping, which is also achieved through the following steps:

[0054] Step 41: Iterate through the center point coordinate parameters of the bounding boxes of the first nodules corresponding to all two-dimensional scan images in sequence, and calculate the distance between the center points of two adjacent first nodal bounding boxes according to the center point coordinate parameters.

[0055] Step 42: Determine whether the distance is not greater than the preset distance threshold. If the distance is greater than the threshold, classify the first bounding box of the first node in the two first bounding boxes currently involved in the calculation into the corresponding first bounding box group.

[0056] Step 43: In the remaining first nodal bounding boxes to be classified, starting from the second first nodal bounding box among the two currently participating first nodal bounding boxes, calculate the distance between the center points of two adjacent first nodal bounding boxes in turn, and select the first nodal bounding box whose distance is not greater than a preset distance threshold, until the distance is greater than the preset distance threshold or the remaining first nodal bounding boxes to be classified are traversed, to obtain another group of first bounding boxes. The second first nodal bounding box among the two currently participating first nodal bounding boxes is the first nodal bounding box of the other group of first bounding boxes.

[0057] In this embodiment, the target nodule identification results of each two-dimensional scan image are traversed along the z-axis. Classification begins with the two-dimensional scan image containing the target nodule, and it is determined whether the next two-dimensional scan image contains the target nodule. If the determination result contains the target nodule, the coordinate parameters of the center point of the first nodule bounding box of each target nodule are compared with the coordinate parameters of the center point of the first nodule bounding box of the target nodule in the previous two-dimensional scan image. If the distance between the two center points is less than or equal to a set threshold, the target nodules identified in the two two-dimensional scan images are classified as the same target nodule in the three-dimensional scan image, and classification continues to the next two-dimensional scan image. If the distance between the two center points is greater than the set threshold, the target nodule identified in the two-dimensional scan image and the lung target nodule identified in the previous two-dimensional scan image are defined as different target nodules in the three-dimensional scan image, and the classification of the target nodule in the previous two-dimensional scan image is stopped. A new classification begins from the target nodule in this two-dimensional scan image. If, before determining the distance, it is determined that the two-dimensional scan image does not contain the target nodule, the classification of the target nodule in the previous two-dimensional scan image is stopped.

[0058] Through steps 41 to 43 above, the merging of first nodule bounding boxes belonging to the same target nodule and the grouping of first nodule bounding boxes belonging to different groups of different target nodules are further realized, thereby realizing the classification and merging of target nodules framed in two-dimensional form and obtaining target nodules at different positions in three-dimensional form.

[0059] In some embodiments, a trained target recognition network is used to detect target nodules in the target organ region of a two-dimensional scan image to obtain lesion information, which is achieved through the following steps:

[0060] Step 51: Using a target recognition network, nodule target detection is performed on the two-dimensional scan image to obtain the label information and second bounding box information corresponding to the candidate target. The label information includes the nodule category corresponding to the candidate target and the target confidence level corresponding to the nodule category. The second bounding box information includes the second bounding box corresponding to the candidate target, the center point coordinate parameters of the second bounding box, and the second bounding box size.

[0061] In this embodiment, when using the target recognition network to perform target recognition, it outputs the corresponding label information and bounding box information of the corresponding two-dimensional scan image. The label information is used to indicate the nodule category of the identified target and the confidence level of the corresponding nodule category. By using the corresponding nodule category and confidence level, it is determined whether there is a target nodule in the corresponding two-dimensional scan image.

[0062] Step 52: Based on the nodule category, select candidate nodules from multiple candidate targets, and determine whether the target confidence corresponding to the candidate nodule is greater than the preset confidence threshold.

[0063] In this embodiment, the identification of a target is first determined based on the nodule category to determine whether it is a set detection target. For example, it is determined whether the target is a lung nodule. If the nodule category is lung nodule, it means that the candidate target can be used as a candidate nodule. Then, through further confidence judgment, it is finally determined whether it is the target nodule.

[0064] Step 53: If the target confidence level corresponding to the candidate nodule is greater than the preset confidence threshold, the candidate nodule is determined as the target nodule, and the second anchor frame information corresponding to the candidate nodule is used as the target information corresponding to the target nodule to obtain the lesion information.

[0065] Step 54: If it is determined that the target confidence scores corresponding to the candidate nodules are all not greater than the preset confidence threshold, then it is determined that there are no target nodules in the two-dimensional scan image.

[0066] Step 55: If it is determined that the target confidence scores corresponding to all candidate nodules are not greater than the preset confidence threshold, the diagnostic result is determined to include the absence of target nodules in the 3D scan image.

[0067] In this embodiment, the expected target nodules are selected by judging the target confidence of the selected candidate nodules.

[0068] The above steps utilize a target recognition network to detect nodules in a 2D scan image, obtaining label information and second bounding box information corresponding to candidate targets. The label information includes the nodule category and the target confidence level corresponding to the nodule category. The second bounding box information includes the second bounding box corresponding to the candidate target, the center point coordinates of the second bounding box, and the second bounding box size. Based on the nodule category, candidate nodules are selected from multiple candidate targets, and it is determined whether the target confidence level corresponding to the candidate nodule is greater than a preset confidence threshold. If the target confidence level of a candidate nodule is greater than the preset confidence threshold, then... Under the threshold condition, candidate nodules are identified as target nodules, and the second anchor box information corresponding to the candidate nodules is used as the target information corresponding to the target nodules to obtain lesion information. If it is determined that the target confidence scores corresponding to the candidate nodules are not greater than the preset confidence threshold, it is determined that there are no target nodules in the two-dimensional scan image. If it is determined that the target confidence scores corresponding to all candidate nodules are not greater than the preset confidence threshold, it is determined that the diagnostic result includes the absence of target nodules in the three-dimensional scan image. This realizes the verification of whether the target identified by the target recognition network is a target nodule, that is, the verification of lesion information.

[0069] In some embodiments, the training steps of the target recognition network include:

[0070] Step 61: Slice the three-dimensional sample scan image according to the preset slicing direction to generate a first sample slice image set corresponding to the two-dimensional sample slice image. The three-dimensional sample scan image includes the sphere center coordinate parameters and sphere size parameters of the target nodule. The first sample slice image set includes a first sphere center sample slice image and a first non-sphere center plane slice image corresponding to the plane where the sphere center of the target nodule is located.

[0071] Step 62: Use Mask R-CNN to segment and extract the target organ region from the first sphere-centered sample slice image and the first non-sphere-centered plane slice image, and determine the corresponding target organ region in the first sphere-centered sample slice image and the first non-sphere-centered plane slice image.

[0072] Step 63: Based on the height of each first non-sphere-centered planar slice image relative to the first sphere-centered sample slice image, the sphere-centered coordinate parameters of the target nodule, and the sphere-sized parameters of the nodule, determine the sphere radius of the nodules present in the first sphere-centered sample slice image and the first non-sphere-centered planar slice image. Based on the determined sphere radius and the sphere-centered coordinate parameters of the target nodule, mark the first sphere-centered sample slice image and the first non-sphere-centered planar slice image of the identified target organ region, and generate the second sphere-centered sample slice image and the second non-sphere-centered planar slice image.

[0073] In this embodiment, the height h of the first non-centered planar slice image relative to the first centered sample slice image is recorded. The spherical radius R of the nodule is obtained from the three-dimensional sample scan image. The spherical radius r of the nodules present in the first centered sample slice image and the first non-centered planar slice image is determined by: r² = R² - h². Then, based on r and the x-axis and y-axis coordinates in the spherical center coordinate parameters, the corresponding circumscribed square of the circle is constructed, and this circumscribed square is used as the corresponding aiming frame in the first centered sample slice image and the first non-centered planar slice image. It can be understood that h is 0 for the first centered sample slice image, therefore the spherical radius r of the nodules present in the first centered sample slice image is equal to R.

[0074] Step 64: Use the actual lesion information corresponding to the three-dimensional sample scan image, the second spherical center sample slice image, and the second non-spherical center plane slice image to train the YOLOv5 algorithm until convergence, and obtain the target recognition network.

[0075] Figure 3 This is a flowchart of an image-assisted diagnostic method based on instance segmentation according to a preferred embodiment of this application, with reference to... Figure 3 The embodiments of this application are further described below:

[0076] The preferred embodiment of the instance segmentation-based image-assisted diagnosis method of this application includes the following steps:

[0077] Step S301: Image data input and preprocessing.

[0078] In this embodiment, after obtaining the three-dimensional scan image of the target thoracic cavity, the three-dimensional scan image is sliced ​​along the z-axis to obtain a two-dimensional slice image of the thoracic cavity. Thus, the three-dimensional scan image of the target thoracic cavity is converted into several two-dimensional slice images, which are used as the processing images for the next step.

[0079] Step S302: Mask R-CNN algorithm segmentation.

[0080] In this embodiment, the trained Mask R-CNN is used to segment each two-dimensional slice image, extract the target organ region of each two-dimensional slice image, and form a two-dimensional scan image containing only the target organ region.

[0081] In some alternative implementations, training the Mask R-CNN and identifying target dirty regions from two-dimensional slice images of the slices include the following steps:

[0082] Step 1: Train Mask-RCNN using existing medical datasets containing pixel-level coordinates of the target organ region (the Mask image size is the same as the 3D chest CT image size, and the Mask image has only two numbers, 0 and 1, which correspond to the pixels in the 3D chest image that are not in the target organ region and the pixels that are in the target organ region, respectively).

[0083] Step 2: Divide the 3D scan image of the target object's chest cavity into multiple 2D slice images in z-axis order.

[0084] Step 3: Input the two-dimensional slice image into the trained Mask-RCNN to obtain a two-dimensional scan image containing only the target organ's dirty region.

[0085] Step S303, YOLOv5 algorithm detection.

[0086] In this embodiment, a trained YOLOv5 target recognition network is used to detect targets in two-dimensional scan images containing only target organ regions and determine lesion information to obtain the confidence level of target nodules in the two-dimensional scan images. The confidence level is then determined to be greater than a preset confidence threshold to obtain a judgment result. If the confidence level is greater than the preset confidence threshold, the recognition result is determined to include the detection of the corresponding target nodule. If the confidence level is not greater than the preset confidence threshold, the recognition result is determined to include the absence of detected target nodules. Then, based on the recognition results of the detected target nodules, the number, size, and location of target nodules in each two-dimensional scan image are recorded.

[0087] Step S304: Merge and classify 2D CT images.

[0088] In this embodiment, the target nodule identification results of each two-dimensional scan image are traversed along the z-axis. Classification begins with the two-dimensional scan image containing the target nodule, and it is determined whether the next two-dimensional scan image contains the target nodule. If the determination result contains the target nodule, the coordinate parameters of the center point of the first nodule bounding box of each target nodule are compared with the coordinate parameters of the center point of the first nodule bounding box of the target nodule in the previous two-dimensional scan image. If the distance between the two center points is less than or equal to a set threshold, the target nodules identified in the two two-dimensional scan images are classified as the same target nodule in the three-dimensional scan image, and classification continues to the next two-dimensional scan image. If the distance between the two center points is greater than the set threshold, the target nodule identified in the two-dimensional scan image and the lung target nodule identified in the previous two-dimensional scan image are defined as different target nodules in the three-dimensional scan image, and the classification of the target nodule in the previous two-dimensional scan image is stopped. A new classification begins from the target nodule in this two-dimensional scan image. If, before determining the distance, it is determined that the two-dimensional scan image does not contain the target nodule, the classification of the target nodule in the previous two-dimensional scan image is stopped.

[0089] Step S305: Analyze the test results.

[0090] In this embodiment, based on the target nodule information in the 3D scan image, corresponding information is returned. Specifically, it is determined whether the number of target nodules in the 3D scan image is greater than 0, and a judgment result is obtained. When the judgment result is that the number of target nodules is greater than 0, the 3D scan image is diagnosed, showing that there are target nodules and displaying specific information about the lesions, such as number, size, and location. Then, diagnostic text is generated according to a pre-defined template. When the judgment result is that the number of target nodules is equal to 0, the 3D scan image is diagnosed, showing that there are no target nodules.

[0091] This embodiment also provides an interactive image restoration device based on deep learning, which is used to implement the above embodiments and preferred embodiments, and will not be repeated as already described. As used below, the terms "module," "unit," "subunit," etc., can refer to a combination of software and / or hardware that performs a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0092] Figure 4 This is a structural block diagram of an image-assisted diagnostic device based on instance segmentation according to an embodiment of this application, as shown below. Figure 4 As shown, the device includes an acquisition module 41, a segmentation module 42, an identification module 43, and a processing module 44, wherein:

[0093] The acquisition module 41 is used to acquire a three-dimensional scan image of the chest cavity of the target object, and to slice the three-dimensional scan image to obtain multiple two-dimensional slice images.

[0094] The segmentation module 42, coupled to the acquisition module 41, is used to segment and extract the dirty region of the target organ from multiple two-dimensional slice images using the trained instance segmentation model Mask R-CNN, to obtain a two-dimensional scan image corresponding to each two-dimensional slice image, wherein the two-dimensional scan image includes the dirty region of the target organ.

[0095] The recognition module 43, coupled to the segmentation module 42, is used to detect target nodules in the target organ region of a two-dimensional scan image using a trained target recognition network to obtain lesion information. The lesion information includes target information of the target nodule. The target recognition network is a neural network based on the YOLOv5 algorithm and trained according to a preset two-dimensional sample scan image dataset and the actual lesion information corresponding to the two-dimensional sample scan images in the dataset. The two-dimensional sample scan image is generated by extracting the target organ region and adding preset bounding boxes in the two-dimensional sample slice image using Mask R-CNN. The two-dimensional sample slice image is generated by slicing the three-dimensional sample scan image. The bounding boxes are used to determine the target nodule.

[0096] The processing module 44, coupled to the identification module 43, is used to merge and classify target nodules based on the target information corresponding to each two-dimensional scan image, so as to determine the diagnostic results corresponding to the three-dimensional scan image. The diagnostic results include the number, size and location of target nodules in the three-dimensional scan image.

[0097] The image-assisted diagnostic device based on instance segmentation according to the embodiments of this application acquires a three-dimensional scan image corresponding to the thoracic cavity of the target object, slices the three-dimensional scan image to obtain multiple two-dimensional slice images; in the multiple two-dimensional slice images, the target organ region is segmented and extracted using a trained instance segmentation model Mask R-CNN to obtain a two-dimensional scan image corresponding to each two-dimensional slice image, and the two-dimensional scan image includes the target organ region; using a trained target recognition network, target nodules are detected in the target organ region of the two-dimensional scan image to obtain lesion information, including target nodule information. The target recognition network is a neural network based on the YOLOv5 algorithm, trained according to a preset two-dimensional sample scan image dataset and the actual diagnostic lesion information corresponding to the two-dimensional sample scan images in the two-dimensional sample scan image dataset. The two-dimensional sample scan images are obtained using Mask R-CNN. R-CNN is generated by extracting target organ regions and labeling pre-defined bounding boxes in two-dimensional sample slice images. Two-dimensional sample slice images are generated by slicing three-dimensional sample scan images, and bounding boxes are used to identify target nodules. Based on the target information corresponding to each two-dimensional scan image, target nodules are merged and categorized to determine the diagnostic results corresponding to the three-dimensional scan image. The diagnostic results include the number, size, and location of target nodules in the three-dimensional scan image. This addresses the low efficiency and accuracy of image-assisted diagnosis methods in related technologies. By combining the target recognition network with Mask R-CNN for instance segmentation, invalid regions where target nodules are impossible are removed, avoiding the detection of invalid regions and improving the efficiency and accuracy of target nodule identification. Furthermore, by merging and categorizing the nodule identification results after target nodule identification, the target nodule information in the corresponding three-dimensional scan image is statistically analyzed, improving the accuracy and practicality of nodule identification.

[0098] In some embodiments, the processing module 44 further includes:

[0099] The first acquisition unit is used to acquire the first bounding box information corresponding to the target nodule from the target information corresponding to each two-dimensional scan image. The first bounding box information includes the corresponding first nodule bounding box and the center point coordinate parameters of the first nodule bounding box.

[0100] The first calculation unit, coupled to the first acquisition unit, is used to classify the bounding boxes of the first nodule according to the center point coordinate parameters to obtain the first bounding box groups, wherein each first bounding box group corresponds to a target nodule.

[0101] The first determining unit, coupled to the first computing unit, is used to determine the number of target nodules based on the number of groups of the first bounding box.

[0102] The second determining unit is coupled to the first calculation unit and the first determining unit respectively. It is used to determine the size of a target nodule corresponding to the first bounding box group based on the target nodule bounding box selected from all the first nodule bounding boxes of the first bounding box group, and to determine the position of a target nodule corresponding to the first bounding box group based on the center point coordinate parameters of the target nodule bounding box.

[0103] In some embodiments, the first calculation unit is further configured to sequentially traverse the center point coordinate parameters of the first nodule bounding boxes corresponding to all two-dimensional scan images, and calculate the distance between the center points of two adjacent first nodule bounding boxes according to the center point coordinate parameters; determine whether the distance is not greater than a preset distance threshold, and if the distance is determined to be not greater than the preset distance threshold, classify the corresponding two first nodule bounding boxes into the corresponding first bounding box group to obtain multiple first bounding box groups.

[0104] In some embodiments, the first calculation unit is further configured to, when it is determined that the distance is greater than a threshold distance, classify the first nodal bounding box that is ranked earlier among the two first nodal bounding boxes currently participating in the calculation into the corresponding first bounding box group; among the remaining first nodal bounding boxes to be classified, starting from the second first nodal bounding box that is ranked later among the two first nodal bounding boxes currently participating in the calculation, calculate the distance between the center points of two adjacent first nodal bounding boxes in turn, and select the first nodal bounding box whose distance is not greater than a preset distance threshold, until the distance is greater than the preset distance threshold or the remaining first nodal bounding boxes to be classified are traversed, to obtain another first bounding box group, wherein the second first nodal bounding box that is ranked later among the two first nodal bounding boxes currently participating in the calculation is the first nodal bounding box of the other first bounding box group.

[0105] In some embodiments, the first aiming frame information also includes the first anchor frame size corresponding to the first nodule bounding box. The second determining unit is further configured to select the first nodule bounding box with the largest first anchor frame size from all the first nodule bounding boxes in the first bounding box group to obtain the target nodule bounding box.

[0106] In some embodiments, the identification module 43 further includes:

[0107] The first detection unit is used to perform nodule target detection on the two-dimensional scan image using a target recognition network, and obtain the label information and second bounding box information corresponding to the candidate target. The label information includes the nodule category corresponding to the candidate target and the target confidence level corresponding to the nodule category. The second bounding box information includes the second bounding box corresponding to the candidate target, the center point coordinate parameters of the second bounding box, and the second bounding box size.

[0108] The first selection unit, coupled to the first detection unit, is used to select candidate nodules from multiple candidate targets according to the nodule category, and to determine whether the target confidence corresponding to the candidate nodule is greater than a preset confidence threshold.

[0109] The first judgment unit, coupled to the first selection unit, is used to determine the candidate nodule as the target nodule when the target confidence corresponding to the candidate nodule is greater than a preset confidence threshold, and to use the second anchor frame information corresponding to the candidate nodule as the target information corresponding to the target nodule to obtain lesion information.

[0110] The second judgment unit, coupled to the first selection unit, is used to determine that there are no target nodules in the two-dimensional scan image when the target confidence scores corresponding to the candidate nodules are all not greater than a preset confidence threshold.

[0111] The third judgment unit, coupled to the first selection unit, is used to determine the diagnostic result, including the absence of target nodules in the three-dimensional scan image, when it is determined that the target confidence scores corresponding to all candidate nodules are not greater than a preset confidence threshold.

[0112] In some embodiments, the image-assisted diagnostic device based on instance segmentation is further used to slice the three-dimensional sample scan image according to a preset slicing direction to generate a first sample slice image set corresponding to the two-dimensional sample slice image. The three-dimensional sample scan image includes the sphere center coordinate parameters and the sphere size parameters of the target nodule. The first sample slice image set includes a first sphere-centered sample slice image and a first non-sphere-centered plane slice image corresponding to the plane containing the sphere center of the target nodule. The first sphere-centered sample slice image and the first non-sphere-centered plane slice image are processed using a Mask... R-CNN is used to segment and extract the target organ region, determining the corresponding target organ region in the first centered sample slice image and the first non-centered planar slice image. Based on the height of each first non-centered planar slice image relative to the first centered sample slice image, the spherical coordinate parameters of the target nodule, and the spherical size parameters of the nodule, the spherical radius of the nodules in the first centered sample slice image and the first non-centered planar slice image is determined. Based on the determined spherical radius and the spherical coordinate parameters of the target nodule, bounding boxes are marked on the first centered sample slice image and the first non-centered planar slice image of the determined target organ region, generating the second centered sample slice image and the second non-centered planar slice image. The YOLOv5 algorithm is trained using the actual lesion information corresponding to the 3D sample scan image, the second centered sample slice image, and the second non-centered planar slice image until convergence, resulting in the target recognition network.

[0113] It should be noted that the above modules can be functional modules or program modules, and can be implemented through software or hardware. For modules implemented through hardware, the above modules can reside in the same processor; or the above modules can be located in different processors in any combination.

[0114] This embodiment also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the steps in any of the above method embodiments.

[0115] Optionally, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor and the input / output device is connected to the processor.

[0116] Optionally, in this embodiment, the processor can be configured to perform the following steps via a computer program:

[0117] S1. Obtain the three-dimensional scan image corresponding to the chest cavity of the target object, and perform slicing processing on the three-dimensional scan image to obtain multiple two-dimensional slice images.

[0118] S2, In multiple two-dimensional slice images, the trained instance segmentation model Mask R-CNN is used to segment and extract the dirty region of the target organ, resulting in a two-dimensional scan image corresponding to each two-dimensional slice image, which includes the dirty region of the target organ.

[0119] S3. Using a trained target recognition network, target nodules are detected in the target organ region of the two-dimensional scan image to obtain lesion information. The lesion information includes the target information of the target nodule. The target recognition network is a neural network based on the YOLOv5 algorithm and trained according to a preset two-dimensional sample scan image dataset and the actual lesion information corresponding to the two-dimensional sample scan images in the dataset. The two-dimensional sample scan image is generated by extracting the target organ region and adding preset bounding boxes in the two-dimensional sample slice image using Mask R-CNN. The two-dimensional sample slice image is generated by slicing the three-dimensional sample scan image. The bounding boxes are used to determine the target nodule.

[0120] S4. Based on the target information corresponding to each two-dimensional scan image, the target nodules are merged and classified to determine the diagnostic results corresponding to the three-dimensional scan image. The diagnostic results include the number, size, and location of the target nodules in the three-dimensional scan image.

[0121] It should be noted that the specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementations, and will not be repeated here.

[0122] Furthermore, in conjunction with the instance-segmentation-based image-assisted diagnostic method in the above embodiments, this application embodiment can provide a storage medium for implementation. This storage medium stores a computer program; when executed by a processor, the computer program implements any of the instance-segmentation-based image-assisted diagnostic methods in the above embodiments.

[0123] Those skilled in the art should understand that the technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments have been described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0124] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.

Claims

1. An image-assisted diagnostic method based on instance segmentation, characterized in that, include: A three-dimensional scan image of the chest cavity of the target object is obtained, and the three-dimensional scan image is sliced ​​to obtain multiple two-dimensional slice images. In multiple two-dimensional slice images, the target organ dirty region is segmented and extracted using the trained instance segmentation model Mask R-CNN to obtain a two-dimensional scan image corresponding to each two-dimensional slice image, wherein the two-dimensional scan image includes the target organ dirty region. Using a trained target recognition network, target nodules are detected in the target organ region of the two-dimensional scan image to obtain lesion information. The lesion information includes target information of the target nodules. The target recognition network is a neural network based on the YOLOv5 algorithm, trained according to a preset two-dimensional sample scan image dataset and the actual lesion information corresponding to the two-dimensional sample scan images in the dataset. The two-dimensional sample scan image is generated by extracting the target organ region and annotating preset anchor boxes in the two-dimensional sample slice image using the Mask R-CNN. The two-dimensional sample slice image is generated by slicing the three-dimensional sample scan image. The anchor boxes are used to determine the target nodules. Based on the target information corresponding to each of the two-dimensional scan images, the target nodules are merged and categorized to determine the diagnostic results corresponding to the three-dimensional scan images. The diagnostic results include the number, size, and location of the target nodules in the three-dimensional scan images, and also include: In the target information corresponding to each of the two-dimensional scan images, the first anchor frame information corresponding to the target nodule is obtained, wherein the first anchor frame information includes the corresponding first nodule bounding box and the center point coordinate parameters of the first nodule bounding box; The center point coordinate parameters of the first nodule bounding boxes corresponding to all the two-dimensional scan images are sequentially traversed, and the distance between the center points of two adjacent first nodule bounding boxes is calculated according to the center point coordinate parameters. Determine whether the distance is not greater than a preset distance threshold, and if the distance is determined to be not greater than the preset distance threshold, classify the corresponding two first nodal bounding boxes into the corresponding first bounding box group to obtain multiple first bounding box groups; The number of target nodules is determined based on the number of groups in the first bounding box group; Based on the target nodule bounding box selected from all the first nodule bounding boxes in the first bounding box group, the size of the target nodule corresponding to the first bounding box group is determined, and based on the center point coordinate parameters corresponding to the target nodule bounding box, the position of the target nodule corresponding to the first bounding box group is determined.

2. The method according to claim 1, characterized in that, If it is determined that the distance is greater than a threshold distance, the method further includes: The first nodule bounding box that appears first in the two first nodule bounding boxes currently being calculated is grouped into the corresponding first bounding box group. In the remaining first nodal bounding boxes to be classified, starting from the second first nodal bounding box among the two currently participating in the calculation, the distance between the center points of two adjacent first nodal bounding boxes is calculated sequentially, and the first nodal bounding box whose distance is not greater than a preset distance threshold is selected, until the distance is greater than the preset distance threshold or all the remaining first nodal bounding boxes to be classified are traversed, to obtain another group of first bounding boxes. The first nodal bounding box that is second among the two currently participating in the calculation is the first nodal bounding box of the other group of first bounding boxes.

3. The method according to claim 1, characterized in that, The first anchor frame information also includes the first anchor frame size corresponding to the first nodule bounding box. The target nodule bounding box selected from all the first nodule bounding boxes in the first bounding box group includes: selecting the first nodule bounding box with the largest first anchor frame size from all the first nodule bounding boxes in the first bounding box group to obtain the target nodule bounding box.

4. The method according to claim 1, characterized in that, Using a trained target recognition network, target nodules are detected in the target organ region of the two-dimensional scan image to obtain lesion information, including: Using the target recognition network, nodule target detection is performed on the two-dimensional scan image to obtain label information and second anchor box information corresponding to the candidate target. The label information includes the nodule category corresponding to the candidate target and the target confidence level corresponding to the nodule category. The second anchor box information includes the second bounding box corresponding to the candidate target, the center point coordinate parameters of the second bounding box, and the second anchor box size. Based on the nodule category, candidate nodules are selected from multiple candidate targets, and it is determined whether the target confidence corresponding to the candidate nodule is greater than a preset confidence threshold. If it is determined that the target confidence level corresponding to the candidate nodule is greater than a preset confidence threshold, the candidate nodule is determined as the target nodule, and the second anchor frame information corresponding to the candidate nodule is used as the target information corresponding to the target nodule to obtain the lesion information.

5. The method according to claim 4, characterized in that, If it is determined that the target confidence scores corresponding to the candidate nodules are all not greater than a preset confidence threshold, then it is determined that the target nodules do not exist in the two-dimensional scan image. If it is determined that the target confidence level corresponding to all the candidate nodules is not greater than a preset confidence threshold, the diagnostic result is determined to include the absence of the target nodule in the three-dimensional scan image.

6. The method according to claim 1, characterized in that, The training steps of the target recognition network include: The three-dimensional sample scan image is sliced ​​according to a preset slicing direction to generate a first sample slice image set corresponding to the two-dimensional sample slice image. The three-dimensional sample scan image includes the sphere center coordinate parameters and the sphere size parameters of the target nodule. The first sample slice image set includes a first sphere center sample slice image and a first non-sphere center plane slice image corresponding to the plane where the sphere center of the target nodule is located. The target organ and dirty region are segmented and extracted using Mask R-CNN on the first sphere-centered sample slice image and the first non-sphere-centered plane slice image to determine the corresponding target organ and dirty region in the first sphere-centered sample slice image and the first non-sphere-centered plane slice image. Based on the height of each of the first non-sphere-centered planar slice images relative to the first sphere-centered sample slice image, the sphere-centered coordinate parameters of the target nodule, and the sphere size parameters of the nodule, the sphere radius of the nodules present in the first sphere-centered sample slice image and the first non-sphere-centered planar slice image is determined. Based on the determined sphere radius and the sphere-centered coordinate parameters of the target nodule, anchor frames are added to the first sphere-centered sample slice image and the first non-sphere-centered planar slice image of the target organ region to generate the second sphere-centered sample slice image and the second non-sphere-centered planar slice image. The YOLOv5 algorithm is trained using the actual lesion information corresponding to the three-dimensional sample scan image, the second sphere-centered sample slice image, and the second non-sphere-centered plane slice image until convergence, thus obtaining the target recognition network.

7. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to run the computer program to perform the steps of the image-assisted diagnostic method based on instance segmentation as described in any one of claims 1 to 6.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the image-assisted diagnosis method based on instance segmentation as described in any one of claims 1 to 6.