Image detection methods, apparatus, equipment and storage media

TWI938350BActive Publication Date: 2026-09-11ALIBABA (CHINA) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
TW111132012
Authority / Receiving Office
TW · TW
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-05-24
Filing Date
2022-08-25
Publication Date
2026-09-11
Estimated Expiration
2042-08-24

AI Technical Summary

Technical Problem

Pancreatic cancer is difficult to detect early due to low contrast in plain CT images, leading to late-stage diagnoses and low survival rates, and enhanced CT scans pose risks and increase costs.

Method used

An image detection method using two neural network models for image classification and segmentation, where the first model identifies the presence of any lesion type and the second model distinguishes specific lesion types, eliminating the need for contrast agents and enhancing detection accuracy.

Benefits of technology

Enables accurate and early detection of pancreatic cancer and other pancreas-related diseases without contrast agents, improving survival rates by identifying lesion types and areas with high precision.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure TWG2TB001909910_001
    Figure TWG2TB001909910_001
  • Figure TWG2TB001909910_002
    Figure TWG2TB001909910_002
  • Figure TWG2TB001909910_003
    Figure TWG2TB001909910_003
Patent Text Reader

Abstract

This invention provides an image detection method, apparatus, device, and storage medium. The method includes: acquiring a detection image obtained by plain CT scan; extracting a target body part image from the detection image; performing a first image classification and segmentation process on the target body part image using a first image detection model to determine the presence of a first target lesion type and a lesion region corresponding to the first target lesion type in the target body part image; and performing a second image classification and segmentation process on the target body part image using a second image detection model to determine a second target lesion type and a lesion region present in the target body part image, wherein the second target lesion type is a subcategory of the first target lesion type. Through the collaboration of the two image detection models, refined detection of images including lesion type and lesion region can be achieved.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This invention relates to the field of artificial intelligence technology, and in particular to an image detection method, apparatus, device and storage medium. [Previous Technology]

[0002] Pancreatic cancer is a highly fatal malignant tumor that is difficult to screen and detect. Once discovered, it is often in an advanced stage, making surgery impossible, and the 5-year survival rate is very low. Currently, for pancreatic cancer and other pancreatic-related diseases, diagnosis is often aided by acquiring computed tomography (CT) images. Commonly used CT scans include contrast-enhanced CT and plain CT. Contrast-enhanced CT requires the injection of contrast agents, which pose a risk of allergic reactions, increase costs, and expose patients to more radiation due to the multi-phase scanning. The CT images shown above contain a wealth of valuable information, making detailed analysis of these images as needed crucial. [Summary of the Invention]

[0003] Embodiments of the present invention provide an image detection method, apparatus, device, and storage medium for achieving refined detection of images containing lesion information. In a first embodiment, the present invention provides an image detection method, the method comprising: acquiring a detection image obtained by plain CT scan; extracting a target body part image corresponding to a target body part from the detection image; performing a first image classification and segmentation process on the target body part image using a first image detection model to determine the existence of a first target lesion type and a lesion region corresponding to the first target lesion type in the target body part image; and performing a second image classification and segmentation process on the target body part image using a second image detection model to determine a second target lesion type and a lesion region in the target body part image, wherein the second target lesion type is a subcategory of the first target lesion type. In a second embodiment of the present invention, an image detection device is provided, comprising: an acquisition module for acquiring detection images obtained by plain scan computed tomography; a segmentation module for extracting target body part images corresponding to target body parts from the detection images; a first detection module for performing first image classification and segmentation processing on the target body part images using a first image detection model to determine the presence of a first target lesion type and a lesion region corresponding to the first target lesion type in the target body part images; and a second detection module for performing second image classification and segmentation processing on the target body part images using a second image detection model to determine a second target lesion type and a lesion region present in the target body part images, wherein the second target lesion type is a subcategory of the first target lesion type. In a third embodiment of the present invention, an electronic device is provided, comprising: a memory, a processor, and a communication interface; wherein the memory stores executable code, and when the executable code is executed by the processor, the processor performs the image detection method as described in the first embodiment. In the fourth embodiment of the present invention, a non-transitory machine-readable storage medium is provided, wherein executable code is stored on the non-transitory machine-readable storage medium, and when the executable code is executed by a processor of an electronic device, the processor is able to implement at least the image detection method as described in the first embodiment.Fifthly, this embodiment of the invention provides an image detection method, the method comprising: receiving a request triggered by a user device calling an image detection service, the request including a detection image obtained by performing a plain CT scan of a human body; and performing the following steps using the processing resources corresponding to the image detection service: receiving a request triggered by a user device calling an image detection service, the request including a detection image obtained by a plain CT scan; and performing the following steps using the processing resources corresponding to the image detection service: extracting a target body part image corresponding to a target body part from the detection image; performing a first image classification and segmentation process on the target body part image using a first image detection model to determine the existence of a first target lesion type and a lesion region corresponding to the first target lesion type in the target body part image; performing a second image classification and segmentation process on the target body part image using a second image detection model to determine a second target lesion type and a lesion region in the target body part image, the second target lesion type being a subcategory of the first target lesion type; and feeding back the detection image marked with the second target lesion type and lesion region to the user device. The sixth embodiment of the present invention provides an image detection method applied to an extended reality device. The method includes: acquiring a detection image obtained by plain CT scan; extracting a target body part image corresponding to a target body part from the detection image; performing a first image classification and segmentation process on the target body part image using a first image detection model to determine the existence of a first target lesion type and a lesion region corresponding to the first target lesion type in the target body part image; performing a second image classification and segmentation process on the target body part image using a second image detection model to determine a second target lesion type and a lesion region in the target body part image, wherein the second target lesion type is a subcategory of the first target lesion type; and displaying the detection image marked with the second target lesion type and lesion region. In this embodiment of the present invention, taking the detection of lesions in the pancreatic region as an example, a plain CT scan of the user's chest or abdomen can be used to obtain the detection image, without relying on enhanced CT, thus eliminating the need for contrast agent injection into the user (non-invasive). Then, the target body part image corresponding to the target body part (such as the pancreatic region) is segmented from the detection image. Next, the first image detection model is used to classify and segment the target body part image to determine whether the first target lesion type and the lesion area corresponding to the first lesion type exist in the target body part image.If present, a second image detection model is used to perform second image classification and segmentation on the target body part image to determine the second target lesion type and lesion region within the target body part image. The second target lesion type is a subcategory of the first target lesion type. Therefore, the first image detection model is used to detect the presence of one or more types of diseases and corresponding lesion regions in the target body part image, with the core function of determining whether the target body part image contains lesions requiring further identification. Based on this, the second image detection model is used to further and more accurately detect the specific lesion types and lesion regions contained in the target body part image. The use of these two image detection models can provide auxiliary image information for the localization of diseases in different parts of the body, such as the pancreas, lungs, and stomach.

Implementation Method

[0005] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are some embodiments of the present invention, but not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. In addition, the sequence of steps in the following method embodiments is only an example and not a strict limitation. In real life, taking pancreatic cancer as an example, there is still no officially recommended screening method for pancreatic cancer. Most pancreatic cancer patients are diagnosed at an advanced stage, losing the opportunity for surgery, and the 5-year survival rate is very low. If pancreatic cancer can be detected early, the 5-year survival rate can be significantly improved after adjuvant chemotherapy. The first-line diagnostic imaging scheme for pancreatic cancer is to use enhanced CT, which requires the injection of contrast agent into the user (the user in this article refers to the person who needs to be tested for the disease, usually a patient). However, contrast agents pose a risk of allergic reactions to patients, increase costs, and expose patients to more radiation due to multi-phase scanning. Chest CT scans are widely used in physical examinations, and their scanning area includes most of the pancreas. However, the low contrast of plain CT images makes it difficult for doctors to visually determine whether there are tumors or cancer on the pancreas. In fact, pancreatic tumors are often missed during physical examinations, which is one of the reasons why they are often discovered at a late stage. In real life, pancreatic-related diseases include not only malignant pancreatic cancers such as pancreatic ductal adenocarcinoma (PDAC), but also pancreatic diseases such as primitive neuroectodermal tumor (PNET), intraductal papillary mucinous neoplasm (IPMN), serous cystic tumor (SCN), mucinous cystic neoplasm (MCN), solid-Pseudopapillary tumor of pancreas (SPT), and chronic pancreatitis (CP).Taking the aforementioned pancreatic-related diseases as an example, in this embodiment of the invention, the detection of common lesion types in the pancreatic region can be divided into two classification and segmentation tasks: the first is to identify three categories based on plain CT images: PDAC, non-PDAC diseases (including subtypes such as PNET, SPT, IPMN, MCN, CP, and SCN), and no disease, as well as the lesion area corresponding to each lesion type; the second is to further classify and identify the subtype based on plain CT images if the output result of the first task is either non-PDAC or PDAC. Distinguishing PDAC from other non-PDAC types is a very important classification method because PDAC accounts for 90% of pancreatic cancers and is the most malignant type of pancreatic cancer. The above is only a brief introduction to the two classification and segmentation tasks using the pancreatic region as an example. In fact, the detection of other body parts is similar. That is to say, the image detection scheme provided by this embodiment of the invention is not limited to the detection of specific diseases in certain specific locations, but provides an image detection scheme applicable to multiple lesion types in many body parts. For example, lung cancer, tuberculosis, pulmonary edema, etc. It should be noted that the image detection scheme provided in this embodiment of the invention only learns the lesion features presented in the image through a neural network model to achieve accurate detection of the lesion type and lesion region that may be contained in the image. The detection result is provided to the doctor as an intermediate result information. The image detection method provided in this embodiment of the invention is described in detail below. Figure 1 is a flowchart of an image detection method provided in this embodiment of the invention. As shown in Figure 1, the method includes the following steps: 101. Obtain the detection image obtained by plain CT scan, and extract the target body part image corresponding to the target body part from the detection image. 102. Perform a first image classification and segmentation process on the target body part image using a first image detection model to determine the existence of a first target lesion type and a lesion region corresponding to the first target lesion type in the target body part image. 103. Perform a second image classification and segmentation process on the target body part image using a second image detection model to determine the existence of a second target lesion type and lesion region in the target body part image, wherein the second target lesion type is a subcategory of the first target lesion type. In real life, the detection of many diseases requires the assistance of medical imaging. CT is a common auxiliary method. In this embodiment of the invention, when a user needs to be detected for a certain disease, only a plain CT scan is required, without the need for an enhanced CT scan. The images obtained by the plain CT scan are called detection images. When acquiring CT images, the user's entire area, such as the chest or abdomen, is often scanned. However, the diagnosis of some diseases often only requires focusing on a specific body part (target body part), such as one or several organs. Taking the screening of pancreatic-related diseases as an example, the target body part is the pancreas.Therefore, after obtaining the detection image, to facilitate the subsequent image recognition process, it is first necessary to extract the image region corresponding to the target body part from the detection image, called the target body part image, such as the pancreas region. In practical applications, a segmentation model can be pre-trained to segment the target body part image from the detection image. Taking the pancreas as an example, simply put, during the training phase of this segmentation model, a large number of training sample images with labeled information can be obtained. The training sample images are images containing the pancreas region, and the labeled information is the location of the pancreas region in the image, usually marked with polygonal boxes, indicating that the pixels within the polygonal boxes correspond to the pancreas region. Based on these training sample images with labeled information, the segmentation model can be trained. Once the training converges, the segmentation model will have the ability to locate the pancreas region. Based on this, after inputting the above-mentioned detection image into the segmentation model, the segmentation model can predict the category label corresponding to each pixel in the detection image: whether it is located within the pancreatic region. Assuming that category label 1 represents being located within the pancreatic region and category label 0 represents not being located within the pancreatic region, then based on the category label corresponding to each pixel, the continuous region formed by the pixels corresponding to category label 1 can be determined as the pancreatic region, that is, the image region corresponding to the target body part. Here, "continuous region" means that even if there are a very small number of pixels with category label 0 among a large number of pixels corresponding to category label 1, the category label 0 of these pixels will be ignored. The specific method for determining the continuous region can refer to the existing related technologies, which will not be elaborated in this embodiment. As mentioned above, taking the pancreatic disease screening scenario as an example, the pancreatic disease screening task is divided into two classification and segmentation tasks. The first classification and segmentation task is used to classify and segment the target body part image for the first group of lesion types; the second classification and segmentation task is used to classify and segment the target body part image for the second group of lesion types. Here, image segmentation refers to locating the corresponding lesion region in the target body part image for the appropriate lesion type, while classification refers to classifying and identifying the lesion type. For pancreatic diseases, the first group of lesion types can include PDAC type, non-PDAC type, and no lesion (i.e., normal). The second group of lesion types can include PNET, SPT, IPMN, MCN, CP, SCN, etc., and may even include PDAC. In an optional embodiment, the second group of lesion types can be considered a subcategory of the non-PDAC type. Therefore, in practical applications, when the first classification segmentation task outputs that the lesion type corresponding to the target body part in the first group of lesion types is a non-PDAC type, the second classification segmentation task can be performed.It should be noted that the system can also be configured to perform a second classification segmentation task when the output of the first classification segmentation task indicates the presence of a PDAC or non-PDAC type (i.e., the presence of disease) in the target body part. In this case, the second classification segmentation task needs to provide classification functionality for PDAC and non-PDAC type subtypes (subcategories). Based on the example of pancreatic diseases mentioned above, to summarize, the first image detection model is used to detect the first group of lesion types, which includes: the third target lesion type, the first target lesion type, and no lesion, which are classified according to the severity of the disease corresponding to the target body part. The first target lesion type refers to the collective term for lesion types other than the third target lesion type. That is to say, the first target lesion type does not point to a specific disease, but rather indicates the presence of an abnormality other than the third target lesion type. In practical applications, the above two classification segmentation tasks can be completed by training two image detection models respectively, namely the first image detection model and the second image detection model. It can be understood that if the detection result of the first image detection model on the target body part image is: the presence of the third target lesion type and its corresponding lesion region. Optionally, the processing of the second image detection model can be omitted. Dividing the image detection task into the two classification and segmentation tasks mentioned above has the following advantages: First, compared to using a single model to complete the image detection task, the two image detection models can each have their own focus, making it easier to train a high-performing model with a more disease-specific focus, reducing the impact of sample imbalance on model performance. Second, classifying diseases involving the same organ into the two groups mentioned above: the first group includes the most severe disease, no disease, and other diseases collectively, ensuring that the first image detection model can fully learn the characteristics of the most severe disease and the characteristics of the organ when it is normal (no disease), ensuring the accuracy of the detection of the most severe disease, thereby ensuring the timely detection of this malignant disease. The first classification and segmentation task is equivalent to a coarse classification and segmentation task, completing the identification of whether the patient has a disease and whether they have certain types of diseases. The second classification and segmentation task is equivalent to a fine classification and segmentation task, and the corresponding second image detection model is trained to learn the subtle differences in features of more types of diseases, ensuring more accurate image detection results. Third, it takes into account the needs of practical applications. Because the first image detection model completes the classification and segmentation of the first group of lesion types, it can determine not only whether the patient has the most severe disease based on the classification results, but also the lesion area based on the segmentation results. Those skilled in the art will understand that the second image detection model provided in the above embodiments can be selected and used based on actual needs.In summary, in practical applications, after locating the target body part image from the detected image, the target body part image is input into the first image detection model. The first image detection model performs classification and image segmentation processing on the target body part image for a first group of lesion types to obtain the lesion types and lesion regions present in the target body part image within the first group of lesion types. If the target body part image contains a first target lesion type (referring to one or more specific lesion types) included in the first group of lesion types, then the second image detection model performs classification and image segmentation processing on the target body part image for a second group of lesion types to determine the second target lesion type and lesion region present in the target body part image within the second group of lesion types. It is understandable that, because the above two image detection models need to provide classification and image segmentation capabilities, during model training, they need to acquire training sample images with two types of annotation information. One type of annotation information is the lesion type contained in the training sample image (included in the corresponding group of lesion types), and the other is the location of the corresponding lesion region in the training sample image for that lesion type. In an optional embodiment, both the first image detection model and the second image detection model described above can be trained using the following cross-validation training method. For ease of description, the first image detection model and the second image detection model will be used as bullseye image detection models, respectively. The training process includes: obtaining a training sample set for training the bullseye image detection models; constructing multiple training sample subsets corresponding to the training sample set; and training multiple bullseye image detection models using the multiple training sample subsets. For example, assuming there are 100 sample images in the training sample set of the bullseye image detection models, a total of 5 bullseye image detection models (different image detection models) are trained. The training sample subset corresponding to each bullseye image detection model contains 80 sample images sampled from these 100 sample images (each bullseye image detection model has a different training sample subset), and the remaining 20 sample images are used as the test set for the corresponding bullseye image detection model. Using the above training method, multiple different first image detection models and multiple different second image detection models can be obtained. Based on this, the target body part image can be classified and segmented using multiple first image detection models to obtain the output results of each of the multiple first image detection models. Then, based on the output results of each of the multiple first image detection models, it can be determined whether the target body part image has a lesion region corresponding to the first lesion type.Optionally, the output with the highest percentage among the outputs of multiple first image detection models can be selected as the final output. For example, if four out of five first image detection models output the classification result PDAC, and the lesion regions they output are the same, then the lesion type in the target body part image is determined to be PDAC, and the lesion region is the lesion region output by these four first image detection models. The same logic applies to determining the second lesion type and lesion region in the target body part image using multiple second image detection models. In summary, combining the image detection scheme provided by the embodiments of the present invention, taking pancreatic disease screening as an example, through the synergistic cooperation of the two image detection models, based on the first image detection model trained to recognize a small number of serious diseases, accurate detection of serious diseases can be achieved. Other types of diseases can be detected based on the second image detection model, realizing comprehensive and accurate identification of multiple lesion types in pancreatic images. Figure 2 illustrates the pancreatic image detection process implemented using the detection scheme provided by the embodiments of the present invention. The structure and working process of the first and second image detection models are described below. Regarding the first image detection model: Structurally, the first image detection model includes a first feature extraction sub-model and a first classification and segmentation sub-model. The first feature extraction sub-model includes a first encoding module, a first decoding module, and a jumper layer between the first encoding module and the first decoding module. From a working process perspective, the process is as follows: The first encoding module extracts a first feature map set corresponding to the target body part image. The first feature map set consists of feature maps at multiple scales. The jumper layer inputs the first feature map set into the first decoding module. The first decoding module obtains a second feature map set corresponding to the target body part image. The second feature map set consists of feature maps at multiple scales. The second feature map set is input into the first classification and segmentation sub-model, which fuses the feature maps contained in the second feature map set. Based on the fused feature maps, it is determined whether a first target lesion type and a lesion region corresponding to the first target lesion type exist in the target body part image. For ease of understanding, Figure 3 is used as an example to illustrate the composition of the first image detection model. In Figure 3, the left branch corresponds to the first feature extraction sub-model, and the right branch corresponds to the first classification and segmentation sub-model. Assuming the first encoding module consists of 6 convolutional blocks, the first decoding module also consists of 6 convolutional blocks. Each convolutional block is represented as 2xConv3D, meaning it consists of two consecutive three-dimensional (3D) convolutional layers. Each convolutional block corresponds to a different scale to extract feature maps of that scale, such as 5x8x5, 10x16x10, 20x32x20, ... as shown in the figure, where C = 320, 256, 128, ..., where C represents the number of channels.The directed connections between convolutional blocks of the same scale in the first encoding module and the first decoding module, as shown in the figure, represent skip layers between convolutional blocks of the same scale. Based on the structure of the first feature extraction sub-model shown in Figure 3, when the target body part image is input into the first encoding module, it can sequentially obtain feature maps of six scales from large to small through convolution and downsampling processing of the above six convolutional blocks. The output of the previous convolutional block is used as the input of the next convolutional block. The output of the last convolutional block in the first encoding module is input to the first decoding module. The first decoding module performs upsampling and deconvolution processing layer by layer through the six convolutional blocks, and outputs feature maps of six scales from small to large in sequence. However, after one convolutional block in the first decoding module upsamples the feature map of a certain scale i input from the previous convolutional block to obtain the feature map of scale j, it needs to use a skip layer to concatenate the feature map of scale j extracted by the first encoding module with the feature map of scale j before performing subsequent convolution and other operations. In other words, the function of the jump-join layer is to concatenate feature maps of the same scale obtained from downsampling (encoding module) and upsampling (decoding module) along the channel dimension. The feature maps of multiple scales extracted by the first encoding module are actually shallow features, forming the first feature map group; while the feature maps of multiple scales extracted by the first decoding module are actually deep features, forming the second feature map group. Feature fusion is achieved through jump-joining. As shown in Figure 3, the first classification and segmentation sub-model may include multiple pooling layers (such as the global max pooling layer shown in the figure) and fully connected layers (FC) connected to each pooling layer, as well as a feature fusion layer. The inputs corresponding to the feature maps of multiple scales contained in the second feature map group are compressed by multiple pooling layers. Finally, the compressed feature maps at multiple scales are input into the feature fusion layer for feature fusion (stitching). Based on the fused feature maps, it is determined whether there is a lesion region in the target body part image that corresponds to the first target lesion type, such as the recognition results of the three categories of PDAC / non-PDAC / normal shown in the figure, as well as the lesion region segmentation results corresponding to PDAC / non-PDAC, to enhance interpretability.Regarding the second image detection model: From a structural perspective, the second image detection model includes a first feature extraction sub-model, a second classification and segmentation sub-model, and a pooling module. The second classification and segmentation sub-model includes a memory unit and an attention module. The memory unit is trained to store the location and visual features of different lesion types within the target body part, and is configured to store the location and visual features using a target number of memory vectors. From a workflow perspective, the workflow is as follows: The second feature extraction sub-model extracts a third feature map group corresponding to the target body part image. The third feature map group consists of feature maps at multiple scales. For the target feature map in the feature maps of the multiple scales, the following steps are performed sequentially: A pooling module is used to pool the target feature map to compress it into a target number of feature vectors; an attention module is used to perform cross-attention processing on the target number of reference vectors and the target number of feature vectors, and self-attention processing on the target number of reference vectors; the cross-attention processing result and the self-attention processing result are summed; wherein, when the target feature map is the first in the feature maps of the multiple scales, the reference vector is the memory vector; when the target feature map is not the first in the feature maps of the multiple scales, the reference vector is the sum of the cross-attention processing result and the self-attention processing result corresponding to the previous target feature map; the target feature map is any one of the third feature map group; based on the sum of the cross-attention processing result and the self-attention processing result corresponding to the last target feature map, the second target lesion type and lesion region existing in the target body part image are determined. In fact, similar to the first feature extraction sub-model in the first image detection model, the second feature extraction sub-model also includes a second encoding module, a second decoding module, and a jumper layer between the second encoding module and the second decoding module. Based on this, optionally, the extraction of a third feature map group corresponding to the target body part image by the second feature extraction sub-model can be implemented as follows: extracting a fourth feature map group corresponding to the target body part image by the second encoding module; inputting the fourth feature map group into the second decoding module by the jumper layer; obtaining a fifth feature map group corresponding to the target body part image by the second decoding module; and determining that a portion of the feature maps contained in the fourth feature map group and a portion of the feature maps contained in the fifth feature map group constitute the third feature map group. Additionally, in an optional embodiment, the second image detection model may further include a location embedding module.Based on this, by using an attention module to perform cross-attention processing on the reference vector and feature vector of the target quantity, it can be achieved as follows: The corresponding position embedding vector is superimposed on the feature vector of the target quantity, where the position embedding vector superimposed on any feature vector represents the positional information of that feature vector within the feature vector of the target quantity; The attention module performs cross-attention processing on the reference vector and feature vector of the target quantity, which are superimposed with their respective corresponding position embedding vectors. The introduction of the position embedding vector essentially introduces the spatial location features of the disease on the target body part, which helps to achieve a more accurate recognition effect. For ease of understanding, Figure 4 is used as an example to illustrate the composition of the second image detection model. In Figure 4, the left branch corresponds to the second feature extraction sub-model, the middle branch corresponds to the position embedding module, and the right branch corresponds to the second classification and segmentation sub-model. Similar to Figure 3, it is assumed that the second encoding module consists of 6 convolutional blocks, and correspondingly, the second decoding module also consists of 6 convolutional blocks. The composition and related parameters of each convolutional block are shown in the figure. The directed connections between convolutional blocks of the same scale in the second encoding module and the second decoding module, as shown in the figure, represent skip layers between convolutional blocks of the same scale. Consistent with the working process of the first feature extraction sub-model described above, based on the structure of the second feature extraction sub-model shown in Figure 4, when the target body part image is input into the second encoding module, it sequentially undergoes convolution and downsampling processing through the aforementioned six convolutional blocks, resulting in six feature maps of varying scales from largest to smallest, forming the fourth feature map group. The second decoding module then sequentially performs upsampling and deconvolution processing through the six convolutional blocks, outputting six feature maps of varying scales from smallest to largest, forming the fifth feature map group. Subsequently, as shown in Figure 4, partial feature maps are selected from the fourth and fifth feature map groups respectively to form the third feature map group. It should be noted that, theoretically, the third feature map group can also be composed of all the feature maps contained in the fourth and fifth feature map groups, but this would increase computational complexity. Furthermore, since the fourth and fifth feature map groups correspond to shallow and deep features respectively, selecting a subset of feature maps from these groups can balance the needs of both shallow and deep features while minimizing computational cost. Next, for each feature map in the third feature map group, a pooling module is used to pool each feature map, compressing it into a target number of feature vectors. The pooling module can provide adaptive average pooling, as illustrated in Figure 4. In Figure 4, the target number is assumed to be 5x8x5=200, matching the size of the 200 320-dimensional vectors stored in the memory unit.In the feature maps of multiple scales illustrated in Figure 4, the smallest scale is 5x8x5, and the other scales are integer multiples of this scale. Therefore, feature maps of other scales can be compressed based on the 5x8x5 scale, and the compression result is represented by 200 feature vectors. In an optional embodiment, after each feature map in the third feature map is compressed to the target number of feature vectors, a corresponding position embedding vector (pos) can be superimposed on each feature vector. The position embedding vector superimposed on any feature vector is used to characterize the position information of the corresponding feature vector in the target number of feature vectors. Then, as shown in Figure 4, the attention module performs cross-attention processing on the target number of reference vectors and the target number of feature vectors superimposed with their respective corresponding position embedding vectors, and performs self-attention processing on the target number of reference vectors. The results of the cross-attention processing and the self-attention processing are summed. The target number of reference vectors originates from the target number of memory vectors stored in the memory unit. The memory vector of the target quantity stored in the memory unit is updated and learned during the training of the second image detection model. When the model converges, the memory vector of the target quantity stored in the memory unit will be finally stored for later use. As shown in Figure 4, each feature map in the third feature map group, after the above compression and position embedding processing, is equivalent to forming a vector sequence. Each vector sequence is input into the attention module in sequence. Thus, the attention module can be regarded as being composed of multiple attention units as shown in the figure. Each attention unit provides the functions of self-attention mechanism and cross-attention mechanism. For ease of description and understanding, the feature vectors of the target quantity after pooling processing of the four feature maps in the third feature map group shown in the figure are represented as C1, C2, C3, and C4, respectively. The corresponding four position embedding vectors are represented as P1, P2, P3, and P4, respectively. The memory vector of the target quantity stored in the memory unit is represented as M0. Based on this, the self-attention processing of the memory vector of the target quantity by the first attention unit is represented as: self-attention(M0) = M01. The self-attention processing of the feature vector C1 of the target quantity memory vector and the target quantity superimposed with their respective position embedding vectors is represented as: cross-attention(C1+P1,M0) = M02, where C1+P1 represents the superposition of the target quantity feature vector C1 and its respective position embedding vector. Let the sum of the self-attention processing result and the cross-attention processing result be: M01 + M02 = M1. Then M1 is the input of the next attention unit: the reference vector of the target quantity.The self-attention processing of the second attention unit on the reference vector M1 of the target quantity at this time is represented as: self-attention(M1) = M11. The self-attention processing of the second attention unit on the feature vector C2 of the target quantity reference vector and the feature vector C2 of the target quantity superimposed on their respective position embedding vectors is represented as: cross-attention(C2+P2,M1) = M12. The sum of the self-attention processing result and the cross-attention processing result is denoted as: M11 + M12 = M2. Similarly, the self-attention processing of the third attention unit on the reference vector M2 of the target quantity at this time is represented as: self-attention(M2) = M21. The self-attention processing of the third attention unit on the feature vector C3 of the target quantity reference vector and the feature vector C3 of the target quantity superimposed on their respective position embedding vectors is represented as: cross-attention(C3+P3,M2) = M22. The sum of the self-attention processing result and the cross-attention processing result is denoted as: M21 + M22 = M3. The self-attention processing of the fourth attention unit on the reference vector M3 for the target quantity at this time is expressed as: self-attention(M3) = M31. The self-attention processing of the fourth attention unit on the feature vector C4, which is the superposition of the reference vector for the target quantity and the corresponding position embedding vector of the target quantity, is expressed as: cross-attention(C4+P4,M3) = M32. The sum of the self-attention processing result and the cross-attention processing result at this time is: M31+M32=M4. Then, based on M4, the type and region of the second target lesion in the target body part image are determined. For example, the response module shown in Figure 4 performs some kind of pooling processing on the 200 320-dimensional vectors corresponding to M4, selects the feature value from each vector, inputs it into the subsequent fully connected layer (FC), and then outputs the prediction results of the lesion type and lesion region. In this embodiment, the memory unit stores a set of globally shared model parameters. During model training, the initial values ​​in the memory unit can be randomly determined, and these parameters are updated each time the model undergoes repeated training. The memory unit is designed to learn global contextual information and location information, such as the relative location of pancreatic tumors within the pancreas, thereby providing distinguishable descriptors for each pancreatic disease type included under the first target lesion type (i.e., the second group of lesion types). In other words, the memory unit aims to store feature information such as the location (spatial) and texture (visual) of different pancreatic diseases. The updating and construction of this feature information requires the use of self-attention and cross-attention mechanisms. The image detection method provided in this embodiment can be executed in the cloud, where several computing nodes can be deployed, each with processing resources such as computing and storage.In the cloud, multiple computing nodes can be organized to provide a certain service; of course, a single computing node can also provide one or more services. The cloud can provide this service through a service interface, which users can call to use the corresponding service. Service interfaces include software development kits (SDKs), application programming interfaces (APIs), and other forms. Regarding the solution provided in this embodiment of the invention, the cloud can provide a service interface with image detection services. Users can call this service interface through their devices to trigger an image detection request to the cloud, which includes detection images obtained through plain CT scans. The cloud determines the computing node that will respond to the request and utilizes the processing resources of that computing node to perform the following steps: Extracting the target body part image corresponding to the target body part from the detected image; Performing a first image classification and segmentation process on the target body part image using a first image detection model to determine the existence of a first target lesion type and a lesion region corresponding to the first target lesion type in the target body part image; Performing a second image classification and segmentation process on the target body part image using a second image detection model to determine the existence of a second target lesion type and a lesion region in the target body part image, wherein the second target lesion type is a subcategory of the first target lesion type; Feeding back the detected image marked with the second target lesion type and lesion region to the user device. The above execution process can be referred to the relevant descriptions in the other embodiments mentioned above, and will not be repeated here. For ease of understanding, Figure 5 is used as an example. The user can use the user device E1 shown in Figure 5 to call the image detection service to upload a service request containing a detected image. In the cloud, as shown in the figure, in addition to deploying several computing nodes, a management node E2 running management services is also deployed. After receiving a service request sent by the user device E1, the management node E2 determines the computing node E3 to respond to the service request. After receiving the service request, the computing node E3 executes the above-mentioned calculation process to obtain a detection image marked with the second target lesion type and lesion area. Then, the computing node E3 sends the detection image with these marking information to the user device E1, and the user device E1 displays the detection image, on which the user can perform further editing and other operations. The image detection device of one or more embodiments of the present invention will be described in detail below. Those skilled in the art will understand that these devices can all be configured using commercially available hardware components by the steps taught in this solution. Figure 6 is a schematic diagram of the structure of an image detection device provided by an embodiment of the present invention. As shown in Figure 6, the device includes: an acquisition module 11, a cutting module 12, a first detection module 13, and a second detection module 14.The acquisition module 11 is used to acquire detection images obtained by plain scan computed tomography. The segmentation module 12 is used to extract the target body part image corresponding to the target body part from the detection images. The first detection module 13 is used to perform a first image classification and segmentation process on the target body part image using a first image detection model to determine the presence of a first target lesion type and a lesion region corresponding to the first target lesion type in the target body part image. The second detection module 14 is used to perform a second image classification and segmentation process on the target body part image using a second image detection model to determine the presence of a second target lesion type and a lesion region in the target body part image, wherein the second target lesion type is a subcategory of the first target lesion type. Optionally, the first image detection model is used to detect a first group of lesion types, which includes: a third target lesion type, the first target lesion type, and no lesion, which are divided according to the severity of the disease corresponding to the target body part. The first target lesion type refers to the collective term for lesion types other than the third target lesion type. Optionally, the first image detection model includes a first feature extraction sub-model and a first classification and segmentation sub-model. The first feature extraction sub-model includes a first encoding module, a first decoding module, and a jumper layer between the first encoding module and the first decoding module. Based on this, the first detection module 13 is used to: extract a first feature map group corresponding to the target body part image using the first encoding module, the first feature map group consisting of feature maps at multiple scales; input the first feature map group to the first decoding module using the jumper layer; obtain a second feature map group corresponding to the target body part image using the first decoding module, the second feature map group consisting of feature maps at multiple scales; input the second feature map group to the first classification and segmentation sub-model, so that the first classification and segmentation sub-model fuses the feature maps contained in the second feature map group, and determines whether a first target lesion type and a lesion region corresponding to the first target lesion type exist in the target body part image based on the fused feature maps. Optionally, the second image detection model includes a second feature extraction sub-model, a second classification and segmentation sub-model, and a pooling module. The second classification and segmentation sub-model includes a memory unit and an attention module. The memory unit is trained to store the location and visual features of different lesion types included under the first target lesion type in the target body part. The memory unit is configured to store the location and visual features in a target number of memory vectors.Based on this, the second detection module 14 is used to: extract a third feature map group corresponding to the target body part image using the second feature extraction sub-model, the third feature map group being composed of feature maps at multiple scales; sequentially, for the target feature map in the multiple scales: perform pooling processing on the target feature map using the pooling module to compress the target feature map into the target number of feature vectors; and perform cross-attention processing on the reference vector of the target number and the feature vector of the target number using the attention module, perform self-attention processing on the reference vector of the target number, and sum the cross-attention processing result and the self-attention processing result; wherein, when the target feature map is the first in the multiple scales of feature maps, the reference vector is the memory vector, and when the target feature map is not the first in the multiple scales of feature maps, the reference vector is the sum of the cross-attention processing result and the self-attention processing result corresponding to the previous target feature map; the target feature map is any one of the multiple scales of feature maps; and determine the second target lesion type and lesion region existing in the target body part image based on the sum of the cross-attention processing result and the self-attention processing result corresponding to the last target feature map. Optionally, the second image detection model includes a location embedding module. Based on this, the second detection module 14 is further configured to: superimpose corresponding location embedding vectors onto the feature vectors of the target quantity, wherein the location embedding vector superimposed on any feature vector is used to characterize the location information corresponding to any feature vector in the feature vectors of the target quantity; and perform cross-attention processing on the reference vectors of the target quantity and the feature vectors of the target quantity superimposed with their respective location embedding vectors using the attention module. Optionally, the second feature extraction sub-model includes a second encoding module, a second decoding module, and a jumper layer between the second encoding module and the second decoding module. Based on this, the second detection module 14 is configured to: extract a fourth feature map group corresponding to the target body part image using the second encoding module; input the fourth feature map group into the second decoding module using the jumper layer; obtain a fifth feature map group corresponding to the target body part image using the second decoding module; and determine that a portion of the feature maps contained in the fourth feature map group and a portion of the feature maps contained in the fifth feature group constitute the third feature map group. Optionally, the first image detection model and the second image detection model are used as bullseye image detection models, respectively. The device further includes: a training module, used to acquire a training sample set for training the bullseye image detection models; construct multiple training sample subsets corresponding to the training sample set; and train multiple bullseye image detection models using the multiple training sample subsets, respectively.Based on this, the first detection module 13 is specifically used to: perform first image classification and segmentation processing on the target body part image using multiple first image detection models to obtain the output results of each of the multiple first image detection models; and determine whether a first target lesion type and a lesion region corresponding to the first target lesion type exist in the target body part image based on the output results of each of the multiple first image detection models. The second classification module detection is specifically used to: perform second image classification and segmentation processing on the target body part image using multiple second image detection models to obtain the output results of each of the multiple second image detection models; and determine the second target lesion type and lesion region present in the target body part image based on the output results of each of the multiple second image detection models. The device shown in Figure 6 can perform the steps in the foregoing embodiments. For detailed execution processes and technical effects, please refer to the description in the foregoing embodiments, which will not be repeated here. In one possible design, the structure of the image detection device shown in Figure 6 can be implemented as an electronic device. As shown in Figure 7, the electronic device may include: a processor 21, a memory 22, and a communication interface 23. Memory 22 stores executable code, which, when executed by processor 21, enables processor 21 to at least implement the image detection method provided in the foregoing embodiments. Additionally, this embodiment provides a non-transitory machine-readable storage medium storing executable code, which, when executed by a processor of an electronic device, enables the processor to at least implement the image detection method provided in the foregoing embodiments. In an optional embodiment, the electronic device used to execute the image detection method provided in this embodiment can be an Extended Reality (XR) device. XR is a general term for various forms such as virtual reality and augmented reality. At this point, the detection images obtained from plain CT scans can be input into the extended reality device. The extended reality device contains the aforementioned first and second image detection models. These two models perform the image classification and segmentation processing described above on the target body part images extracted from the detection images, ultimately determining the type and region of the second target lesion present in the target body part image, and displaying the detection image marked with the second target lesion type and region. In this way, doctors and users can see this marked information through the extended reality device. In practical applications, for ease of viewing, optionally, after receiving the initially acquired detection images, or detection images with the aforementioned marked information, the extended reality device can generate a virtual environment for clearer viewing of the detection images, rendering and displaying the detection images within this virtual environment. Furthermore, doctors and users can also input interactive operations into the extended reality device during viewing, such as rotating or zooming in on the detection images.The device embodiments described above are merely illustrative. The units described as separate elements may or may not be physically separate. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without creative effort. From the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of necessary general-purpose hardware platforms, or by a combination of hardware and software. Based on this understanding, the above technical solutions, in essence or in terms of their contribution to the prior art, can be embodied in the form of a computer product. This invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code. Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention. [Simplified Explanation of the Diagram]

[0004] To more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. [Figure 1] is a flowchart of an image detection method provided in an embodiment of the present invention; [Figure 2] is an application diagram of a pancreatic disease screening process provided in an embodiment of the present invention; [Figure 3] is a schematic diagram of the composition of a first image detection model provided in an embodiment of the present invention; [Figure 4] is a schematic diagram of the composition of a second image detection model provided in an embodiment of the present invention; [Figure 5] is an application diagram of an image detection method provided in an embodiment of the present invention; [Figure 6] is a structural schematic diagram of an image detection device provided in an embodiment of the present invention; [Figure 7] is a structural schematic diagram of an electronic device provided in this embodiment.

Claims

1. An image detection method, executed by a processor, characterized in that it comprises: Acquire detection images obtained by plain CT scan; extract target body part images corresponding to the target body part from the detection images; perform first image classification and segmentation processing on the target body part images using a first image detection model to determine the presence of a first target lesion type and the lesion region corresponding to the first target lesion type in the target body part images; perform second image classification and segmentation processing on the target body part images using a second image detection model to determine the presence of a second target lesion type and the lesion region in the target body part images, wherein the second target lesion type is a subcategory of the first target lesion type, wherein the first image detection model is used to detect a first group of lesion types, which includes: a third target lesion type divided according to the severity of the disease corresponding to the target body part, the first target lesion type, and no lesion, wherein the first target lesion type refers to the collective term for lesion types other than the third target lesion type.

2. The method according to request item 1, wherein, The first image detection model includes a first feature extraction sub-model and a first classification and segmentation sub-model. The first feature extraction sub-model includes a first encoding module, a first decoding module, and a jumper layer between the first encoding module and the first decoding module. The first image detection model is used to perform first image classification and segmentation processing on the target body part image, including: extracting a first feature map group corresponding to the target body part image using the first encoding module, the first feature map group consisting of feature maps at multiple scales; inputting the first feature map group into the first decoding module using the jumper layer; obtaining a second feature map group corresponding to the target body part image using the first decoding module, the second feature map group consisting of feature maps at multiple scales; inputting the second feature map group into the first classification and segmentation sub-model, so that the first classification and segmentation sub-model fuses the feature maps contained in the second feature map group, and determines whether a first target lesion type and a lesion region corresponding to the first target lesion type exist in the target body part image based on the fused feature maps.

3. The method according to request item 1, wherein, The second image detection model includes a second feature extraction sub-model, a second classification and segmentation sub-model, and a pooling module. The second classification and segmentation sub-model includes a memory unit and an attention module. The memory unit is trained to store the location and visual features of different lesion types within the first target lesion type in the target body part. The memory unit is configured to store the location and visual features using a memory vector of the target number. The second image classification and segmentation process performed on the target body part image using the second image detection model includes: extracting a third feature map group corresponding to the target body part image using the second feature extraction sub-model. The third feature map group consists of feature maps at multiple scales. For the target feature map in the feature maps of the multiple scales, the pooling module performs pooling processing on the target feature map to compress it into feature vectors of the target number; and the attention module performs cross-attention processing on the reference vectors of the target number and the feature vectors of the target number, and performs self-attention processing on the reference vectors of the target number, summing the cross-attention processing results and the self-attention processing results; wherein, when the target feature map is the first in the feature maps of the multiple scales, the reference vector is the memory vector; when the target feature map is not the first in the feature maps of the multiple scales, the reference vector is the sum of the cross-attention processing results and the self-attention processing results corresponding to the previous target feature map; the target feature map is any one of the feature maps of the multiple scales; based on the sum of the cross-attention processing results and the self-attention processing results corresponding to the last target feature map, the second target lesion type and lesion region existing in the target body part image are determined.

4. The method according to request item 3, wherein, The second image detection model includes a location embedding module; The attention module performs cross-attention processing on the reference vector and feature vector of the target quantity, including: superimposing corresponding position embedding vectors on the feature vector of the target quantity, wherein the position embedding vector superimposed on any feature vector is used to represent the position information of any feature vector in the feature vector of the target quantity; and performing cross-attention processing on the reference vector and feature vector of the target quantity superimposed on their respective position embedding vectors.

5. The method according to request item 3, wherein, The second feature extraction sub-model includes a second encoding module, a second decoding module, and a jumper layer between the second encoding module and the second decoding module. Extracting the third feature map group corresponding to the target body part image using the second feature extraction sub-model includes: extracting the fourth feature map group corresponding to the target body part image using the second encoding module; inputting the fourth feature map group into the second decoding module using the jumper layer; obtaining the fifth feature map group corresponding to the target body part image using the second decoding module; and determining that a portion of the feature maps contained in the fourth feature map group and a portion of the feature maps contained in the fifth feature map group constitute the third feature map group.

6. The method according to request item 1, wherein, Using the first image detection model and the second image detection model as bullseye image detection models respectively, the method further includes: obtaining a training sample set for training the bullseye image detection model; constructing multiple training sample subsets corresponding to the training sample set; and training multiple bullseye image detection models using the multiple training sample subsets respectively.

7. The method according to claim 6, wherein, The method of performing first image classification and segmentation processing on the target body part image using a first image detection model to determine the presence of a lesion region corresponding to a first target lesion type in the target body part image includes: performing first image classification and segmentation processing on the target body part image using multiple first image detection models to obtain the output results of each of the multiple first image detection models; determining whether the target body part image contains a first target lesion type and a lesion region corresponding to the first target lesion type based on the output results of each of the multiple first image detection models; and performing second image classification and segmentation processing on the target body part image using a second image detection model to determine the second target lesion type and lesion region present in the target body part image includes: performing second image classification and segmentation processing on the target body part image using multiple second image detection models to obtain the output results of each of the multiple second image detection models; determining the second target lesion type and lesion region present in the target body part image based on the output results of each of the multiple second image detection models.

8. An image detection device, characterized in that it comprises: The system comprises: an acquisition module for acquiring detection images obtained through plain CT scans; a segmentation module for extracting target body part images corresponding to the target body part from the detection images; a first detection module for performing first image classification and segmentation processing on the target body part images using a first image detection model to determine the presence of a first target lesion type and the corresponding lesion region in the target body part images; and a second detection module for performing second image classification and segmentation processing on the target body part images using a second image detection model to determine the presence of a second target lesion type and the lesion region in the target body part images, wherein the second target lesion type is a subcategory of the first target lesion type. The first image detection model is used to detect a first group of lesion types, which includes: a third target lesion type classified according to the severity of the disease corresponding to the target body part, the first target lesion type, and no lesion. The first target lesion type refers to all lesion types except the third target lesion type.

9. An electronic device, characterized in that it comprises: The memory, processor, and communication interface are provided; wherein the memory stores executable code that, when executed by the processor, causes the processor to perform the image detection method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Generation system and generation method for perspective images

    TW202215374A

  • Fine-tuning a generic model via a multi-model medical scan analysis system

    US20200160520A1

  • System, method, and computer-accessible medium for virtual pancreatography

    US20200226748A1