Image Detection Method, Apparatus, Device, and Storage Medium
Through plain scanning CT and image detection models, the risk and cost problems of enhancing CT scans in the prior art are solved, and efficient and safe diagnosis of pancreatic cancer is achieved.
Patent Information
- Application Number
- CN202210575258.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-24
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2042-05-24
AI Technical Summary
The prior art In the diagnosis of pancreatic-related diseases such as pancreatic cancer, relying on enhanced CT scans, there are problems with the risk of contrast agent allergy, increased medical expenses, and patients are exposed to more radiation.
Detection images are acquired through flat-scan electronic computed tomography (CT), and the lesions at the pancreatic site are refined using the first and second image detection models, including classification of lesions and segmentation of lesions, avoiding contrast agent injection and high radiation exposure to the patient.
The refined detection of pancreatic lesions has been achieved, which improves the accuracy and safety of diagnosis, and reduces medical costs and radiation exposure risks.
Smart Images

Figure CN115018775B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to an image detection method, device, equipment and storage medium. Background Art
[0002] Pancreatic cancer is a highly lethal malignant tumor, which is not easy to screen and detect. Once detected, it is often in the advanced stage of cancer, losing the opportunity for surgery and having a very low 5-year survival rate.
[0003] For pancreatic-related diseases such as pancreatic cancer, currently, the collection of Computed Tomography (CT) images of patients is often carried out to assist in diagnosis. Commonly used CT includes enhanced CT and plain scan CT. Among them, enhanced CT requires injection of contrast agents, which poses a risk of allergy to patients, increases the cost, and exposes patients to more radiation due to multi-phase image scanning.
[0004] The above CT images will contain a lot of valuable information, and it is of great significance to perform refined detection on these images as needed. Summary of the Invention
[0005] Embodiments of the present invention provide an image detection method, device, equipment and storage medium for realizing refined detection of images containing lesion information.
[0006] In a first aspect, an embodiment of the present invention provides an image detection method, the method comprising:
[0007] Obtaining a detection image obtained by plain scan Computed Tomography;
[0008] Extracting a target body part image corresponding to a target body part from the detection image;
[0009] Performing a first image classification and segmentation process on the target body part image through a first image detection model to determine a first target lesion type and a lesion area corresponding to the first target lesion type in the target body part image;
[0010] Performing a second image classification and segmentation process on the target body part image through a second image detection model to determine a second target lesion type and a lesion area existing in the target body part image, where the second target lesion type is a subcategory of the first target lesion type.
[0011] In a second aspect, an embodiment of the present invention provides an image detection device, the device comprising:
[0012] An acquisition module, configured to obtain a detection image obtained by plain scan Computed Tomography;
[0013] A cutting module, configured to extract a target body part image corresponding to a target body part from the detection image;
[0014] A first detection module, configured to perform a first image classification and segmentation process on the target body part image through a first image detection model, so as to determine a first target lesion type and a lesion area corresponding to the first target lesion type existing in the target body part image;
[0015] A second detection module, configured to perform a second image classification and segmentation process on the target body part image through a second image detection model, so as to determine a second target lesion type and a lesion area existing in the target body part image, where the second target lesion type is a subcategory of the first target lesion type.
[0016] In a third aspect, an embodiment of the present invention provides an electronic device, including: a memory, a processor, and a communication interface; wherein, an executable code is stored on the memory, and when the executable code is executed by the processor, the processor is caused to execute the image detection method as described in the first aspect.
[0017] In a fourth aspect, an embodiment of the present invention provides a non-transitory machine-readable storage medium, on which an executable code is stored, and when the executable code is executed by a processor of an electronic device, the processor can at least implement the image detection method as described in the first aspect.
[0018] In a fifth aspect, an embodiment of the present invention provides an image detection method, where the method includes:
[0019] Receiving a request triggered by a user device by invoking an image detection service, where the request includes a detection image obtained by performing a plain scan computed tomography on a human body;
[0020] Using processing resources corresponding to the image detection service to perform the following steps:
[0021] Receiving a request triggered by a user device by invoking an image detection service, where the request includes a detection image obtained by a plain scan computed tomography;
[0022] Using processing resources corresponding to the image detection service to perform the following steps:
[0023] Extracting a target body part image corresponding to a target body part from the detection image;
[0024] Perform first image classification and segmentation processing on the target body part image through a first image detection model to determine a first target lesion type and a lesion area corresponding to the first target lesion type in the target body part image;
[0025] Perform second image classification and segmentation processing on the target body part image through a second image detection model to determine a second target lesion type and a lesion area existing in the target body part image, where the second target lesion type is a subcategory of the first target lesion type;
[0026] Feed back the detection image marked with the second target lesion type and the lesion area to the user device.
[0027] In a sixth aspect, an embodiment of the present invention provides an image detection method applied to an extended reality device, and the method includes:
[0028] Obtain a detection image obtained by plain scan computed tomography;
[0029] Extract a target body part image corresponding to a target body part from the detection image;
[0030] Perform first image classification and segmentation processing on the target body part image through a first image detection model to determine a first target lesion type and a lesion area corresponding to the first target lesion type in the target body part image;
[0031] Perform second image classification and segmentation processing on the target body part image through a second image detection model to determine a second target lesion type and a lesion area existing in the target body part image, where the second target lesion type is a subcategory of the first target lesion type;
[0032] Display the detection image marked with the second target lesion type and the lesion area.
[0033] In an embodiment of the present invention, taking the detection of lesions in the pancreas as an example, a plain CT scan can be performed on the chest or abdomen of a user to obtain a detection image, without relying on enhanced CT, and thus there is no need to inject a contrast agent into the user (non-invasive). Then, the target body part image corresponding to the target body part (such as the pancreatic region) is segmented from the detection image. Then, the target body part image is first subjected to image classification and segmentation processing by a first image detection model to determine whether there is a first target lesion type and the lesion region corresponding to the first lesion type in the target body part image. If so, the target body part image is subjected to second image classification and segmentation processing by a second image detection model to determine the second target lesion type and the lesion region existing in the target body part image, where the second target lesion type is a subcategory of the first target lesion type.
[0034] It can be seen from this that the above first image detection model is used to detect whether there is a certain type or several types of diseases and the corresponding lesion regions in the target body part image, and the core is to determine whether the target body part image contains lesions that need to be further identified; based on this, the second image detection model is used to further accurately detect the specific lesion types and lesion regions contained in the target body part image. Based on the use of the above two image detection models, auxiliary image information can be provided for the localization of diseases related to different parts such as the pancreas, lungs, and stomach. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0036] Figure 1 It is a flowchart of an image detection method provided by an embodiment of the present invention;
[0037] Figure 2 It is an application schematic diagram of a pancreatic disease screening process provided by an embodiment of the present invention;
[0038] Figure 3 It is a composition schematic diagram of a first image detection model provided by an embodiment of the present invention;
[0039] Figure 4 It is a composition schematic diagram of a second image detection model provided by an embodiment of the present invention;
[0040] Figure 5 It is an application schematic diagram of an image detection method provided by an embodiment of the present invention;
[0041] Figure 6 The structural schematic diagram of an image detection device provided by an embodiment of the present invention;
[0042] Figure 7 The structural schematic of an electronic device provided by this embodiment. Specific implementation manners
[0043] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0044] In addition, the step timings in the following method embodiments are only examples and are not strictly limited.
[0045] In real life, taking pancreatic cancer as an example, there is no officially recommended screening method for pancreatic cancer so far. Most patients with pancreatic cancer are found to be in the advanced stage once diagnosed, losing the opportunity for surgery, and the 5-year survival rate is very low. If pancreatic cancer can be detected early and adjuvant chemotherapy is given after surgery, the 5-year survival rate is expected to increase significantly.
[0046] The first-line diagnostic imaging scheme for pancreatic cancer is to use enhanced CT, and a contrast agent needs to be injected into the user (the user in this article refers to the person who needs to undergo disease detection, usually the patient). However, the contrast agent poses a risk of allergy to the patient, increases the cost, and exposes the patient to more radiation due to the multi-phase imaging scan.
[0047] Plain chest CT is widely used in physical examinations, and its scanning range already includes most of the pancreas. However, the contrast of the plain CT image is relatively low, and it is very difficult for doctors to judge with the naked eye whether there is a tumor on the pancreas and whether it is cancer. In fact, it often happens that pancreatic tumors are missed during physical examinations, which is also one of the reasons why most patients are found to be in the advanced stage once diagnosed.
[0048] In real life, pancreatic-related diseases include not only malignant pancreatic cancer such as pancreatic ductal adenocarcinoma (PDAC), but also other pancreatic diseases such as primitive neuroectodermal tumor (PNET), intraductal papillary mucinous neoplasm (IPMN), serous cystadenoma (SCN), mucinous cystic neoplasm (MCN), solid-pseudopapillary tumor of the pancreas (SPT), and chronic pancreatitis (CP).
[0049] Taking the above-mentioned pancreatic-related diseases as an example, in the embodiments of the present invention, the detection of common lesion types in the pancreatic region can be considered to be divided into two classification and segmentation tasks: the first is to identify three categories, namely PDAC, non-PDAC diseases (including subtypes such as PNET, SPT, IPMN, MCN, CP, SCN, etc.), and disease-free, and the lesion regions corresponding to each lesion type based on non-contrast CT images; the second is to further perform subtype classification and identification based on non-contrast CT images if the output result of the first task is non-PDAC type or PDAC.
[0050] The classification of PDAC and other non-PDAC type diseases is a very important classification because PDAC accounts for 90% of pancreatic cancer and is the most malignant type of pancreatic cancer.
[0051] The above only gives a simple introduction to two classification and segmentation tasks taking the pancreatic region as an example. In fact, the detection of other body parts is the same. That is to say, the image detection scheme provided by the embodiments of the present invention is not limited to the detection of specific diseases in certain specific parts, but provides an image detection scheme applicable to a variety of lesion types in many body parts. For example, lung cancer, pulmonary tuberculosis, pulmonary edema, etc.
[0052] It should be noted that the image detection scheme provided by the embodiments of the present invention only realizes the accurate detection of the lesion categories and lesion regions that may be included in the image through the learning of the lesion features presented in the image by the neural network model, and the detection result is provided to the doctor as an intermediate result information.
[0053] The following specifically introduces the image detection method provided by the embodiments of the present invention.
[0054] Figure 1 The flowchart of an image detection method provided by an embodiment of the present invention is as follows. As Figure 1 shown, the method includes the following steps:
[0055] 101. Obtain a detection image obtained by plain scan CT, and extract a target body part image corresponding to the target body part from the detection image.
[0056] 102. Perform a first image classification and segmentation process on the target body part image through a first image detection model to determine the first target lesion type and the lesion area corresponding to the first target lesion type in the target body part image.
[0057] 103. Perform a second image classification and segmentation process on the target body part image through a second image detection model to determine the second target lesion type and the lesion area existing in the target body part image, where the second target lesion type is a subcategory of the first target lesion type.
[0058] In real life, the detection of many diseases requires the assistance of medical images. CT is a common auxiliary method. In the embodiment of the present invention, when a certain user needs to detect a certain disease, only a plain scan CT scan of the user is required, and an enhanced CT is not required. The image obtained by plain scan CT is called a detection image.
[0059] When collecting CT images, the whole area of the user, such as the chest, abdomen, etc., is often scanned, and the judgment of some diseases often only needs to focus on a certain body part (target body part) among them, such as a certain or several organs. Taking the screening of pancreatic-related diseases as an example, the target body part to be concerned about is the pancreas. Therefore, after obtaining the detection image, in order to facilitate the subsequent image recognition process, it is first necessary to extract the image area corresponding to the target body part from the detection image, which is called the target body part image, such as the pancreas area.
[0060] In practical applications, a segmentation model can be pre-trained to achieve the segmentation of the target body part image from the detection image.
[0061] Taking the pancreas as an example, simply speaking, in the training stage of the segmentation model, a large number of training sample images with annotation information can be obtained. Among them, the training sample images are images containing the pancreas area, and the annotation information is the position corresponding to the pancreas area in the image, usually marked in the form of a polygon box, indicating that the pixels within the polygon box are all corresponding to the pancreas area. Based on these training sample images with annotation information, the training of the segmentation model can be realized, and the trained segmentation model with convergence has the ability to locate the pancreas area.
[0062] Based on this, after inputting the above detection image into the segmentation model, the segmentation model can predict the class labels corresponding to each pixel in the detection image: whether it is within the pancreatic region. Assuming that the class label 1 is used to represent being within the pancreatic region and the class label 0 is used to represent not being within the pancreatic region, then based on the class labels corresponding to each pixel, the continuous region composed of the pixels corresponding to the class label 1 can be determined as the pancreatic region, that is, the image region corresponding to the target body part. Herein, the "continuous region" means that even if there are a particularly small number of pixels with the class label 0 among a large number of pixels corresponding to the class label 1, these pixels with the class label 0 will be ignored. Specifically, the method for determining the continuous region can be implemented with reference to the existing related technologies and will not be elaborated in this embodiment.
[0063] As described above, taking the pancreatic disease screening scenario as an example, the screening task of pancreatic diseases is divided into two classification and segmentation tasks. Among them, the first classification and segmentation task is used to perform the classification processing and image segmentation processing of the first group of lesion types on the target body part image; the second classification and segmentation task is used to perform the classification processing and image segmentation processing of the second group of lesion types on the target body part image.
[0064] Among them, the image segmentation processing here refers to locating the lesion region corresponding to the corresponding lesion type in the target body part image, and the classification processing is to perform the classification and recognition of the lesion type.
[0065] For pancreatic diseases, the first group of lesion types may include PDAC type, non-PDAC type, and no lesion (i.e., normal), and the second group of lesion types may include types such as PNET, SPT, IPMN, MCN, CP, SCN, etc., and may even include PDAC.
[0066] In an alternative embodiment, the second group of lesion types can be considered as subcategories under the non-PDAC type. Therefore, in practical applications, when the first classification and segmentation task outputs that the lesion type corresponding to the target body part in the first group of lesion types is the non-PDAC type, the second classification and segmentation task can be executed.
[0067] It should be noted that it can also be configured such that when the output result of the first classification and segmentation task indicates that there is PDAC or non-PDAC type (i.e., there is a disease) in the target body part, the second classification and segmentation task is executed. At this time, the second classification and segmentation task needs to provide the classification function for the subtypes (subcategories) of PDAC and non-PDAC types.
[0068] Based on the above examples of pancreatic diseases, generally speaking, the first image detection model is used to detect the first group of lesion types, which includes: the third target lesion type, the first target lesion type, and no lesion, which are sequentially divided according to the disease severity corresponding to the target body part. Among them, the first target lesion type is a general term for lesion types other than the third target lesion type. That is to say, the first target lesion type does not refer to a specific disease, but indicates an abnormal situation of a disease other than the third target lesion type.
[0069] In practical applications, the above two classification and segmentation tasks can be completed by training two image detection models respectively, namely the above-mentioned first image detection model and the second image detection model.
[0070] It can be understood that if the detection result of the first image detection model for the target body part image is that there is the third target lesion type and its corresponding lesion area. At this time, optionally, the processing process of the second image detection model can be not executed.
[0071] Dividing the image detection task into the above two classification and segmentation tasks to execute has the following advantages:
[0072] First, compared with using one model to complete the image detection task, the two image detection models can have their own focuses respectively, and it is easy to train a model with better performance and more disease focus, reducing the impact of sample imbalance on the model performance.
[0073] Second, the diseases related to the same organ are divided into the above two categories: the first group includes the most serious disease, no disease, and a general term for other diseases, ensuring that the first image detection model can fully learn the characteristics of this most serious disease and the characteristics of the organ in the normal (disease-free) state, and can ensure the detection accuracy of this most serious disease, thus ensuring the timely discovery of this malignant disease. The first classification and segmentation task is equivalent to a rough classification and segmentation task, which completes the identification process of whether the patient has a disease and whether the patient has certain types of diseases. The second classification and segmentation task is equivalent to a fine classification and segmentation task, and its corresponding second image detection model is trained to be able to learn the fine-grained difference characteristics of more types of diseases, ensuring the greater accuracy of the image detection results.
[0074] Third, it takes into account the actual application requirements. Since the first image detection model will complete the classification and segmentation processing of the first group of lesion types, not only can it be known whether the patient has the above-mentioned most serious disease based on the classification processing result, but also the lesion area of the patient can be known based on the segmentation processing result.
[0075] Those skilled in the art can understand that the second image detection model provided in the above embodiments can be selected and used according to actual needs.
[0076] In summary, in practical applications, after the target body part image is located from the detection image, the target body part image is input into the first image detection model, and the first image detection model performs classification processing and image segmentation processing on the target body part image for the first set of lesion types to obtain the lesion types and lesion regions existing in the target body part image in the first set of lesion types. If there is a first target lesion type (referring to a specific one or several lesion types) included in the first set of lesion types in the target body part image, then the second image detection model performs classification processing and image segmentation processing on the target body part image for the second set of lesion types to determine the second target lesion type and lesion region existing in the target body part image in the second set of lesion types.
[0077] It can be understood that since the above two image detection models need to provide classification and image segmentation capabilities, during the model training process, it is necessary to obtain training sample images with two types of annotation information. One type of annotation information is the lesion type included in the training sample image (included in the corresponding set of lesion types), and the other is the position of the lesion region corresponding to the lesion type in the training sample image.
[0078] In an optional embodiment, whether it is the above first image detection model or the second image detection model, the following cross-validation training method can be adopted for training. For ease of description, the first image detection model and the second image detection model will be used as the target image detection models respectively, and the training process includes:
[0079] Obtain a training sample set for training the target image detection model;
[0080] Construct multiple training sample subsets corresponding to the training sample set;
[0081] Train multiple target image detection models respectively through multiple training sample subsets.
[0082] For example, assume that there are 100 sample images in the training sample set of the target image detection model, and a total of 5 target image detection models (different image detection models) are trained. Among them, each training sample subset corresponding to the target image detection model contains 80 sample images sampled from these 100 sample images (each training sample subset corresponding to the target image detection model is different), and the remaining 20 sample images are used as the test set for the corresponding target image detection model.
[0083] Through the above training method, multiple different first image detection models and multiple different second image detection models can be obtained. Based on this, the first image classification and segmentation processing can be respectively performed on the target body part image through multiple first image detection models to obtain the output results of each of the multiple first image detection models. Then, according to the output results of each of the multiple first image detection models, it is determined whether there is a lesion area corresponding to the first lesion type in the target body part image.
[0084] Optionally, among the output results of each of the multiple first image detection models, the output result with a high proportion can be selected as the final output result. For example, if 4 out of 5 first image detection models all output the classification result of PDAC, and the lesion areas output by each are the same, then it is determined that the lesion type existing in the target body part image is PDAC, and the lesion area is the lesion area output by these 4 first image detection models.
[0085] Similarly, the second lesion type and the lesion area existing in the target body part image are determined through multiple second image detection models.
[0086] In summary, taking the pancreatic disease screening scenario as an example, in combination with the image detection solution provided in the embodiments of the present invention, through the collaborative cooperation of the above two image detection models, based on the first image detection model trained to have the recognition ability of a small number of several serious diseases, the accurate detection of serious diseases can be achieved. For some other types of diseases, they can be detected based on the second image detection model, realizing the comprehensive and accurate recognition of the pancreatic part image for multiple lesion types. Figure 2 Schematically shows the pancreatic image detection process implemented in combination with the detection solution provided in the embodiments of the present invention.
[0087] The structures and working processes of the first image detection model and the second image detection model are introduced below.
[0088] Regarding the first image detection model:
[0089] From a structural perspective, the first image detection model includes a first feature extraction sub-model and a first classification and segmentation sub-model. Among them, the first feature extraction sub-model includes a first encoding module, a first decoding module, and a skip connection layer between the first encoding module and the first decoding module.
[0090] From a working process perspective, the working process is as follows:
[0091] The first feature map group corresponding to the target body part image is extracted through the first encoding module, and the first feature map group is composed of feature maps of multiple scales;
[0092] The first feature map group is input into the first decoding module through the skip connection layer;
[0093] Obtain a second set of feature maps corresponding to the target body part image through the first decoding module, where the second set of feature maps is composed of feature maps of multiple scales;
[0094] Input the second set of feature maps into the first classification and segmentation sub-model to fuse the feature maps included in the second set of feature maps through the first classification and segmentation sub-model, and determine whether there is a first target lesion type and the lesion area corresponding to the first target lesion type in the target body part image based on the fused feature maps.
[0095] For ease of understanding, in combination with Figure 3 to exemplarily illustrate the composition of the first image detection model.
[0096] In Figure 3 , the left branch corresponds to the first feature extraction sub-model, and the right branch corresponds to the first classification and segmentation sub-model. Among them, it is assumed that the first encoding module is composed of 6 convolutional blocks. Correspondingly, the first decoding module is also composed of 6 convolutional blocks. Among them, each convolutional block is represented as: 2xConv3D, which means it is composed of 2 serial three-dimensional (3D) convolutional layers. Among them, the scale corresponding to each convolutional block is different to extract feature maps of corresponding scales, such as 5x8x5, 10x16x10, 20x32x20,... shown in the figure, where C = 320, 256, 128..., and C represents the number of channels. The directed connection lines between the convolutional blocks with the same scale in the first encoding module and the first decoding module shown in the figure represent skip connections between the convolutional blocks of the same scale.
[0097] Based on Figure 3 the structure of the first feature extraction sub-model shown in, when the target body part image is input into the first encoding module, through the convolution and downsampling processing of the above 6 convolutional blocks in sequence, 6 scales of feature maps with gradually decreasing scales can be obtained in sequence, where the output of the previous convolutional block is used as the input of the next convolutional block. The output of the last convolutional block in the first encoding module will be input into the first decoding module, and the first decoding module will perform upsampling and transposed convolution processing layer by layer through 6 convolutional blocks in sequence, and will output 6 scales of feature maps with gradually increasing scales in sequence. Only after a convolutional block in the first decoding module upsamples the feature map of a certain scale i input by the previous convolutional block to obtain a feature map of scale j, it is necessary to splice the feature map of scale j extracted from the first encoding module with this feature map through a skip connection layer before performing subsequent convolution and other operations. That is to say, the role of the skip connection layer is: the splicing of feature maps of the same scale obtained by downsampling (encoding module) and upsampling (decoding module) in the channel dimension.
[0098] The feature maps of multiple scales extracted by the first encoding module are actually shallow features, constituting the first feature map group; while the feature maps of multiple scales extracted by the first decoding module are actually deep features, constituting the second feature map group. Feature fusion is achieved through skip connections.
[0099] As Figure 3 shown, the first classification and segmentation sub-model may include multiple pooling layers (such as the global maximum pooling layer shown in the figure: global pooling) and fully connected layers (FC) connected to each pooling layer, and also includes a feature fusion layer. The feature maps of multiple scales included in the second feature map group are correspondingly input into the multiple pooling layers for compression processing.
[0100] Finally, the feature maps of multiple scales after compression processing are input into the feature fusion layer for feature fusion (concatenation), and based on the fused feature maps, it is determined whether there is a lesion area corresponding to the first target lesion type in the target body part image, such as the recognition results of the three categories of PDAC / non-PDAC / normal shown in the figure, and the lesion area segmentation results corresponding to PDAC / non-PDAC respectively, to enhance interpretability.
[0101] For the second image detection model:
[0102] From a structural perspective, the second image detection model includes a first feature extraction sub-model, a second classification and segmentation sub-model, and a pooling module. Among them, the second classification and segmentation sub-model includes a memory unit and an attention module; among them, the memory unit is trained to store the positions and visual features corresponding to different lesion types included in the first target lesion type in the target body part, and the memory unit is configured to store the positions and visual features with a target number of memory vectors.
[0103] From a working process perspective, the working process is as follows:
[0104] The third feature map group corresponding to the target body part image is extracted through the second feature extraction sub-model, and the third feature map group is composed of feature maps of multiple scales.
[0105] For each target feature map in the feature maps of the multiple scales in sequence: performing pooling processing on the target feature map through a pooling module to compress the target feature map into a target number of feature vectors; and performing cross-attention processing on the target number of reference vectors and the target number of feature vectors through an attention module, performing self-attention processing on the target number of reference vectors, and adding the cross-attention processing result and the self-attention processing result; wherein, when the target feature map is the first one in the feature maps of the multiple scales, the reference vector is the memory vector, and when the target feature map is not the first one in the feature maps of the multiple scales, the reference vector is the addition result of the cross-attention processing result and the self-attention processing result corresponding to the previous target feature map; the target feature map is any one in the third feature map group;
[0106] Determine the second target lesion type and lesion area existing in the target body part image according to the addition result of the cross-attention processing result and the self-attention processing result corresponding to the last target feature map.
[0107] Actually, similar to the first feature extraction sub-model in the first image detection model, the second feature extraction sub-model also includes a second encoding module, a second decoding module, and a skip connection layer between the second encoding module and the second decoding module. Based on this, optionally, extracting the third feature map group corresponding to the target body part image through the second feature extraction sub-model can be implemented as:
[0108] Extracting a fourth feature map group corresponding to the target body part image through the second encoding module;
[0109] Inputting the fourth feature map group into the second decoding module through the skip connection layer;
[0110] Obtaining a fifth feature map group corresponding to the target body part image through the second decoding module;
[0111] Determine that the partial feature maps included in the fourth feature map group and the partial feature maps included in the fifth feature map group constitute the third feature map group.
[0112] In addition, in an optional embodiment, a position embedding module may further be included in the second image detection model. Based on this, performing cross-attention processing on the target number of reference vectors and the target number of feature vectors through the attention module can be implemented as:
[0113] Respectively add the corresponding position embedding vectors to the target number of feature vectors, wherein the position embedding vector added to any one feature vector is used to represent the position information corresponding to the any one feature vector in the target number of feature vectors;
[0114] Cross-attention processing is performed on the feature vectors obtained by superimposing the reference vector of the target quantity and the respective position embedding vectors of the target quantity through the attention module.
[0115] The introduction of the position embedding vector essentially introduces the spatial position characteristics of the disease on the target body part, which helps to achieve a more accurate recognition effect.
[0116] For ease of understanding, in combination with Figure 4 to exemplarily illustrate the composition of the second image detection model.
[0117] In Figure 4 the left branch corresponds to the second feature extraction sub-model, the middle branch corresponds to the position embedding module, and the right branch corresponds to the second classification and segmentation sub-model.
[0118] Among them, similar to Figure 3 assuming that the second encoding module consists of 6 convolutional blocks, correspondingly, the second decoding module also consists of 6 convolutional blocks, and the composition and related parameters of each convolutional block are as shown in the figure. The directed connection lines between the convolutional blocks corresponding to the same scale in the second encoding module and the second decoding module schematically shown in the figure represent the skip connections between the convolutional blocks of the same scale.
[0119] Consistent with the working process of the first feature extraction sub-model introduced above, based on Figure 4 the structure of the second feature extraction sub-model schematically shown in the figure, when the target body part image is input into the second encoding module, through the convolution and downsampling processing of the above 6 convolutional blocks in sequence, 6 feature maps with gradually decreasing scales can be obtained in sequence, constituting the fourth feature map group. The second decoding module performs upsampling and transposed convolution processing layer by layer through 6 convolutional blocks in sequence, and will output 6 feature maps with gradually increasing scales in sequence, constituting the fifth feature map group.
[0120] After that, as Figure 4 shown in the figure, some feature maps are respectively selected from the fourth feature map group and the fifth feature map group to constitute the third feature map group.
[0121] It should be noted that theoretically, the third feature map group can also be composed of all the feature maps included in the fourth feature map group and the fifth feature map group, but this will increase the computational complexity. In addition, since the fourth feature map group and the fifth feature map group correspond to shallow features and deep features respectively, selecting some feature maps from the fourth feature map group and the fifth feature map group respectively can take into account both shallow features and deep features as well as the computational amount.
[0122] After that, for each feature map in the third group of feature maps, each feature map is pooled through a pooling module to compress the feature map into a target number of feature vectors. Among them, the pooling module can provide the adaptive average pooling process (Adaptive Average Pooling) as shown in the figure. In Figure 4 it is assumed that the target number is 5x8x5 = 200, which matches the size of the 200 320-dimensional vectors stored in the memory unit.
[0123] Among them, in Figure 4 the feature maps of multiple scales shown, the smallest scale is 5x8x5, and other scales are integer multiples of this scale. Therefore, feature maps of other scales can all be compressed based on this 5x8x5 scale, and the compression results are represented by 200 feature vectors.
[0124] In an optional embodiment, after each feature map in the above-mentioned third feature map is compressed into the above-mentioned target number of feature vectors, for each of the feature vectors, a corresponding position embedding vector (pos) can be superimposed. Among them, the position embedding vector superimposed on any feature vector is used to represent the position information corresponding to the any feature vector in the target number of feature vectors.
[0125] After that, as Figure 4 shown in it, cross-attention processing is performed on the target number of reference vectors and the feature vectors with their respective corresponding position embedding vectors superimposed through an attention module, and self-attention processing is performed on the target number of reference vectors, and the cross-attention processing result and the self-attention processing result are added.
[0126] Among them, the target number of reference vectors originate from the target number of memory vectors stored in the memory unit. The target number of memory vectors stored in this memory unit are updated and learned during the training process of the second image detection model. When the model converges, the target number of memory vectors stored in the memory unit will be finally stored for later use.
[0127] As Figure 4 shown in it, after each feature map in the third group of feature maps undergoes the above-mentioned compression and position embedding processing, it is equivalent to forming a vector sequence, and each vector sequence will be input into the attention module in turn. Thus, the attention module can be regarded as being composed of multiple attention units shown in the figure, and each attention unit provides the functions of self-attention mechanism and cross-attention mechanism.
[0128] For ease of description and understanding, the feature vectors of the target quantities after pooling the 4 feature maps in the third feature map group shown in the figure are respectively denoted as C1, C2, C3, and C4, the corresponding 4 position embedding vectors are respectively denoted as P1, P2, P3, and P4, and the memory vector of the target quantity stored in the memory unit is denoted as M0. Based on this, the self-attention processing of the memory vector of the target quantity by the first attention unit is expressed as: self-attention(M0) = M 01 , the self-attention processing of the memory vector of the target quantity by the first attention unit and the feature vector C1 of the target quantity with its corresponding position embedding vector added to it is expressed as: cross-attention(C1 + P1, M0) = M 02 , where C1 + P1 represents the superposition of the feature vector C1 of the target quantity and its corresponding position embedding vector.
[0129] Denote the sum of the self-attention processing result and the cross-attention processing result as: M 01 +M 02 = M1, then M1 is the input of the next attention unit: the reference vector of the target quantity.
[0130] The self-attention processing of the reference vector M1 of the target quantity by the second attention unit is expressed as: self-attention(M1) = M 11 , the self-attention processing of the reference vector of the target quantity by the second attention unit and the feature vector C2 of the target quantity with its corresponding position embedding vector added to it is expressed as: cross-attention(C2 + P2, M1) = M 12 , denote the sum of the self-attention processing result and the cross-attention processing result at this time as: M 11 +M 12 = M2.
[0131] And so on, the self-attention processing of the reference vector M2 of the target quantity by the third attention unit is expressed as: self-attention(M2) = M 21 , the self-attention processing of the reference vector of the target quantity by the third attention unit and the feature vector C3 of the target quantity with its corresponding position embedding vector added to it is expressed as: cross-attention(C3 + P3, M2) = M 22 , denote the sum of the self-attention processing result and the cross-attention processing result at this time as: M 21 +M 22 = M3. The self-attention processing of the reference vector M3 of the target quantity by the fourth attention unit is expressed as: self-attention(M3) = M 31, The self-attention processing of the reference vector of the target quantity and the superimposed eigenvector C4 of the position embedding vector corresponding to the target quantity by the fourth attention unit is expressed as: cross-attention(C4 + P4, M3) = M 32 , Denote the sum of the self-attention processing result and the cross-attention processing result at this time as: M 31 +M 32 = M4.
[0132] After that, determine the second target lesion type and lesion area existing in the target body part image according to M4. For example, through Figure 4 The response module shown in performs a certain pooling process on the 200 320-dimensional vectors corresponding to M4, selects the eigenvalue from each vector, inputs it to the subsequent fully connected layer (FC), and then outputs the prediction results of the lesion type and lesion area.
[0133] In this embodiment, the memory unit is used to store a set of globally shared model parameters. In the model training stage, the initial value in the memory unit can be randomly determined. After that, this parameter will be updated every time the model is iteratively trained. The memory unit is designed to learn global context information and position information, such as the relative position of pancreatic tumors in the pancreas, so as to provide distinguishable descriptors for each pancreatic disease type included in the first target lesion type (i.e., the second group of lesion types). That is to say, the memory unit aims to store the location (space) and texture (vision) and other feature information of different pancreatic diseases. The update and construction of this feature information need to be realized through the self-attention mechanism and the cross-attention mechanism.
[0134] The image detection method provided by the embodiments of the present invention can be executed in the cloud. There can be several computing nodes deployed in the cloud, and each computing node has processing resources such as computing and storage. In the cloud, a service can be organized to be provided by multiple computing nodes. Of course, a single computing node can also provide one or more services. The way the cloud provides this service can be to provide a service interface externally, and users call this service interface to use the corresponding service. The service interface includes forms such as a Software Development Kit (SDK) and an Application Programming Interface (API).
[0135] For the solution provided by the embodiments of the present invention, the cloud can provide a service interface for image detection services. Users call this service interface through the user device to trigger an image detection request to the cloud. This request includes the detection image obtained by plain scan computed tomography. The cloud determines the computing node that responds to this request and uses the processing resources in this computing node to execute the following steps:
[0136] Extract a target body part image corresponding to a target body part from the detected image;
[0137] Perform a first image classification and segmentation process on the target body part image through a first image detection model to determine a first target lesion type and a lesion area corresponding to the first target lesion type in the target body part image;
[0138] Perform a second image classification and segmentation process on the target body part image through a second image detection model to determine a second target lesion type and a lesion area existing in the target body part image, where the second target lesion type is a subcategory of the first target lesion type;
[0139] Feed back the detected image marked with the second target lesion type and the lesion area to the user device.
[0140] The above execution process can refer to the relevant descriptions in the foregoing other embodiments and will not be elaborated herein.
[0141] For ease of understanding, in combination with Figure 5 for exemplary illustration. The user can call an image detection service through the user device E1 shown in Figure 5 to upload a service request including a detected image. In the cloud, as shown in the figure, in addition to deploying several computing nodes, a management node E2 running a control service is also deployed. After receiving the service request sent by the user device E1, the management node E2 determines a computing node E3 that responds to the service request. After receiving the service request, the computing node E3 executes the above computing process to obtain a detected image marked with the second target lesion type and the lesion area. After that, the computing node E3 sends the detected image with these marked information to the user device E1, and the user device E1 displays the detected image, and the user can perform further editing and other operations on this basis.
[0142] The following will describe in detail an image detection device according to one or more embodiments of the present invention. Those skilled in the art can understand that these devices can all be configured by using commercially available hardware components through the steps taught by this solution.
[0143] Figure 6 As shown in Figure 6 is a schematic structural diagram of an image detection device provided by an embodiment of the present invention. The device includes: an acquisition module 11, a cutting module 12, a first detection module 13, and a second detection module 14.
[0144] The acquisition module 11 is used to obtain a detected image obtained by plain scan computed tomography.
[0145] The cutting module 12 is configured to extract a target body part image corresponding to the target body part from the detected image.
[0146] The first detection module 13 is configured to perform a first image classification and segmentation process on the target body part image through a first image detection model, so as to determine a first target lesion type and a lesion area corresponding to the first target lesion type in the target body part image.
[0147] The second detection module 14 is configured to perform a second image classification and segmentation process on the target body part image through a second image detection model, so as to determine a second target lesion type and a lesion area existing in the target body part image, and the second target lesion type is a subcategory of the first target lesion type.
[0148] Optionally, the first image detection model is used to detect a first group of lesion types, and the first group of lesion types includes: a third target lesion type sequentially divided according to the disease severity corresponding to the target body part, the first target lesion type, and no lesion, where the first target lesion type is a general term for lesion types other than the third target lesion type.
[0149] Optionally, the first image detection model includes a first feature extraction sub-model and a first classification and segmentation sub-model, and the first feature extraction sub-model includes a first encoding module, a first decoding module, and a skip connection layer between the first encoding module and the first decoding module. Based on this, the first detection module 13 is configured to: extract a first feature map group corresponding to the target body part image through the first encoding module, and the first feature map group is composed of feature maps of multiple scales; input the first feature map group into the first decoding module through the skip connection layer; obtain a second feature map group corresponding to the target body part image through the first decoding module, and the second feature map group is composed of feature maps of multiple scales; input the second feature map group into the first classification and segmentation sub-model, so as to fuse the feature maps included in the second feature map group through the first classification and segmentation sub-model, and determine whether a first target lesion type and a lesion area corresponding to the first target lesion type exist in the target body part image based on the fused feature maps.
[0150] Optionally, the second image detection model includes a second feature extraction sub-model, a second classification and segmentation sub-model, and a pooling module. The second classification and segmentation sub-model includes a memory unit and an attention module. Among them, the memory unit is trained to store the positions and visual features corresponding to different lesion types included in the first target lesion type in the target body part. The memory unit is configured to store the positions and visual features with a target number of memory vectors. Based on this, the second detection module 14 is configured to: extract a third feature map group corresponding to the target body part image through the second feature extraction sub-model, and the third feature map group is composed of feature maps of multiple scales; sequentially for the target feature map among the feature maps of multiple scales: perform pooling processing on the target feature map through the pooling module to compress the target feature map into the target number of feature vectors; and perform cross-attention processing on the target number of reference vectors and the target number of feature vectors through the attention module, perform self-attention processing on the target number of reference vectors, and sum the cross-attention processing result and the self-attention processing result; where when the target feature map is the first one among the feature maps of multiple scales, the reference vector is the memory vector, and when the target feature map is not the first one among the feature maps of multiple scales, the reference vector is the sum result of the cross-attention processing result and the self-attention processing result corresponding to the previous target feature map; the target feature map is any one of the feature maps of multiple scales; determine the second target lesion type and the lesion area existing in the target body part image according to the sum result of the cross-attention processing result and the self-attention processing result corresponding to the last target feature map.
[0151] Optionally, the second image detection model includes a position embedding module. Based on this, the second detection module 14 is further configured to: respectively add corresponding position embedding vectors to the target number of feature vectors, where the position embedding vector added to any one of the feature vectors is used to represent the position information corresponding to the any one of the feature vectors in the target number of feature vectors; perform cross-attention processing on the target number of reference vectors and the target number of feature vectors with their respective corresponding position embedding vectors added through the attention module.
[0152] Optionally, the second feature extraction sub-model includes a second encoding module, a second decoding module, and a skip connection layer between the second encoding module and the second decoding module. Based on this, the second detection module 14 is configured to: extract a fourth feature map group corresponding to the target body part image through the second encoding module; input the fourth feature map group into the second decoding module through the skip connection layer; obtain a fifth feature map group corresponding to the target body part image through the second decoding module; and determine that partial feature maps included in the fourth feature map group and partial feature maps included in the fifth feature map group constitute the third feature map group.
[0153] Optionally, using the first image detection model and the second image detection model as target image detection models respectively, the apparatus further includes: a training module, configured to obtain a training sample set for training the target image detection model; construct a plurality of training sample subsets corresponding to the training sample set; and train a plurality of target image detection models respectively through the plurality of training sample subsets.
[0154] Based on this, the first detection module 13 is specifically configured to: perform first image classification and segmentation processing on the target body part image respectively through a plurality of first image detection models to obtain output results of the plurality of first image detection models respectively; and determine whether there is a first target lesion type and a lesion area corresponding to the first target lesion type in the target body part image according to the output results of the plurality of first image detection models respectively. The second classification module is specifically configured to: perform second image classification and segmentation processing on the target body part image respectively through a plurality of second image detection models to obtain output results of the plurality of second image detection models respectively; and determine a second target lesion type and a lesion area existing in the target body part image according to the output results of the plurality of second image detection models respectively.
[0155] Figure 6 The illustrated apparatus may execute the steps in the foregoing embodiments. For the detailed execution process and technical effects, refer to the descriptions in the foregoing embodiments, which will not be elaborated herein.
[0156] In a possible design, the structure of the foregoing Figure 6 illustrated image detection apparatus may be implemented as an electronic device. As Figure 7 illustrated, the electronic device may include: a processor 21, a memory 22, and a communication interface 23. Wherein, executable code is stored on the memory 22, and when the executable code is executed by the processor 21, the processor 21 can at least implement the image detection method provided in the foregoing embodiments.
[0157] In addition, an embodiment of the present invention provides a non-transitory machine-readable storage medium, on which executable code is stored. When the executable code is executed by a processor of an electronic device, the processor can at least implement the image detection method provided in the foregoing embodiments.
[0158] In an alternative embodiment, the electronic device for executing the image detection method provided in the embodiment of the present invention may be an Extended Reality (XR) device. XR is a general term for various forms such as virtual reality and augmented reality. At this time, the detection image obtained by plain CT can be input into the extended reality device. The first image detection model and the second image detection model are set in the extended reality device, and the two models perform the above-mentioned image classification and segmentation processing on the target body part image extracted from the detection image, and finally determine the second target lesion type and lesion area existing in the target body part image, and display the detection image marked with the second target lesion type and lesion area. In this way, doctors and users can both see these marked information through the extended reality device.
[0159] In practical applications, for the convenience of viewing, optionally, after receiving the initially collected detection image or the detection image with the above-mentioned marked information, the extended reality device can generate a virtual environment for more clearly viewing the detection image, and render and display the detection image in the virtual environment. In addition, doctors and users can also input interaction operations to the extended reality device during viewing, such as rotating the detection image, magnifying the detection image, etc. The device embodiments described above are merely illustrative, and the units described as separate components may or may not be physically separated. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0160] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of adding a necessary general hardware platform, and of course, it can also be implemented by a combination of hardware and software. Based on such an understanding, the essence of the above technical solution or the part that contributes to the prior art can be embodied in the form of a computer product. The present invention can be implemented in the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0161] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An image detection method, characterized in that, comprising: obtaining a detection image obtained by plain scan computed tomography; extracting a target body part image corresponding to a target body part from the detection image; performing a first image classification and segmentation process on the target body part image through a first image detection model to determine a first target lesion type and a lesion area corresponding to the first target lesion type in the target body part image; extracting a third feature map group corresponding to the target body part image through a second image detection model, and determining a target number of feature vectors corresponding to a target feature map, where the target feature map is any one of the feature maps of multiple scales included in the third feature map group; performing cross-attention processing on the feature vectors and the target number of reference vectors, and performing self-attention processing on the reference vectors; wherein, when the target feature map is the first feature map in the third feature map group, the reference vector is a memory vector pre-storing the positions and visual features corresponding to different lesion types included in the first target lesion type in the target body part; when the target feature map is not the first feature map, the reference vector is the sum result of the cross-attention processing result and the self-attention processing result corresponding to the previous target feature map; determining a second target lesion type and a lesion area existing in the target body part image according to the sum result of the cross-attention processing result and the self-attention processing result corresponding to the last target feature map in the third feature map group, where the second target lesion type is a subcategory of the first target lesion type.
2. The method according to claim 1, characterized in that, the first image detection model is used for detecting a first group of lesion types, and the first group of lesion types includes: a third target lesion type, the first target lesion type, and no lesion, which are sequentially divided according to the disease severity corresponding to the target body part, where the first target lesion type is a general term for lesion types other than the third target lesion type.
3. The method according to claim 1, characterized in that, the first image detection model includes a first feature extraction sub-model and a first classification and segmentation sub-model, and the first feature extraction sub-model includes a first encoding module, a first decoding module, and a skip connection layer between the first encoding module and the first decoding module; the performing a first image classification and segmentation process on the target body part image through the first image detection model includes: extracting a first feature map group corresponding to the target body part image through the first encoding module, and the first feature map group is composed of feature maps of multiple scales; inputting the first feature map group into the first decoding module through the skip connection layer; obtaining a second feature map group corresponding to the target body part image through the first decoding module, and the second feature map group is composed of feature maps of multiple scales; Input the second feature map group into the first classification and segmentation sub-model, so as to fuse the feature maps included in the second feature map group through the first classification and segmentation sub-model, and determine whether there is a first target lesion type and the lesion area corresponding to the first target lesion type in the target body part image based on the fused feature maps.
4. The method according to claim 1, wherein, the second image detection model includes a second feature extraction sub-model, a second classification and segmentation sub-model and a pooling module, and the second classification and segmentation sub-model includes a memory unit and an attention module; wherein, the memory unit is trained to store the target number of memory vectors; The extracting the third feature map group corresponding to the target body part image through the second image detection model and determining the target number of feature vectors corresponding to the target feature map includes: extracting the third feature map group corresponding to the target body part image through the second feature extraction sub-model; performing pooling processing on the target feature map through the pooling module to compress the target feature map into the target number of feature vectors; The performing cross-attention processing on the feature vectors and the target number of reference vectors and performing self-attention processing on the reference vectors includes: performing cross-attention processing on the target number of reference vectors and the target number of feature vectors through the attention module, performing self-attention processing on the target number of reference vectors, and adding the cross-attention processing result and the self-attention processing result.
5. The method according to claim 4, wherein, the second image detection model includes a position embedding module; the performing cross-attention processing on the target number of reference vectors and the target number of feature vectors through the attention module includes: respectively adding the corresponding position embedding vectors to the target number of feature vectors, wherein the position embedding vector added to any one of the feature vectors is used to represent the position information corresponding to the any one of the feature vectors in the target number of feature vectors; performing cross-attention processing on the target number of reference vectors and the target number of feature vectors added with their respective corresponding position embedding vectors through the attention module.
6. The method according to claim 4, wherein, the second feature extraction sub-model includes a second encoding module, a second decoding module and a skip connection layer between the second encoding module and the second decoding module; The extracting the third feature map group corresponding to the target body part image through the second feature extraction sub-model includes: extracting the fourth feature map group corresponding to the target body part image through the second encoding module; inputting the fourth feature map group into the second decoding module through the skip connection layer; obtaining the fifth feature map group corresponding to the target body part image through the second decoding module; determining that a partial feature map included in the fourth feature map group and a partial feature map included in the fifth feature map group constitute the third feature map group.
7. The method according to claim 1, It is characterized in that using the first image detection model and the second image detection model as the target image detection models respectively, the method further includes: obtaining a training sample set for training the target image detection model; constructing a plurality of training sample subsets corresponding to the training sample set; training a plurality of target image detection models respectively through the plurality of training sample subsets.
8. The method according to claim 7, it is characterized in that the first image classification and segmentation processing of the target body part image by the first image detection model to determine a lesion area corresponding to the first target lesion type in the target body part image includes: performing first image classification and segmentation processing on the target body part image respectively through a plurality of first image detection models to obtain output results of the plurality of first image detection models respectively; determining whether there is a first target lesion type and a lesion area corresponding to the first target lesion type in the target body part image according to the output results of the plurality of first image detection models respectively; the determination of the second target lesion type and the lesion area existing in the target body part image includes: outputting the second target lesion type and the lesion area existing in the target body part image respectively through a plurality of second image detection models; determining the second target lesion type and the lesion area existing in the target body part image according to the output results of the plurality of second image detection models respectively.
9. An image detection device, it is characterized in that including: a collection module for obtaining a detection image obtained by plain scan computed tomography; a cutting module for extracting a target body part image corresponding to a target body part from the detection image; a first detection module for performing first image classification and segmentation processing on the target body part image by a first image detection model to determine a first target lesion type and a lesion area corresponding to the first target lesion type in the target body part image; A second detection module, configured to extract a third feature map group corresponding to the target body part image through a second image detection model, and determine a target number of feature vectors corresponding to a target feature map, where the target feature map is any one of feature maps of multiple scales included in the third feature map group; perform cross-attention processing on the feature vectors and the target number of reference vectors, and perform self-attention processing on the reference vectors; wherein, when the target feature map is the first feature map in the third feature map group, the reference vector is a memory vector that pre-stores the positions and visual features corresponding to different lesion types included in the first target lesion type in the target body part; when the target feature map is not the first feature map, the reference vector is the sum result of the cross-attention processing result and the self-attention processing result corresponding to the previous target feature map; determine the second target lesion type and the lesion area existing in the target body part image according to the sum result of the cross-attention processing result and the self-attention processing result corresponding to the last target feature map in the third feature map group, and the second target lesion type is a sub-category of the first target lesion type.
10. An electronic device, characterized in that it includes: a memory, a processor, and a communication interface; wherein, an executable code is stored on the memory, and when the executable code is executed by the processor, the processor is caused to execute the image detection method according to any one of claims 1 to 8.
11. A non-transitory machine-readable storage medium, characterized in that an executable code is stored on the non-transitory machine-readable storage medium, and when the executable code is executed by a processor of an electronic device, the processor is caused to execute the image detection method according to any one of claims 1 to 8.
12. An image detection method, characterized in that it includes: receiving a request triggered by a user device by invoking an image detection service, where the request includes a detection image obtained by plain scan computed tomography; using processing resources corresponding to the image detection service to perform the following steps: extracting a target body part image corresponding to a target body part from the detection image; performing first image classification and segmentation processing on the target body part image through a first image detection model to determine a first target lesion type and a lesion area corresponding to the first target lesion type existing in the target body part image; extracting a third feature map group corresponding to the target body part image through a second image detection model, and determining a target number of feature vectors corresponding to a target feature map, where the target feature map is any one of feature maps of multiple scales included in the third feature map group; Perform cross-attention processing on the feature vector and the target number of reference vectors, and perform self-attention processing on the reference vectors. Wherein, when the target feature map is the first feature map in the third feature map group, the reference vector is a memory vector that pre-stores the positions and visual features corresponding to different lesion types included in the first target lesion type in the target body part; when the target feature map is not the first feature map, the reference vector is the sum result of the cross-attention processing result and the self-attention processing result corresponding to the previous target feature map; Determine the second target lesion type and the lesion area existing in the target body part image according to the sum result of the cross-attention processing result and the self-attention processing result corresponding to the last target feature map in the third feature map group, and the second target lesion type is a sub-category of the first target lesion type; Feed back the detection image marked with the second target lesion type and the lesion area to the user device.
13. An image detection method, Characterized in that, Applied to an extended reality device, including: Obtain a detection image obtained by plain scan computed tomography; Extract a target body part image corresponding to the target body part from the detection image; Perform first image classification and segmentation processing on the target body part image through a first image detection model to determine the first target lesion type and the lesion area corresponding to the first target lesion type existing in the target body part image; Extract a third feature map group corresponding to the target body part image through a second image detection model, and determine the target number of feature vectors corresponding to the target feature map, where the target feature map is any one of the feature maps of multiple scales included in the third feature map group; Perform cross-attention processing on the feature vector and the target number of reference vectors, and perform self-attention processing on the reference vectors. Wherein, when the target feature map is the first feature map in the third feature map group, the reference vector is a memory vector that pre-stores the positions and visual features corresponding to different lesion types included in the first target lesion type in the target body part; when the target feature map is not the first feature map, the reference vector is the sum result of the cross-attention processing result and the self-attention processing result corresponding to the previous target feature map; Determine the second target lesion type and the lesion area existing in the target body part image according to the sum result of the cross-attention processing result and the self-attention processing result corresponding to the last target feature map in the third feature map group, and the second target lesion type is a sub-category of the first target lesion type; Display the detection image marked with the second target lesion type and the lesion area.
Citation Information
Patent Citations
Lung focus detection method and device and training method of image detection model
CN111325739A
Thyroid image processing method and device, electronic equipment and storage medium
CN113689412A