Image processing-based anomaly detection method and device, electronic equipment and medium
By acquiring medical images of target blood vessels and bones, extracting image feature data using a deep learning semantic segmentation model, and combining anatomical image information for feature extraction, the accuracy problem of blood vessel abnormality detection is solved, and more efficient abnormality localization is achieved.
Patent Information
- Application Number
- CN202311174630.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-12
- Publication Date
- 2025-12-16
- Estimated Expiration
- 2043-09-12
AI Technical Summary
The accuracy of vascular abnormality detection in existing technologies is not high, resulting in poor detection results, especially in cases of vascular variations where it is difficult to accurately locate the occlusion.
By acquiring medical images of the target blood vessels and target bones, image feature data is extracted using a deep learning semantic segmentation model. This is combined with anatomical image information for feature extraction, and abnormal information of the target blood vessels is detected. An anomaly detection is performed using a multi-task UNet network structure.
It improves the accuracy and effectiveness of vascular abnormality detection, enabling more precise location of abnormal vascular positions, especially in cases of vascular variations, thus reducing false positives and false negatives.
Smart Images

Figure CN117237291B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of image processing, deep learning, and the like, and in particular to an abnormality detection method and device based on image processing, an electronic device, and a storage medium. BACKGROUND
[0002] Blood vessels are an important part of the human body and play a crucial role in human health. Therefore, detecting abnormalities and pathological changes in blood vessels is of great significance to human health. However, the accuracy of blood vessel abnormality detection technology in the related art is not high, and the abnormality detection effect is not good. SUMMARY
[0003] The present application aims to at least partially solve one of the technical problems in the related art. To this end, the present application aims to provide an abnormality detection method and device based on image processing, an electronic device, a storage medium, and a program product.
[0004] The present application provides an abnormality detection method based on image processing, comprising: acquiring a first medical image for a target blood vessel and a target bone and an anatomical image for the target bone; performing feature extraction on the first medical image and the anatomical image to obtain image feature data, wherein the image feature data is used to represent the features of the target blood vessel and the relative position relationship between the target blood vessel and the target bone; and detecting abnormal information of the target blood vessel based on the image feature data.
[0005] Another embodiment of the present application provides an abnormality detection device based on image processing, comprising: an acquisition module, an extraction module, and a detection module. The acquisition module is configured to acquire a first medical image for a target blood vessel and a target bone and an anatomical image for the target bone; the extraction module is configured to perform feature extraction on the first medical image and the anatomical image to obtain image feature data, wherein the image feature data is used to represent the features of the target blood vessel and the relative position relationship between the target blood vessel and the target bone; and the detection module is configured to detect abnormal information of the target blood vessel based on the image feature data.
[0006] Another embodiment of the present application provides an electronic device comprising a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the method according to any one of the above embodiments.
[0007] Another embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program is executed by a processor to implement the steps of the method according to any one of the above embodiments.
[0008] Another embodiment of the present application provides a computer program product comprising instructions, which, when executed by a processor of a computer device, enable the computer device to perform the steps of the method according to any one of the embodiments described above.
[0009] In the above embodiments, by acquiring the first medical image for the target blood vessel and the target bone and the anatomical image for the target bone, image feature data is obtained by performing feature extraction on the first medical image and the anatomical image, and abnormal information of the target blood vessel is detected based on the image feature data, thereby improving the accuracy of abnormal detection. BRIEF DESCRIPTION OF DRAWINGS
[0010] Figure 1 A flowchart of an abnormality detection method based on image processing provided by an embodiment of the present application is shown in the figure;
[0011] Figure 2 A schematic diagram of image preprocessing provided by an embodiment of the present application is shown in the figure;
[0012] Figure 3 A schematic diagram of image segmentation provided by an embodiment of the present application is shown in the figure;
[0013] Figure 4 A schematic diagram of detecting abnormalities using an abnormality detection model provided by an embodiment of the present application is shown in the figure;
[0014] Figure 5 A schematic diagram of an abnormality detection device based on image processing provided by an embodiment of the present application is shown in the figure. DETAILED DESCRIPTION
[0015] Embodiments of the present application are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals represent the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by reference to the accompanying drawings are exemplary and are intended to explain the present application, and cannot be understood as limiting the present application.
[0016] Vessels are important components of the human body and play a crucial role in the health of the human body, so detecting abnormalities and pathological changes of vessels is of great significance to the health of the human body. However, the accuracy of the vessel abnormality detection technology in the related art is not high, and the effect of abnormality detection is not good.
[0017] Vessels include, for example, vertebral arteries. The vertebral artery is an important part of the initial stage of the posterior circulation system and supplies blood to an important area of the posterior circulation of the brain. Developmental abnormalities and pathological changes of the vertebral artery can cause posterior circulation ischemia, causing symptoms such as dizziness, and in severe cases, posterior circulation ischemia can also cause cerebral infarction. Therefore, early detection of abnormalities and pathological changes of the vertebral artery has important clinical significance for preventing cardiovascular and cerebrovascular diseases.
[0018] Human cervical spine includes C1 ~ C7 seven blocks, vertebral artery originates from the subclavian artery, divided into left and right two branches, each branch passes through the sixth cervical vertebra C6 to atlas C1, through the foramen magnum into the cranial fossa, at the pontomedullary junction into a basilar artery. Vertebral artery is divided into four segments, the first segment V1 from the subclavian starting point to the sixth cervical vertebra C6 transverse foramen; the second segment V2 passes through (penetrates) the sixth cervical vertebra C6 to the second cervical vertebra C2 transverse foramen; the third segment V3 from the first cervical vertebra C1 transverse foramen to the atlantooccipital membrane; the fourth segment V4 is the intracranial segment, through the atlantooccipital membrane into the skull and merges into the basilar artery. Accurate identification of the position of each segment of the vertebral artery can quickly locate the abnormality.
[0019] As the starting and important part of the posterior circulation, the shape and position of the vertebral artery have particularity, and are prone to congenital deficiency and acquired lesions. Common lesions include developmental failure (thin), abnormal transverse foramen, abnormal origin, etc.
[0020] Most people's vertebral arteries pass through the transverse foramen of the sixth cervical vertebra C6, run between the transverse foramen of each cervical vertebra, and a small number of people have variations that may pass through the transverse foramen of the seventh cervical vertebra C7, the fifth cervical vertebra C5, the fourth cervical vertebra C4, and the third cervical vertebra C3. The abnormal transverse foramen of the vertebral artery belongs to congenital developmental deficiency, and has no obvious clinical signs and symptoms, but when the anterior circulation is affected by lesions affecting blood flow, the abnormal vertebral artery cannot play a compensatory role and cause basilar artery ischemia, which may cause serious consequences. The diameter of the vertebral artery lumen is small, on the one hand due to congenital developmental thinness, and on the other hand due to atherosclerosis and small intima tube diameter. Patients with vertebral artery hypoplasia need to bear greater blood flow pressure and compensation, which is due to their vertebral artery is usually too narrow, and vertebral artery stenosis and occlusion is the main cause of posterior circulation ischemic stroke, so the thinness and stenosis of the vertebral artery are of great significance to the clinic.
[0021] In an example, the occlusion state of a conventional blood vessel needs to be detected manually, which is limited by professional degree and measurement error, and is prone to missed detection and inaccurate detection.
[0022] In an example, the intracranial CTA medical image can be acquired by a Computed Tomography angiography (CTA) technology, and the occlusion state of the blood vessel can be detected based on the intracranial CTA medical image. For example, first, a segmentation model is used to segment the blood vessel to obtain a blood vessel mask image, then a segmentation model is used to segment the segmented blood vessel mask image, and finally, the occlusion state is detected by intercepting the blood vessel mask image segment of interest. However, this method directly segments the blood vessel mask image for detection, which is prone to errors when there are blood vessel variations, resulting in deviation in the positioning of the occlusion position.
[0023] In an example, whether the blood vessel structure or morphology is abnormal can be detected based on the blood vessel image. For example, first, the blood vessel in the image is segmented to obtain each blood vessel segment in the blood vessel, and the blood vessel segment is input into a classification model for abnormality detection. However, this method only inputs the mask image of the blood vessel when extracting features and extracts features from the blood vessel mask image, which has strict requirements for the generation of the blood vessel mask image. When the blood vessel mask image extraction effect is poor, the abnormality detection is prone to errors.
[0024] Therefore, the embodiments of the present application provide an optimized image processing-based abnormality detection method.
[0025] Figure 1 A flowchart of the image processing-based abnormality detection method provided by the embodiments of the present application is shown.
[0026] As shown in Figure 1 The image processing-based abnormality detection method 100 provided by the embodiments of the present application includes steps S110-S130.
[0027] In step S110, a first medical image for a target blood vessel and a target bone and an anatomical image for the target bone are acquired.
[0028] Exemplarily, the first medical image includes the target blood vessel and the target bone, and the first medical image can be a three-dimensional image. Specifically, the first medical image can be a CTA medical image acquired by a Computed Tomography angiography (CTA) technology. The CTA technology is a widely used, minimally invasive, and high-yield blood vessel disease diagnosis technology, which is suitable for targeted diagnosis of important information of cardiovascular and cerebrovascular diseases. When the CTA medical image is collected, a contrast agent that makes the blood vessel visible is injected into the vein of the subject, and when the contrast agent passes through the detection position, a three-dimensional tomographic scan of a certain thickness is performed on the detection position to obtain the CTA medical image.
[0029] In an example, the target blood vessel includes a vertebral artery, and the target bone includes a cervical vertebra. Specifically, the cervical vertebra includes seven pieces of C1-C7, and the target bone can be any one of the cervical vertebra C1-C7, for example, the sixth cervical vertebra C6. Of course, in the example, the target blood vessel can be other types of blood vessels in addition to the vertebral artery, and the target bone can be other types of bones in addition to the cervical vertebra, as long as there is a specific relative positional relationship between the target blood vessel and the target bone so that the abnormality of the target blood vessel can be determined with the target bone as a reference.
[0030] In step S120, feature extraction is performed on the first medical image and the anatomical image to obtain image feature data.
[0031] Exemplarily, the image feature data is used to represent the features of the target blood vessel and the relative positional relationship between the target blood vessel and the target bone. The features of the target blood vessel, for example, are the morphology of the target blood vessel, which represents whether the morphology of the target blood vessel is abnormal, thin, or narrow, and the like. The relative positional relationship between the target blood vessel and the target bone, for example, includes whether the target blood vessel passes through the target bone, where the target blood vessel enters the target bone, where the target blood vessel exits the target bone, and the like.
[0032] In step S130, based on the image feature data, abnormal information of the target blood vessel is detected.
[0033] After the image feature data is extracted, the abnormal information of the target blood vessel can be determined based on the image feature data. The abnormal information, for example, includes whether the target blood vessel is abnormal, the type of abnormality of the target blood vessel, and the like. For example, the abnormal information can specifically include that the target blood vessel passes through the target bone, the target blood vessel does not pass through the target bone, the target blood vessel is narrow, the target blood vessel is infarcted, the morphology of the target blood vessel is normal, and the like.
[0034] It can be understood that the embodiments of the present application obtain the image feature data representing the internal relationship between the target blood vessel and the target bone by performing feature extraction on the first medical image and the anatomical image, and determine the abnormality of the target blood vessel based on the image feature data, which realizes the detection of the abnormality of the target blood vessel with the anatomical information of the target bone as a reference, so that the accuracy of the abnormality detection is higher, and the effect of the abnormality detection is better.
[0035] In another example, before step S110 is performed, an anatomical image for the target bone needs to be obtained. Next, how to obtain the anatomical image for the target bone will be described.
[0036] For example, at least one of the first medical image and the second medical image is processed to obtain an anatomical image of the target bone. The first medical image is a CTA medical image, and the second medical image is, for example, a CT (Computed Tomography) plain scan image. Unlike the CTA medical image, the CT plain scan image does not require any contrast agent when obtained, and can be directly scanned by a CT device to obtain the CT plain scan image. The CT plain scan image can be used for disease screening, and the image scanning range needs to cover the cervical spine of the human body.
[0037] The anatomical image of the target bone is obtained without injecting contrast agent into the vein, so in addition to being able to obtain the anatomical image based on the first medical image, the anatomical image can also be obtained based on the second medical image, and of course the anatomical image can also be obtained based on the first medical image and the second medical image.
[0038] Figure 2 The schematic diagram of image preprocessing provided for the embodiments of the present application.
[0039] After scanning the image by CTA or CT technology, the scanned image can be positioned based on the key anatomical position of the cervical spine to obtain a cervical bone region image, and the cervical bone region image is taken as the first medical image or the second medical image.
[0040] Before obtaining the anatomical image of the target bone based on the first medical image and the second medical image, the first medical image and the second medical image can also be preprocessed, and subsequent analysis can be performed based on the preprocessed first medical image and the second medical image. The preprocessing includes converting the format of the image, adjusting the window width and window level based on the window technology (the window technology is to select the CT value range of interest by using the window level and the window width), data resampling, normalization, image cropping and the like. The preprocessed first medical image or second medical image is as shown in Figure 2 .
[0041] Figure 3 The schematic diagram of image segmentation provided for the embodiments of the present application.
[0042] As Figure 3The anatomical images of the target bones can be obtained by the trained semantic segmentation model. For example, at least one of the preprocessed first medical image and the second medical image is input into the trained semantic segmentation model. The semantic segmentation model is used to extract features of at least one of the first medical image and the second medical image, to obtain position information 301 of a plurality of bones in the image and category information of each bone. When the bones include cervical vertebrae, the plurality of bones include seven cervical vertebrae (first cervical vertebra to seventh cervical vertebra), the position information 301 includes coordinates of the center of each bone in the image, and the category information of the bone represents the category of the corresponding bone, with the category C1 representing the first cervical vertebra, the category C2 representing the second cervical vertebra, the category C3 representing the third cervical vertebra, and so on. Of course, if the first medical image and the second medical image also include information of other types of bones (for example, spinal column), the semantic segmentation model can segment to obtain position information of other types of bones in addition to the position information of the cervical vertebrae in the image.
[0043] After obtaining the position information 301 of each bone, at least one of the first medical image and the second medical image can be cropped based on the position information 301 to obtain a plurality of anatomical images corresponding to the plurality of bones one by one. Each anatomical image is a partial region of the first medical image or the second medical image, and each anatomical image includes at least the corresponding bone (cervical vertebra). The first medical image and the second medical image can be two-dimensional images or three-dimensional images, and can be three-dimensional head and neck images. Taking one bone as an example, after obtaining the center coordinates of the bone, the center coordinates are taken as the starting point to extend in each dimension and the extension length does not exceed other adjacent bones, and after the extension, a partial region containing the bone but not containing other bones is cropped to obtain an anatomical image of the bone. The plurality of anatomical images can be distinguished by using the category information of the bone as a label.
[0044] After obtaining the anatomical image of each bone, for a target bone in the plurality of bones, an anatomical image for the target bone can be determined from the plurality of anatomical images based on the category information.
[0045] In an example, the semantic segmentation model includes at least a first encoding network, a first decoding network, and a first feature concatenation network.
[0046] The first encoding network includes a plurality of first network levels connected in sequence, the number of channels of the feature maps extracted by the plurality of first network levels increases gradually, and the size of the feature maps extracted by the plurality of first network levels decreases gradually.
[0047] The first decoding network is connected with the first encoding network, the first decoding network includes a plurality of second network levels connected in sequence, the number of channels of the feature maps extracted by the plurality of second network levels decreases gradually, and the size of the feature maps extracted by the plurality of second network levels increases gradually.
[0048] The first feature splicing network is connected with the first encoding network and the first decoding network, and is configured to splice the feature map output by the first encoding network and the same channel feature map output by the first decoding network.
[0049] In an example, the semantic segmentation model is a deep learning model, which can be a U-shaped residual network structure, for example, and can be a Res-UNet network. The first encoding network includes an Encoder end, and the first decoding network includes a Decoder end.
[0050] Taking an example in which the Encoder end includes four first network levels A1, A2, A3, and A4, the four first network levels A1, A2, A3, and A4 are connected in sequence. Each first network level can be composed of a feature extraction layer and a down-sampling network, the feature extraction layer can include a convolutional neural network layer (CNN), and the down-sampling network can include a max pooling layer (Max Pool). Each of the four first network levels can generate a feature map with a different number of channels. The Encoder stage is a feature extraction stage, in which the number of channels increases step by step, the size of the feature map decreases step by step, and the feature strength of the obtained bone (cervical vertebra) increases step by step.
[0051] Taking an example in which the Decoder end includes four second network levels B1, B2, B3, and B4, the four second network levels B1, B2, B3, and B4 are connected in sequence. Each second network level can be composed of a feature extraction layer and an up-sampling network, and the feature extraction layer can include a convolutional neural network layer (CNN). Each of the four second network levels can generate a feature map with a different number of channels, in which the number of channels decreases step by step, the size of the feature map increases step by step, and the feature strength of the obtained bone (cervical vertebra) increases step by step. Through the Decoder stage, the image information is gradually recovered, so that the number of channels decreases and the size of the feature map increases after each level.
[0052] The output of the Encoder end is channel spliced with the output of the Decoder end to realize the fusion of shallow features and deep features through channel splicing, providing more semantic information for the Decoder stage. For example, the number of channels of the feature map output by the A2 level of the Encoder end is consistent with the number of channels of the feature map output by the B3 level of the Decoder end, and the feature map output by the A2 level is spliced with the feature map output by the B3 level and input into the B4 level. For example, the number of channels of the feature map output by the A2 level is 2, and the number of channels of the feature map output by the B3 level is 2, so that the feature map output by the A2 level is spliced with the feature map output by the B3 level to obtain a 4-channel feature map, and the 4-channel feature map is input into the B4 level. The number of channels of the feature map output by the A3 level of the Encoder end is consistent with the number of channels of the feature map output by the B2 level of the Decoder end, and the feature map output by the A3 level is spliced with the feature map output by the B2 level and input into the B3 level, and so on.
[0053] In addition, the first network level and / or the second network level includes a first residual network. For example, the first network level includes a first residual network, and each of the plurality of first network levels can include a first residual network. Or the second network level includes a first residual network, and each of the plurality of second network levels can include a first residual network. Or the first network level and the second network level both include a first residual network, and each of the plurality of first network levels and each of the plurality of second network levels can include a first residual network. It can be understood that the first residual network included in each first network level can be a different residual network, and the first residual network included in each second network level can be a different residual network.
[0054] The first residual network includes a skip connection structure for transmitting feature information between different first network levels and / or for transmitting feature information between different second network levels.
[0055] For example, for the case that the first network level includes a first residual network, the first residual network is used to transmit shallow feature information of a low-level first network level to deep feature information of a high-level first network level, realizing the interaction of information. For example, the first network levels A1 and A2 both include a first residual network, and the first residual network transmits shallow feature information of the first network level A1 to deep feature information of the first network level A2.
[0056] Similarly, for the case that the second network level includes the first residual network, the first residual network is used to pass the shallow feature information of the second network level at a low level to the deep feature information of the second network level at a high level, to realize the interaction of information. For example, the second network levels B1 and B2 both include the first residual network, and the first residual network passes the shallow feature information of the second network level B1 to the deep feature information of the second network level B2.
[0057] Similarly, for the case that the first network level and the second network level both include the first residual network, for example, the first residual network passes the shallow feature information of the first network level A1 to the deep feature information of the first network level A2, and the first residual network passes the shallow feature information of the second network level B1 to the deep feature information of the second network level B2.
[0058] The semantic segmentation model can solve the problems of gradient explosion and gradient disappearance in the model training process by increasing the first residual network, so as to more effectively extract high-level semantic features of the region of interest using a shallower network structure, and enhance the feature expression capability.
[0059] The deep learning semantic segmentation model is used to segment the cervical vertebrae C1-C7 of a three-dimensional head and neck image (CAT medical image or CT plain scan image), to determine the center point information of the cervical vertebrae C1-C7, which contains the spatial position coordinates of the center point and the category to which the center point belongs. This process is divided into a training phase and a prediction phase. The training phase can use supervised learning, and the gold standard (label) can be annotated by experienced doctors. When training the model, the first medical image (CTA medical image) and / or the second medical image (CT plain scan image) are input as training samples, and the loss function is iteratively optimized until the network is optimal, and the optimal network model parameters are saved. In order to improve the robustness of the network model and increase the diversity of the samples, random enhancement processing can be performed on the training samples, such as adding random noise, random rotation, random flipping, random mirroring, random cropping, etc. to the samples, so as to expand the number of samples and improve the robustness of the network model. In the prediction process, the first medical image (CTA medical image) and / or the second medical image (CT plain scan image) to be segmented are input, and the trained model parameters are loaded, and the segmentation result of the cervical vertebrae is output.
[0060] In an example, the loss function value of the semantic segmentation model is a weighted value between the cross entropy value and the set similarity measure value. The semantic segmentation model is a model suitable for multi-class segmentation tasks, and its loss function can use the Cross Entropy Loss loss function and the Dice Loss loss function for set similarity measure, and be combined with different weights.
[0061] For example, the cross-entropy loss function is shown in formula (1):
[0062]
[0063] where C is the number of classes, [y i ,L,y C-1 ] is the onehot encoding of the sample label, p i represents the probability that the sample belongs to the i-th class, and the i-th class represents any one of the cervical vertebrae C1-C7.
[0064] The Dice Loss loss function is defined based on the Dice coefficient and can be used to evaluate the similarity between the predicted value and the true value. The Dice Loss loss function is shown in formula (2):
[0065]
[0066] where |X∩Y| represents the intersection between X and Y; |X| and |Y| represent the true label and the prediction result in the segmentation process, respectively.
[0067] The total loss function is obtained based on the cross-entropy loss function and the Dice Loss loss function, and the total loss function is shown in formula (3):
[0068] Loss 总 =α*Loss ce +(1-α)*Loss dice (0<α<1) (3)
[0069] When training the semantic segmentation model, as the loss function is iterated continuously, the model parameters become more and more stable, and the role of the cross-entropy loss function may become smaller and smaller as the training goes on, which may cause the model to fall into a local optimum. Therefore, the embodiment of the present application can improve the accuracy of model training by combining the dice Loss loss function for model training.
[0070] In addition to obtaining the anatomical image of the target bone through the semantic segmentation model described above, the embodiment of the present application also proposes another image processing method to obtain the anatomical image of the target bone.
[0071] For example, for at least one of the first medical image and the second medical image, an edge pixel point can be determined from the image based on a preset pixel difference value, in view of the fact that there is a certain difference between the pixel points in the region where the bone is located and the pixel points outside the region. For example, if the pixel difference between a certain pixel point and its surrounding pixel points is greater than the preset pixel difference value, the pixel point is determined as an edge pixel point, thereby determining a plurality of edge pixel points. The image region surrounded by the plurality of edge pixel points is a plurality of initial image regions. In addition to including the region where the bone is located, the plurality of initial image regions can also include regions where other human tissues are located, and therefore the plurality of initial image regions need to be further screened in order to screen a plurality of target image regions corresponding to the plurality of bones from the plurality of initial image regions. For example, if the plurality of bones includes 7 bones, the screened target image regions are also 7 image regions.
[0072] After the plurality of target image regions are screened, screening can be performed based on at least one of the following four conditions: a size condition of the region, a shape condition of the region, a distance condition between adjacent regions, and a relative position relationship condition between regions. Taking screening based on the four conditions as an example, the screened target image regions satisfy the following conditions: the size of each target image region is within a preset region size range, the preset region size range representing the size of the bone; the shape of each target image region is a preset shape, the preset shape representing the shape of the bone; the distance between two adjacent target image regions is less than a preset distance, the preset distance representing the distance between two adjacent bones; and the relative position relationship between the plurality of target image regions is a preset relative position relationship, the preset relative position relationship representing the relative position between the plurality of bones, for example, the relative position relationship between the plurality of target image regions is that they are arranged in sequence along a curve with a certain curvature.
[0073] After the plurality of target image regions corresponding to the plurality of bones are screened, in an example, the plurality of target image regions can be directly used as the anatomical images for the plurality of bones, or the plurality of target image regions can be respectively image-extended, and the plurality of larger image regions obtained after the image extension can be used as the anatomical images for the plurality of bones. Then, based on the relative position relationship between the plurality of anatomical images, an anatomical image for a target bone can be determined from the plurality of anatomical images. For example, the relative position relationship can include that the plurality of anatomical images are arranged in sequence from top to bottom, and the plurality of anatomical images correspond to the plurality of bones C1-C7 in sequence from top to bottom. Based on the relative position relationship between the plurality of anatomical images, an anatomical image for a target bone can be determined from the plurality of anatomical images.
[0074] In another example, after obtaining the plurality of target image regions, the region centers of the plurality of target image regions can be determined as the position information of the plurality of bones in the image respectively. Then, based on the relative positional relationship between the position information of the plurality of bones in the image, the category information C1-C7 of each bone is determined, for example, the relative positional relationship can indicate that the positions of the plurality of bones in the image are arranged from top to bottom in sequence, and the category information of the bones corresponding to the region centers of the plurality of target image regions from top to bottom in sequence is C1-C7. After obtaining the position information of the plurality of bones in the image and the category information of each bone, at least one of the first medical image and the second medical image is cropped based on the position information to obtain a plurality of anatomical images corresponding to the plurality of bones one by one, and the anatomical image for the target bone is determined from the plurality of anatomical images based on the category information, and the specific process is as follows Figure 3 The description of the embodiments is not repeated here.
[0075] It can be understood that in the process of identifying edge pixel points from the image based on the preset pixel difference to obtain the initial image region, the identification accuracy of the edge pixel points has certain limitations, and therefore directly taking the target image region as the anatomical image may have the case that the accuracy of the anatomical image is not high enough, and the way of cropping the image based on the position information to obtain the anatomical image makes the accuracy of the anatomical image higher.
[0076] After obtaining the anatomical image for the target bone through the semantic segmentation model as described above, the first medical image and the anatomical image need to be further processed to obtain abnormal information of the target blood vessel.
[0077] For example, the first medical image is segmented to obtain a mask image of the target blood vessel. Another semantic segmentation model can be used to segment the first medical image to obtain the mask image. The other semantic segmentation model is similar to the semantic segmentation model used to segment to obtain the anatomical image as described above, and is not repeated here.
[0078] Then, an anomaly detection model is used to extract features from the first medical image, the mask image, and the anatomical image to obtain image feature data, and the first medical image can be a preprocessed image. The image feature data includes: global feature data of the target blood vessel and the target bone represented by the first medical image, morphological feature data of the target blood vessel represented by the mask image, and local feature data of the target bone represented by the anatomical image.
[0079] Figure 4 An illustrative diagram of detecting abnormalities using an anomaly detection model is provided for the embodiments of the present application.
[0080] As Figure 4As shown, the anomaly detection model 420 can be a deep learning model, which is similar to the semantic segmentation model above, but different. Specifically, the anomaly detection model 420 includes a first classification sub-model 421 and a second classification sub-model 422.
[0081] The image feature data includes first feature data for characterizing target vessel abnormalities, and the first classification sub-model 421 can be used to extract features from the first medical image, the mask image, and the anatomical image to obtain the first feature data.
[0082] The image feature data also includes second feature data for characterizing target bone abnormalities, and the second classification sub-model 422 can be used to extract features from the first medical image, the mask image, and the anatomical image to obtain the second feature data.
[0083] In an example, the first classification sub-model 421 and the second classification sub-model 422 can share a feature extraction layer.
[0084] Specifically, the anomaly detection model 420 includes at least a second encoding network, a second decoding network, a second feature concatenation network, a first classification network, a second classification network, and an attention network.
[0085] The second encoding network includes a plurality of third network levels connected in sequence, the channel of the feature map extracted by the plurality of third network levels increases gradually, and the size of the feature map extracted by the plurality of third network levels decreases gradually.
[0086] The second decoding network is connected with the second encoding network, and the second decoding network includes a plurality of fourth network levels connected in sequence, the channel of the feature map extracted by the plurality of fourth network levels decreases gradually, and the size of the feature map extracted by the plurality of fourth network levels increases gradually.
[0087] The third network level and / or the fourth network level includes a second residual network, and the second residual network is used to transmit feature information between different third network levels and / or used to transmit feature information between different fourth network levels.
[0088] The second feature concatenation network is connected with the second encoding network and the second decoding network, and the second feature concatenation network concatenates the feature map output by the second encoding network with the same channel feature map output by the second decoding network.
[0089] The anomaly detection model 420 can be a U-shaped residual network structure, specifically a Res-UNet network. The second encoding network is similar to the first encoding network in the semantic segmentation model described above, the second decoding network is similar to the first decoding network in the semantic segmentation model described above, the second feature concatenation network is similar to the first feature concatenation network in the semantic segmentation model described above, and the second residual network is similar to the first residual network in the semantic segmentation model described above, and thus will not be described again here.
[0090] The first classification network and the second classification network are both connected to the second decoding network, the first classification network is configured to output a first classification result for representing a target blood vessel anomaly, and the second classification network is configured to output a second classification result for representing a target bone anomaly. The first classification network and the second classification network can each include one or more fully connected layers. For example, the second decoding network includes four network levels connected in sequence, each network level can be composed of a feature extraction layer and an up-sampling network, and the first classification network and the second classification network can be connected as two branches after the up-sampling network of the fourth network level (the last level of the second decoding network).
[0091] The present embodiment takes the target bone as a cervical vertebra and the target blood vessel as a vertebral artery as an example. The first classification result includes at least one of the following, for example: a probability that the vertebral artery passes through the transverse foramen of the cervical vertebra, a probability that the vertebral artery exists but does not pass through the transverse foramen of the cervical vertebra, a probability of vertebral artery infarction, a probability of vertebral artery thinness, and a probability of normal vertebral artery morphology. The second classification result includes at least one of the following, for example: a probability of cervical intervertebral foramen stenosis, a probability of intervertebral disc herniation, and a probability of vertebral fracture. It should be understood that the first classification result and the second classification result can serve as reference information to assist doctors in making anomaly judgments.
[0092] The attention network includes a spatial attention network for the image and a channel attention network for the feature map, and the spatial attention network and the channel attention mechanism are both connected to the second encoding network. For example, the second encoding network includes four network levels connected in sequence, each network level can be composed of a feature extraction layer and a down-sampling network, and the spatial attention network and the channel attention network can be added to the down-sampling network of the fourth network level (the last level of the second encoding network). The attention network includes an attention mechanism, which can be used to score each dimension of the input feature, and then the feature is weighted according to the score to highlight the important features to the downstream network in the model, thereby improving the detection effect of the model.
[0093] The first classification sub-model 421 in the above can include a second encoding network, a second decoding network, a second feature splicing network, and a first classification network. The second classification sub-model 422 can include a second encoding network, a second decoding network, a second feature splicing network, and a second classification network. The second encoding network, the second decoding network, and the second feature splicing network can constitute a feature extraction layer, and the first classification sub-model 421 and the second classification sub-model 422 can share the feature extraction layer.
[0094] Specifically, referring to Figure 4 The first medical image 411, the mask image 412, and the anatomical image 413 are input into the abnormality detection model 420 for processing, and the first medical image 411, the mask image 412, and the anatomical image 413 can each be a three-dimensional image. The anatomical image 413 is an anatomical image of any one of the cervical vertebrae C1-C7, and can be input according to actual conditions. If it is necessary to detect the cervical vertebrae C1-C7, the anatomical images of the cervical vertebrae C1-C7 need to be input in turn, that is, an anatomical image, the first medical image 411, and the mask image 412 are input each time, and the model output result is waited for, and then the next anatomical image, the first medical image 411, and the mask image 412 are input.
[0095] The abnormality detection model of the embodiment of the present application includes a multi-task UNet network structure, the model contains multiple inputs and multiple outputs and has a spatial attention mechanism and a channel attention mechanism, and the multiple tasks are jointly trained and mutually assisted. The model learns the target region in the image comprehensively and from multiple directions based on global deep semantic information, local detail information, anatomical position information, and morphological features of the image. The first classification sub-model 421 can predict the probability of the vertebral artery penetrating each cervical vertebra, and the second classification sub-model 422 can judge whether the cervical vertebra has an abnormality and a lesion, and the two sub-models share convolution layers. The CTA medical image input into the model can enable the network to learn global information, the mask image of the vertebral artery can assist the model to learn the morphology of the vertebral artery, and the anatomical image of the cervical vertebra can enable the network to focus on local information, thereby realizing deep expression of features while keeping the model relatively light.
[0096] The vertebral artery includes left and right vertebral arteries, and each cervical vertebra includes left and right transverse foramina. The first classification sub-model 421 outputs a first classification result [P li ,P ri ,P lo ,P ro ,P lc ,P rc ,P ln ,P rn ,P ls ,P rs ] representing a blood vessel abnormality. Wherein, P li and P riP and P represent the probability of the left vertebral artery and the right vertebral artery passing through the left transverse foramen and the right transverse foramen of the cervical vertebra, respectively; P lo and P ro represent the probability of the left vertebral artery and the right vertebral artery existing but not passing into the left transverse foramen and the right transverse foramen of the cervical vertebra, respectively; P lc and P rc represent the probability of the left vertebral artery and the right vertebral artery infarction, respectively; P ls and P rs represent the probability of the left vertebral artery and the right vertebral artery being thin, respectively; P ln and P rn represent the probability of the left vertebral artery and the right vertebral artery being normal, respectively.
[0097] The second classification sub-model 422 outputs the second classification result [P vs ,P sd ,P vf ] representing the bone abnormalities. Among them, P vs represents the probability of the intervertebral foramen stenosis of the cervical vertebra; P sd represents the probability of the intervertebral disc herniation, P vf represents the probability of the vertebral fracture.
[0098] It can be seen that the above abnormality detection model can at least predict the following abnormalities and pathological conditions:
[0099] The left / right vertebral artery enters the transverse foramen from the cervical vertebra C7, and the characteristic expression is that the vertebral artery runs in the transverse foramen of the cervical vertebra C7, and passes through the cervical vertebra C6 to the cervical vertebra C1;
[0100] The left / right vertebral artery enters the transverse foramen from the cervical vertebra C5, and the characteristic expression is that the vertebral artery does not pass through the transverse foramen of the cervical vertebra C7 and the cervical vertebra C6, but runs in the transverse foramen of the cervical vertebra C5 and passes through the cervical vertebra C1;
[0101] The left / right vertebral artery enters the transverse foramen from the cervical vertebra C4, and the characteristic expression is that the vertebral artery does not pass through the transverse foramen of the cervical vertebra C7-C5, but runs in the transverse foramen of the cervical vertebra C4 and passes through the cervical vertebra C1;
[0102] The left / right vertebral artery enters the transverse foramen from the cervical vertebra C3, and the characteristic expression is that the vertebral artery does not pass through the transverse foramen of the cervical vertebra C7-C4, but runs in the transverse foramen of the cervical vertebra C3 and passes through the cervical vertebra C1;
[0103] It can also predict that the left / right vertebral artery enters the transverse foramen from the cervical vertebra C2 or C1, etc.
[0104] The left / right vertebral artery is infarcted at the cervical vertebra C i (i=1, 2, …, 7);
[0105] The left / right vertebral artery is infarcted at the cervical vertebra C i(i = 1, 2, …, 7) cervical vertebrae are thin.
[0106] It should be understood that the abnormal information of the blood vessel (vertebral artery) and the abnormal information of the bone (cervical vertebra) predicted by the abnormality detection model in the embodiments of the present application are only for reference, and the abnormality detection result can assist doctors in medical diagnosis.
[0107] The deep learning model of the embodiments of the present application can simultaneously detect the abnormality of the vertebral artery and the abnormality of the cervical vertebra, and comprehensively learn the global deep semantic information of the CTA medical image, the morphological information of the vertebral artery in the mask image, the local detail information of the anatomical image, and the like in multiple directions. It can be seen that by combining the anatomical structure of the cervical vertebra and the morphological information of the vertebral artery, the internal relationship between the relevant features and the learning goal can be more easily found. The model uses a multi-task learning method to supervise and assist learning in different tasks to obtain better classification results, so that the model can simultaneously judge multiple abnormal types such as vertebral artery morphological abnormalities, stenosis, and thinness, and calculate the specific position of the abnormality, and more targetedly learn and capture various variation conditions and locate lesions. It can be seen that through the embodiments of the present application, the doctor can be quickly assisted to locate the abnormality and the lesion site, improve the work efficiency of the doctor, and play an early prevention role for ischemic posterior circulation and infarction.
[0108] Figure 5 A schematic diagram of an image processing-based abnormality detection device provided by the embodiments of the present application.
[0109] The embodiments of the present application provide an image processing-based abnormality detection device 500, please refer to Figure 5 The image processing-based abnormality detection device 500 includes an acquisition module 510, an extraction module 520, and a detection module 530.
[0110] Illustratively, the acquisition module 510 is configured to acquire a first medical image for a target blood vessel and a target bone, and an anatomical image for the target bone.
[0111] Illustratively, the extraction module 520 is configured to perform feature extraction on the first medical image and the anatomical image to obtain image feature data, wherein the image feature data is used to represent the features of the target blood vessel and the relative position relationship between the target blood vessel and the target bone.
[0112] Illustratively, the detection module 530 is configured to detect abnormal information of the target blood vessel based on the image feature data.
[0113] It can be understood that the specific description of the image processing-based abnormality detection device 500 can be referred to the description of the image processing-based abnormality detection method in the above.
[0114] Exemplarily, the feature extraction on the first medical image and the anatomical image to obtain the image feature data comprises: performing segmentation processing on the first medical image to obtain a mask image of the target blood vessel; and performing feature extraction on the first medical image, the mask image and the anatomical image by using the anomaly detection model to obtain the image feature data.
[0115] Exemplarily, the image feature data comprises at least one of: global feature data of the target blood vessel and the target bone represented by the first medical image; morphological feature data of the target blood vessel represented by the mask image; and local feature data of the target bone represented by the anatomical image.
[0116] Exemplarily, the anomaly detection model comprises a first classification sub-model, and the image feature data comprises first feature data for representing the target blood vessel anomaly; wherein the feature extraction on the first medical image, the mask image and the anatomical image by using the anomaly detection model to obtain the image feature data comprises: performing feature extraction on the first medical image, the mask image and the anatomical image by using the first classification sub-model to obtain the first feature data.
[0117] Exemplarily, the anomaly detection model further comprises a second classification sub-model, and the image feature data further comprises second feature data for representing the target bone anomaly; wherein the anomaly detection apparatus 500 further comprises: a second feature data extraction module configured to perform feature extraction on the first medical image, the mask image and the anatomical image by using the second classification sub-model to obtain the second feature data.
[0118] Exemplarily, the first classification sub-model and the second classification sub-model share a feature extraction layer.
[0119] Exemplarily, the anomaly detection apparatus 500 further comprises: a processing module configured to perform processing on at least one of the first medical image and the second medical image to obtain the anatomical image for the target bone.
[0120] Exemplarily, the processing on at least one of the first medical image and the second medical image to obtain the anatomical image for the target bone comprises: performing processing on at least one of the first medical image and the second medical image to obtain position information of a plurality of bones in the image and category information of each bone; performing cropping on at least one of the first medical image and the second medical image based on the position information to obtain a plurality of anatomical images corresponding to the plurality of bones one by one; and determining the anatomical image for the target bone from the plurality of anatomical images based on the category information.
[0121] Exemplarily, the processing of at least one of the first medical image and the second medical image to obtain the position information of the plurality of bones in the image and the category information of each bone comprises: performing feature extraction on at least one of the first medical image and the second medical image by using a semantic segmentation model to obtain the position information of the plurality of bones in the image and the category information of each bone.
[0122] Exemplarily, the processing of at least one of the first medical image and the second medical image to obtain the position information of the plurality of bones in the image and the category information of each bone comprises: determining, based on a preset pixel difference value, an edge pixel point from at least one of the first medical image and the second medical image, wherein an image region surrounded by the edge pixel point is a plurality of initial image regions; screening a plurality of target image regions corresponding to the plurality of bones from the plurality of initial image regions; determining the region centers of the plurality of target image regions as the position information of the plurality of bones in the image, respectively; and determining the category information of each bone based on the relative positional relationship between the position information of the plurality of bones in the image.
[0123] Exemplarily, the plurality of target image regions satisfy at least one of the following conditions: the region size of each target image region is within a preset region size range, the preset region size range representing the size of the bone; the region shape of each target image region is a preset shape, the preset shape representing the shape of the bone; the distance between two adjacent target image regions is less than a preset distance, the preset distance representing the distance between two adjacent bones; and the relative positional relationship between the plurality of target image regions is a preset relative positional relationship, the preset relative positional relationship representing the relative positions of the plurality of bones.
[0124] Exemplarily, the semantic segmentation model is a loss function value that is a weighted value between a cross-entropy value and a set similarity measure value.
[0125] Exemplarily, the semantic segmentation model comprises: a first encoding network, the first encoding network comprising a plurality of first network levels connected in sequence, the channel of the feature map extracted by the plurality of first network levels increasing gradually, and the size of the feature map extracted by the plurality of first network levels decreasing gradually; a first decoding network connected with the first encoding network, the first decoding network comprising a plurality of second network levels connected in sequence, the channel of the feature map extracted by the plurality of second network levels decreasing gradually, and the size of the feature map extracted by the plurality of second network levels increasing gradually; and a first feature splicing network connected with the first encoding network and the first decoding network, the first feature splicing network splicing the feature map output by the first encoding network with the same channel feature map output by the first decoding network, wherein the first network level and / or the second network level comprises a first residual network, the first residual network being used for transmitting feature information between different first network levels and / or being used for transmitting feature information between different second network levels.
[0126] Exemplarily, the anomaly detection model comprises: a second encoding network, the second encoding network comprising a plurality of third network layers connected in sequence, the plurality of third network layers extracting feature maps with gradually increasing channels and gradually decreasing sizes; a second decoding network connected with the second encoding network, the second decoding network comprising a plurality of fourth network layers connected in sequence, the plurality of fourth network layers extracting feature maps with gradually decreasing channels and gradually increasing sizes; a second feature splicing network connected with the second encoding network and the second decoding network, the second feature splicing network splicing the feature maps output by the second encoding network with the same channel feature maps output by the second decoding network; a first classification network connected with the second decoding network, the first classification network being configured to output a first classification result for representing a target blood vessel anomaly; a second classification network connected with the second decoding network, the second classification network being configured to output a second classification result for representing a target bone anomaly; a spatial attention network for the image and a channel attention network for the feature maps, the spatial attention network and the channel attention network both being connected with the second encoding network, wherein the third network layers and / or the fourth network layers comprise a second residual network, the second residual network being configured to transmit feature information between different third network layers and / or transmit feature information between different fourth network layers.
[0127] Exemplarily, the target blood vessel comprises a vertebral artery, and the target bone comprises a cervical vertebra.
[0128] An electronic device is provided in an embodiment of the present application, comprising a memory and a processor, the memory storing a computer program, and the processor implementing the steps of the method in any of the above embodiments when executing the computer program.
[0129] A computer readable storage medium is provided in an embodiment of the present application, the computer readable storage medium storing a computer program, and the computer program implementing the steps of the method in any of the above embodiments when executed by a processor.
[0130] An embodiment of the present application provides a computer program product, the computer program product comprising instructions, the instructions being executed by a processor of a computer device to enable the computer device to perform the steps of the method in any of the above embodiments.
[0131] It should be noted that the logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequential list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this application, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: electrical connections (electronic devices) having one or more wires, portable computer disk drives (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Furthermore, computer-readable media can even be paper or other suitable media on which programs can be printed, because programs can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpreting, or otherwise processing as necessary, and then stored in computer memory.
[0132] It should be understood that various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0133] In the description of this application, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this application, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0134] In the description of the present application, it needs to be understood that the orientation or positional relationship indicated by the terms "center", "longitudinal", "lateral", "length", "width", "thickness", "upper", "lower", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", "axial", "radial", "circumferential" and the like is based on the orientation or positional relationship shown in the drawings, and is only for the purpose of facilitating the description of the present application and simplifying the description, and does not indicate or imply that the device or element referred to must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as a limitation on the present application.
[0135] In addition, the terms "first", "second", and the like used in the embodiments of the present application are only for the purpose of description, and cannot be understood as indicating or implying relative importance, or implicitly indicating the number of technical features referred to in the embodiments. Therefore, the features defined with the terms "first", "second" and the like in the embodiments of the present application can be explicitly or implicitly indicated to include at least one of the features. In the description of the present application, the meaning of the word "plurality" is at least two or two or more, such as two, three, four, etc., unless otherwise specifically limited in the embodiments.
[0136] In the present application, unless otherwise specifically defined or limited in the embodiments, the terms "mounting", "connecting", "connecting" and "fixing" and the like appearing in the embodiments should be understood broadly, for example, the connection can be a fixed connection, or a detachable connection, or integrated, which can be understood, or can be a mechanical connection, an electrical connection, etc. Of course, it can also be directly connected, or indirectly connected through an intermediate medium, or it can be the internal communication of two elements, or the interaction relationship between two elements. For those skilled in the art, the specific meaning of the above terms in the present application can be understood according to the specific implementation situation.
[0137] In the present application, unless otherwise specifically defined or limited, the first feature "on" or "under" the second feature can be that the first and second features are in direct contact, or the first and second features are indirectly in contact through an intermediate medium. Moreover, the first feature "above", "above" and "above" the second feature can be that the first feature is directly above or obliquely above the second feature, or only indicates that the horizontal height of the first feature is higher than that of the second feature. The first feature "below", "below" and "below" the second feature can be that the first feature is directly below or obliquely below the second feature, or only indicates that the horizontal height of the first feature is less than that of the second feature.
[0138] Although the embodiments of the present application have been shown and described above, it is understood that the above-described embodiments are exemplary and are not to be construed as limiting the present application, and that changes, modifications, substitutions and variations can be made by those skilled in the art without departing from the scope of the present application.
Claims
1. An image processing-based anomaly detection method, characterized by, The method comprises: obtaining a first medical image for a target blood vessel and a target bone and an anatomical image for the target bone; performing feature extraction on the first medical image and the anatomical image to obtain image feature data, comprising: performing segmentation processing on the first medical image to obtain a mask image of the target blood vessel, and performing feature extraction on the first medical image, the mask image and the anatomical image by using an anomaly detection model to obtain the image feature data, wherein the image feature data is used to represent characteristics of the target blood vessel and a relative position relationship between the target blood vessel and the target bone, and the image feature data comprises at least one of: global feature data of the target blood vessel and the target bone represented by the first medical image, morphological feature data of the target blood vessel represented by the mask image, and local feature data of the target bone represented by the anatomical image; and detecting abnormal information of the target blood vessel and abnormal information of the target bone based on the image feature data; wherein the anomaly detection model comprises a first classification sub-model and a second classification sub-model, the image feature data comprises first feature data used to represent the target blood vessel anomaly and second feature data used to represent the target bone anomaly; the first classification sub-model is used to perform feature extraction on the first medical image, the mask image and the anatomical image to obtain the first feature data; and the second classification sub-model is used to perform feature extraction on the first medical image, the mask image and the anatomical image to obtain the second feature data; the first classification sub-model and the second classification sub-model share a feature extraction layer, and the anomaly detection model comprises a multi-task UNet network structure.
2. The method of claim 1, wherein, The method further comprises: processing at least one of the first medical image and a second medical image to obtain an anatomical image for the target bone.
3. The method of claim 2, wherein, The processing at least one of the first medical image and the second medical image to obtain the anatomical image for the target bone comprises: processing at least one of the first medical image and the second medical image to obtain position information of a plurality of bones in the image and category information of each bone; based on the position information, cropping at least one of the first medical image and the second medical image to obtain a plurality of anatomical images corresponding to the plurality of bones one by one; and based on the category information, determining the anatomical image for the target bone from the plurality of anatomical images.
4. The method of claim 3, wherein, The processing at least one of the first medical image and the second medical image to obtain the position information of the plurality of bones in the image and the category information of each bone comprises: performing feature extraction on at least one of the first medical image and the second medical image by using a semantic segmentation model to obtain the position information of the plurality of bones in the image and the category information of each bone.
5. The method of claim 3, wherein, The processing of at least one of the first medical image and the second medical image to obtain the position information of the plurality of bones in the image and the category information of each bone comprises: Based on a preset pixel difference value, edge pixel points are determined from at least one of the first medical image and the second medical image, wherein an image region surrounded by the edge pixel points is a plurality of initial image regions; A plurality of target image regions corresponding one-to-one to the plurality of bones are selected from the plurality of initial image regions; The region centers of the plurality of target image regions are determined as the position information of the plurality of bones in the image, respectively; and Based on the relative positional relationship between the position information of the plurality of bones in the image, the category information of each bone is determined.
6. The method of claim 5, wherein, The plurality of target image regions satisfy at least one of the following conditions: The region size of each target image region is within a preset region size range, and the preset region size range represents the size of the bone; The region shape of each target image region is a preset shape, and the preset shape represents the shape of the bone; The distance between two adjacent target image regions is less than a preset distance, and the preset distance represents the distance between two adjacent bones; The relative positional relationship between the plurality of target image regions is a preset relative positional relationship, and the preset relative positional relationship represents the relative position between the plurality of bones.
7. The method of claim 4, wherein, The loss function value of the semantic segmentation model is a weighted value between a cross-entropy value and a set similarity measure value.
8. The method according to claim 4 or 7, characterized in that, The semantic segmentation model comprises: A first encoding network, the first encoding network comprising a plurality of first network levels connected in sequence, the number of channels of the feature maps extracted by the plurality of first network levels increasing gradually, and the size of the feature maps extracted by the plurality of first network levels decreasing gradually; A first decoding network connected to the first encoding network, the first decoding network comprising a plurality of second network levels connected in sequence, the number of channels of the feature maps extracted by the plurality of second network levels decreasing gradually, and the size of the feature maps extracted by the plurality of second network levels increasing gradually; and A first feature splicing network connected to the first encoding network and the first decoding network, the first feature splicing network splicing the feature maps output by the first encoding network and the same channel feature maps output by the first decoding network, Wherein, the first network level and / or the second network level comprises a first residual network, the first residual network being used for transmitting feature information between different first network levels and / or being used for transmitting feature information between different second network levels.
9. The method of claim 1, wherein, The anomaly detection model comprises: A second encoding network, the second encoding network comprising a plurality of third network levels connected in sequence, the number of channels of the feature maps extracted by the plurality of third network levels increasing gradually, and the size of the feature maps extracted by the plurality of third network levels decreasing gradually; A second decoding network connected to the second encoding network, the second decoding network comprising a plurality of fourth network levels connected in sequence, the number of channels of the feature maps extracted by the plurality of fourth network levels decreasing gradually, and the size of the feature maps extracted by the plurality of fourth network levels increasing gradually; a second feature concatenation network connected with the second encoding network and the second decoding network, the second feature concatenation network concatenating the feature map output by the second encoding network and the same-channel feature map output by the second decoding network; a first classification network connected with the second decoding network, the first classification network being configured to output a first classification result for representing the target blood vessel abnormality; a second classification network connected with the second decoding network, the second classification network being configured to output a second classification result for representing the target bone abnormality; and a spatial attention network for the image and a channel attention network for the feature map, the spatial attention network and the channel attention network both being connected with the second encoding network, wherein the third network level and / or the fourth network level comprises a second residual network configured to pass feature information between different third network levels and / or pass feature information between different fourth network levels.
10. The method of claim 1, wherein, The target blood vessel comprises a vertebral artery, and the target bone comprises a cervical vertebra.
11. An abnormality detection apparatus based on image processing, characterized by comprising: The apparatus comprises: an acquisition module configured to acquire a first medical image for a target blood vessel and a target bone, and an anatomical image for the target bone; an extraction module configured to perform feature extraction on the first medical image and the anatomical image to obtain image feature data, including: performing segmentation processing on the first medical image to obtain a mask image of the target blood vessel, and performing feature extraction on the first medical image, the mask image, and the anatomical image by using an abnormality detection model to obtain the image feature data, wherein the image feature data is configured to represent features of the target blood vessel and a relative positional relationship between the target blood vessel and the target bone, and the image feature data comprises at least one of: global feature data of the target blood vessel and the target bone represented by the first medical image, morphological feature data of the target blood vessel represented by the mask image, and local feature data of the target bone represented by the anatomical image; and a detection module configured to detect abnormal information of the target blood vessel and abnormal information of the target bone based on the image feature data. The abnormality detection model comprises a first classification sub-model and a second classification sub-model, and the image feature data comprises first feature data for representing the target blood vessel abnormality and second feature data for representing the target bone abnormality; the first classification sub-model is configured to perform feature extraction on the first medical image, the mask image, and the anatomical image to obtain the first feature data; and the second classification sub-model is configured to perform feature extraction on the first medical image, the mask image, and the anatomical image to obtain the second feature data. The first classification sub-model and the second classification sub-model share a feature extraction layer, and the abnormality detection model comprises a multi-task UNet network structure.
12. An electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and wherein the computer program is configured to perform the method according to any one of claims 1-11. The processor executes the computer program to implement the steps of the method of any one of claims 1-10.
13. A computer readable storage medium having stored thereon a computer program, characterized in that The computer program, which is executed by a processor, implements the steps of the method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Method and apparatus for positioning interested area in medical image, and system thereof
CN107545584A
Image segmentation method, radiotherapy system, computer equipment and storage medium
CN112001925A
Micro-hemorrhage focus segmentation method based on convolutional neural network
CN112927243A
Vascular system variation detection method and device and storage medium
CN114913174A
Vascular abnormality analysis method and device, storage medium and electronic equipment
CN116363070A