Model training and image processing method, device and equipment for blood vessel tree segmentation
Patent Information
- Application Number
- CN202310556309.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-17
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2043-05-17
AI Technical Summary
具体而言,由于末端血管的比例很小,会使得末端血管部分的分割效果很难在学习过程中表现出来,这会导致图像分割模型无法针对末端血管的分割得以充分的学习,进而使得训练后的图像分割模型在实际应用时,得到的最终分割结果中末端部分容易发生漏检与断裂等问题
[0025]本申请实施例中,在进行血管树分割模型的训练时,结合各个样本图像的标注结果中,各个血管分支的粗细程度以及各像素点在血管分支中的边缘程度,来生成各个像素点对应的训练权重,且该训练权重与粗细程度以及边缘程度均呈负相关,也就是说,每个像素点所在的血管分支越细,或者像素点的在血管分支中的边缘程度越低,则该像素点所对应的训练权重越高,从而使得样本图像中末端血管的训练权重更高,增加了此类血管分支在整体训练数据中的比重,从而使得训练过程针对此类血管分支的感知程度更强,则最终得到的血管树分割模型应用到实际场景时,则针对末端血管等小血管的感知能力得以显著提升,以达到提升血管树的末端血管部分的分割效果,提升血管树分割的整体准确性的目的。
Smart Images

Figure CN116958539B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence (AI) technology, and more particularly to the field of vascular tree segmentation technology, providing a model training and image processing method, apparatus and device for vascular tree segmentation. Background Technology
[0002] With the development of computer vision (CV) technology, its applications in medicine are becoming increasingly widespread. For example, image segmentation techniques can be used to segment vascular tree images from wide-angle fundus photographs, thereby assisting in the diagnosis of diseases such as diabetic retinopathy. Peripheral lesions observed in a wide field of view can provide more sufficient diagnostic evidence for the staging of diabetic retinopathy. Alternatively, combined with echocardiographic images, cardiac vascular tree image segmentation can be performed to assist in the diagnosis of heart-related diseases. It is evident that the accuracy of vascular tree segmentation directly affects the outcome of disease diagnosis; therefore, accurate vascular tree segmentation is fundamental to accurate disease diagnosis.
[0003] In vascular tree image segmentation, terminal vessels constitute a very small proportion of the entire vascular tree, making their segmentation more challenging than that of main vessels. Specifically, the small proportion of terminal vessels makes it difficult for them to be effectively segmented during the learning process. This results in the image segmentation model failing to adequately learn how to segment terminal vessels, leading to issues such as missed detections and fragmentation of terminal segments in the final segmentation results obtained during practical applications.
[0004] Therefore, how to achieve more accurate vascular tree segmentation based on vascular tree images is a technical problem that urgently needs to be solved. Summary of the Invention
[0005] This application provides a model training and image processing method, apparatus, and device for vascular tree segmentation, which is used to relatively increase the training weight of terminal blood vessels during the training process, so as to improve the accuracy of vascular tree segmentation.
[0006] On the one hand, a model training method for vascular tree segmentation is provided, the method comprising:
[0007] A sample image set for model training is obtained, wherein each sample image in the sample image set is an image containing a vascular tree and is associated with annotation results representing the true location of the vascular tree in the sample image;
[0008] Based on the annotation results of each sample image, the attribute recognition processing of blood vessel branches is performed to obtain the attribute information of each blood vessel branch in each sample image. The attribute information includes: the thickness of the corresponding blood vessel branch, and the edge degree of the pixels included in the corresponding blood vessel branch in the blood vessel branch.
[0009] Based on the attribute information corresponding to each sample image, the training weights corresponding to each pixel in the corresponding sample image are obtained respectively. The training weights are negatively correlated with the thickness and the edge degree.
[0010] During the training of the blood vessel tree segmentation model, the model loss value is determined based on the prediction results, annotation results, and corresponding training weights of the blood vessel trees in each sample image, and the model parameters are adjusted based on the model loss value.
[0011] On the one hand, an image processing method for vascular tree segmentation is provided, the method comprising:
[0012] Obtain the target image for vascular tree segmentation;
[0013] Based on the vascular tree segmentation model obtained through the above model training method, the target image is subjected to vascular tree segmentation processing based on image visual detection, and the corresponding vascular tree segmentation result is output; wherein, the vascular tree segmentation result is used to indicate the position of the vascular tree in the target image.
[0014] On the one hand, a model training device for vascular tree segmentation is provided, the device comprising:
[0015] The sample acquisition unit is used to obtain a set of sample images for model training. Each sample image in the set is an image containing a vascular tree and is associated with a labeling result representing the true location of the vascular tree in the sample image.
[0016] The attribute recognition unit is used to perform attribute recognition processing of blood vessel branches based on the annotation results of each sample image, and to obtain the attribute information of each blood vessel branch in each sample image. The attribute information includes: the thickness of the corresponding blood vessel branch, and the edge degree of the pixels included in the corresponding blood vessel branch in the blood vessel branch.
[0017] The weight determination unit is used to obtain the training weights corresponding to each pixel in the corresponding sample image based on the attribute information corresponding to each sample image. The training weights are negatively correlated with the thickness and the edge degree.
[0018] The training unit is used to determine the model loss value based on the prediction results, annotation results and corresponding training weights of the blood vessel tree segmentation model during the training process, and to adjust the model parameters based on the model loss value.
[0019] On one hand, an image processing apparatus for vascular tree segmentation is provided, the apparatus comprising:
[0020] An image acquisition unit is used to acquire the target image to be segmented into a blood vessel tree.
[0021] The segmentation unit is used to perform image visual detection-based vascular tree segmentation processing on the target image based on the vascular tree segmentation model obtained by the above-described model training method, and output the corresponding vascular tree segmentation result; wherein, the vascular tree segmentation result is used to indicate the position of the vascular tree in the target image.
[0022] On one hand, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of any of the above methods.
[0023] On the one hand, a computer storage medium is provided, on which a computer program is stored, which, when executed by a processor, implements the steps of any of the above methods.
[0024] On one hand, a computer program product is provided, comprising a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium, and executes the computer program, causing the computer device to perform the steps of any of the methods described above.
[0025] In this embodiment, during the training of the vascular tree segmentation model, the thickness of each vascular branch and the edge degree of each pixel within the vascular branch are combined with the annotation results of each sample image to generate training weights corresponding to each pixel. These training weights are negatively correlated with both thickness and edge degree; that is, the thinner the vascular branch containing each pixel, or the lower the edge degree of the pixel within the vascular branch, the higher the training weight corresponding to that pixel. This results in higher training weights for terminal vascular branches in the sample images, increasing the proportion of such vascular branches in the overall training data. Consequently, the training process becomes more sensitive to these vascular branches. Therefore, when the final vascular tree segmentation model is applied to real-world scenarios, its ability to perceive small vascular branches such as terminal vascular branches is significantly improved, thereby enhancing the segmentation effect of the terminal vascular portion of the vascular tree and improving the overall accuracy of vascular tree segmentation. Attached Figure Description
[0026] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0027] Figure 1 This is a schematic diagram illustrating an application scenario provided in the embodiments of this application;
[0028] Figure 2 A schematic diagram of the architecture of a blood vessel tree segmentation system based on fundus images provided in an embodiment of this application;
[0029] Figure 3 A schematic flowchart illustrating the model training method for vascular tree segmentation provided in this application embodiment;
[0030] Figure 4 A schematic diagram illustrating the calculation process of training weights provided in an embodiment of this application;
[0031] Figure 5 A schematic diagram of the skeletal lines provided in the embodiments of this application;
[0032] Figure 6 A schematic diagram of a training process for a blood vessel tree segmentation model provided in an embodiment of this application;
[0033] Figure 7 A schematic diagram of another training process for the blood vessel tree segmentation model provided in this application embodiment;
[0034] Figure 8 A schematic diagram of a partial region provided in an embodiment of this application;
[0035] Figure 9 A schematic diagram of a sliding window input image provided in an embodiment of this application;
[0036] Figure 10 This is a schematic flowchart of an image processing method for vascular tree segmentation provided in an embodiment of this application;
[0037] Figure 11 This is a schematic diagram of the fundus vascular tree in a wide-angle fundus photograph provided in an embodiment of this application;
[0038] Figure 12a and Figure 12b This is a schematic diagram of the segmentation results provided in an embodiment of this application;
[0039] Figure 13 A schematic diagram of a model training device for vascular tree segmentation provided in an embodiment of this application;
[0040] Figure 14 A schematic diagram of an image processing apparatus for vascular tree segmentation provided in an embodiment of this application;
[0041] Figure 15 This is a schematic diagram of the composition structure of a computer device provided in an embodiment of this application;
[0042] Figure 16 This is a schematic diagram of the composition structure of another computer device using an embodiment of this application. Detailed Implementation
[0043] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application. Unless otherwise specified, the embodiments and features in the embodiments of this application can be arbitrarily combined with each other. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that shown here.
[0044] It is understood that the following specific embodiments of this application involve vascular tree image data, such as sample images used in the training phase or vascular tree images collected in the application phase. When the various embodiments of this application are applied to specific products or technologies, relevant licenses or consents are required, and the collection, use, and processing of relevant data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, when relevant data is needed, it can be obtained by recruiting relevant volunteers and signing relevant agreements authorizing the use of their data; or, it can be implemented within an authorized organization, using data from internal members to implement the following implementation methods to make relevant recommendations to internal members; or, in specific implementations, all relevant data used are simulated data.
[0045] To facilitate understanding of the technical solutions provided in the embodiments of this application, some key terms used in the embodiments of this application will be explained below:
[0046] Vascular Tree: Blood vessels are the conduits through which blood is transported in living organisms. In the human body, blood vessels are distributed throughout the entire body. Arteries carry blood from the heart to the body tissues, veins carry blood from between the tissues back to the heart, and capillaries connect arteries and veins, serving as the primary site for the exchange of substances between blood and tissues. Different organisms possess different vascular morphologies. A vascular tree is a tree-like structure formed by interconnected blood vessels. In other words, a vascular tree is a way of representing the flow and shape of blood vessels in a tree-like structure. Of course, given the vast number of blood vessels in the body, vascular trees are usually drawn based on a specific organ or local area of the body, depicting the vascular pathways within that organ or region. Taking an organ as an example, the starting point of the vascular tree is the bifurcation point of the blood vessels flowing into that organ. From this bifurcation point, the blood vessel may continue to branch out further.
[0047] Vascular tree image: This includes the sample images involved in the embodiments of this application and the target images in the actual segmentation process. The vascular tree image can be an image containing a vascular tree collected in a real scene. It usually refers to medical images collected by medical means, such as wide-angle fundus photographs taken by a wide-angle fundus camera, computed tomography (CT) medical images, magnetic resonance imaging (MRI) medical images, or ultrasound medical images, etc.
[0048] Wide-angle fundus photographs: These are fundus photographs with a wide field of view, such as fundus photographs with a field of view of 200 degrees. They can observe more than 80% of the fundus retinal manifestations, which can help in the diagnosis of retinal-related diseases.
[0049] Image visual detection refers to the use of artificial intelligence technology to simulate human vision and perform visual detection on vascular tree images in order to extract information related to vascular trees from the images. In other words, it is a technology that uses artificial intelligence to distinguish which parts of a vascular tree image belong to blood vessels and which parts belong to non-blood vessels.
[0050] Vascular tree segmentation refers to the process of obtaining the location of vascular trees in a vascular tree image from an input image. It can also be described as the process of segmenting the pixels that belong to the vascular tree part of the vascular tree image. Therefore, the vascular tree segmentation process can also be called the vascular tree segmentation process. For example, after inputting a vascular tree image, the result is an image that only contains the pixels that belong to the vascular tree in the vascular tree image. Other non-vascular pixels are only the background part in the obtained image.
[0051] Terminal vessels: In this embodiment, terminal vessels refer to vessels with relatively small diameters. In a vascular tree, vessels closer to the end generally have smaller diameters. Therefore, terminal vessels can be defined according to their branching level. For example, terminal vessels can refer to smaller third-order vessels and vessels below the third order. The branching level refers to the number of bifurcations in a vessel; each time a vessel bifurcates, the branching level increases by one level. Alternatively, terminal vessels can be defined according to their diameter. For example, terminal vessels can refer to vascular branches with a diameter not exceeding a certain diameter threshold.
[0052] Skeleton extraction algorithms, also known as binary image thinning, refer to removing a portion of points from a binary image while retaining the original shape of the remaining points. This algorithm simplifies a planar region into a graph's structural shape representation, thinning a connected region to a width of one pixel for feature extraction and target topological representation. In the resulting graph, the pixel corresponding to each connected region can be called a skeleton point, and the line connecting these skeleton points can be called a skeleton line.
[0053] The embodiments of this application relate to artificial intelligence and machine learning (ML) technologies, and are primarily designed based on machine learning in artificial intelligence.
[0054] Artificial intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess the functions of perception, reasoning, and decision-making.
[0055] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning.
[0056] Machine learning is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory, among others. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and instructional learning.
[0057] Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence. Its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, and inductive learning. Artificial Neural Networks (ANNs) abstract the neural network of the human brain from an information processing perspective, establishing a simple model and forming different networks with different connection methods. A neural network is a computational model composed of a large number of interconnected nodes (or neurons). Each node represents a specific output function called the activation function. The connection between any two nodes represents a weighted value for the signal passing through that connection, called a weight. This is equivalent to the memory of the artificial neural network. The network's output varies depending on the connection method, weight values, and activation functions. The network itself is usually an approximation of a certain algorithm or function in nature, or it may be an expression of a logical strategy.
[0058] This application involves the process of image segmentation for images with a vein-like structure, such as vascular trees. The segmentation results can then be applied to downstream applications. For example, segmenting vascular trees in medical images can assist medical personnel in diagnosing related diseases. When segmenting vascular tree images, a deep learning-based vascular tree segmentation model is required. Specifically, this application uses machine learning to obtain a vascular tree segmentation model that performs the segmentation process. This model, based on machine learning, enables the processing and understanding of whether each pixel in the input vascular tree image is a vascular pixel. Specifically, the vascular tree segmentation model can comprehensively evaluate the probability that each pixel in the vascular tree image is a vascular pixel based on its trained vascular tree recognition capabilities and combined with image features of different dimensions and scales in the vascular tree image. This probability represents the likelihood that a pixel is identified as a vascular pixel by artificial intelligence.
[0059] Specifically, the vascular tree segmentation process in this embodiment can be divided into two parts: a training part and an application part. The training part involves machine learning, where an artificial neural network model (the vascular tree segmentation model mentioned later) is trained using machine learning techniques. This model is trained based on a set of sample images with labeled vascular tree locations provided in this embodiment, and the model parameters are continuously adjusted through optimization algorithms until the model converges. The application part uses the artificial neural network model trained in the training part to segment vascular tree images acquired during actual use, thereby segmenting the pixels belonging to blood vessels in the vascular tree image. It should also be noted that the artificial neural network model in this embodiment can be trained online or offline, without specific limitations. This paper uses offline training as an example.
[0060] Among related technologies, vascular tree segmentation technology is increasingly widely used in the field of medical diagnosis. However, the accuracy of vascular tree segmentation is the foundation for accurate disease diagnosis, and it directly affects the results of disease diagnosis.
[0061] Currently, the following methods are used in related technologies to segment the vascular tree:
[0062] The first approach, based on image semantic segmentation, transforms the binary labels of foreground and background of blood vessels into five categories: thick blood vessels, small blood vessels, edges of thick blood vessels, edges of small blood vessels, and background. This aims to increase the focus on small blood vessels during training. However, because this approach involves subdividing the categories, a corresponding weight hyperparameter is required for each category, resulting in high model training complexity and making the training results highly sensitive to the hyperparameters. Furthermore, small blood vessels can vary in thickness and background complexity; treating them all as small blood vessels reduces segmentation accuracy.
[0063] The first approach, based on skeletal line prediction, treats the prediction of the expanded skeletal lines as an auxiliary task. This method primarily alleviates the problem of uneven area distribution between large and small blood vessels in vascular labeling, thereby improving the model's sensitivity to small vessels. However, for vascular tree images such as those of the fundus vessels, even considering only skeletal lines, there is a significant difference in the number of skeletal points belonging to large and small blood vessels, thus limiting the effectiveness of this method.
[0064] The present application's embodiments take into account the high difficulty of accurately segmenting terminal blood vessels, primarily because these vessels constitute only a very small proportion of the entire vascular tree. Therefore, during model training, the segmentation performance of terminal blood vessels is difficult to reflect in the model loss. Thus, to improve the segmentation performance of terminal blood vessels, it is necessary to enhance the learning ability of the training process for terminal blood vessels. Therefore, the present application's embodiments, based on the unique tree structure of the vascular tree, improve the loss function by assigning different weights to different foreground positions when calculating the model loss.
[0065] Specifically, this application provides a model training and image processing method for vascular tree segmentation. In this method, during the training of the vascular tree segmentation model, the thickness of each vascular branch and the edge degree of each pixel within the vascular branch are combined with the annotation results of each sample image to generate training weights corresponding to each pixel. These training weights are negatively correlated with both thickness and edge degree; that is, the thinner the vascular branch containing each pixel, or the lower the edge degree of the pixel within the vascular branch, the higher the training weight corresponding to that pixel. This results in higher training weights for terminal vascular branches in the sample image, increasing the proportion of such vascular branches in the overall training data. Consequently, the training process becomes more sensitive to these vascular branches. Therefore, when the final vascular tree segmentation model is applied to real-world scenarios, its ability to perceive small vascular branches such as terminal vascular branches is significantly improved, thereby enhancing the segmentation effect of the terminal vascular portion of the vascular tree and improving the overall accuracy of vascular tree segmentation.
[0066] Furthermore, by adjusting the weight distribution, the proportion of the segmentation effect of the terminal part in the total loss is increased. This allows for the detection of more terminal blood vessels. However, the diameter of individual blood vessel branches in the prediction results may be significantly larger than the labeled results, and even incorrect connections between blood vessel branches may occur. Therefore, in order to improve the segmentation effect at the terminal blood vessels, local statistical feature constraints for the terminal blood vessels are added to the loss function. This combines the single-point prediction loss (i.e., the weighted prediction loss based on a single pixel) with the neighborhood prediction loss, thereby reducing the model's sensitivity to a small amount of labeled noise. This helps the model focus more on segmenting error-prone locations during training, thus improving the segmentation accuracy at the terminal blood vessels.
[0067] The following is a brief introduction to the application scenarios to which the technical solutions of the embodiments of this application are applicable. It should be noted that the application scenarios described below are only for illustrating the embodiments of this application and are not intended to limit the scope. In specific implementation, the technical solutions provided by the embodiments of this application can be flexibly applied according to actual needs.
[0068] The solution provided in this application can be applied to vascular tree segmentation scenarios, such as vascular tree segmentation based on fundus images, vascular tree segmentation based on echocardiogram images, or vascular tree segmentation based on other types of medical images. However, it should be noted that although this application primarily uses vascular tree segmentation as an example, the method is equally applicable to image types similar to vascular trees. In other words, this application can perform image segmentation on images with tree-like or vein-like structures to extract the corresponding tree-like or vein-like structures. For example, segmenting mountains based on remote sensing images, segmenting river veins based on aerial images, or segmenting leaf veins based on leaf images.
[0069] like Figure 1 The diagram shown is an application scenario provided by an embodiment of this application. In this scenario, a terminal device 101 and a server 102 may be included.
[0070] Terminal device 101 can be, for example, a mobile phone, tablet computer (PAD), laptop computer, desktop computer, smart TV, smart in-vehicle device, and smart wearable device. Terminal device 101 can be used to provide a target image to be segmented into a vascular tree and upload the target image to server 102 for vascular tree segmentation processing.
[0071] In one possible implementation, the terminal device 101 may provide a front-end application for providing target images. The front-end application has the functions of acquiring and uploading target images and presenting vascular tree segmentation results. The application may be a software client, or a web page, applet, or other client, without limiting the specific type of client.
[0072] Server 102 is used to implement the image segmentation process and provide backend services for front-end applications. For example, it can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, i.e., content delivery network (CDN), as well as big data and artificial intelligence platforms, but it is not limited to these.
[0073] It should be noted that the model training method or image processing method for vascular tree segmentation in this embodiment can be executed by the terminal device 101 alone, or by the server 102 and the terminal device 101 together. For example, when the processing capability of the terminal device 101 meets the requirements of model training or image processing, the aforementioned model training method for vascular tree segmentation can be executed to train the vascular tree segmentation model and obtain a trained vascular tree segmentation model, or the aforementioned image processing method for vascular tree segmentation can be executed to segment and obtain the vascular trees in the image. Alternatively, the above process can be implemented entirely by the server 102. Furthermore, for the model training process, the terminal device 101 can determine the training weights for each sample image and pass the obtained training weights to the server 102 so that the vascular tree segmentation model can be trained based on these training weights. This application does not make specific limitations here; the following mainly uses the example of the server 102 and the terminal device 101 jointly executing the process for illustration.
[0074] Both server 102 and terminal device 101 may include one or more processors, memory, and I / O interfaces for interaction. Furthermore, server 102 may be configured with a database to store sample image data required for the training process and model parameters obtained from training. When server 102 and terminal device 101 implement the above processes, their memory may also store the program instructions required for execution in the model training and image processing methods for vascular tree segmentation provided in this application embodiment. These program instructions, when executed by the processor, can be used to implement the model training and image processing processes for vascular tree segmentation provided in this application embodiment.
[0075] In this embodiment, the terminal device 101 and the server 102 can communicate directly or indirectly through one or more networks 103. The network 103 can be a wired network or a wireless network. For example, the wireless network can be a mobile cellular network or a Wireless-Fidelity (WIFI) network. Of course, it can also be other possible networks, and this embodiment does not limit them.
[0076] The vascular tree segmentation method provided in this application is mainly applied in the field of medical diagnosis. Here, vascular tree segmentation based on fundus images is used as an example; of course, vascular tree segmentation for other medical images can refer to this process. See also... Figure 2 The diagram shown is an architectural schematic of a vascular tree segmentation system based on fundus images provided in an embodiment of this application. The system includes a first terminal device 201, a server 202, a fundus camera 203, and a second terminal device 204.
[0077] The first terminal device 201 is the terminal device corresponding to the personnel deploying the vascular tree segmentation model. It is used to deploy the trained vascular tree segmentation model to the server 202, or to deploy the initial model of the vascular tree segmentation model to the server 202, and to provide sample images or specify sample images in the server 202 to initiate the training of the initial model and save the training results to the server 202, so that the server 202 can perform the vascular tree segmentation process based on the vascular tree segmentation model. The sample images involved here can be fundus photographs with the actual locations of the vascular trees already labeled.
[0078] Server 202 is used to provide backend services for vascular tree segmentation. After the vascular tree segmentation model is trained and deployed on server 202, fundus photos of patients can be taken by fundus camera 203. The fundus photos can be sent to server 202 so that server 202 can perform vascular tree segmentation of the fundus photos based on its own deployed vascular tree segmentation model, and transmit the obtained vascular tree segmentation results to the second terminal device 204 used by the doctor so that the doctor can combine the vascular tree segmentation results to diagnose diseases such as diabetic retinopathy.
[0079] Among them, fundus photographs can be, for example, wide-angle fundus photographs. Wide-angle fundus photographs have been widely used in the diagnosis of diabetic retinopathy. The peripheral lesions observed in the wide field of view can provide more sufficient evidence for the staging of diabetic retinopathy. Lesions on fundus vessels, such as intraretinal microvascular abnormality (IRMA) and new vessels (NV), are the most important features for distinguishing between the proliferative and non-proliferative phases. Therefore, the method provided by the embodiments of this application can achieve more accurate segmentation of terminal vessels, thereby accurately quantifying the morphology of terminal vessels, further assisting in screening branches of potential lesions, and realizing the detection function of vascular lesions.
[0080] It should be noted that in real-world scenarios, if the computing power of the device allows, some of the aforementioned devices can be implemented using the same device. For example, the functions of server 202 and fundus camera 203 can be implemented using the same device, or the functions of fundus camera 203 and second terminal device 204 can be implemented using the same device. This application embodiment does not impose any restrictions on this.
[0081] The following describes the model training and image processing method for blood vessel tree segmentation provided by the exemplary embodiments of this application, in conjunction with the application scenarios described above and with reference to the accompanying drawings. It should be noted that the above application scenarios are only shown to facilitate understanding of the spirit and principles of this application, and the embodiments of this application are not limited in any way in this respect.
[0082] Since the vascular tree segmentation process in this application embodiment needs to be based on the vascular tree segmentation model, and the vascular tree segmentation model needs to be trained before it can be used, the training process of the vascular tree segmentation model used in this application embodiment will be introduced before introducing the image processing method for vascular tree segmentation in this application embodiment.
[0083] See Figure 3 The diagram shown is a flowchart illustrating a model training method for blood vessel tree segmentation provided in an embodiment of this application. The process includes the following steps:
[0084] Step 301: Obtain a set of sample images for model training. Each sample image in the set contains a vascular tree and is associated with a label that represents the true location of the vascular tree in the sample image.
[0085] In this embodiment of the application, the sample image set contains multiple sample images required for model training, and each sample image is associated with a corresponding annotation result.
[0086] In practical scenarios, the annotation results can be implemented using different representation methods, and this application embodiment does not limit the representation method. For example, the annotation result can be a mask image of the vascular tree in the sample image, that is, an image that only contains pixels at the location of the vascular tree. Alternatively, the annotation result can also be represented using positional information, that is, indicating which pixels in the sample image belong to the vascular tree. Or, the annotation result can also be in the form of a numerical matrix of the same size as the sample image, where each value in the numerical matrix corresponds to a pixel in the sample image, and the value represents the true probability value of the pixel being a vascular pixel. For example, when the value is 1, it represents that the pixel is a vascular pixel, and when the value is 0, it represents that the pixel is a non-vascular pixel. Of course, other values can also be used for representation, and this application embodiment does not limit this.
[0087] Specifically, sample images can be obtained from existing public datasets. Sample images in public datasets are usually images that have already been annotated at the pixel level. By obtaining sample images from public datasets, there is no need to perform the annotation process again, thus saving time and costs.
[0088] In addition, considering that model training requires a large number of training samples to support the accuracy of the trained model, if the number of samples in the public dataset is insufficient for training, images from real-world scenarios can be obtained and annotated with vascular trees to serve as sample images for model training. For example, with the authorization of the dataset owner, images from a private dataset in a real-world scenario can be used for annotation to obtain sample images.
[0089] Specifically, the annotation process for each image can be combined with computer vision capabilities. Computer vision technology can be used to identify different image regions within each image. Each image region can be a connected region composed of identical or similar pixels. Annotators can then selectively choose image regions belonging to the vascular tree, eliminating the need for pixel-by-pixel selection and improving the efficiency of sample image annotation. Alternatively, computer vision technology can be used to initially identify vascular tree regions in each image. Annotators can then refine these identifications to obtain more accurate annotation results. Similarly, this method also eliminates the need for pixel-by-pixel selection, further improving the efficiency of sample image annotation.
[0090] In this embodiment of the application, considering that labeling sample images collected in real-world scenarios is extremely time-consuming, it may take a lot of time to prepare the model in order to obtain a large number of sample images. Therefore, in order to obtain new sample images by transforming the labeled sample images, the workload of labeling can be reduced and the time spent on preliminary preparation can be increased.
[0091] Specifically, image transformation methods include, but are not limited to, perspective transformation, affine transformation, image movement, scaling, rotation, adding noise, or image stitching. Of course, other possible image transformation methods can also be included, thereby quickly and easily increasing the number of samples required for training.
[0092] In this embodiment, based on the application scenario of the vascular tree segmentation model, sample images in the corresponding scenario can be collected specifically. For example, when applying vascular tree segmentation to a specific body part, images of that body part are collected for training. For instance, when performing vascular tree segmentation on fundus images, fundus images can be collected specifically; or, when performing segmentation of cardiac vascular trees, cardiac images can be collected specifically. Furthermore, this embodiment does not limit the image acquisition method; for example, it can be a wide-angle fundus photograph taken by a wide-angle fundus camera, computed tomography (CT) medical images, MRI medical images, or ultrasound medical images, etc.
[0093] Step 302: Based on the annotation results of each sample image, perform attribute recognition processing of blood vessel branches to obtain the attribute information of each blood vessel branch in each sample image. The attribute information includes: the thickness of the corresponding blood vessel branch, and the edge degree of the pixels included in the corresponding blood vessel branch in the blood vessel branch.
[0094] In this embodiment, vascular tree segmentation is a process of dividing pixels in a sample image into vascular pixels and non-vascular pixels. Vascular pixels are the foreground parts to be segmented in the sample image, while non-vascular pixels are the corresponding background parts. In other words, vascular tree segmentation can also be considered a binary classification process. During the specific segmentation process, considering that the recognition level of vascular branches of different thicknesses is different, thicker vascular branches are obviously easier to segment due to their larger area, while segmenting small vessels such as terminal vessels is relatively difficult. Therefore, to improve the learning level of small vessels, different weights can be assigned to different foreground positions to increase the training weight of small vessels.
[0095] For details, see Figure 4 The diagram illustrates the calculation process of training weights provided in this embodiment. Factors influencing the identification of small blood vessels can include the thickness of the blood vessel branches and the distance between pixels within the small blood vessel and the edge of its skeletal line. The edge distance measures the distance of a pixel from the skeletal line. Therefore, different training weights can be assigned to different pixels based on these two factors. Of course, other possible influencing factors may also be included in real-world scenarios, and this embodiment does not impose any limitations on this.
[0096] Therefore, since the annotation results of the sample images truly reflect the information of the location of the vascular tree, after annotating the sample images, the attribute recognition processing of the vascular branches can be performed on each sample image based on the annotation results of each sample image, thereby obtaining the attribute information of each vascular branch in each sample image, namely the thickness of the vascular branch mentioned above, and the edge degree of the pixels included in the vascular branch, etc.
[0097] Since the process of determining attribute information is similar for each sample image, we will use one sample image as an example for explanation.
[0098] In this embodiment of the application, the attribute information can be determined based on the annotation results of the sample image. The following is an example of the attribute information including the two parameters of thickness and edge degree mentioned above, which will be described separately.
[0099] In one possible implementation, the thickness of a vascular branch can be measured by the diameter of that branch, thereby determining the diameter of each branch by using the vascular tree marked in the annotation results, thus measuring the thickness of the vascular branch.
[0100] Specifically, for each blood vessel branch in each sample image, theoretically, when the lengths are equal, the larger the area occupied, the larger the diameter of the blood vessel branch; similarly, when the area occupied is equal, the shorter the length, the larger the diameter of the blood vessel branch. Therefore, the diameter of the blood vessel branch is positively correlated with the area and negatively correlated with the length, so its diameter can be determined based on the area and length occupied by the blood vessel branch.
[0101] In practical applications, one approach is to classify the thickness levels of blood vessel branches, which is equivalent to classifying them according to different diameters. Different thickness levels correspond to different areas and length ranges. Therefore, based on the range of area and length occupied by each blood vessel branch in the annotation results, a blood vessel branch can be classified into a thickness level to measure the thickness of different blood vessel branches.
[0102] Alternatively, the diameter of a blood vessel branch can be measured by the ratio of its area to its length. Therefore, for a given sample image, the diameter of each blood vessel branch can be obtained based on the size of the image region occupied by each branch and the length of its skeletal line, thus characterizing the thickness of the branch.
[0103] In this context, skeletal lines represent the orientation of corresponding vascular branches. They are lines with the same shape as the branch after thinning a vascular branch. Skeletal lines can be extracted using skeleton extraction algorithms. Specifically, these algorithms extract the skeleton of the annotated vascular tree in the sample image to obtain the skeleton of the vascular tree, which contains the skeletal lines of each vascular branch. See also... Figure 5 The diagram shown is a schematic diagram of the skeletal lines provided in the embodiment of this application. Here, only a part of the blood vessel tree is used as an example. The black and white boxes each represent a blood vessel pixel. The white boxes represent the skeletal points of the blood vessel tree. The connection of each skeletal point forms the skeletal lines of the blood vessel branches. The combination of all the skeletal lines is the skeleton of the blood vessel tree. It can be seen that in this local blood vessel tree, there are two blood vessel bifurcation points. The skeleton of this local blood vessel tree is consistent with the overall direction of the blood vessel tree.
[0104] Therefore, when determining the diameter of a blood vessel branch, it can be based on the size of the area occupied by the blood vessel branch in the image, combined with the length of the skeletal line of the blood vessel branch.
[0105] The size of the area occupied by a blood vessel branch in an image can be characterized by the total number of branch pixels it contains, referred to here as the first total number. Similarly, the length of a skeletal line can also be determined by the total number of skeletal line pixels contained within the skeletal line, referred to here as the second total number. The diameter of the blood vessel branch can then be obtained based on the ratio between the first total number and the second total number.
[0106] In one possible implementation, the diameter of the vascular branch can be expressed as:
[0107]
[0108] c i N represents the average diameter of the blood vessel branch containing the i-th pixel. pixel N represents the total number of branch pixels contained in this blood vessel branch. skel This represents the total number of skeletal line pixels contained in the skeletal line of the blood vessel branch.
[0109] In one possible implementation, in real-world scenarios, blood vessels typically become progressively smaller with increasing branching levels. Therefore, the thickness of a blood vessel branch can also be measured by its branching level. This allows the branching level of each blood vessel branch to be determined using the annotated blood vessel tree in the annotation results, thus measuring the thickness of the branch. The branching level can be determined based on the number of times the blood vessel branch bifurcates; more bifurcations result in a higher score, indicating a lower thickness, i.e., a thinner blood vessel branch.
[0110] In this embodiment, in addition to the diameter of the blood vessel, the weight can also be attenuated based on the distance from each foreground point (i.e., the pixel point belonging to the blood vessel branch) to the bone point, that is, based on the edge degree of each foreground point in the blood vessel branch.
[0111] Specifically, the edge-ness of each pixel within a blood vessel branch can be characterized by the distance of each pixel from the center of the blood vessel branch; the greater the distance, the closer the pixel is to the edge of the blood vessel branch. Therefore, for a blood vessel branch in a sample image, the skeletal points of the blood vessel tree can be extracted based on the annotation results of the sample image. Then, based on the positional information of each pixel within each blood vessel branch in the sample image, and the positional information of the skeletal line of the corresponding blood vessel branch, the distance between each pixel and the skeletal line of the corresponding blood vessel branch can be obtained.
[0112] In one possible implementation, the distance between a pixel and the skeletal line of the corresponding blood vessel branch can be expressed as:
[0113]
[0114] Where, d i This represents the distance between the i-th pixel and the nearest bone point in the skeletal line of the blood vessel branch to which it belongs, (x i y i Let (x0, y0) represent the coordinates of the i-th pixel, and (x0, y0) represent the coordinates of the nearest bone point to the i-th pixel. It should be noted that d here... i Specifically, the example uses Euclidean distance, but in real-world scenarios, other possible methods can also be used for distance calculation, and this application does not limit this approach.
[0115] Step 303: Based on the attribute information corresponding to each sample image, obtain the training weights corresponding to each pixel in the corresponding sample image. The training weights are negatively correlated with the thickness and edge degree.
[0116] In this embodiment, attribute information obtained from the annotation results of sample images is used to assign different weights to each pixel in the sample image. The main purpose is to improve the learning degree of terminal blood vessels during the training process, thereby making the trained blood vessel tree segmentation model more sensitive to the segmentation of terminal blood vessels and improving the segmentation effect of terminal blood vessels. Specifically, in order to improve the blood vessel tree segmentation model's attention to terminal blood vessels, higher weights can be assigned to finer blood vessel branches. At the same time, in order to reduce the impact of area imbalance caused by the diameter of the vessel, the weights at different positions are further attenuated along the direction from the center line to the edge. This method can, to a certain extent, ensure the segmentation effect of the main blood vessel part. In short, the finer the blood vessel, the higher its training weight, and the higher the edge degree, the lower the training weight. That is, the training weight is negatively correlated with the thickness and edge degree.
[0117] Therefore, in this embodiment of the application, for each pixel in each sample image, taking the diameter of a blood vessel branch as an example to represent its thickness, a first sub-weight based on the diameter of the blood vessel branch can be assigned to a pixel. Specifically, for a pixel, a weight mapping process based on the diameter of the blood vessel branch to which the pixel belongs and a preset diameter threshold can be performed to obtain the first sub-weight of the pixel based on the diameter. The idea behind the weight mapping process is to make the obtained first sub-weight negatively correlated with the diameter.
[0118] Since the embodiments of this application primarily aim to enhance the weight of terminal blood vessels, the preset diameter threshold can be the upper limit of the diameter of terminal blood vessels, encompassing finer tertiary blood vessels and those below tertiary level, which is also the range of small blood vessels referred to in the embodiments of this application. This upper limit of the diameter can be set based on experience, through experimental procedures, or flexibly based on the diameter of each blood vessel branch in each sample image. For example, it can be the upper limit of the diameter of tertiary blood vessel branches in the sample image, or set according to the proportion of the diameter values of blood vessel branches, etc. Furthermore, the preset diameter threshold is also directly related to the resolution of the sample image. The higher the resolution of the sample image, the more pixels the blood vessel branches should contain; therefore, the preset diameter threshold will increase accordingly with higher resolution.
[0119] Specifically, this application embodiment designs a linearly decreasing weight function based on the diameter of the blood vessel branches. This function assigns higher weights to finer branches, and the function is expressed as follows:
[0120]
[0121] in, c is the first sub-weight of the i-th pixel. set This is the preset pipe diameter threshold.
[0122] In this embodiment, considering that the main trunk vessels account for the majority of the proportion, no additional weight is required. Therefore, the category of vessel branches can be considered when assigning weights. The main trunk vessels can refer to other vessels that are not terminal vessels, or vessels with a diameter greater than a certain threshold (i.e., thicker vessels), or they can be defined according to branch level (i.e., vessels corresponding to smaller branch levels).
[0123] Specifically, when calculating the first sub-weight for a single pixel, the blood vessel type can be determined based on the diameter or branch level of the blood vessel branch containing the pixel, and then a corresponding first sub-weight can be assigned based on the blood vessel type. See also Figure 4 As shown, if the blood vessel type is a terminal blood vessel, the first sub-weight of the pixel is obtained based on the ratio between the diameter of the blood vessel branch where the pixel is located and a preset diameter threshold, i.e., the corresponding first sub-weight is calculated according to the weight function described above; if the blood vessel type is a non-terminal blood vessel, the first sub-weight of the pixel is set to zero, i.e., the first sub-weight of the large blood vessel is zero. Set it to 0.
[0124] In this embodiment, in addition to the diameter of the blood vessel, the weight attenuation is further set based on the distance from the foreground point to the bone point. That is, for each pixel in each sample image, taking the distance between the pixel and the nearest bone point to represent its edge degree as an example, the pixel can be given a second sub-weight based on the distance between the pixel and the nearest bone point.
[0125] Specifically, for a pixel, a distance-based weight mapping can be performed on the distance between the pixel and the skeletal line of the blood vessel branch to which it belongs, i.e., the distance between the pixel and the nearest skeletal point, and a preset distance threshold to obtain the second sub-weight of the pixel based on distance.
[0126] In one possible implementation, the second sub-weight can be calculated using the following formula:
[0127]
[0128] in, The second sub-weight representing the i-th pixel, d max A preset distance threshold is used, which is not less than the maximum distance in the sample image. For example, it can be the maximum distance in the current sample image, thus mapping the second sub-weight to the range of 0 to 1. The weight decay in this way is mainly located at the edge of the main blood vessel, while the weights at small blood vessels are fully retained. This method suppresses the edge of large blood vessels to improve the attention to terminal blood vessels during training, thereby solving the problem of class imbalance between large and small blood vessels.
[0129] See Figure 4 As shown, after obtaining the first sub-weight Second sub-weight Then, a weighted process can be performed based on the two weights, so that the training weight of each pixel can be obtained by superimposing the two weight decay methods, which is used for the weight of each pixel when calculating the segmentation loss of the subsequent blood vessel tree segmentation model.
[0130] In one possible implementation, the two weights can be combined to obtain the training weights for each pixel in the following way:
[0131] w i =(kw diam +1)w skel
[0132] Among them, w i The training weights represent the i-th pixel, and k is the first sub-weight. The weight parameters.
[0133] It should be noted that during training, the training weights can be applied only to the foreground points, while the background points do not need to be given much attention. Therefore, for the background points, their training weights can be directly set to the initial weights, such as 1 or other possible values. Thus, the above process of calculating training weights can be performed only on the foreground points, thereby improving the efficiency of training weight calculation.
[0134] Step 304: During the training of the blood vessel tree segmentation model, the model loss value is determined based on the prediction results, annotation results and corresponding training weights of the blood vessel trees in each sample image, and the model parameters are adjusted based on the model loss value.
[0135] Specifically, the model training process is an iterative optimization process. The purpose of iterative optimization is to make the model's prediction results gradually approach the true values. The model loss value is an indicator parameter that measures the difference between the current predicted value and the true value. In this embodiment, the loss of the terminal blood vessel portion is increased during loss calculation to enhance the blood vessel tree segmentation model's focus on the terminal blood vessel portion. Therefore, in each training process, when obtaining the prediction results of each sample image through the blood vessel tree segmentation model, the current model loss value can be determined based on the prediction results, annotation results, and corresponding training weights, and then model parameter tuning and optimization can be performed based on the model loss value.
[0136] In one possible implementation, the model loss value based on the training weights can be calculated using the following loss function:
[0137]
[0138] Among them, l seg p represents the single-point prediction loss value for a single pixel. i Let y be the predicted value of the i-th pixel in the prediction result of the sample image. i Let r be the annotation value of the i-th pixel in the annotation results of the sample image, where N is the total number of pixels, and r is the value of the pixel. l It is a hyperparameter that helps the training process focus more on the misclassified foreground parts.
[0139] It should be noted that the above loss function can also obtain similar results based on the weighting of other basic loss functions. For example, the training weights of this application embodiment can be added to the cross-entropy loss function or the focal loss function. This application embodiment does not limit this.
[0140] See Figure 6The diagram shown illustrates a training process for the vascular tree segmentation model provided in this embodiment. This training process involves multiple rounds of iterative training of the vascular tree segmentation model based on the aforementioned sample image set. Each round of training is similar; therefore, this section primarily uses a single round of training as an example.
[0141] Step 601: Using the vascular tree segmentation model used in this round, perform vascular tree segmentation processing based on image visual detection on each of the obtained sample images to obtain the prediction results representing the predicted positions of vascular trees in the sample images.
[0142] In this embodiment, during each training round, the blood vessel tree segmentation model used in this round can be used to perform forward prediction on each input sample image to obtain the corresponding prediction result. Specifically, in the first training round, the blood vessel tree segmentation model used in this round is the initial model; in subsequent training rounds, the blood vessel tree segmentation model used in this round is the model after parameter adjustment in the previous round.
[0143] Specifically, for each sample image, the vascular tree segmentation model uses an artificial neural network to simulate human vision to perform image visual detection on the sample image, in order to identify which parts of the sample image belong to the vascular part and which parts belong to the non-vascular part, so as to achieve vascular tree segmentation of the sample image and obtain the image segmentation result of each sample image.
[0144] In practical applications, any possible deep neural network model can be selected as the blood vessel tree segmentation model in this embodiment of the application. For example, WingsNet can be selected as the network used for blood vessel segmentation. The WingsNet structure was originally used for 3D tracheal tree segmentation. To make it more suitable for 2D image segmentation, in practical applications, the overall structure of WingsNet can be kept unchanged, only replacing all 3D kernels with 2D kernels. Of course, other possible image segmentation models can also be used in real-world scenarios, and this embodiment of the application does not impose any limitations on this.
[0145] The prediction results are similar to the annotation results, and can be implemented using different representation methods. This application does not limit the representation methods used. For example, the prediction result can be a mask image of the vascular tree in the sample image, i.e., an image containing only the pixels at the location of the vascular tree. Alternatively, the prediction result can be represented using location information, indicating which pixels in the sample image belong to the vascular tree. Or, the prediction result can be in the form of a numerical matrix of the same size as the sample image, where each value in the numerical matrix corresponds to a pixel in the sample image, and the value represents the predicted probability that the pixel is a vascular pixel. Unlike the annotation value, which has two possibilities (vascular pixel or non-vascular pixel), the predicted probability value represents the likelihood that a pixel is a vascular pixel; for example, 90% indicates a high probability that it is a vascular pixel.
[0146] Step 602: Based on the training weights of each pixel in each sample image, perform weighted processing on the labeled values of each pixel in the annotation results and the predicted values of the corresponding pixels in the prediction results to obtain the model loss value.
[0147] The process of step 602 is similar to the process of calculating the model loss value in step 304 above. Therefore, please refer to the description of the corresponding part in step 304 above, and it will not be repeated here.
[0148] Step 603: Determine whether the iteration termination condition is met.
[0149] Specifically, the iteration termination condition may include at least one of the following conditions:
[0150] (1) The model loss value is not greater than the preset loss value threshold. The model loss value represents the accuracy of the model to a certain extent. That is, when the accuracy of the model is sufficient, iterative training can be stopped.
[0151] (2) The number of training sessions has reached the maximum number of training sessions.
[0152] If the result of step 603 is yes, that is, the training of the representation model is complete, the training process ends, and the vascular tree segmentation model can be used in the subsequent vascular tree segmentation process.
[0153] Step 604: If the result of step 603 is negative, then the parameters of the blood vessel tree segmentation model used in this round are tuned based on the model loss value, and the next round of training process is started, i.e., jump to step 601 to continue execution.
[0154] Specifically, the model parameter tuning process can be carried out using optimization algorithms, such as batch stochastic gradient descent, or other possible model optimization methods, such as gradient descent or Newton's method.
[0155] In this embodiment, during the above process, the weight distribution of each pixel in the image is adjusted to increase the proportion of the terminal blood vessel segmentation effect in the total loss. To improve the sensitivity to terminal blood vessels, it is necessary to increase the weight. However, this method often reduces accuracy to some extent while improving sensitivity. This means that while detecting more terminal branches, the diameter of the predicted blood vessel branches may be thicker than the labeled result, and even incorrect connections between blood vessel branches may occur. Incorrect connections will directly change the local topology and have a significant impact on subsequent quantitative analysis. Therefore, to improve the segmentation effect at terminal blood vessels, this embodiment also adds a constraint term for terminal blood vessels to the loss function. Through region constraints, more attention is paid to segmenting error-prone locations during training.
[0156] See Figure 7 The diagram shown illustrates another training process for the blood vessel tree segmentation model provided in this embodiment. Similarly, this training process will primarily focus on a single training round as an example.
[0157] Step 701: Obtain the sample image set for model training.
[0158] Step 702: Based on the annotation results of each sample image, the pixels located on the skeleton line in each sample image are used as reference points.
[0159] In this embodiment of the application, in order to reduce the probability of blood vessel segmentation errors, region constraints are applied to the sample images. In order to improve the prediction results of the blood vessel branches becoming thicker or even having incorrect connections due to the increase in sensitivity, this embodiment of the application selects the pixel points on the skeletal line, that is, the skeletal points of the blood vessel branches, as reference points to apply region constraints to these reference points.
[0160] The location of the vascular tree has been marked in the annotation results. Therefore, a skeleton extraction operation can be performed on the annotation results of each sample image to obtain the skeletal line of the vascular tree in the sample image. Then, each pixel point located on the skeletal line can be used as a reference point.
[0161] In one possible implementation, since the terminal blood vessels are more prone to errors in practical applications, the region constraint can be applied specifically to the terminal blood vessels. Furthermore, applying region constraints only to the terminal blood vessels can further reduce the computational load required for training and improve model training efficiency. Therefore, for each sample image involved in this round, the pixels located on the skeletal lines of the terminal blood vessels in each sample image can be used as reference points; that is, only the skeletal points of the terminal blood vessels need to be selected as reference points for region constraint.
[0162] It should be noted that in this embodiment, skeletal points are mainly selected as reference points. This is based on the consideration of segmenting edges that are prone to errors, and on the other hand, it can reduce the number of selected skeletal points and improve the efficiency of the entire training process. However, in practical applications, other pixels or more pixels can also be selected as reference points. This embodiment does not limit this.
[0163] Step 703: Extract regional features from the local area where each benchmark point is located to obtain the labeled neighborhood features corresponding to each benchmark point.
[0164] In this embodiment, the local region where each reference point is located can refer to a local region extracted with respect to the reference point and a kernel of a fixed size. Since this local region contains only the neighboring pixels of the reference point, it can also be called the neighborhood of the reference point. The kernel of a fixed size can be set according to the needs of the actual scene, such as the resolution of the sample image or the granularity requirements of the constraints. This embodiment does not limit this.
[0165] See Figure 8 The diagram shown is a schematic representation of a local area provided in an embodiment of this application. Each white square represents a skeletal point. The diagram uses three skeletal points as an example to illustrate their corresponding local areas. For example, skeletal point 1, with a fixed-size kernel of 9x9, and R... s Representing a neighborhood centered on a bone point, we can extract a 9x9 local region Rs1 centered on bone point 1. Similarly, we can extract 9x9 local regions Rs2 and Rs3 centered on bone point 2 and bone point 3 respectively.
[0166] Specifically, for each sample image, after determining the reference point in a sample image, the local region corresponding to each reference point can be determined in the sample image based on a preset local region size, with each reference point as the center. Based on the annotation values of the pixels included in each local region in the annotation results of the sample image, the annotation neighborhood features of the reference point corresponding to each local region are obtained. The annotation neighborhood features represent the distribution in the neighborhood of a reference point. Subsequently, after obtaining the prediction results through the model, the accuracy of the model prediction can be measured by comparing with the neighborhood region, thereby achieving regional constraints.
[0167] In practical applications, the labeled neighborhood features can include the first mean and first standard deviation of the labeled values of the pixels within the local region. It should be noted that the first and second values mentioned here are only used to represent different feature values in the labeling and prediction stages, and are not limited by any specific order. Therefore, for each skeletal point of the terminal blood vessel in the sample image, its neighborhood is extracted using a kernel of fixed size, and the mean μ and standard deviation σ are calculated within the neighborhood.
[0168] In one possible implementation, the mean and standard deviation of each local region can be achieved using an average pooling layer. Then, a label value matrix composed of the label values of pixels within each local region can be extracted from the labeling results of each local region. Average pooling is then performed on each label value matrix to obtain its first mean. Furthermore, average pooling is performed on each label value matrix based on its own basic product operation result. Finally, based on the pooling results and the difference between the corresponding matrix means, the first standard deviation of each label value matrix is obtained.
[0169] Specifically, the mean μ and standard deviation σ can be calculated as follows:
[0170] μ = AvgPool(R) s )
[0171]
[0172] Among them, R s AvgPool represents the labeled value matrix that constitutes a local region, and represents the average pooling layer. It is the basic product, also known as the Hadamard product.
[0173] See Figure 8 As shown, extracting Rs1 yields the labeled value matrix shown on the right, where 0 represents non-vascular regions and 1 represents vascular regions. The mean μ of Rs1 is 0.44, and the corresponding standard deviation σ can be obtained similarly.
[0174] Step 704: Using the vascular tree segmentation model used in this round, perform vascular tree segmentation processing based on image visual detection on each of the obtained sample images to obtain the prediction results representing the predicted positions of vascular trees in the sample images.
[0175] The process of step 704 is similar to that of step 601 above, so please refer to the description of step 601 above, and it will not be repeated here.
[0176] In this embodiment, considering that the resolution of the input image may be large, and in order to improve the accuracy of the blood vessel tree segmentation model, for the image input to the blood vessel tree segmentation model (including sample images and target images in the application stage), local images can be sampled from the sample images in sequence according to the preset sliding window size and sliding step size, and the obtained local images are sequentially input into the blood vessel tree segmentation model to obtain the sub-prediction results of each local image output by the blood vessel tree segmentation model. Based on the overlapping parts in each local image, the obtained sub-prediction results are merged to obtain the prediction result of the sample image.
[0177] See Figure 9 As shown, taking the sliding window method as an example, with an input size of 1024×1024 and a sliding distance of 512, the first sampling can obtain a local image of size 1024×1024, that is... Figure 9 As shown in section 1 of the sliding window, the second sample can obtain a local image of size 1024×1024, i.e. Figure 9 The sliding window 2 shown, the overlapping part between the two, see [reference needed]. Figure 9 As shown, after inputting sliding window 1 and sliding window 2 into the blood vessel tree segmentation model, we can obtain their respective sub-prediction results. Furthermore, for the overlapping part of the two, we can obtain two prediction values, or even more prediction values. Then, we can merge the multiple prediction values of the overlapping part to obtain the final prediction result.
[0178] The method of merging multiple predicted values can be, for example, by averaging or by selecting the median or maximum value among the multiple predicted values. This application does not limit this method.
[0179] Step 705: Based on the training weights of each pixel in each sample image, perform weighted processing on the labeled values of each pixel in the annotation results and the predicted values of the corresponding pixels in the prediction results to obtain the single-point prediction sub-loss.
[0180] The content of step 705 can be found in the description of the corresponding part in step 304 above, and will not be repeated here.
[0181] Step 706: Based on the prediction results of each sample image, extract the regional features of the local area where each reference point is located, and obtain the predicted neighborhood features corresponding to each reference point in each sample image.
[0182] Specifically, similar to the process of obtaining the labeled neighborhood features, the predicted neighborhood features are obtained from the corresponding predicted values of the local regions where each reference point is located in the prediction results. Therefore, the calculation process will not be elaborated further.
[0183] It should be noted that the local region here refers to the neighborhood extracted based on the annotation results, and the neighborhood will not be extracted again during the prediction stage.
[0184] Step 707: Accumulate the difference between the predicted neighborhood features and the corresponding labeled neighborhood features of each reference point in each sample image to obtain the region prediction sub-loss.
[0185] After obtaining the predicted neighborhood features of each benchmark point, the differences in the features of the remaining labeled neighborhoods can be compared to measure the accuracy of the current model.
[0186] Specifically, the predicted neighborhood features can also include the second mean and second standard deviation of the predicted values of the pixels included in the local region. Then, a weighted summation process can be performed based on the differences between the first and second means of each reference point in each sample image, and the differences between the first and corresponding second standard deviations, to obtain the region prediction sub-loss.
[0187] In one possible implementation, the region prediction sub-loss can be expressed as:
[0188]
[0189] Where, N s This represents the number of skeletal points belonging to the terminal blood vessels in a sample image. The mean obtained during the prediction phase, i.e., the second mean mentioned above, represents the value of the mean. The mean obtained during the characterization and annotation phase, i.e., the first mean mentioned above, The standard deviation obtained during the characterization and prediction phase, i.e., the second standard deviation mentioned above, The standard deviation obtained during the characterization and labeling stage is the first standard deviation mentioned above. γ is a model hyperparameter used to balance the weights of the mean and standard deviation.
[0190] It should be noted that the embodiments of this application mainly take the L1 loss of prediction and labeling of all end bone point positions as an example, but in practical applications, other loss functions can also be used, and the embodiments of this application do not limit this.
[0191] Step 708: Perform weighted processing based on the single-point prediction sub-loss and the regional prediction sub-loss to obtain the model loss value.
[0192] In this embodiment, the total loss value of the model in this training round is obtained by combining the single-point prediction sub-loss and the region prediction sub-loss. The single-point prediction sub-loss tends to measure the segmentation effect of a single pixel, while the region prediction sub-loss tends to represent the segmentation quality of each bone point by the comprehensive segmentation performance of the neighborhood, thus dealing with the noise introduced during the annotation process. Here, because many terminal blood vessels cannot be clearly observed during the annotation process due to the influence of image quality and surrounding lesions, the labeling of these parts based on past experience inevitably introduces some noise. If the loss is calculated only on a single-point basis, it is easily affected by annotation noise. In contrast, the statistical features within the neighborhood are not sensitive to a small amount of annotation noise, which can help the network focus more on the location of segmentation errors rather than the location of labeling errors during training, thereby improving the accuracy of the model.
[0193] In one possible implementation, the final joint loss function is:
[0194] l Total =l seg +αl reg
[0195] Here, α is a weighted hyperparameter used to control the degree of influence of the regional prediction sub-loss.
[0196] Step 709: Determine whether the iteration termination condition is met. If it is met, the training process is complete.
[0197] If the result of step 709 is yes, that is, the training of the representation model is complete, the training process ends, and the vascular tree segmentation model can be used in subsequent vascular tree segmentation processes. Since the above process adds regional constraints for bone points, it can effectively focus more on the location of segmentation errors during training, thereby improving the accuracy of the vascular tree segmentation model.
[0198] Step 710: If the result of step 709 is negative, then the parameters of the blood vessel tree segmentation model used in this round are tuned based on the model loss value, and the next round of training process is started, i.e., jump to step 704 to continue execution.
[0199] Through the above training process, a blood vessel tree segmentation model that can be used in real-world scenarios can be obtained. Therefore, the following describes the blood vessel tree segmentation process based on the blood vessel tree segmentation model obtained by the above training method, as provided in the embodiments of this application. See also Figure 10 The diagram shown is a flowchart illustrating the image processing method for vascular tree segmentation provided in an embodiment of this application.
[0200] Step 1001: Obtain the target image for vascular tree segmentation.
[0201] Step 1002: Based on the blood vessel tree segmentation model obtained by the above model training method, perform blood vessel tree segmentation processing on the target image based on image visual detection, and output the corresponding blood vessel tree segmentation results; wherein, the blood vessel tree segmentation results are used to indicate the position of the blood vessel tree in the target image.
[0202] The vascular tree segmentation process of this application embodiment can be applied to any scenario involving vascular tree segmentation, such as fundus vascular tree segmentation or cardiac vascular tree segmentation. Taking fundus vascular tree segmentation as an example, wide-angle fundus images have been widely used in the diagnosis of diabetic retinopathy. Peripheral lesions observed in a wide field of view can provide more sufficient evidence for the staging of diabetic retinopathy. Furthermore, lesions on fundus vessels, such as IRMA, are the most important feature distinguishing between proliferative and non-proliferative phases. Therefore, the precise terminal vessel segmentation method provided in this application embodiment helps to quantitatively analyze the morphology of terminal branches, further screen out branches with potential lesions, and realize the function of vascular lesion detection. See also Figure 11 The image shown is a schematic diagram of the fundus vascular tree in a wide-angle fundus photograph provided in this application embodiment. Segmenting the vascular tree can help doctors more easily observe lesion areas, such as... Figure 11 The area marked by the middle circle is where IRMA lesions appear, allowing doctors to determine the stage of diabetic retinopathy by combining the vascular tree extracted from wide-angle fundus images.
[0203] In summary, in the embodiments of this application, different training weights are assigned to different foreground positions when calculating the model loss. These training weights are directly related to the diameter of the blood vessel branches and their distance from the skeletal line. Higher weights are assigned to finer branches, and the weights for different positions are further attenuated along the direction from the center line to the edge. This can improve the learning degree of terminal blood vessels during the training process to a certain extent, while also ensuring the segmentation effect of the main blood vessel section.
[0204] In addition, additional constraints were imposed on the segmentation performance of terminal blood vessels. These constraints mainly focus on the statistical features of the region adjacent to the terminal blood vessels. The constraints on the statistical features provide a general direction for the optimization of this part. Compared with constraining each point, constraining the overall performance within a neighborhood has a certain ability to resist noise in the label. This allows the training process to focus more on those regions with poor overall performance, while the error caused by label noise will be ignored to a certain extent, thereby improving the detection length and segmentation accuracy of the terminal blood vessel part in blood vessel segmentation.
[0205] In this embodiment of the application, the effectiveness of the method provided by this embodiment of the application is verified by a specific application in the scenario of fundus vascular tree segmentation. The final conclusion is that, compared with the solutions of related technologies, the method provided by this embodiment of the application can significantly improve the segmentation effect of vascular tree.
[0206] Specifically, in practical applications, two datasets were used for model training. One was the public dataset PRIME, which contains multiple wide-angle fundus images and corresponding pixel-level annotations. Each fundus image was taken with a wide-angle fundus color camera, with an image resolution of 4000×4000. The other dataset was BUVS, which contains 80 fundus images collected from 70 authorized patients, with an image resolution of 3072×3900. The vascular portions of the images were delineated by an experienced ophthalmologist and then reviewed by another ophthalmologist, who supplemented any missing terminal vessels. In the BUVS dataset, c set Setting it to 4, in the PRIME dataset, due to the presence of a large amount of black background and the lower resolution of the fundus region compared to BUVS, c set Set it to 3.
[0207] During model training, WingsNet was chosen as the network for blood vessel segmentation, with all 3D kernels replaced by 2D kernels. Since the PRIME dataset contains only 15 examples, it was used as the test set, while the 80 images from BUVS were divided into a training set (50 examples), a validation set (10 examples), and a test set (20 examples). The training set was used for model training, and the hyperparameters were determined based on the results from the validation set. Specifically, R... s The model is represented by an 11×11 convolutional kernel with α = 0.3. The input to the vascular tree segmentation model is the green channel of a wide-angle fundus image. The network is trained for 80 rounds. In each round, 40 512×512 slices are sampled from each image. The sampling uses a hard sample sampling strategy for training. The optimization method adopts the stochastic gradient descent algorithm with an initial learning rate of 0.05, which is reduced to 0.01 and 0.001 in the 40th and 70th rounds, respectively. During model inference, a sliding window is used, with each input size of 1024×1024 and a sliding distance of 512. The model is trained on a cloud hardware environment.
[0208] See Figure 12a The diagram shows the segmentation results of the method based on the embodiments of this application on the BUVS dataset. Segmentation based on weighted loss refers to segmentation without adding local constraints, while segmentation based on joint loss refers to segmentation with local constraints included. Both methods were tested here. Figure 12aThe two examples shown correspond to vessel segmentation under simple and complex backgrounds, respectively. The first column of images is free from background noise and large retinal lesions, and both loss functions provided in this application's embodiments can achieve good segmentation results. However, the second column of images contains reflections, large hemorrhages, and hard effusion lesions, all of which affect the visibility of terminal vessels. In areas affected by background factors, see [reference needed]. Figure 12a As shown in the labeled box, the segmentation results based on weighted loss exhibit fragmentation, while the joint loss result maintains the connectivity of the branch. Furthermore, the joint loss segmentation is more refined than the weighted loss segmentation, effectively mitigating oversegmentation and improving segmentation accuracy.
[0209] See Figure 12b The diagram shown illustrates the segmentation results of the method based on this application embodiment on the PRIME dataset. Similar to the BUVS dataset described above, two methods were used for testing. It can be seen that... Figure 12b In the two examples shown, the predicted pipes are finer in the joint loss-based segmentation results than in the weighted loss-based segmentation results. The branches in the weighted loss-based segmentation results are significantly coarser than the labeled results, while the joint loss-based segmentation results are closer to the manually labeled results.
[0210] Table 1 shows a comparison of the accuracy of segmentation based on various methods provided in this application embodiment. The evaluation metrics selected are the Dice coefficient, sensitivity, accuracy, and the ratio of detected length. The ratio of detected length can be calculated by dividing the number of detected bone points by the total number of bone points. Furthermore, the segmentation performance of all blood vessels and the terminal blood vessel portion were compared separately. To calculate the segmentation performance near the terminal blood vessel, the terminal blood vessel region was first extracted. Specifically, a vascular tree dilation operation with a kernel of size 5 was performed to obtain region R1. Then, each foreground point was assigned to the nearest bone point based on its coordinate distance, and the pixels belonging to the terminal branch bone points constituted region R2. The intersection of regions R1 and R2 is the region near the terminal blood vessel. The four metrics mentioned above were then recalculated within this region as an evaluation of the segmentation effect of the terminal portion.
[0211]
[0212] Table 1
[0213] As shown in Table 1, semantic segmentation can detect 72.3% of small blood vessels in the BUVS dataset and 79.0% in the PRIME dataset. Skeletal point prediction focuses on the skeletal point representation of the tracheal tree, achieving 68.9% of the terminal blood vessel detection length in BUVS and 76.2% in PRIME. In the method proposed in this application, we first weight the segmentation loss based on the blood vessel diameter and distance from the skeletal line. Compared to the semantic segmentation results, all four metrics are improved in the small blood vessel region, with sensitivity and detection length ratio showing the most significant improvement. Meanwhile, the overall segmentation metrics do not change significantly. Generally, the metrics at small blood vessels and the overall metrics are mutually exclusive; an improvement in one often leads to a decrease in the other. However, based on a reasonable weight design, the segmentation accuracy of terminal blood vessels can be further improved while maintaining the overall segmentation metrics. Furthermore, this application further introduces constraints on the statistical characteristics of the terminal portion, which can effectively improve the segmentation effect of the terminal portion. Compared to weighted loss, joint loss improved accuracy from 74.9% to 78.9% on the BUVS dataset and from 60.5% to 64.5% on the PRIME dataset. This improved accuracy also led to improvements in Dice performance, achieving a 1.9% improvement on both datasets. Simultaneously, joint loss maintained very similar overall segmentation results and vessel detection lengths, achieving optimal performance on both datasets.
[0214] Please see Figure 13 Based on the same inventive concept, embodiments of this application also provide a model training device 130 for vascular tree segmentation, the device comprising:
[0215] The sample acquisition unit 1301 is used to acquire a set of sample images for model training. Each sample image in the sample image set is an image containing a vascular tree and is associated with a labeling result representing the true location of the vascular tree in the sample image.
[0216] The attribute recognition unit 1302 is used to perform attribute recognition processing of blood vessel branches based on the annotation results of each sample image, and to obtain the attribute information of each blood vessel branch in each sample image. The attribute information includes: the thickness of the corresponding blood vessel branch, and the edge degree of the pixels included in the corresponding blood vessel branch in the blood vessel branch.
[0217] The weight determination unit 1303 is used to obtain the training weights corresponding to each pixel in the corresponding sample image based on the attribute information of each sample image. The training weights are negatively correlated with the thickness and edge degree.
[0218] Training unit 1304 is used to determine the model loss value based on the prediction results, annotation results and corresponding training weights of the blood vessel tree segmentation model during the training process, and to adjust the model parameters based on the model loss value.
[0219] In one possible implementation, the attribute recognition unit 1302 is specifically used for:
[0220] For each sample image, perform the following processing:
[0221] For a given sample image, based on the size of the image region occupied by each blood vessel branch in the sample image and the length of the skeletal line of the corresponding blood vessel branch, the diameter of each blood vessel branch is obtained. The diameter represents the thickness of the corresponding blood vessel branch, and the skeletal line is used to represent the direction of the corresponding blood vessel branch.
[0222] Based on the positional information of each pixel in each blood vessel branch and the positional information of the skeletal line of the corresponding blood vessel branch, the distance between each pixel and the skeletal line of the corresponding blood vessel branch is obtained. The distance represents the degree of the pixel's position on the edge of the corresponding blood vessel branch.
[0223] In one possible implementation, the attribute recognition unit 1302 is specifically used for:
[0224] The sample image is processed by skeleton extraction to obtain the skeleton of the blood vessel tree in the sample image. The skeleton contains the skeletal lines of each blood vessel branch.
[0225] For each vascular branch, the following treatments were performed:
[0226] For a blood vessel branch, obtain the first total number of branch pixels contained in the sample image of the blood vessel branch, and obtain the second total number of skeletal line pixels contained in the sample image of the skeletal line of the blood vessel branch.
[0227] The diameter of the vascular branch is obtained based on the ratio between the first total number and the second total number.
[0228] In one possible implementation, the weight determination unit 1303 is specifically used for:
[0229] For each pixel in the sample image, perform the following processing:
[0230] For a single pixel, a weight mapping process based on the diameter of the blood vessel branch to which the pixel is located and a preset diameter threshold is performed to obtain the first sub-weight of the pixel based on the diameter. The preset diameter threshold is the upper limit of the diameter of the terminal blood vessel.
[0231] Based on the distance between a pixel and the skeletal line of the corresponding blood vessel branch, a distance-based weight mapping process is performed with a preset distance threshold to obtain a second sub-weight of the pixel based on distance. The preset distance threshold is not less than the maximum distance in the sample image.
[0232] The training weights of the pixels are obtained by weighting the first sub-weights and the second sub-weights.
[0233] In one possible implementation, the weight determination unit 1303 is specifically used for:
[0234] For a given pixel, the vessel type is determined based on the diameter or branch level of the vessel branch to which the pixel is located.
[0235] If the blood vessel type is terminal blood vessel, the first sub-weight of the pixel is obtained based on the ratio between the diameter of the blood vessel branch where the pixel is located and the preset diameter threshold.
[0236] If the blood vessel type is a non-terminal blood vessel, then the first sub-weight of the pixel is set to zero.
[0237] In one possible implementation, the training unit 1304 is specifically used for:
[0238] Based on the sample image set, the blood vessel tree segmentation model is trained in multiple rounds of iterations. In each round of training, the following steps are performed:
[0239] Using the vascular tree segmentation model used in this round, the obtained sample images are processed by vascular tree segmentation based on image visual detection to obtain the prediction results representing the predicted positions of vascular trees in the sample images.
[0240] Based on the training weights of each pixel in each sample image, the labeled values of each pixel in the annotation results and the predicted values of the corresponding pixels in the prediction results are weighted to obtain the model loss value.
[0241] The model parameters of the vascular tree segmentation model used in this round were adjusted based on the model loss value.
[0242] In one possible implementation, the device further includes a neighborhood feature extraction unit 1305, for:
[0243] Based on the annotation results of each sample image, the pixels located on the skeleton line in each sample image are taken as reference points, and the local regions where each reference point is located are extracted to obtain the annotation neighborhood features corresponding to each reference point.
[0244] The training unit 1304 is specifically used to perform weighted processing on the labeled values of each pixel in the annotation results and the predicted values of the corresponding pixels in the prediction results based on the training weights of each pixel in each sample image, to obtain a single-point prediction sub-loss; and to extract regional features from the local region where each reference point is located based on the prediction results of each sample image, to obtain the prediction neighborhood features corresponding to each reference point in each sample image; and to accumulate the difference between the prediction neighborhood features and the corresponding labeled neighborhood features of each reference point in each sample image to obtain a region prediction sub-loss, and to perform weighted processing on the single-point prediction sub-loss and the region prediction sub-loss to obtain the model loss value.
[0245] In one possible implementation, the neighborhood feature extraction unit 1305 is specifically used for:
[0246] For each sample image, perform the following steps:
[0247] For a given sample image, the pixels located on the skeletal line of the terminal blood vessel are used as reference points.
[0248] Using each reference point as the center, and based on a preset local region size, the local region corresponding to each reference point is determined in the sample image;
[0249] Based on the annotation values of pixels in each local region of the sample image, the annotation neighborhood features of the reference point corresponding to each local region are obtained.
[0250] In one possible implementation, the labeled neighborhood features include a first mean and a first standard deviation of the labeled values of the pixels included in the local region, and the predicted neighborhood features include a second mean and a second standard deviation of the predicted values of the pixels included in the local region; then the training unit 1304 is specifically used for:
[0251] Based on the difference between the first mean and the second mean of each reference point in each sample image, and the difference between the first standard deviation and the corresponding second standard deviation, a weighted summation process is performed to obtain the regional prediction sub-loss.
[0252] In one possible implementation, the training unit 1304 is specifically used for:
[0253] For each sample image, perform the following operations:
[0254] For a sample image, local images are sampled sequentially from the sample image according to a preset sliding window size and sliding step size, and the obtained local images are sequentially input into the blood vessel tree segmentation model to obtain the sub-prediction results of each local image output by the blood vessel tree segmentation model.
[0255] Based on the overlapping parts in each local image, the obtained sub-prediction results are merged to obtain the prediction result of the sample image.
[0256] This device can be used to execute the model training methods shown in the various embodiments of this application. Therefore, the functions that each functional module of this device can achieve can be referred to the description of the foregoing embodiments, and will not be repeated here.
[0257] Please see Figure 14 Based on the same inventive concept, embodiments of this application also provide an image processing apparatus 140 for vascular tree segmentation, the apparatus comprising:
[0258] Image acquisition unit 1401 is used to acquire the target image to be segmented into a blood vessel tree;
[0259] The segmentation unit 1402 is used to perform image visual detection-based vascular tree segmentation processing on the target image based on the vascular tree segmentation model obtained by the above-described model training method, and output the corresponding vascular tree segmentation result; wherein, the vascular tree segmentation result is used to indicate the position of the vascular tree in the target image.
[0260] This device can be used to execute the image processing methods shown in the various embodiments of this application. Therefore, the functions that each functional module of this device can achieve can be referred to the description of the foregoing embodiments, and will not be repeated here.
[0261] Using the aforementioned device, during the training of the blood vessel tree segmentation model, the thickness of each blood vessel branch and the edge degree of each pixel within the blood vessel branch are combined with the annotation results of each sample image to generate the training weights corresponding to each pixel, thereby obtaining the model's single-point prediction loss. Furthermore, local statistical feature constraints for terminal blood vessels are added to the loss function. Based on the aforementioned single-point prediction loss, the neighborhood prediction loss is combined to reduce the model's sensitivity to a small amount of annotation noise, assisting the model to focus more on the location of segmentation errors during training and improving the segmentation accuracy of terminal blood vessels.
[0262] Please see Figure 15 Based on the same technical concept, embodiments of this application also provide a computer device. In one embodiment, the computer device can be... Figure 1 The server shown or Figure 2 The cloud-related device shown is a computer device such as... Figure 15 As shown, it includes a memory 1501, a communication module 1503, and one or more processors 1502.
[0263] The memory 1501 is used to store computer programs executed by the processor 1502. The memory 1501 may mainly include a program storage area and a data storage area. The program storage area may store the operating system and programs required to run instant messaging functions, etc.; the data storage area may store various instant messaging information and operation instruction sets, etc.
[0264] Memory 1501 may be volatile memory, such as random-access memory (RAM); memory 1501 may also be non-volatile memory, such as read-only memory, flash memory, hard disk drive (HDD), or solid-state drive (SSD); or memory 1501 may be any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but is not limited thereto. Memory 1501 may be a combination of the above-described memories.
[0265] Processor 1502 may include one or more central processing units (CPUs) or digital processing units, etc. Processor 1502 is used to implement the above-described model training and image processing methods for vascular tree segmentation when calling computer programs stored in memory 1501.
[0266] The communication module 1503 is used to communicate with terminal devices and other servers.
[0267] This application embodiment does not limit the specific connection medium between the memory 1501, communication module 1503, and processor 1502. This application embodiment... Figure 15 The memory 1501 and the processor 1502 are connected via a bus 1504, and the bus 1504 is in Figure 15 The diagram uses thick lines to describe the connections between other components; these are for illustrative purposes only and should not be considered limiting. The 1504 bus can be divided into address bus, data bus, control bus, etc. For ease of description, Figure 15 It is described using only a thick line, but does not indicate that there is only one bus or one type of bus.
[0268] The memory 1501 stores a computer storage medium, which stores computer-executable instructions. The computer-executable instructions are used to implement the model training and image processing method for vascular tree segmentation according to the embodiments of this application. The processor 1502 is used to execute the model training and image processing method for vascular tree segmentation according to the above embodiments.
[0269] In another embodiment, the computer device can also be a terminal device, such as... Figure 1 The terminal device shown. In this embodiment, the structure of the computer device can be as follows. Figure 16 As shown, it includes components such as: communication component 1610, memory 1620, display unit 1630, camera 1640, sensor 1650, audio circuit 1660, Bluetooth module 1670, processor 1680, etc.
[0270] The communication component 1610 is used to communicate with the server. In some embodiments, it may include a Circuit-Based Wireless Fidelity (WiFi) module, which is a short-range wireless transmission technology. Computer devices can use WiFi modules to help users send and receive information.
[0271] The memory 1620 can be used to store software programs and data. The processor 1680 executes various functions of the terminal device and data processing by running the software programs or data stored in the memory 1620. The memory 1620 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. The memory 1620 stores an operating system that enables the terminal device to run. In this application, the memory 1620 may store the operating system and various application programs, and may also store code that executes the model training and image processing methods for vascular tree segmentation according to the embodiments of this application.
[0272] The display unit 1630 can also be used to display information input by the user or information provided to the user, as well as various menus of the terminal device, forming a graphical user interface (GUI). Specifically, the display unit 1630 may include a display screen 1632 disposed on the front of the terminal device. The display screen 1632 may be configured as a liquid crystal display, a light-emitting diode, or the like. The display unit 1630 can be used to display various input image pages or segmentation result pages in the embodiments of this application.
[0273] The display unit 1630 can also be used to receive input digital or character information and generate signal inputs related to user settings and function control of the terminal device. Specifically, the display unit 1630 may include a touch screen 1631 disposed on the front of the terminal device, which can collect touch operations of the user on or near it, such as clicking buttons, dragging scroll boxes, etc.
[0274] The touchscreen 1631 can be placed on top of the display screen 1632, or the touchscreen 1631 and the display screen 1632 can be integrated to realize the input and output functions of the terminal device. After integration, it can be referred to as a touch display screen. In this application, the display unit 1630 can display the application and the corresponding operation steps.
[0275] Camera 1640 can be used to capture still images, which users can then post comments on via the application. There can be one or multiple cameras 1640. An object is projected onto a photosensitive element through a lens, generating an optical image. This photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then transmitted to the processor 1680 for conversion into a digital image signal.
[0276] The terminal device may also include at least one sensor 1650, such as an accelerometer 1651, a proximity sensor 1652, a fingerprint sensor 1653, and a temperature sensor 1654. The terminal device may also be equipped with other sensors such as a gyroscope, barometer, hygrometer, thermometer, infrared sensor, light sensor, and motion sensor.
[0277] Audio circuitry 1660, speaker 1661, and microphone 1662 provide an audio interface between the user and the terminal device. Audio circuitry 1660 converts received audio data into electrical signals, which are then transmitted to speaker 1661, where they are converted into sound signals for output. The terminal device can also be equipped with volume buttons for adjusting the volume of the sound signal. Conversely, microphone 1662 converts collected sound signals into electrical signals, which are then received by audio circuitry 1660, converted back into audio data, and output to communication component 1610 for transmission to, for example, another terminal device, or to memory 1620 for further processing.
[0278] The Bluetooth module 1670 is used to interact with other Bluetooth devices that also have a Bluetooth module via the Bluetooth protocol. For example, a terminal device can establish a Bluetooth connection with a wearable computer device (such as a smartwatch) that also has a Bluetooth module through the Bluetooth module 1670, thereby exchanging data.
[0279] The processor 1680 is the control center of the terminal device, connecting various parts of the terminal through various interfaces and lines. It executes various functions and processes data by running or executing software programs stored in the memory 1620 and calling data stored in the memory 1620. In some embodiments, the processor 1680 may include one or more processing units; the processor 1680 may also integrate an application processor and a baseband processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the baseband processor mainly handles wireless communication. It is understood that the baseband processor may not be integrated into the processor 1680. In this application, the processor 1680 can run an operating system, applications, user interface display and touch response, as well as the model training and image processing methods for vascular tree segmentation according to embodiments of this application. Furthermore, the processor 1680 is coupled to the display unit 1630.
[0280] Based on the same inventive concept, embodiments of this application also provide a storage medium storing a computer program that, when run on a computer, causes the computer to perform the steps in the model training and image processing methods for vascular tree segmentation according to various exemplary embodiments of this application described above.
[0281] In some possible implementations, various aspects of the model training and image processing methods for vascular tree segmentation provided in this application can also be implemented in the form of a computer program product, which includes a computer program that, when run on a computer device, causes the computer device to perform the steps in the model training and image processing methods for vascular tree segmentation according to various exemplary embodiments of this application described above. For example, the computer device can perform the steps of the various embodiments.
[0282] The program product may employ any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples (a non-exhaustive list) of readable storage media include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0283] The program product of the embodiments of this application may employ a portable compact disc read-only memory (CD-ROM) and include a computer program, and may run on a computer device. However, the program product of this application is not limited thereto. In this application, the readable storage medium may be any tangible medium that contains or stores a program, and the computer program included therein may be used by or in conjunction with a command execution system, apparatus, or device.
[0284] A readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying a readable computer program. This propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium, capable of sending, propagating, or transmitting a program for use by or in conjunction with a command execution system, apparatus, or device.
[0285] Computer programs contained on readable media may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination thereof.
[0286] Computer programs for performing the operations of this application can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages.
[0287] It should be noted that although several units or sub-units of the device have been mentioned in the detailed description above, this division is merely exemplary and not mandatory. In fact, according to embodiments of this application, the features and functions of two or more units described above can be embodied in one unit. Conversely, the features and functions of one unit described above can be further divided and embodied by multiple units.
[0288] Furthermore, although the operations of the method of this application are described in a specific order in the accompanying drawings, this does not require or imply that these operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps.
[0289] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0290] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application. Clearly, those skilled in the art can make various alterations and variations to this application without departing from its spirit and scope. Thus, if such alterations and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such alterations and variations.
Claims
1. A model training method for vascular tree segmentation, characterized in that, The method includes: A sample image set for model training is obtained, wherein each sample image in the sample image set is an image containing a vascular tree and is associated with annotation results representing the true location of the vascular tree in the sample image; Based on the annotation results of each sample image, attribute recognition processing of blood vessel branches is performed to obtain the attribute information of each blood vessel branch in each sample image: For a sample image, based on the size of the image area occupied by each blood vessel branch in the sample image and the length of the skeletal line of the corresponding blood vessel branch, the diameter of each blood vessel branch is obtained. The diameter represents the thickness of the corresponding blood vessel branch, and the skeletal line is used to represent the direction of the corresponding blood vessel branch. Based on the position information of each pixel point included in each blood vessel branch and the position information of the skeletal line of the corresponding blood vessel branch, the distance between each pixel point and the skeletal line of the corresponding blood vessel branch is obtained. The distance represents the degree to which the pixel point is on the edge of the corresponding blood vessel branch. Based on the attribute information corresponding to each sample image, the training weights corresponding to each pixel in the corresponding sample image are obtained respectively. The training weights are negatively correlated with the thickness and the edge degree. During the training of the blood vessel tree segmentation model, the model loss value is determined based on the prediction results, annotation results, and corresponding training weights of the blood vessel trees in each sample image, and the model parameters are adjusted based on the model loss value.
2. The method as described in claim 1, characterized in that, The method of obtaining the diameter of each blood vessel branch based on the size of the image region occupied by each blood vessel branch in the sample image and the length of the skeletal line of the corresponding blood vessel branch includes: The sample image is subjected to skeleton extraction processing to obtain the skeleton of the blood vessel tree in the sample image, the skeleton containing the skeletal lines of each blood vessel branch; For each of the aforementioned vascular branches, the following processing is performed respectively: For a blood vessel branch, a first total number of branch pixels contained in the sample image of the blood vessel branch is obtained, and a second total number of skeletal line pixels contained in the sample image of the skeletal line of the blood vessel branch are obtained. The diameter of the blood vessel branch is obtained based on the ratio between the first total number and the second total number.
3. The method as described in claim 1, characterized in that, Based on the attribute information corresponding to each sample image, the training weights corresponding to each pixel in the corresponding sample image are obtained, including: For each pixel in the sample image, the following processing is performed: For a single pixel, a weight mapping process based on the diameter of the blood vessel branch to which the pixel is located and a preset diameter threshold is performed to obtain the first sub-weight of the pixel based on the diameter. The preset diameter threshold is the upper limit value of the diameter of the terminal blood vessel. Based on the distance between the pixel and the skeletal line of the corresponding blood vessel branch, a distance-based weight mapping process is performed with a preset distance threshold to obtain the second sub-weight of the pixel based on distance. The preset distance threshold is not less than the maximum distance in the sample image. The training weights of the pixels are obtained by weighting the first sub-weights and the second sub-weights.
4. The method as described in claim 3, characterized in that, For a given pixel, a diameter-based weight mapping process is performed based on the diameter of the blood vessel branch to which the pixel is located and a preset diameter threshold to obtain the first sub-weight of the pixel based on the diameter, including: For a given pixel, the blood vessel type is determined based on the diameter or branch level of the blood vessel branch to which the pixel is located. If the blood vessel type is a terminal blood vessel, then the first sub-weight of the pixel is obtained based on the ratio between the diameter of the blood vessel branch where the pixel is located and a preset diameter threshold. If the blood vessel type is a non-terminal blood vessel, then the first sub-weight of the pixel is determined to be zero.
5. The method according to any one of claims 1 to 4, characterized in that, During the training process of the blood vessel tree segmentation model, the model loss value is determined based on the prediction results, annotation results, and corresponding training weights of the blood vessel trees in each sample image, and the model parameters are adjusted based on the model loss value, including: Based on the sample image set, the blood vessel tree segmentation model is trained through multiple rounds of iterative training. In each round of training, the following steps are performed: Using the vascular tree segmentation model used in this round, the obtained sample images are processed by vascular tree segmentation based on image visual detection to obtain the prediction results representing the predicted positions of vascular trees in the sample images. The labeled values of each pixel in the annotation result and the predicted values of the corresponding pixels in the prediction result are weighted based on the training weights of each pixel in each sample image to obtain the model loss value. The model parameters of the vascular tree segmentation model used in this round are adjusted based on the model loss value.
6. The method as described in claim 5, characterized in that, After obtaining the sample image set for model training, the method further includes: Based on the annotation results of each sample image, the pixels located on the skeleton line in each sample image are taken as reference points, and the local regions where each reference point is located are extracted to obtain the annotation neighborhood features corresponding to each reference point. Then, based on the training weights of each pixel in each sample image, the labeled values of each pixel in the annotation result and the predicted values of the corresponding pixels in the prediction result are weighted to obtain the model loss value, including: Based on the training weights of each pixel in each sample image, the labeled values of each pixel in the labeled result and the predicted values of the corresponding pixels in the predicted result are weighted to obtain the single-point prediction sub-loss. Based on the prediction results of each sample image, regional features are extracted from the local area where each reference point is located, and the predicted neighborhood features corresponding to each reference point in each sample image are obtained respectively. The difference between the predicted neighborhood features and the corresponding labeled neighborhood features of each reference point in each sample image is accumulated to obtain the region prediction sub-loss. The single-point prediction sub-loss and the region prediction sub-loss are then weighted to obtain the model loss value.
7. The method as described in claim 6, characterized in that, The annotation results based on each sample image are used to select pixels located on the skeletal lines in each sample image as reference points. Regional features are extracted from the local areas where each reference point is located to obtain the annotation neighborhood features corresponding to each reference point, including: For each sample image, perform the following steps: For a given sample image, the pixels located on the skeletal line of the terminal blood vessel are used as reference points. Using each reference point as the center, and based on a preset local region size, the local region corresponding to each reference point is determined in the sample image. Based on the annotation values of the pixels included in each local region in the annotation results of the sample image, the annotation neighborhood features of the reference point corresponding to each local region are obtained.
8. The method as described in claim 7, characterized in that, The labeled neighborhood features include the first mean and the first standard deviation of the labeled values of the pixels included in the local region, and the predicted neighborhood features include the second mean and the second standard deviation of the predicted values of the pixels included in the local region. The difference between the predicted neighborhood features and the corresponding labeled neighborhood features of each reference point in each sample image is accumulated to obtain the region prediction sub-loss, including: Based on the difference between the first mean and the second mean of each reference point in each sample image, and the difference between the first standard deviation and the corresponding second standard deviation, a weighted summation process is performed to obtain the region prediction sub-loss.
9. The method as described in claim 5, characterized in that, The blood vessel tree segmentation model used in this round performs blood vessel tree segmentation processing based on image visual detection on each obtained sample image to obtain prediction results representing the predicted positions of blood vessel trees in the sample image, including: For each sample image, perform the following operations: For a sample image, local images are sampled sequentially from the sample image according to a preset sliding window size and sliding step size, and the obtained local images are sequentially input into the blood vessel tree segmentation model to obtain the sub-prediction results of each local image output by the blood vessel tree segmentation model; Based on the overlapping parts in each local image, the obtained sub-prediction results are merged to obtain the prediction result of the sample image.
10. An image processing method for vascular tree segmentation, characterized in that, The method includes: Obtain the target image for vascular tree segmentation; Based on the vascular tree segmentation model obtained by the model training method according to any one of claims 1 to 9, the target image is subjected to vascular tree segmentation processing based on image visual detection, and the corresponding vascular tree segmentation result is output; wherein, the vascular tree segmentation result is used to indicate the position of the vascular tree in the target image.
11. A model training device for vascular tree segmentation, characterized in that, The device includes: The sample acquisition unit is used to obtain a set of sample images for model training. Each sample image in the set is an image containing a vascular tree and is associated with a labeling result representing the true location of the vascular tree in the sample image. The attribute recognition unit is used to perform attribute recognition processing of blood vessel branches based on the annotation results of each sample image, and to obtain the attribute information of each blood vessel branch in each sample image: For a sample image, based on the size of the image area occupied by each blood vessel branch in the sample image and the length of the skeletal line of the corresponding blood vessel branch, the diameter of each blood vessel branch is obtained, wherein the diameter represents the thickness of the corresponding blood vessel branch, and the skeletal line is used to represent the direction of the corresponding blood vessel branch; Based on the position information of each pixel point included in each blood vessel branch and the position information of the skeletal line of the corresponding blood vessel branch, the distance between each pixel point and the skeletal line of the corresponding blood vessel branch is obtained, wherein the distance represents the degree of the pixel point on the edge of the corresponding blood vessel branch; The weight determination unit is used to obtain the training weights corresponding to each pixel in the corresponding sample image based on the attribute information corresponding to each sample image. The training weights are negatively correlated with the thickness and the edge degree. The training unit is used to determine the model loss value based on the prediction results, annotation results and corresponding training weights of the blood vessel tree segmentation model during the training process, and to adjust the model parameters based on the model loss value.
12. An image processing apparatus for vascular tree segmentation, characterized in that, The device includes: An image acquisition unit is used to acquire the target image to be segmented into a blood vessel tree. The segmentation unit is used to perform image visual detection-based vascular tree segmentation processing on the target image based on the vascular tree segmentation model obtained by the model training method according to any one of claims 1 to 9, and output the corresponding vascular tree segmentation result; wherein the vascular tree segmentation result is used to indicate the position of the vascular tree in the target image.
13. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 9 or 10.
14. A computer storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the steps of the method described in any one of claims 1 to 9 or 10.
15. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the steps of the method described in any one of claims 1 to 9 or 10.
Citation Information
Patent Citations
Training method and device and application method and device of blood vessel segmentation model
CN113066090A