Blood vessel segmentation method and device for OCTA image, equipment and medium

By performing feature encoding and branching processing of the volume data and projection map of OCTA images, integrating global and local information, the problem of low vascular segmentation accuracy in OCTA technology is solved, and higher segmentation accuracy and efficiency are achieved, which is suitable for ophthalmic disease diagnosis.

CN120298423APending Publication Date: 2025-07-11LIAONING MOBILE COMM +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510322645.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

When the existing OCTA technology is manually segmented in retinal images, the accuracy is low, the texture is complex, and the clarity of the blood vessel boundary is limited, resulting in large errors.

Method used

The vascular segmentation method of OCTA images is adopted. By inputting the volume data and projection maps into the feature encoder respectively, different projection learning branches are used to process features, fuse global and local information, and the feature decoder is used to restore blood vessel information to improve segmentation accuracy.

Benefits of technology

It improves the accuracy and efficiency of vascular segmentation, enhances the connectivity between blood vessels, and can better assist in the diagnosis of ophthalmic diseases.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298423A_ABST
    Figure CN120298423A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a blood vessel segmentation method and device for an OCTA image, equipment and a medium, and belongs to the field of image processing. The method comprises the following steps: inputting volume data of an obtained optical OCTA image into a first feature encoder, and processing the volume data by using a plurality of residual convolution down-sampling layers in the first feature encoder to obtain first features output by the plurality of residual convolution down-sampling layers; respectively inputting each first feature into a first projection learning branch and a second projection learning branch for processing to obtain a second feature corresponding to each first feature; inputting the obtained projection image of the OCTA image into a second feature encoder, and processing the projection image by using a plurality of residual convolution down-sampling layers in the second feature encoder to obtain third features output by the plurality of residual convolution down-sampling layers; and obtaining blood vessel segmentation data corresponding to the OCTA image based on the second feature, the third feature and a feature decoder. According to the embodiment of the invention, the accuracy of blood vessel segmentation can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of image processing, and particularly relates to a method, device, equipment and medium for vascular segmentation of OCTA images. Background Art

[0002] The retinal vascular system is an important structure in the fundus of the eye. The vascular information of the fundus can be obtained by using retinal imaging technology, and retinal diseases can be diagnosed based on the vascular information. Common retinal imaging technologies may include color fundus imaging technology, fluorescein fundus angiography technology, optical coherence tomography angiography (OCTA) technology, etc. The OCTA technology can obtain retinal images under non-invasive conditions and is more widely used. In the actual operation process, professional doctors need to manually mark the blood vessels in the retinal images obtained by using the OCTA technology, so as to segment the blood vessels from the retinal images. However, the texture of the retinal images is complex and the clarity of the blood vessel boundaries is limited, and there are large errors in the accuracy of manually marking the blood vessels from the retinal images, and the accuracy of blood vessel segmentation is relatively low. Summary of the Invention

[0003] The embodiments of this application provide a method, device, equipment and medium for vascular segmentation of OCTA images, which can improve the accuracy of vascular segmentation.

[0004] In a first aspect, the embodiments of this application provide a method for vascular segmentation of OCTA images, including: inputting the volume data of the obtained optical coherence tomography angiography (OCTA) images into a first feature encoder, and using multiple residual convolutional downsampling layers in the first feature encoder to process the volume data to obtain first features output by the multiple residual convolutional downsampling layers; respectively inputting each first feature into a first projection learning branch and a second projection learning branch for processing to obtain second features corresponding to each first feature, and the processing of the first projection learning branch is different from that of the second projection learning branch; inputting the projection map of the obtained OCTA image into a second feature encoder, and using multiple residual convolutional downsampling layers in the second feature encoder to process the projection map to obtain third features output by the multiple residual convolutional downsampling layers; and obtaining vascular segmentation data corresponding to the OCTA image based on the second features, the third features and a feature decoder.

[0005] In some possible embodiments, in the first feature encoder, the input data of the first residual convolutional downsampling layer includes volume data, and the input data of the subsequent residual convolutional downsampling layer includes the first features output by the previous residual convolutional downsampling layer; in the second feature encoder, the input data of the first residual convolutional downsampling layer includes projection maps, and the input data of the subsequent residual convolutional downsampling layer includes the third features output by the previous residual convolutional downsampling layer.

[0006] In some possible embodiments, the processing performed by the residual convolutional downsampling layer includes residual convolution processing and downsampling processing, and the input data of the downsampling processing includes the output data of the residual convolution processing; the downsampling processing includes: performing average pooling processing and max pooling processing on the output data of the residual convolution processing respectively, combining the data after average pooling processing and the data after max pooling processing to obtain first combined data, and performing convolution processing on the first combined data to obtain the output data of the downsampling processing.

[0007] In some possible embodiments, the first projection learning branch includes 3D convolutional layers and unidirectional convolutional layers, and the second projection learning branch includes 3D convolutional layers and unidirectional pooling layers; inputting each first feature into the first projection learning branch and the second projection learning branch respectively for processing to obtain second features corresponding to each first feature, including: inputting the first feature into the first projection learning branch and processing it sequentially through 3D convolutional layers, 3D convolutional layers, unidirectional convolutional layers, 3D convolutional layers, 3D convolutional layers, unidirectional convolutional layers to obtain first sub-features; inputting the first feature into the second projection learning branch and processing it sequentially through 3D convolutional layers, 3D convolutional layers, unidirectional pooling layers, 3D convolutional layers, 3D convolutional layers, unidirectional pooling layers to obtain second sub-features; combining the first sub-features and the second sub-features to obtain second features.

[0008] In some possible embodiments, obtaining the vascular segmentation data corresponding to the OCTA image based on the second features, the third features, and the feature decoder includes: performing matrix transformation processing including two types of pooling processing on both the second features and the third features respectively to obtain a first channel attention matrix corresponding to the second features and a second channel attention matrix corresponding to the third features; obtaining multiple fused features according to the first channel attention matrix, the second channel attention matrix, the second features, and the third features; and using multiple transposed convolutional upsampling layers in the feature decoder to process the multiple fused features to obtain the vascular segmentation data.

[0009] In some possible embodiments, matrix transformation processing including two types of pooling processing is respectively performed on the second feature and the third feature to obtain a first channel attention matrix corresponding to the second feature and a second channel attention matrix corresponding to the third feature, including: respectively performing max pooling processing and average pooling processing on the second feature, inputting the data obtained by the max pooling processing and the data obtained by the average pooling processing into a multi-layer perceptron respectively, combining the data output by the multi-layer perceptron to obtain second combined data, and processing the second combined data through an activation function to obtain the first channel attention matrix; respectively performing max pooling processing and average pooling processing on the third feature, inputting the data obtained by the max pooling processing and the data obtained by the average pooling processing into a multi-layer perceptron respectively, combining the data output by the multi-layer perceptron to obtain third combined data, and processing the third combined data through an activation function to obtain the second channel attention matrix.

[0010] In some possible embodiments, multiple fused features are obtained according to the first channel attention matrix, the second channel attention matrix, the second feature, and the third feature, including:

[0011] Processing the first channel attention matrix and the second channel attention matrix using a max function to obtain a comprehensive channel attention matrix; enhancing the second feature and the third feature respectively according to the comprehensive channel attention matrix to obtain a first enhanced feature and a second enhanced feature; performing convolution processing on the sum of the first enhanced feature and the second enhanced feature to obtain a fused feature.

[0012] In some possible embodiments, the transposed convolution upsampling layer in the feature decoder corresponds one-to-one with the residual convolution downsampling layer in the first feature encoder and the residual convolution downsampling layer in the second feature encoder; in the feature decoder, the input data of the last transposed convolution upsampling layer includes the fused feature corresponding to the last residual convolution downsampling layer in the first feature encoder and the last residual convolution downsampling layer in the second feature encoder, the output data of the first transposed convolution upsampling layer includes vascular segmentation data, and the input data of a previous transposed convolution upsampling layer includes the combined data of the fused feature corresponding to the previous residual convolution downsampling layer and the output data of the next transposed convolution upsampling layer.

[0013] In some possible embodiments, the method further includes: training a neural network model according to the sample volume data and the sample projection map of the OCTA image, where the neural network model includes a first feature encoder, a second feature encoder, a first projection learning branch, a second projection learning branch, and a feature decoder; obtaining a loss function of the neural network model, where the loss function is obtained by using a weight algorithm through a Dice loss function and a cross-entropy loss function; and in the case where the loss function does not meet the training cut-off condition, adjusting the model parameters of the neural network model, and training the neural network after adjusting the model parameters until the latest obtained loss function meets the training cut-off condition.

[0014] In a second aspect, an embodiment of the present application provides a vascular segmentation device for OCTA images, including: a first encoding processing module, configured to input the volume data of the acquired optical coherence tomography angiography (OCTA) image into a first feature encoder, and process the volume data by using a plurality of residual convolutional downsampling layers in the first feature encoder to obtain a first feature output by the plurality of residual convolutional downsampling layers; a projection learning branch module, configured to input each first feature into a first projection learning branch and a second projection learning branch respectively for processing to obtain a second feature corresponding to each first feature, where the processing of the first projection learning branch is different from the processing of the second projection learning branch; a second encoding processing module, configured to input the projection map of the acquired OCTA image into a second feature encoder, and process the projection map by using a plurality of residual convolutional downsampling layers in the second feature encoder to obtain a third feature output by the plurality of residual convolutional downsampling layers; and a fusion decoding processing module, configured to obtain vascular segmentation data corresponding to the OCTA image based on the second feature, the third feature, and the feature decoder.

[0015] In a third aspect, an embodiment of the present application provides a vascular segmentation device for OCTA images, the device including: a processor, and a memory storing computer program instructions; the processor reads and executes the computer program instructions to implement the vascular segmentation method for OCTA images in the first aspect.

[0016] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, where computer program instructions are stored on the computer storage medium, and when the computer program instructions are executed by a processor, the vascular segmentation method for OCTA images in the first aspect is implemented.

[0017] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, the vascular segmentation method for OCTA images in the first aspect is implemented.

[0018] The embodiments of the present application provide a method, apparatus, device, and medium for vascular segmentation of OCTA images. The volume data of the OCTA image is input into the first feature encoder, and the volume data is processed by multiple residual convolutional downsampling layers in the first feature encoder to obtain the first feature learned from the volume data. The projection map of the OCTA image is input into the second feature encoder, and the projection map is processed by multiple residual convolutional downsampling layers in the second feature encoder to obtain the third feature learned from the projection map. Different two projection learning branches are used to process the first feature to obtain the second feature that can better retain the spatial information in the depth direction. By fusing the second feature and the third feature, the global information and local information of the blood vessels can be integrated, so that the advantageous information of the volume data and the advantageous information of the projection map are complementary, and the feature decoder is used to more fully restore the blood vessel information in the decoding stage, thereby more accurately segmenting the blood vessels from the OCTA image and improving the accuracy of vascular segmentation. Description of the Drawings

[0019] To more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required to be used in the embodiments of the present application. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.

[0020] Figure 1 It is a schematic diagram of an example of the volume data and the projection map of the OCTA image provided by the embodiments of the present application;

[0021] Figure 2 It is a flowchart of the method for vascular segmentation of the OCTA image provided by an embodiment of the present application;

[0022] Figure 3 It is a schematic diagram of an example of the first feature encoder provided by the embodiments of the present application;

[0023] Figure 4 It is a schematic logical diagram of an example of the residual convolution processing provided by the embodiments of the present application;

[0024] Figure 5 It is a schematic logical diagram of an example of the downsampling processing provided by the embodiments of the present application;

[0025] Figure 6 It is a schematic comparison diagram of an example of the vascular segmentation results of multiple methods provided by the embodiments of the present application;

[0026] Figure 7 It is a schematic diagram of an example of the first projection learning branch and the second projection learning branch provided by the embodiments of the present application;

[0027] Figure 8A logic diagram showing an example of obtaining a fused feature provided by an embodiment of the present application;

[0028] Figure 9 A schematic diagram showing an example of a neural network model provided by an embodiment of the present application;

[0029] Figure 10 A schematic structural diagram of a vascular segmentation device for OCTA images provided by an embodiment of the present application;

[0030] Figure 11 A schematic structural diagram of a vascular segmentation device for OCTA images provided by an embodiment of the present application. Detailed implementation manners

[0031] The features and exemplary embodiments of various aspects of the present application will be described in detail below. To make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain the present application, rather than to limit the present application. For those skilled in the art, the present application can be implemented without some of these specific details. The following description of the embodiments is only intended to provide a better understanding of the present application by showing examples of the present application.

[0032] It should be noted that in this document, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover a non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the element.

[0033] The retinal vascular system is an important structure in the fundus of the eye. The vascular information of the fundus can be obtained by using retinal imaging technology, and retinal diseases can be diagnosed based on the vascular information. Common retinal imaging technologies may include color fundus imaging technology, fluorescein fundus angiography technology, OCTA technology, etc. OCTA technology can obtain retinal images under non-invasive conditions and is more widely used. In the actual operation process, professional doctors need to manually mark the blood vessels in the retinal images obtained by using OCTA technology, so as to segment the blood vessels from the retinal images. However, the texture of retinal images is complex, and the clarity of blood vessel boundaries is limited. There are large errors in the accuracy of manually marking blood vessels from retinal images, and the accuracy of blood vessel segmentation is relatively low. For example, Figure 1 is a schematic diagram of an example of the volume data and projection map of the OCTA image provided by the embodiment of the present application. As Figure 1 shown, the projection of the volume data of the OCTA image in the vertical direction can obtain the projection map of the OCTA image. There is noise interference in the volume data, and the signal-to-noise ratio is relatively low. The overall contrast of the projection map generated by the vertical projection of the volume data is relatively low. The projection map shows the situation of the retina. In addition to the blood vessel area, the projection map also shows the macular area, the lesion area, etc. The macular area, the lesion area, etc. will interfere with the segmentation of blood vessels. It is very difficult to manually mark blood vessels in the projection map, which is time-consuming and laborious, and the accuracy will also be adversely affected.

[0034] In order to reduce the marking error of blood vessels, there is also a method of using a deep learning OCTA blood vessel segmentation algorithm to segment blood vessels from the OCTA image showing the retina. For example, 2D-to-2D methods and 3D-to-2D methods can be used; the 2D-to-2D method can use the vertical projection map of a certain layer in the volume data of the OCTA image as the input, so as to output a two-dimensional blood vessel prediction map. However, due to the selection of the vertical projection maps of different layers and the influence of the projection technology, the blood vessel information contained in the vertical projection map used by the 2D-to-2D method may be incomplete, resulting in mis-segmentation or wrong segmentation of blood vessels, and reducing the accuracy of blood vessel segmentation; the 3D-to-2D method can use the volume data of the OCTA image as the input and output a two-dimensional blood vessel prediction map. However, the 3D-to-2D method uses a convolutional neural network to extract blood vessel features, and there are problems such as limited receptive field of the convolutional kernel and lack of global context information, resulting in poor blood vessel segmentation results and relatively low accuracy of blood vessel segmentation.

[0035] The present application provides a method, apparatus, device, medium, and program product for vascular segmentation of OCTA images. The volume data and projection map of the OCTA image can be used as inputs to complement each other by leveraging the characteristics of the volume data and the projection map. The volume data and the projection map are respectively input into corresponding feature encoders for processing, and different two projection learning branches are used to process the features corresponding to the volume data output by the feature encoder differently, so as to better retain the spatial information in the depth direction. The features obtained after processing by the projection learning branches are fused with the features corresponding to the projection map output by the feature encoder, and a feature decoder is used to restore the vascular information to a greater extent during the decoding stage, thereby segmenting the blood vessels more accurately from the OCTA image.

[0036] The method, apparatus, device, medium, and program product for vascular segmentation of OCTA images provided by the present application will be described separately below.

[0037] The present application provides a method for vascular segmentation of OCTA images. The method for vascular segmentation of OCTA images can be executed by a device, equipment, etc. for vascular segmentation of OCTA images, and is not limited herein. Figure 2 For the flowchart of the method for vascular segmentation of OCTA images provided by an embodiment of the present application, as Figure 2 shown, the method for vascular segmentation of OCTA images may include steps S101 to S104.

[0038] In step S101, the volume data of the acquired OCTA image is input into the first feature encoder, and the volume data is processed by multiple residual convolutional downsampling layers in the first feature encoder to obtain the first features output by the multiple residual convolutional downsampling layers.

[0039] The volume data of the OCTA image is a data set of the OCTA image in three-dimensional space, which can reflect the situation of the retinal image shown in the OCTA image in three-dimensional space. The retinal image includes a vascular image, and the volume data of the OCTA image also contains the data of the vascular image in three-dimensional space.

[0040] The first feature encoder can extract the features of the volume data. The first feature encoder includes multiple residual convolutional downsampling layers, and each residual convolutional downsampling layer has a residual convolutional processing function and a downsampling processing function. Each residual convolutional downsampling layer can output a first feature according to the input data. The number of the first features is consistent with the number of residual convolutional downsampling layers in the first feature encoder. In some examples, in the first feature encoder, the input data of the first residual convolutional downsampling layer includes the volume data, and the input data of the subsequent residual convolutional downsampling layer includes the first feature output by the previous residual convolutional downsampling layer. For example, Figure 3 For a schematic diagram of an example of the first feature encoder provided by an embodiment of the present application, asFigure 3 As shown in the figure, the first feature encoder 21 may include three residual convolutional downsampling layers, namely the residual convolutional downsampling layer 211, the residual convolutional downsampling layer 212, and the residual convolutional downsampling layer 213. The residual convolutional downsampling layer 211 is the first residual convolutional downsampling layer. The volume data is input into the residual convolutional downsampling layer 211. After being processed by the residual convolutional downsampling layer 211, the first feature C1 is output. The first feature C1 is used as the input data and input into the residual convolutional downsampling layer 212. After being processed by the residual convolutional downsampling layer 212, the first feature C2 is output. The first feature C2 is used as the input data and input into the residual convolutional downsampling layer 213. After being processed by the residual convolutional downsampling layer 213, the first feature C3 is output. If the size of the volume data is W×H×D×C, where W represents the width, specifically the number of pixels of the volume data in the horizontal direction, H represents the height, specifically the number of pixels of the volume data in the vertical direction, D represents the depth, specifically the number of volume data, and C represents the number of channels, the size of the first feature C1 output by the residual convolutional downsampling layer 211 is W / 2×H / 2×D / 2×C, the size of the first feature C2 output by the residual convolutional downsampling layer 212 is W / 4×H / 4×D / 4×C, and the size of the first feature C3 output by the residual convolutional downsampling layer 213 is W / 16×H / 16×D / 16×4C.

[0041] In some examples, the processing performed by the residual convolutional downsampling layer may include residual convolution processing and downsampling processing. The input data of the downsampling processing includes the output data of the residual convolution processing. The residual convolution processing may adopt the residual convolution processing proposed by ResNet-50, but is not limited thereto. The residual convolution processing may specifically include convolution processing, batch normalization processing, activation function processing, etc. For example, Figure 4 is a logical schematic diagram of an example of the residual convolution processing provided by the embodiments of the present application. As Figure 4 shown, first, perform 1×1 convolution processing on the data, then perform batch normalization processing and activation function processing. The activation function processing may specifically be ReLU (Rectified Linear Unit) processing; then sequentially perform 3×3 convolution processing, batch standard processing and activation function processing, 3×3 convolution processing and batch normalization processing. The output data can be merged with the input data through residual connection and then perform activation function processing again; the batch standard processing can accelerate the training process and improve the stability of the neural network model. The ReLU function processing can introduce non-linearity, enabling the neural network model to learn more complex features. The downsampling processing may specifically include: respectively performing average pooling processing and maximum pooling processing on the output data of the residual convolution processing, merging the data after average pooling processing and the data after maximum pooling, obtaining the first merged data, and performing convolution processing on the first merged data to obtain the output data of the downsampling processing. For example, Figure 5A logical schematic diagram of an example of downsampling processing provided by an embodiment of the present application, as Figure 5 shown, average pooling is performed on the input data for downsampling processing, that is, the output data of residual convolution processing, and max pooling is performed on the input data for downsampling processing, that is, the output data of residual convolution processing. Indicates the merging process. The data after the merging process, that is, the first merged data, is subjected to 1×1 convolution processing to obtain the first feature. Through this downsampling processing method, the information emphasized by average pooling is different from the information emphasized by max pooling. The information emphasized by average pooling and the information emphasized by max pooling can complement each other, retain the feature information of blood vessels to a greater extent, and minimize or even avoid the loss of the structural information of blood vessels.

[0042] In step S102, each first feature is respectively input into the first projection learning branch and the second projection learning branch for processing to obtain a second feature corresponding to each first feature.

[0043] The measurement of blood vessels in the retina is quantified on the projection map rather than in three-dimensional space. In the embodiment of the present application, the first projection learning branch and the second projection learning branch can reduce the depth of the first feature along the projection direction to 1, so that when training the neural network model used in the present application, the two-dimensional blood vessel segmentation label can supervise the first feature of the three-dimensional volume data. The processing of the first projection learning branch is different from the processing of the second projection learning branch. The processing of the two different projection learning branches can complement each other, and the fusion of the data processed by the two different projection learning branches can learn more diverse features in the depth direction while reducing the depth of the first feature. After being processed by the first projection learning branch and the second projection learning branch, a second feature can be obtained, and each first feature corresponds to a first feature.

[0044] In step S103, the projection map of the obtained OCTA image is input into the second feature encoder, and the projection map is processed by multiple residual convolution downsampling layers in the second feature encoder to obtain a third feature output by the multiple residual convolution downsampling layers.

[0045] The projection image of the OCTA image is a two-dimensional image. The second feature encoder can extract the features of the projection image. The second feature encoder may include a plurality of residual convolutional downsampling layers, and each residual convolutional downsampling layer has a residual convolutional processing function and a downsampling processing function. Each residual convolutional downsampling layer can output a second feature according to the input first feature, and each first feature corresponds to a second feature. The number of second features is consistent with the number of residual convolutional downsampling layers in the second feature encoder. The number of residual convolutional downsampling layers in the second feature encoder is consistent with the number of residual convolutional downsampling layers in the first feature encoder. In some examples, in the second feature encoder, the input data of the first residual convolutional downsampling layer includes the projection image, and the input data of the subsequent residual convolutional downsampling layer includes the third feature output by the previous residual convolutional downsampling layer. The structure of the second feature encoder can refer to the structure of the first feature encoder in the above embodiments, which will not be elaborated here. For example, if the first feature encoder includes three residual convolutional downsampling layers, then the second feature encoder includes three residual convolutional downsampling layers. The projection image is input into the first residual convolutional downsampling layer in the second feature encoder, and the third feature T1 is output. The third feature T1 is input into the second residual convolutional downsampling layer in the second feature encoder, and the third feature T2 is output. The third feature T2 is input into the third residual convolutional downsampling layer in the second feature encoder, and the third feature T3 is output. If the size of the projection image is W×H×1×C, the size of the third feature T1 output by the first residual convolutional downsampling layer is W / 2×H / 2×1×C, the size of the third feature T2 output by the second residual convolutional downsampling layer is W / 4×H / 4×1×C, and the size of the third feature T3 output by the third residual convolutional downsampling layer is W / 16×H / 16×1×4C.

[0046] In some examples, the processing performed by the residual convolutional downsampling layer may include residual convolutional processing and downsampling processing, and the input data of the downsampling processing includes the output data of the residual convolutional processing. The processing method in the residual convolutional downsampling layer in the second feature encoder is basically the same as the processing method in the residual convolutional downsampling layer in the first feature encoder. Refer to the relevant description above, which will not be elaborated here.

[0047] It should be noted that the execution order of step S103 and step S101 is not limited here. Step S103 may be executed before step S101, may be executed after step S101, or may be executed synchronously with step S101.

[0048] In step S104, based on the second feature, the third feature, and the feature decoder, the vascular segmentation data corresponding to the OCTA image is obtained.

[0049] After obtaining the second feature and the third feature, the second feature and the third feature can be combined to perform attention cross-feature fusion processing, fusing the features from the volume data and the features from the projection map, so as to obtain a richer feature representation of blood vessels during the decoding stage using the feature decoder, thereby improving the accuracy of blood vessel segmentation. The attention cross-feature fusion processing can obtain fused features, and the feature decoder decodes the fused features to obtain blood vessel segmentation data, which may include a blood vessel segmentation map, and the blood vessel segmentation map can clearly and accurately show the blood vessels segmented from the retinal image.

[0050] The feature decoder may include multiple transposed convolution upsampling layers, and the number of transposed convolution upsampling layers in the feature decoder is the same as the number of residual convolution downsampling layers in the first feature encoder and the number of residual convolution downsampling layers in the second feature encoder. The fused features are processed layer by layer through the transposed convolution upsampling layers in the feature decoder, and the feature decoder outputs blood vessel segmentation data.

[0051] In the embodiments of the present application, the volume data of the OCTA image is input into the first feature encoder, and the volume data is processed by multiple residual convolution downsampling layers in the first feature encoder to obtain the first feature learned from the volume data; the projection map of the OCTA image is input into the second feature encoder, and the projection map is processed by multiple residual convolution downsampling layers in the second feature encoder to obtain the third feature learned from the projection map; different two projection learning branches are used to perform different processes on the first feature to obtain the second feature that can better retain the spatial information in the depth direction. Fusing the second feature and the third feature can integrate the global information and local information of blood vessels, making the advantageous information of the volume data complementary to the advantageous information of the projection map, and using the feature decoder to restore the blood vessel information to a greater extent during the decoding stage, so as to more accurately segment the blood vessels from the OCTA image, improving the accuracy of blood vessel segmentation and also improving the efficiency of blood vessel segmentation.

[0052] To verify the effectiveness of the vascular segmentation method for OCTA images provided in the embodiments of the present application, the public dataset can be used to conduct experimental comparisons on the UNet method, the CS-Net method, and the vascular segmentation method for OCTA images provided in the embodiments of the application. The OCTA-500 dataset can be used as the dataset. The OCTA-500 dataset includes multimodal retinal image data of 500 subjects. According to different fields of view, the OCTA-500 dataset can be divided into two subsets, namely the OCTA-3M subset and the OCTA-6M subset. The OCTA-3M subset contains retinal images of 200 subjects, with a field of view of 3mm×3mm, and the size of each image is 304×304. The OCTA-6M subset includes retinal images of 300 subjects, with a field of view of 6mm×6mm, and the size of each image is 400×400. Figure 6 It is a comparative schematic diagram of an example of the vascular segmentation results of multiple methods provided in the embodiments of the present application. Figure 6 The present solution in it is the vascular segmentation method for OCTA images provided in the embodiments of the present application. The ground truth is the ground truth of the input image, and it can be obtained from Figure 6 It can be seen that compared with the UNet method and the CS-Net method, the vascular segmentation map obtained by the vascular segmentation method for OCTA images provided in the embodiments of the present application enhances the connectivity between blood vessels, and the overall segmentation is relatively accurate, which can better assist doctors in diagnosing ophthalmic diseases.

[0053] According to the OCTA-3M subset and the OCTA-6M subset, by using the UNet method, the CS-Net method, and the vascular segmentation method for OCTA images provided in the embodiments of the application, indicators such as the average similarity coefficient (i.e., DICE), Jaccard coefficient (i.e., JAC), and balanced accuracy (i.e., BACC) quantitative analysis of the experiments on the OCTA-3M subset and the OCTA-6M subset can be obtained. The definitions of the average similarity coefficient, Jaccard coefficient, and balanced accuracy quantitative analysis can be shown as the following formulas (1) to (3):

[0054]

[0055] Among them, DICE is the average similarity coefficient; TP represents true positive; FP represents false positive; FN represents false negative; JAC is the Jaccard coefficient; BACC is the balanced accuracy quantitative analysis; SE represents the proportion of pixels predicted as positive examples among all positive example pixels; SP represents the proportion of pixels predicted as negative examples among all negative example pixels.

[0056] The average similarity coefficient, Jaccard coefficient, and balanced accuracy obtained from experiments on the OCTA-3M subset and the OCTA-6M subset of the UNet method, CS-Net method, and the vascular segmentation method for OCTA images provided by the application embodiment can be shown separately in Table 1 and Table 2 for quantitative analysis.

[0057] Table 1

[0058]

[0059] Table 2

[0060]

[0061] It can be seen from Table 1 and Table 2 that whether it is the OCTA-3M subset or the OCTA-6M subset, compared with the UNet method and the CS-Net method, the quantitative analysis of the average similarity coefficient, Jaccard coefficient, and balanced accuracy of the present solution, that is, the vascular segmentation method for OCTA images provided by the application embodiment, is higher, the vascular segmentation effect is better, and the accuracy is higher.

[0062] In some embodiments, the first projection learning branch includes a 3D convolutional layer and a unidirectional convolutional layer, and the second projection learning branch includes a 3D convolutional layer and a unidirectional pooling layer. Figure 7 It is a schematic diagram of an example of the first projection learning branch and the second projection learning branch provided by the application embodiment, as Figure 7As shown, the first projection learning branch 231 may include 3D convolutional layers, 3D convolutional layers, unidirectional convolutional layers, 3D convolutional layers, 3D convolutional layers, and unidirectional convolutional layers arranged in sequence. The second projection learning branch 232 may include 3D convolutional layers, 3D convolutional layers, unidirectional pooling layers, 3D convolutional layers, 3D convolutional layers, and unidirectional pooling layers arranged in sequence. The 3D convolutional layer can perform convolutional processing and activation function processing, and can learn planar features in volumetric data. The unidirectional convolutional layer can perform projection direction convolutional processing and activation function processing, the unidirectional pooling layer can perform projection direction pooling processing and activation function processing, and the unidirectional convolutional layer and the unidirectional pooling layer can learn diverse features of volumetric data in the depth direction while reducing the depth of the first feature. The activation function processing can be implemented as ReLU function processing. The first feature can be input into the first projection learning branch and processed sequentially through 3D convolutional layers, 3D convolutional layers, unidirectional convolutional layers, 3D convolutional layers, 3D convolutional layers, and unidirectional convolutional layers to obtain a first sub-feature; the first feature can be input into the second projection learning branch and processed sequentially through 3D convolutional layers, 3D convolutional layers, unidirectional pooling layers, 3D convolutional layers, 3D convolutional layers, and unidirectional pooling layers to obtain a second sub-feature; the first sub-feature and the second sub-feature are combined to obtain a second feature. The first sub-feature is the output data of the first projection learning branch, and the second sub-feature is the output data of the second projection learning branch. Specifically, the first sub-feature and the second sub-feature can be added to obtain the second feature. The second feature can be obtained according to the following equations (4) to (6):

[0063]

[0064] y = F up (x) = UDConv(Conv3(Conv3(x)))(5)

[0065] y = F low (x) = UDPool(Conv3(Conv3(x)))(6)

[0066] where, F up (x) is the mapping function synthesized by 3D convolutional layers, 3D convolutional layers, and unidirectional convolutional layers in the first projection learning branch; F low (x) is the mapping function synthesized by 3D convolutional layers, 3D convolutional layers, and unidirectional pooling layers in the second projection learning branch; Conv3 is the calculation identifier of the 3D convolutional layer, and the kernel size can be 3×3×3; UDConv is the unidirectional convolutional layer, and the kernel size can be 1×1×D / 2; UDPool is the unidirectional pooling layer, and the kernel size can be 1×1×D / 2; is the i-th second feature, and the size can be W / 4×H / 4×1×C.

[0067] In some embodiments, the channel attention matrices corresponding to the second feature and the third feature may be obtained first, and then the fused feature may be obtained based on the channel attention matrices, and the vascular segmentation data may be obtained according to the fused feature. Specifically, matrix transformation processing including two pooling processes is respectively performed on the second feature and the third feature to obtain a first channel attention matrix corresponding to the second feature and a second channel attention matrix corresponding to the third feature; multiple fused features are obtained according to the first channel attention matrix, the second channel attention matrix, the second feature, and the third feature; and the multiple fused features are processed by multiple deconvolution upsampling layers in the feature decoder to obtain the vascular segmentation data.

[0068] The matrix transformation processing includes two different pooling processes, and the two different pooling processes may include max pooling processing and average pooling processing. The max pooling processing and the average pooling processing can obtain the global information of the second feature and the third feature. The matrix transformation processing may further include processing by a multilayer perceptron (MLP), merging processing, and activation function processing, and the activation function processing may be implemented as Sigmoid function processing. Matrix transformation processing is performed on the second feature to obtain the first channel attention matrix, and matrix transformation processing is performed on the third feature to obtain the second channel attention matrix.

[0069] In some examples, max pooling processing and average pooling processing are respectively performed on the second feature, the data obtained by the max pooling processing and the data obtained by the average pooling processing are respectively input into the multilayer perceptron, the data output by the multilayer perceptron is merged to obtain the second merged data, and the first channel attention matrix is obtained by processing the second merged data through the activation function; max pooling processing and average pooling processing are respectively performed on the third feature, the data obtained by the max pooling processing and the data obtained by the average pooling processing are respectively input into the multilayer perceptron, the data output by the multilayer perceptron is merged to obtain the third merged data, and the second channel attention matrix is obtained by processing the third merged data through the activation function. The max function is used to process the first channel attention matrix and the second channel attention matrix to obtain the comprehensive channel attention matrix; the second feature and the third feature are respectively enhanced according to the comprehensive channel attention matrix to obtain the first enhanced feature and the second enhanced feature; and convolution processing is performed on the sum of the first enhanced feature and the second enhanced feature to obtain the fused feature.

[0070] Figure 8 A logical schematic diagram of an example for obtaining the fused feature provided by an embodiment of the present application is shown in Figure 8 As shown, for the second feature Perform max pooling and average pooling respectively. Process the data after max pooling with a multi-layer perceptron, and process the data after average pooling with a multi-layer perceptron. The data output by the two multi-layer perceptrons can be added to obtain the second merged data. The second merged data is processed by the Sigmoid function to obtain the first channel attention matrix For the third feature T i Perform max pooling and average pooling respectively. Process the data after max pooling with a multi-layer perceptron, and process the data after average pooling with a multi-layer perceptron. The data output by the two multi-layer perceptrons can be added to obtain the third merged data. The third merged data is processed by the Sigmoid function to obtain the second channel attention matrix Use the max function to map the first channel attention matrix and the second channel attention matrix to determine the larger value among them as the comprehensive channel attention matrix CT i Multiply the second feature by the comprehensive channel attention matrix CT i to enhance the second feature and obtain the first enhanced feature, where the first enhanced feature is the product of the second feature and the comprehensive channel attention matrix CT i Multiply the third feature T i by the comprehensive channel attention matrix CT i to enhance the third feature T i and obtain the second enhanced feature, where the second enhanced feature is the product of the third feature T i and the comprehensive channel attention matrix CT i The matrix multiplication used for enhancement can enhance the valuable channel features in the second feature and the third feature T i and suppress the worthless channel features. Perform 1×1 convolution processing on the sum of the first enhanced feature and the second enhanced feature to obtain the fused feature FusCT i The fused feature can be obtained according to the following equations (7) to (10):

[0071]

[0072] where is the first channel attention matrix; is the second channel attention matrix; Sigmoid() is the Sigmoid function; MLP is the processing of the multi-layer perceptron; GMP() is the max pooling processing; GAP() is the average pooling processing; is the second feature; T iis the third feature; Max() is the maximum value function; is matrix addition; is matrix multiplication; CT i is the comprehensive channel attention matrix; FusCT i is the fused feature; 1×1Conv() is the 1×1 convolution process.

[0073] The fused feature is input to the feature decoder. The feature decoder includes multiple deconvolution upsampling layers, and the deconvolution upsampling layer can have deconvolution processing function and upsampling processing function. The deconvolution processing function can use the residual convolution structure provided by ResNet-50 to design the deconvolution. The deconvolution upsampling layers in the feature decoder correspond one-to-one with the residual convolution downsampling layers in the first feature encoder and the residual convolution downsampling layers in the second feature encoder. In the feature decoder, the input data of the last deconvolution upsampling layer includes the fused feature corresponding to the last residual convolution downsampling layer in the first feature encoder and the last residual convolution downsampling layer in the second feature encoder. The output data of the first deconvolution upsampling layer includes the vascular segmentation data, and the input data of the previous deconvolution upsampling layer includes the merged data of the fused feature corresponding to the previous residual convolution downsampling layer and the output data of the next deconvolution upsampling layer.

[0074] Figure 9 is a schematic diagram of an example of the neural network model provided by the embodiment of the present application, Figure 9 The processing of the volume data 22 in the first feature encoder 21 can be referred to the relevant description above about Figure 3 and will not be elaborated here. As Figure 9 shown, the first feature C1 output by the residual convolution downsampling layer 211 can be input to the dual-branch projection learning layer 291, the first feature C2 output by the residual convolution downsampling layer 212 can be input to the dual-branch projection learning layer 292, the first feature C3 output by the residual convolution downsampling layer 213 can be input to the dual-branch projection learning layer 293. Each of the dual-branch projection learning layer 291, the dual-branch projection learning layer 292, and the dual-branch projection learning layer 293 includes a first projection learning branch and a second projection learning branch. The processing of the first projection learning branch and the second projection learning branch can be referred to the relevant content above and will not be elaborated here. The dual-branch projection learning layer 291 outputs the second feature , the dual-branch projection learning layer 292 outputs the second feature , the dual-branch projection learning layer 293 outputs the second feature 。The projection image 25 is input into the residual convolutional downsampling layer 241 in the second feature encoder 24. The residual convolutional downsampling layer 241 outputs the third feature T1. The third feature T1 is input into the residual convolutional downsampling layer 242. The residual convolutional downsampling layer 242 outputs the third feature T2. The third feature T2 is input into the residual convolutional downsampling layer 243. The residual convolutional downsampling layer 243 outputs the third feature T3. The second feature and the third feature T1 are input into the feature fusion layer 261. The second feature and the third feature T2 are input into the feature fusion layer 262. The second feature and the third feature T3 are input into the feature fusion layer 263. Each of the feature fusion layer 261, the feature fusion layer 262, and the feature fusion layer 263 can perform the function of obtaining the fusion feature according to the second feature and the third feature. For the specific content, refer to the above text and will not be elaborated here. The feature fusion layer 261 outputs the fusion feature FusCT1. The feature fusion layer 262 outputs the fusion feature FusCT2. The feature fusion layer 263 outputs the fusion feature FusCT3. The fusion feature FusCT3 is input into the transposed convolutional upsampling layer 273 in the feature decoder 27. The fusion feature FusCT2 and the output data of the transposed convolutional upsampling layer 273 are merged and then input into the transposed convolutional upsampling layer 272 in the feature decoder 27. The fusion feature FusCT1 and the output data of the transposed convolutional upsampling layer 272 are merged and then input into the transposed convolutional upsampling layer 271 in the feature decoder 27. The transposed convolutional upsampling layer 271 outputs the vascular segmentation data 28. Among them, the upsampling process in the transposed convolutional upsampling layer 273 and the transposed convolutional upsampling layer 272 can be 2-fold upsampling. The upsampling process in the transposed convolutional upsampling layer 271 can be 4-fold upsampling. The size of the output data of the transposed convolutional upsampling layer 273 can be W / 8×H / 8×2C. The size of the output data of the transposed convolutional upsampling layer 272 can be W / 4×H / 4×C. The size of the output data of the transposed convolutional upsampling layer 271 can be W×H×1.

[0075] In some embodiments, the neural network model in the embodiments of the present application can be trained in an end-to-end manner, and the loss function of the neural network model is obtained through the ground truth of the OCTA image samples, and the training of the neural network model is supervised through the loss function until a neural network model that meets the training requirements is obtained. Specifically, the neural network model can be trained according to the sample volume data and sample projection map of the OCTA image; the loss function of the neural network model is obtained; when the loss function does not meet the training cut-off condition, the model parameters of the neural network model are adjusted, and the neural network after adjusting the model parameters is trained until the latest obtained loss function meets the training cut-off condition. The neural network model includes a first feature encoder, a second feature encoder, a first projection learning branch, a second projection learning branch, and a feature decoder. In some examples, the neural network model may further include a function of obtaining fused features, that is, the feature fusion layer in the above embodiments. The training cut-off condition may include that the loss function is less than a preset loss threshold, or the loss function reaches the minimum, and the training cut-off condition is not limited here. The loss function is obtained by using a weight algorithm for the Dice loss function and the cross-entropy loss function. The Dice loss function has a corresponding weight coefficient, and the cross-entropy loss function has a corresponding weight coefficient. For example, the Dice loss function, the cross-entropy loss function, and the loss function can be according to the following formulas (11) to (13):

[0076]

[0077] L total =αL Dice +βL BCE (13)

[0078] where L Dice is the Dice loss function; G i is the ground truth of the sample; Y i is the predicted data output by the neural network model; N is the number of samples; L BCE is the cross-entropy loss function; α is the weight coefficient of the Dice loss function; β is the weight coefficient of the cross-entropy loss function; L total is the loss function.

[0079] The vascular segmentation method of the OCTA image provided by the embodiments of the present application integrates the features of the retinal blood vessels generated by different modalities, has higher overall segmentation and stronger robustness, improves the consistency and accuracy of the vascular segmentation results, and effectively solves the technical problem of vascular structure segmentation.

[0080] The present application also provides a vascular segmentation device for OCTA images, Figure 10 is a schematic structural diagram of the vascular segmentation device for OCTA images provided by an embodiment of the present application, as Figure 10As shown, the vascular segmentation device 300 for the OCTA image may include a first encoding processing module 301, a projection learning branch module 302, a second encoding processing module 303, and a fusion decoding processing module 304.

[0081] The first encoding processing module 301 may be configured to input the volume data of the acquired optical coherence tomography angiography (OCTA) image into a first feature encoder, and process the volume data by using a plurality of residual convolutional downsampling layers in the first feature encoder to obtain first features output by the plurality of residual convolutional downsampling layers.

[0082] The projection learning branch module 302 may be configured to input each first feature into a first projection learning branch and a second projection learning branch respectively for processing to obtain second features corresponding to each first feature, and the processing of the first projection learning branch is different from the processing of the second projection learning branch.

[0083] The second encoding processing module 303 may be configured to input the projection map of the acquired OCTA image into a second feature encoder, and process the projection map by using a plurality of residual convolutional downsampling layers in the second feature encoder to obtain third features output by the plurality of residual convolutional downsampling layers.

[0084] The fusion decoding processing module 304 may be configured to obtain vascular segmentation data corresponding to the OCTA image based on the second features, the third features, and a feature decoder.

[0085] In some embodiments, in the first feature encoder, the input data of the first residual convolutional downsampling layer includes volume data, and the input data of a subsequent residual convolutional downsampling layer includes the first features output by the previous residual convolutional downsampling layer. In the second feature encoder, the input data of the first residual convolutional downsampling layer includes the projection map, and the input data of a subsequent residual convolutional downsampling layer includes the third features output by the previous residual convolutional downsampling layer.

[0086] In some embodiments, the processing performed by the residual convolutional downsampling layer includes a residual convolution process and a downsampling process, and the input data of the downsampling process includes the output data of the residual convolution process. The downsampling process includes: respectively performing average pooling processing and maximum pooling processing on the output data of the residual convolution process, combining the data after the average pooling processing and the data after the maximum pooling, obtaining first combined data, and performing convolution processing on the first combined data to obtain the output data of the downsampling process.

[0087] In some embodiments, the first projection learning branch includes a 3D convolutional layer and a unidirectional convolutional layer, and the second projection learning branch includes a 3D convolutional layer and a unidirectional pooling layer.

[0088] The projection learning branch module 302 can be specifically used for: inputting the first feature into the first projection learning branch, and processing it successively through a 3D convolutional layer, a 3D convolutional layer, a unidirectional convolutional layer, a 3D convolutional layer, a 3D convolutional layer, and a unidirectional convolutional layer to obtain a first sub-feature; inputting the first feature into the second projection learning branch, and processing it successively through a 3D convolutional layer, a 3D convolutional layer, a unidirectional pooling layer, a 3D convolutional layer, a 3D convolutional layer, and a unidirectional pooling layer to obtain a second sub-feature; merging the first sub-feature and the second sub-feature to obtain a second feature.

[0089] In some embodiments, the fusion decoding processing module 304 can be used for: respectively performing matrix transformation processing including two types of pooling processing on the second feature and the third feature to obtain a first channel attention matrix corresponding to the second feature and a second channel attention matrix corresponding to the third feature; obtaining a plurality of fusion features according to the first channel attention matrix, the second channel attention matrix, the second feature, and the third feature; using a plurality of deconvolutional upsampling layers in the feature decoder to process the plurality of fusion features to obtain vascular segmentation data.

[0090] In some examples, the fusion decoding processing module 304 can be specifically used for: respectively performing maximum pooling processing and average pooling processing on the second feature, inputting the data obtained by the maximum pooling processing and the data obtained by the average pooling processing into a multi-layer perceptron respectively, merging the data output by the multi-layer perceptron, obtaining second merged data, and processing the second merged data through an activation function to obtain a first channel attention matrix; respectively performing maximum pooling processing and average pooling processing on the third feature, inputting the data obtained by the maximum pooling processing and the data obtained by the average pooling processing into a multi-layer perceptron respectively, merging the data output by the multi-layer perceptron, obtaining third merged data, and processing the third merged data through an activation function to obtain a second channel attention matrix.

[0091] In some examples, the fusion decoding processing module 304 can be specifically used for: processing the first channel attention matrix and the second channel attention matrix by using a maximum value function to obtain a comprehensive channel attention matrix; enhancing the second feature and the third feature respectively according to the comprehensive channel attention matrix to obtain a first enhanced feature and a second enhanced feature; performing convolutional processing on the sum of the first enhanced feature and the second enhanced feature to obtain a fusion feature.

[0092] In some examples, the deconvolution upsampling layers in the feature decoder correspond one-to-one to the residual convolution downsampling layers in the first feature encoder and the residual convolution downsampling layers in the second feature encoder. In the feature decoder, the input data of the last deconvolution upsampling layer includes the fusion features corresponding to the last residual convolution downsampling layer in the first feature encoder and the last residual convolution downsampling layer in the second feature encoder. The output data of the first deconvolution upsampling layer includes blood vessel segmentation data. The input data of a previous deconvolution upsampling layer includes the combined data of the fusion features corresponding to the previous residual convolution downsampling layer and the output data of the subsequent deconvolution upsampling layer.

[0093] In some embodiments, the blood vessel segmentation device 300 for OCTA images may further include a training module. The training module can be used to: train a neural network model according to the sample volume data and sample projection maps of OCTA images, where the neural network model includes a first feature encoder, a second feature encoder, a first projection learning branch, a second projection learning branch, and a feature decoder; obtain the loss function of the neural network model, where the loss function is obtained by using a weight algorithm through a Dice loss function and a cross-entropy loss function; and adjust the model parameters of the neural network model and train the neural network after adjusting the model parameters until the latest obtained loss function meets the training cut-off condition when the loss function does not meet the training cut-off condition.

[0094] It should be noted that the blood vessel segmentation device 300 for OCTA images is a device corresponding to the above-mentioned blood vessel segmentation method for OCTA images. All implementation manners in the above method embodiments are applicable to the embodiments of this device and can achieve the same technical effects.

[0095] The present application also provides a model publishing device in a distributed system. Figure 11 The structural schematic diagram of the blood vessel segmentation device for OCTA images provided by an embodiment of the present application is as Figure 11 shown. The blood vessel segmentation device 400 for OCTA images includes a memory 401, a processor 402, and a computer program stored on the memory 401 and executable on the processor 402.

[0096] In some examples, the above-mentioned processor 402 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.

[0097] The memory 401 may include a Read-Only Memory (ROM), a Random Access Memory (RAM), a magnetic disk storage medium device, an optical storage medium device, a flash memory device, an electrical, optical, or other physical / tangible memory storage device. Thus, generally, the memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the method for vascular segmentation of OCTA images according to embodiments of the present application.

[0098] The processor 402 runs a computer program corresponding to the executable program code by reading the executable program code stored in the memory 401, so as to implement the method for vascular segmentation of OCTA images in the above embodiments.

[0099] In some examples, the vascular segmentation device 400 for OCTA images may further include a communication interface 403 and a bus 404. Among them, as Figure 11 shown, the memory 401, the processor 402, and the communication interface 403 are connected through the bus 404 and complete communication with each other.

[0100] The communication interface 403 is mainly used to implement communication between various modules, devices, units, and / or devices in embodiments of the present application. The input device and / or output device may also be accessed through the communication interface 403.

[0101] The bus 404 includes hardware, software, or both, and couples the components of the OCTA image vascular segmentation device 400 to each other. By way of example and not limitation, the bus 404 may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-E) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local Bus (VLB) bus, or other suitable buses, or a combination of two or more of these. Where appropriate, the bus 404 may include one or more buses. Although embodiments of the present application describe and illustrate specific buses, the present application contemplates any suitable bus or interconnect.

[0102] The present application also provides a computer-readable storage medium, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the OCTA image vascular segmentation method in the above embodiments can be implemented, and the same technical effects can be achieved. To avoid repetition, it will not be described in detail here. Among them, the above computer-readable storage medium may include a non-transitory computer-readable storage medium, such as a Read-Only Memory (ROM), a Random Access Memory (RAM), a magnetic disk, or an optical disc, etc., which is not limited herein.

[0103] The present application also provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the OCTA image vascular segmentation method in the above embodiments can be implemented, and the same technical effects can be achieved. To avoid repetition, it will not be described in detail here.

[0104] It should be clear that the various embodiments in this specification are described in a progressive manner. For the parts that are the same or similar among the various embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. For the device embodiments, equipment embodiments, and computer-readable storage medium embodiments, the relevant parts can refer to the description part of the method embodiments. This application is not limited to the specific steps and structures described above and shown in the figures. Those skilled in the art can make various changes, modifications, and additions, or change the order of the steps after understanding the spirit of this application. And, for the sake of simplicity, the detailed description of known method technologies is omitted here.

[0105] The aspects of this application have been described above with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of this application. It should be understood that each block in the flowchart and / or block diagram, as well as the combinations of blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, or other programmable data processing devices to generate a machine, such that these instructions executed by the processor of the computer or other programmable data processing devices enable the implementation of the functions / actions specified in one or more blocks of the flowchart and / or block diagram. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field-programmable logic circuit. It should also be understood that each block in the block diagram and / or flowchart, as well as the combinations of blocks in the block diagram and / or flowchart, can also be implemented by dedicated hardware that performs the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0106] Those skilled in the art should be able to understand that the above embodiments are all exemplary rather than restrictive. Different technical features that appear in different embodiments can be combined to achieve beneficial effects. Those skilled in the art should be able to understand and implement other variations of the disclosed embodiments based on the study of the drawings, the specification, and the claims. In the claims, the term "comprising" does not exclude other devices or steps; the quantifier "one" does not exclude a plurality; the terms "first" and "second" are used to label names rather than to represent any specific order. Any reference signs in the claims should not be construed as limiting the scope of protection. The functions of multiple parts in the claims can be implemented by a single hardware or software module. The fact that certain technical features appear in different dependent claims does not mean that these technical features cannot be combined to achieve beneficial effects.

Claims

1. A method for vascular segmentation of OCTA images, characterized in that Including: Inputting the volume data of the obtained optical coherence tomography angiography (OCTA) image into a first feature encoder, and processing the volume data by using a plurality of residual convolutional downsampling layers in the first feature encoder to obtain first features output by the plurality of residual convolutional downsampling layers; Respectively inputting each of the first features into a first projection learning branch and a second projection learning branch for processing to obtain second features corresponding to each of the first features, wherein the processing of the first projection learning branch is different from the processing of the second projection learning branch; Inputting the projection map of the obtained OCTA image into a second feature encoder, and processing the projection map by using a plurality of residual convolutional downsampling layers in the second feature encoder to obtain third features output by the plurality of residual convolutional downsampling layers; Based on the second features, the third features, and a feature decoder, obtaining vascular segmentation data corresponding to the OCTA image.

2. The method according to claim 1, wherein In the first feature encoder, the input data of the first residual convolutional downsampling layer includes the volume data, and the input data of a subsequent residual convolutional downsampling layer includes the first feature output by the previous residual convolutional downsampling layer; In the second feature encoder, the input data of the first residual convolutional downsampling layer includes the projection map, and the input data of a subsequent residual convolutional downsampling layer includes the third feature output by the previous residual convolutional downsampling layer.

3. The method according to claim 1 or 2, characterized in that, The processing performed by the residual convolutional downsampling layer includes residual convolution processing and downsampling processing, and the input data of the downsampling processing includes the output data of the residual convolution processing; The downsampling processing includes: respectively performing average pooling processing and maximum pooling processing on the output data of the residual convolution processing, combining the data after the average pooling processing and the data after the maximum pooling to obtain first combined data, and performing convolution processing on the first combined data to obtain the output data of the downsampling processing.

4. The method according to claim 1, wherein The first projection learning branch includes a 3D convolutional layer and a unidirectional convolutional layer, and the second projection learning branch includes a 3D convolutional layer and a unidirectional pooling layer; The step of respectively inputting each of the first features into the first projection learning branch and the second projection learning branch for processing to obtain second features corresponding to each of the first features includes: Inputting the first feature into the first projection learning branch and sequentially processing it through a 3D convolutional layer, a 3D convolutional layer, a unidirectional convolutional layer, a 3D convolutional layer, a 3D convolutional layer, and a unidirectional convolutional layer to obtain a first sub-feature; Inputting the first feature into the second projection learning branch and sequentially processing it through a 3D convolutional layer, a 3D convolutional layer, a unidirectional pooling layer, a 3D convolutional layer, a 3D convolutional layer, and a unidirectional pooling layer to obtain a second sub-feature; Combining the first sub-feature and the second sub-feature to obtain the second feature.

5. The method according to claim 1, characterized in that The step of obtaining the vascular segmentation data corresponding to the OCTA image based on the second features, the third features, and the feature decoder includes: Perform matrix transformation processing including two pooling processes on both the second feature and the third feature respectively to obtain a first channel attention matrix corresponding to the second feature and a second channel attention matrix corresponding to the third feature; Obtain a plurality of fused features according to the first channel attention matrix, the second channel attention matrix, the second feature, and the third feature; Process the plurality of fused features by using a plurality of deconvolution upsampling layers in the feature decoder to obtain the vascular segmentation data.

6. The method according to claim 5, wherein The performing matrix transformation processing including two pooling processes on both the second feature and the third feature respectively to obtain a first channel attention matrix corresponding to the second feature and a second channel attention matrix corresponding to the third feature includes: Perform maximum pooling processing and average pooling processing on the second feature respectively, input the data obtained by the maximum pooling processing and the data obtained by the average pooling processing into a multi-layer perceptron respectively, merge the data output by the multi-layer perceptron to obtain second merged data, and process the second merged data through an activation function to obtain the first channel attention matrix; Perform maximum pooling processing and average pooling processing on the third feature respectively, input the data obtained by the maximum pooling processing and the data obtained by the average pooling processing into a multi-layer perceptron respectively, merge the data output by the multi-layer perceptron to obtain third merged data, and process the third merged data through an activation function to obtain the second channel attention matrix.

7. The method according to claim 5, characterized in that, The obtaining a plurality of fused features according to the first channel attention matrix, the second channel attention matrix, the second feature, and the third feature includes: Process the first channel attention matrix and the second channel attention matrix by using a maximum value function to obtain a comprehensive channel attention matrix; Enhance the second feature and the third feature respectively according to the comprehensive channel attention matrix to obtain a first enhanced feature and a second enhanced feature; Perform convolution processing on the sum of the first enhanced feature and the second enhanced feature to obtain the fused feature.

8. The method according to claim 5, wherein The deconvolution upsampling layers in the feature decoder correspond one-to-one to the residual convolution downsampling layers in the first feature encoder and the residual convolution downsampling layers in the second feature encoder; In the feature decoder, the input data of the last deconvolution upsampling layer includes the fused features corresponding to the last residual convolution downsampling layer in the first feature encoder and the last residual convolution downsampling layer in the second feature encoder, the output data of the first deconvolution upsampling layer includes the vascular segmentation data, and the input data of the previous deconvolution upsampling layer includes the merged data of the fused features corresponding to the previous residual convolution downsampling layer and the output data of the next deconvolution upsampling layer.

9. The method according to claim 1, wherein Further includes: Train a neural network model based on the sample volume data and sample projection map of the OCTA image, where the neural network model includes the first feature encoder, the second feature encoder, the first projection learning branch, the second projection learning branch, and the feature decoder; Obtain the loss function of the neural network model, where the loss function is obtained by using a weight algorithm with the Dice loss function and the cross-entropy loss function; In the case where the loss function does not meet the training cut-off condition, adjust the model parameters of the neural network model, and train the neural network after adjusting the model parameters until the latest obtained loss function meets the training cut-off condition.

10. An apparatus for vascular segmentation of OCTA images, characterized in that, The device includes: A first encoding processing module, configured to input the volume data of the obtained optical coherence tomography angiography (OCTA) image into the first feature encoder, and process the volume data by using multiple residual convolutional downsampling layers in the first feature encoder to obtain first features output by the multiple residual convolutional downsampling layers; A projection learning branch module, configured to input each of the first features into the first projection learning branch and the second projection learning branch for processing respectively to obtain second features corresponding to each of the first features, where the processing of the first projection learning branch is different from the processing of the second projection learning branch; A second encoding processing module, configured to input the projection map of the obtained OCTA image into the second feature encoder, and process the projection map by using multiple residual convolutional downsampling layers in the second feature encoder to obtain third features output by the multiple residual convolutional downsampling layers; A fusion decoding processing module, configured to obtain the vascular segmentation data corresponding to the OCTA image based on the second features, the third features, and the feature decoder.

11. An apparatus for vascular segmentation of OCTA images, characterized in that, The device includes: a processor and a memory storing computer program instructions; the processor reads and executes the computer program instructions to implement the method for segmenting blood vessels in an OCTA image according to any one of claims 1 to 9.

12. A computer-readable storage medium, characterized in that, Computer program instructions are stored on the computer storage medium, and when the computer program instructions are executed by a processor, the method for segmenting blood vessels in an OCTA image according to any one of claims 1 to 9 is implemented.

13. A computer program product, characterized in that, Including a computer program, where when the computer program is executed by a processor, the method for segmenting blood vessels in an OCTA image according to any one of claims 1 to 9 is implemented.