Medical image segmentation method based on multi-axis surface feature fusion and two-dimensional convolutional neural network

By using a two-dimensional convolutional neural network with multi-axial feature fusion in medical image segmentation, the problem of insufficient accuracy of existing algorithms in brain tissue image segmentation is solved, and efficient and accurate medical image segmentation effect is achieved.

CN114581453BActive Publication Date: 2025-05-20CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210251549.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-15
Publication Date
2025-05-20
Estimated Expiration
2042-03-15

AI Technical Summary

Technical Problem

Existing medical image segmentation algorithms are difficult to meet the high accuracy requirements when processing complex brain tissue images. Due to the problems of hardware computing power limitations and high data labeling costs, deep learning algorithms have problems such as training difficulties and slow segmentation speed.

Method used

A two-dimensional convolutional neural network based on multi-axis surface feature fusion is adopted. By performing grayscale normalization and 64*64*64 block random sampling on the image, a multi-axis surface feature fusion network model is built, and three U-shaped network branches are used for feature extraction and fusion, combining residual bottleneck modules and bilinear interpolation for feature transmission and resolution reduction.

Benefits of technology

It realizes high-precision segmentation of three-dimensional medical images, reduces the amount of network parameters, solves the problems of hardware computing power limitation and feature fitting difficulties, and expands to other three-dimensional medical image segmentation tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114581453B_ABST
    Figure CN114581453B_ABST
Patent Text Reader

Abstract

The present invention relates to a medical image segmentation method based on a multi-axis surface feature fusion two-dimensional convolutional neural network, and belongs to the field of image processing. The method comprises: grayscale normalization and block random sampling of training set data; building a multi-axis surface feature fusion network model; rotating the same image block by 90° along the X-axis and Y-axis respectively, obtaining three input image blocks of the same area, slicing the three image blocks, and then inputting the corresponding network branches for feature fusion; performing probability map fusion after the last bilinear interpolation operation of upsampling; after the last convolution operation of the three branches, averaging the three probability maps, and then obtaining the final segmentation result map through an activation function. The present invention makes full use of the image space features, so that the segmentation result is more accurate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image processing and relates to a medical image segmentation method based on a two-dimensional convolutional neural network with multi-axis plane fusion. Background Art

[0002] Magnetic resonance images are widely used in brain imaging due to their high contrast and spatial recognition rate for soft tissues. Since quantitative analysis of the brain tissues to be segmented helps to evaluate the development of various diseases, such as Alzheimer's disease, epilepsy, and schizophrenia, etc. The segmentation and quantitative analysis of brain tissues are the premise for judging the development of diseases. Therefore, the segmentation of brain magnetic resonance images into gray matter, white matter, and cerebrospinal fluid is a very important research topic. Currently, manual segmentation of brain magnetic resonance images (Magnetic Resonance Image, MRI) by professional physicians is still the gold standard. However, the manual pixel-level annotation by professional doctors is a time-consuming and energy-consuming process with a high time cost and is prone to some subjective segmentation errors. Thus, a large number of researchers are engaged in the research of automated segmentation algorithms.

[0003] In the early stage of the research on segmentation algorithms, some researchers proposed a series of traditional methods, including deformation model-based segmentation algorithms such as the Snake model and level set method, as well as single-atlas and multi-atlas registration algorithms based on atlas registration. However, due to the complex boundary texture and small gray-scale variation of brain tissues, traditional algorithms are difficult to meet the high-precision requirements of medical image segmentation. In recent years, segmentation methods based on deep learning have shown their superiority. Deep learning algorithms are data-driven segmentation algorithms. Facing complex segmentation tasks, deep learning algorithms are driven by sufficient data, and artificial neurons learn the abstract features of the target, and finally excellent segmentation results can be obtained. However, due to problems such as the privacy of medical images themselves and high annotation costs, the publicly available datasets are few and the image quality is uneven, which poses high requirements for the design of the network. In medical image segmentation tasks, deep learning algorithms can be mainly divided into two categories: 2D (2-Dimension) segmentation networks and 3D (3-Dimension) segmentation networks. Among them, the number of parameters of 3D segmentation networks increases sharply compared with 2D segmentation networks. Due to hardware computing power limitations, there are problems such as difficult network training and slow segmentation speed, while 2D segmentation networks lack slice context information and have difficulty in feature fitting. Therefore, it is very necessary to add sufficient spatial information to the 2D network. Summary of the Invention

[0004] In view of this, the purpose of the present invention is to provide a medical image segmentation method based on a two-dimensional convolutional neural network with multi-axis plane feature fusion, so as to make full use of the spatial features of the image and make the segmentation result more accurate.

[0005] To achieve the above object, the present invention provides the following technical solutions:

[0006] A medical image segmentation method based on a multi-axis plane feature fusion two-dimensional convolutional neural network, specifically including the following steps:

[0007] S1: Perform gray normalization processing on the training set data, so that the gray intensity of the data is distributed in the range of 0 to 1, enabling the network to focus on the features of the segmentation target itself, strengthening the generalization of the network, and then perform 64*64*64 block random sampling;

[0008] S2: Build a multi-axis plane feature fusion network model, and the model mainly includes three U-shaped network branches.

[0009] S3: Rotate the same image block 90° along the X-axis and Y-axis respectively to obtain three input image blocks of the same area, and slice the three image blocks;

[0010] S4: Input the three image block slices obtained in step S3 into the corresponding network branches respectively. Before the first pooling operation of the three network branches, perform feature map fusion. Rotate the feature maps of the two network branches to the same axial plane as the input image block of the branch to be fused, and then perform feature fusion;

[0011] S5: Perform probability map fusion consistent with step S4 after the last bilinear interpolation operation during upsampling;

[0012] S6: After the last convolution operation of the three branches, average the three probability maps to obtain a value, and then pass through an activation function to obtain the final segmentation result map;

[0013] S7: Adjust the parameters of the network, save the model with the best verification data effect, verify each data, and select the optimal model through multiple cross-validations.

[0014] Further, in step S2, the structure of each network branch includes a downsampling path and an upsampling path; in the downsampling path, a residual bottleneck module is used to extract features, mainly using a 3*3 two-dimensional convolution kernel, and the pooling layer uses a 2*2 maximum pooling kernel three times; the upsampling path includes three upsamplings and a residual bottleneck module, and bilinear interpolation is used for upsampling to restore the resolution. The loss function of each network branch uses a combined loss function of Focal loss and lovász-softmax loss.

[0015] The beneficial effects of the present invention are as follows: The present invention conducts two-dimensional convolution training on three-dimensional images. On the basis of learning sufficient spatial features, it reduces the number of network parameters, and introduces a residual bottleneck module to optimize the feature transfer and gradient transfer of the network. The entire network can be extended to other three-dimensional medical image segmentation tasks. Specifically, the beneficial effects include the following:

[0016] (1) The present invention performs block training on training samples and uses the method of randomly taking blocks to solve the problem of hardware computing power limitation.

[0017] (2) In the downsampling path of the present invention, max-pooling kernels are used for downsampling operations, and effective features are activated through max-pooling.

[0018] (3) In the upsampling path of the network of the present invention, bilinear interpolation is used to restore the resolution, avoiding the checkerboard effect of transposed convolution.

[0019] (4) The present invention extracts features from the input slices through three different axial planes. Each axial plane corresponding to the branch network restores features through three times of downsampling and three times of upsampling, and uses a 2D segmentation network to solve the three-dimensional medical image segmentation problem.

[0020] (5) In the network model with multi-axial plane fusion of the present invention, a feature fusion module is added to fuse the features of different axial planes in the same stage, solving the problem of lack of spatial information in 2D networks.

[0021] Other advantages, objectives, and features of the present invention will be described to some extent in the subsequent specification, and to some extent, will be obvious to those skilled in the art based on the study of the following text, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the following specification. Brief Description of the Drawings

[0022] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be described in detail preferably with reference to the accompanying drawings, where:

[0023] Figure 1 is the flowchart of the three-dimensional brain magnetic resonance image segmentation method based on multi-axial plane fusion UNet in this embodiment;

[0024] Figure 2 is the structural diagram of the multi-axial plane feature fusion network model;

[0025] Figure 3 is the single-branch U-shape network;

[0026] Figure 4 is the multi-axial plane feature fusion module;

[0027] Figure 5This is the overall structural diagram for network training and testing. Specific implementation manners

[0028] The following uses specific specific examples to illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0029] Among them, the attached drawings are only for illustrative purposes, showing only schematic diagrams, not physical diagrams, and should not be construed as limiting the present invention; in order to better illustrate the embodiments of the present invention, some components in the attached drawings will be omitted, enlarged or reduced, which does not represent the size of the actual product; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the attached drawings may be omitted.

[0030] In the attached drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", "front", "rear", etc. indicating the orientation or positional relationship, they are based on the orientation or positional relationship shown in the attached drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, the terms describing the positional relationship in the attached drawings are only for illustrative purposes and should not be construed as limiting the present invention. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific circumstances.

[0031] Please refer to Figures 1 to 5 , Figure 1 This is the flowchart of the three-dimensional brain magnetic resonance image segmentation method based on multi-axis plane fusion UNet provided in this embodiment, which specifically includes the following steps:

[0032] Step 1: Take blocks of the training samples according to a pixel size of 64*64, and use the random sampling method. Each time, take 64 samples and input them into the network for training.

[0033] In this embodiment, the data set is adult brain MR images, and the image resolution is 256*128*256. Due to the hardware computing power limitation problem, the entire image cannot be directly put into the model for training. Therefore, it is necessary to cut the image into blocks for sampling. In the training method, the image is taken in a size of 64*64 to match the current hardware computing power.

[0034] Step 2: Build a multi-axis feature fusion network model.

[0035] As Figure 2 shown, the network mainly consists of three branches, corresponding to the cross-section, coronal plane, and sagittal plane of the image for image input respectively. The network structures of the three branches are the same as Figure 3 shown, both contain a downsampling path and an upsampling path. The downsampling path is mainly composed of residual bottleneck modules and max pooling. A total of three max poolings with a 2*2 pooling kernel and a stride of 2 are performed to ensure that the network can learn the deep abstract features of the image. The corresponding upsampling path also contains three upsamplings and residual bottleneck modules. The upsampling is mainly achieved through bilinear interpolation. The loss function uses a combined loss function of Focal loss and lovász-softmax loss.

[0036] At the same time, each of the downsampling and upsampling paths of the network contains a multi-axis feature fusion module. This module is mainly composed of feature maps at the same level. Since the input image axes of the three branch networks are different, the feature maps need to be rotated to the same axis to be fused before fusion. As Figure 4 shown, taking the fusion module of the cross-section branch as an example, to fuse the features of the coronal plane and sagittal plane into the cross-section, the coronal plane feature map needs to be rotated 90° along the Y axis, and the sagittal plane feature map needs to be rotated 90° along the X axis. The feature maps of the three branches are all in the perspective of the cross-section, and then the feature regions correspond to each other for feature fusion, aiming to supplement the shallow texture features in the branch network and strengthen the shallow feature fitting ability of the network. Similarly, a voxel probability fusion is performed in the upsampling path, and the specific operation is the same as that of the feature fusion module in the downsampling, aiming to increase the fault tolerance rate of the single-branch voxel probability prediction.

[0037] Step 4: Use the multi-axis feature fusion network model to train the training set and tune the parameters of the network model. When verifying the model effect on the validation set, one data is taken as the validation set in each round. Taking the IBSR dataset used as an example, IBSR contains 18 data, 10 of which are taken as the training set, and the other 8 data are taken as the test set. Among the 10 training sets, one data is taken as the validation data, and the validation data does not participate in the training of the network. Since each of the 10 data is used as the validation data once, a total of 10 rounds of training are performed for cross-validation. Finally, the network model saves the version with the best validation data segmentation result;

[0038] Step 5: The entire fully convolutional neural network process is as Figure 5As shown, it mainly includes two stages: the training stage and the testing stage. In the training stage, the network is mainly trained, and the reverse gradient is passed through the loss function to update the weight values in the network. When the loss of the network drops to the minimum value, the network model is saved. The saved network is used to segment the test data. After performing the same normalization process on the test data, it is input into the network for segmentation, and finally the segmentation result is obtained.

[0039] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions, and they should all be covered by the scope of the claims of the present invention.

Claims

1. A medical image segmentation method based on multi-axis surface feature fusion two-dimensional convolutional neural network, characterized in that: The method specifically comprises the following steps: S1: Perform grayscale normalization on the training set data, and then perform random sampling in blocks; S2: Build a multi-axis surface feature fusion network model, which includes three U-shaped network branches; the structure of each network branch includes a downsampling path and an upsampling path; the downsampling path and the upsampling path each contain a multi-axis surface feature fusion module; the multi-axis surface feature fusion module is composed of feature maps of the same level. Since the input image axes of the three branch networks are different, the feature maps need to be rotated to the same axis to be fused before fusion; the downsampling path uses a residual bottleneck module to extract features; the upsampling path includes three upsamplings and a residual bottleneck module; S3: Rotate the same image block by 90° along the X-axis and the Y-axis respectively to obtain three input image blocks of the same area, and slice the three image blocks; S4: The three image block slices obtained in step S3 are respectively input into the corresponding network branches, and feature map fusion is performed before the first pooling operation of the three network branches. The feature maps of the two network branches are respectively rotated to the axial plane consistent with the input image block of the branch to be fused, and then feature fusion is performed; S5: after the last bilinear interpolation operation of upsampling, perform probability map fusion consistent with step S4; S6: After the last convolution operation of the three branches, the three probability maps are averaged and then passed through the activation function to obtain the final segmentation result map; S7: Adjust the parameters of the network, save the model with the best verification data, and verify each data. After multiple cross-validations, the optimal model is obtained.

2. The medical image segmentation method based on multi-axis surface feature fusion two-dimensional convolutional neural network according to claim 1 is characterized in that: In step S2, the downsampling path uses a residual bottleneck module to extract features, uses a 3*3 two-dimensional convolution kernel, and the pooling layer uses a three-time 2*2 maximum pooling kernel; the upsampling path includes three upsampling and a residual bottleneck module, and the upsampling uses bilinear interpolation for resolution restoration.

3. The medical image segmentation method based on multi-axis surface feature fusion two-dimensional convolutional neural network according to claim 1 is characterized in that: In step S2, the loss function of each network branch uses a combination of Focal loss and lovász-softmax loss.