Training method, segmentation method and training device for brain CT image segmentation model

By combining encoding networks and decoding networks in the brain CT image segmentation model to perform spatial and channel information fusion, the problem of low segmentation accuracy in the prior art is solved, and more efficient cerebral ischemia detection is achieved.

CN113808085BActive Publication Date: 2025-05-02SHENZHEN INST OF ADVANCED TECH +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110996998.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-08-27
Publication Date
2025-05-02
Estimated Expiration
2041-08-27

AI Technical Summary

Technical Problem

The prior art fails to pay enough attention to the spatial information of the image and the connection between channels in the segmentation of brain CT images, resulting in low segmentation accuracy.

Method used

The combination method of encoding network and decoding network is adopted to extract and aggregate rich feature information through spatial information fusion and channel information fusion, and then train the segmentation model.

Benefits of technology

It improves the segmentation accuracy of brain CT images, enhances the representation ability of the model, and significantly improves the detection performance and efficiency of large-area cerebral ischemia.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113808085B_ABST
    Figure CN113808085B_ABST
Patent Text Reader

Abstract

The present invention discloses a training method, a segmentation method and a training device for a brain CT image segmentation model. The training method comprises: obtaining an input feature map obtained by a coding network to be trained according to a brain CT sample image with label information, obtaining a segmentation result obtained by a decoding network to be trained according to the input feature map, the decoding network performs spatial information fusion processing and channel information fusion processing on feature maps with different channel numbers, and aggregates the obtained spatial fusion features and channel fusion features to obtain a segmentation result; calculating the difference between the segmentation result and the label information of the brain CT sample image, and updating the model parameters of the coding network and the decoding network according to the difference. In the training process, the rich spatial information in the coding stage is extracted by spatial information fusion to improve the accuracy of segmentation, and at the same time, a dynamic non-dependent relationship between channels is established by channel information fusion, the learning process is simplified, and the representation ability of the model is significantly enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of medical image processing, and specifically relates to a training method for a segmentation model for CT images, a segmentation method, a training device, a segmentation device, a computer-readable storage medium, and a computer device. Background Art

[0002] Cerebral infarction is a cerebrovascular disease with a high clinical incidence. This condition results in obstructed cerebral blood circulation, leading to localized brain tissue ischemia and hypoxia, which in turn leads to softening and necrosis. Large-area cerebral infarction in the hyperacute phase is a more severe form of cerebral infarction, posing a serious threat to the patient's life if timely diagnosis and treatment are not provided. Computed tomography (CT) is widely used in the clinic for the rapid diagnosis of ischemic cerebral infarction due to its rapid imaging speed and low cost. Currently, CT image analysis is primarily performed by physicians based on experience. Well-trained radiologists can identify lesions well, but consistency in their assessment of the extent of ischemia is poor. Furthermore, clinically, it is difficult for physicians to determine the extent of early ischemic changes, especially during the hyperacute phase. Furthermore, ischemic lesions are currently often segmented manually, but manual segmentation is time-consuming and relies heavily on subjective judgment by the operator. However, the existing image processing of hyperacute cerebral ischemia has low precision, large errors, and inaccurate detection. At the same time, if the risk of stroke cannot be assessed in a timely manner, timely treatment will not be possible, delaying the disease.

[0003] It is worth noting that in the existing research on stroke CT images, including related research papers and patents, most cerebral ischemia segmentation and detection methods are based on traditional image processing algorithms. Traditional image processing methods often require a lot of computing power to calculate the shape, grayscale, texture features, etc. of the image, and the detection speed and accuracy are not high.

[0004] Some researchers have also proposed using deep learning methods, using convolutional neural networks for segmentation tasks, to overcome the initial difficulties in extracting image features, improve segmentation speed, and achieve good segmentation accuracy. However, the convolutional neural networks currently used in deep learning are relatively simple, and the inclusion of fully connected layers results in a large number of overall network training parameters, complex computations, large amounts of information, long network training times, and poor segmentation accuracy. The improved fully convolutional networks based on this approach still have low overall segmentation accuracy, and pixel-based classification fails to consider relationships between pixels and lacks spatial consistency. Summary of the Invention

[0005] (1) Technical Problems to be Solved by the Present Invention

[0006] The technical problem solved by the present invention is: how to provide a segmentation model that can fully pay attention to the spatial information of the image and the connection between channels.

[0007] (2) Technical solution adopted by the present invention

[0008] A method for training a brain CT image segmentation model, wherein the segmentation model to be trained includes an encoding network and a decoding network, and the training method includes:

[0009] Obtaining an input feature map obtained by the encoding network to be trained based on a brain CT sample image with labeled information, wherein the input feature map includes feature maps with different numbers of channels;

[0010] Obtaining a segmentation result obtained by a decoding network to be trained based on the input feature map, wherein the decoding network performs spatial information fusion processing and channel information fusion processing on feature maps with different numbers of channels, and aggregates the obtained spatial fusion features and channel fusion features to obtain a segmentation result;

[0011] The difference between the segmentation result and the label information of the brain CT sample image is calculated, and the model parameters of the encoding network and the decoding network are updated according to the difference.

[0012] Preferably, the method for obtaining the input feature map of the encoding network to be trained based on the brain CT sample image with label information includes:

[0013] Performing convolution processing on the brain CT sample image with label information to obtain an underlying feature map;

[0014] The bottom layer feature map is sequentially subjected to several convolution and pooling processes to obtain multiple intermediate layer feature maps with increasing number of channels, wherein the input of the first convolution and pooling process is the bottom layer feature map, and after each convolution and pooling process, an intermediate layer feature map is output and used as the input of the next convolution and pooling process;

[0015] Non-local attention processing is performed on the intermediate layer feature map output after the last convolution and pooling processing to obtain a high-level feature map. The bottom layer feature map, other intermediate layer feature maps except the intermediate layer feature map output after the last convolution and pooling processing, and the high-level feature map constitute the input feature map.

[0016] Preferably, the convolution pooling processing method includes:

[0017] Perform two convolution processes and one maximum pooling process on the input to obtain the output features;

[0018] The number of channels of the feature to be output is doubled to obtain an intermediate layer feature map.

[0019] Preferably, the decoding network includes several fusion modules cascaded in sequence from high level to low level, each of the fusion modules includes a spatial fusion unit, a channel fusion unit, an aggregation unit and an upsampling convolution unit, the spatial fusion unit is used to perform weighted processing on the spatial information of the feature map, the channel fusion unit is used to perform weighted processing on the channel information of the feature map, the aggregation unit is used to aggregate the output data of the spatial fusion unit and the channel fusion unit, the upsampling convolution unit is used to upsample, deconvolve and convolve the data output by the aggregation unit of the fusion module of the previous level, and use the obtained data as input data of the channel fusion unit, wherein the input data of the upsampling convolution unit of the highest-level fusion module is the high-level feature map, the input data of the spatial fusion unit of each fusion module is the other feature maps in the input feature map except the high-level feature map, and the number of channels of the input data of the spatial fusion unit decreases with the level, and the output data of the aggregation unit of the lowest-level fusion module is the segmentation result.

[0020] Preferably, the method for the spatial fusion unit to perform weighted processing on the spatial information of the feature map includes:

[0021] Calculate the average value and maximum value of the spatial information set of each pixel of the input feature map respectively to obtain the average value feature map and the maximum value feature map;

[0022] Convolution processing is performed on the average feature map and the maximum feature map respectively, and activated by a PReLu function to obtain a spatial information weight;

[0023] Output data of the spatial fusion unit is obtained by performing a matrix multiplication operation on the spatial information weight and the input feature map.

[0024] Preferably, the method for the channel fusion unit to perform weighted processing on the channel information of the feature map includes:

[0025] Perform global maximization and global average pooling on the input feature maps respectively, and perform matrix addition operation on the results obtained by global maximization and global average pooling to obtain the channel information weight;

[0026] Perform a sigmoid transformation on the input feature map according to the channel information weight to obtain output data of the channel fusion unit.

[0027] The present application also discloses a brain CT image segmentation method, the segmentation method comprising:

[0028] Acquiring a brain CT image to be tested;

[0029] The brain CT image is input into a brain CT image segmentation model trained according to the above training method, and the segmentation model outputs a detection result.

[0030] The present application also discloses a training device for a brain CT image segmentation model, the training device comprising:

[0031] A first acquisition unit is used to obtain an input feature map obtained by the encoding network to be trained based on the brain CT sample image with label information, wherein the input feature map includes a plurality of feature maps with different channel numbers;

[0032] a second acquisition unit, configured to acquire a segmentation result obtained by a decoding network to be trained based on the input feature map, wherein the decoding network performs spatial information fusion processing and channel information fusion processing on feature maps with different numbers of channels, and aggregates the obtained spatial fusion features and channel fusion features to obtain a segmentation result;

[0033] A training unit is configured to calculate a difference between the segmentation result and label information of a brain CT sample image, and update model parameters of the encoding network and the decoding network according to the difference.

[0034] The present application also discloses a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, any of the above methods is implemented.

[0035] The present application also discloses a computer device, which includes a computer-readable storage medium, a processor, and a computer program stored in the computer-readable storage medium. When the computer program is executed by the processor, any one of the above methods is implemented.

[0036] (3) Beneficial effects

[0037] The present invention discloses a training method and a segmentation method for a brain CT image segmentation model. Compared with existing methods, the present invention has the following technical effects:

[0038] During the training process, the spatial fusion unit is used to extract the rich spatial information in the encoding stage, so that the decoding layer can also utilize the rich spatial information in the shallow layer, improving the segmentation accuracy. At the same time, the channel fusion unit is used to establish a dynamic non-dependent relationship between channels, simplifying the learning process and significantly enhancing the representation ability of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 This is a flowchart of a method for training a brain CT image segmentation model according to the first embodiment of the present invention;

[0040] Figure 2This is a flowchart of extracting an input feature map of the encoding network according to the first embodiment of the present invention;

[0041] Figure 3 Schematic diagram of the data processing process of the segmentation model of the first embodiment of the present invention;

[0042] Figure 4 This is a data fusion diagram of a spatial fusion unit according to the first embodiment of the present invention;

[0043] Figure 5 This is a data fusion diagram of a channel fusion unit according to the first embodiment of the present invention;

[0044] Figure 6 Schematic diagram of a training device for a brain CT image segmentation model according to a third embodiment of the present invention;

[0045] Figure 7 Schematic diagram of a computer device according to a fourth embodiment of the present invention. DETAILED DESCRIPTION

[0046] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0047] Before describing the various embodiments of the present application in detail, the inventive concept of the present application is briefly described first: when using deep learning methods to segment brain CT images in the prior art, the spatial information of the image and the channel connection between pixels are not fully considered, and the segmentation accuracy of the model is not high. To this end, the present application provides a training method for a brain CT image segmentation model, which first uses an encoding network to extract several feature maps with rich image information from brain CT sample images, then uses a decoding network to perform spatial information fusion processing and channel information fusion processing on several feature maps with different numbers of channels, and aggregates the obtained spatial fusion features and channel fusion features to obtain a segmentation result, and finally adjusts the model parameters according to the difference between the segmentation result and the label information of the brain CT sample image.

[0048] Specifically, if Figure 1 As shown, the brain CT image segmentation model of the first embodiment includes two parts: an encoding network and a decoding network. The training method of the brain CT image segmentation model includes the following steps:

[0049] Step S10: obtaining an input feature map obtained by the encoding network to be trained based on the brain CT sample image with label information, wherein the input feature map includes feature maps with different numbers of channels;

[0050] Step S20: obtaining a segmentation result obtained by the decoding network to be trained based on the input feature map, wherein the decoding network performs spatial information fusion processing and channel information fusion processing on feature maps with different numbers of channels, and aggregates the obtained spatial fusion features and channel fusion features to obtain a segmentation result;

[0051] Step S30: Calculate the difference between the segmentation result and the label information of the brain CT sample image, and update the model parameters of the encoding network and the decoding network according to the difference.

[0052] The brain CT images of the first embodiment are taken as an example of CT images of large-area cerebral ischemia in the hyperacute phase. Before the training method of the first embodiment is performed, data processing is first performed, including the following steps:

[0053] 1. Data collection and preprocessing: Collect CT image data of large-area cerebral ischemia in the hyperacute phase and annotate them; since the sizes of CT images collected by different machines are inconsistent, the CT image slices need to be cropped to the same size (512×512). At the same time, since the Hounsfield Unit (HU) of each tissue in the CT image varies greatly, it is necessary to select a window to better represent the ischemic area in the brain parenchyma. In this embodiment 1, we selected a window value of -30-100HU (image pixel values ​​greater than 100Hu are set to 100, those less than -30 are set to -30, and the rest remain unchanged) to better display the diseased tissue. The data set is then divided into corresponding training set, validation set, and test set.

[0054] 2. Data Augmentation: The original training data is relatively monotonous, lacking the information richness required for the network. Insufficient and monotonous data can reduce the generalization of network learning, so data augmentation and enhancement are necessary. We randomly crop, rotate, and shift each image slice, performing geometric transformations such as random cropping, rotation, and translation. We also randomly blur, sharpen, distort, perform edge detection, and add noise to the slices with a 50% probability. Furthermore, the large number of healthy tissue slices in the case results in a significant data imbalance. Removing all healthy tissue further reduces the network's generalization. To mitigate these issues, we perform targeted data augmentation, augmenting the slices with disease once every five times and the non-lesioned areas once.

[0055] 3. Data normalization: To facilitate network training, the original brain CT images and the gold standard of brain ischemic areas need to be normalized. Linear normalization is used here to normalize the grayscale data to the [0, 1] interval. The formula is:

[0056]

[0057] where X norm is the normalized data, X is the original data, and X max 、X min are the maximum and minimum values ​​of the original data set, respectively. Before entering the segmentation model, the brain ischemic area data with a grayscale of 0 / 255 is normalized from 0 to 1 as the gold standard of the ischemic area. The data is divided by 255 and 0.5 is used as the threshold. Values ​​above 0.5 are set to 1, and values ​​below 0.5 are set to 0.

[0058] Furthermore, in step S10, the method for obtaining the input feature map of the encoding network to be trained based on the brain CT sample image with label information includes the following steps:

[0059] Step S101: performing convolution processing on the brain CT sample image with label information to obtain an underlying feature map;

[0060] Step S102: performing convolution and pooling processing on the underlying feature map several times in sequence to obtain multiple intermediate layer feature maps with increasing number of channels, wherein the input of the first convolution and pooling processing is the underlying feature map, and after each convolution and pooling processing, an intermediate layer feature map is output and used as the input of the next convolution and pooling processing;

[0061] Step S103: Perform non-local attention processing on the intermediate layer feature map output after the last convolution pooling process to obtain a high-level feature map. The bottom layer feature map, other intermediate layer feature maps except the intermediate layer feature map output after the last convolution pooling process, and the high-level feature map constitute the input feature map.

[0062] For example, Figure 3As shown, the encoding network of this embodiment is based on the U-Net deep network, which has been expanded and improved. First, a 3×3 convolution operation is performed on the brain CT sample image to obtain a 32-channel feature image Conv1 / 32. Subsequently, four convolution and pooling operations are performed to extract image features. Each convolution and pooling process consists of two repeated 3×3 convolutions. Each convolution layer is followed by a batch normalization layer and a nonlinear activation function PReLu. After two convolutions, a 2×2 maximum pooling operation is performed. After each pooling operation, the number of feature channels is doubled to extract richer image features. Four intermediate layer feature maps Conv2 / 64, Conv3 / 128, Conv4 / 256, and Conv1 / 512 are obtained, with channel numbers of 64, 128, 256, and 512, respectively. After the fourth convolution and pooling process, a non-local attention module (Non-local) is introduced to obtain a high-level feature map. In this way, global context information can be used to increase feature extraction. When calculating the response at a certain location, the module considers the weighted features of all channel positions and spatial positions to improve the detection of cerebral ischemic areas and suppress false positives.

[0063] In step S20, Figure 3 As shown in the figure, the decoding network includes several fusion modules cascaded from high level to low level, each fusion module includes a spatial fusion unit SIF, a channel fusion unit CIF, an aggregation unit CAT and an upsampling unit UWC, the spatial fusion unit is used to perform weighted processing on the spatial information of the feature map, the channel fusion unit CIF is used to perform weighted processing on the channel information of the feature map, the aggregation unit CAT is used to aggregate the output data of the spatial fusion unit SIF and the channel fusion unit CIF, the upsampling convolution unit UWC is used to upsample, deconvolve and convolve the data output by the aggregation unit CAT of the previous level fusion module, and use the obtained data as input data of the channel fusion unit CIF, wherein the input data of the upsampling convolution unit UWC of the highest level fusion module is the high-level feature map, the input data of the spatial fusion unit SIF of each fusion module is the other feature maps in the input feature map except the high-level feature map, and the number of channels of the input data of the spatial fusion unit SIF decreases with the level, and the output data of the aggregation unit CAT of the lowest level fusion module is the segmentation result.

[0064] For example, the number of fusion modules is four, such as Figure 4As shown in FIG, the method for the spatial fusion unit to perform weighted processing on the spatial information of the feature map includes: respectively calculating the average value and maximum value of the spatial information set of each pixel of the input feature map to obtain the average value feature map and the maximum value feature map; respectively performing convolution processing on the average value feature map and the maximum value feature map, and activating them through the PReLu function to obtain the spatial information weight; and performing matrix multiplication operation on the spatial information weight and the input feature map to obtain the output data of the spatial fusion unit. Where X is the input feature of the current spatial fusion unit, F max and F avg They are the maximum and average operations respectively, W is the spatial information weight, and Y is the output feature of the current spatial fusion unit. The channel information at the same position point on the feature map is compressed and fused to the same spatial position.

[0065] For example, the encoded feature data of the same cascade is used as the input of the spatial fusion unit, compressed in the channel dimension, and the average and maximum values ​​of the spatial information sets of each pixel in the feature map are calculated, respectively, to obtain two single-channel two-dimensional feature maps. These two feature maps are then subjected to a 1*1 convolution and activated by the PReLu function to obtain the final spatial information weights. Finally, the obtained spatial information weights are weighted to the input of the spatial fusion unit through a simple matrix multiplication, and the weighted spatial information is used as the output of the spatial fusion unit.

[0066] like Figure 5 As shown, the method for the channel fusion unit CIF to perform weighted processing on the channel information of the feature map includes: performing global maximization processing and global average pooling processing on the input feature map respectively, and performing matrix addition operation on the results obtained by the global maximization processing and the global average pooling processing to obtain the channel information weight; performing sigmoid transformation on the input feature map according to the channel information weight to obtain the output data of the channel fusion unit.

[0067] For example, the spatial information of each channel on the feature map is compressed and fused onto the same channel. The decoded feature information obtained from the lower layer serves as the input to this module, where it is compressed in the spatial dimension to increase the sensitivity of each channel to the effective channel information. Global max pooling and global average pooling are used to fuse the spatial information of each channel to obtain weighted coefficients. The obtained channel information weights are then activated using the PReLu function to control the excitation of each channel. However, information between channels does not exist in isolation and nonlinear interactions exist between them. To simultaneously capture global information while focusing on multiple channel information and strengthen channel interdependence, the channel information weights can be mapped to [0, 1] using a sigmoid transform to capture the correlation of channel information. By establishing inter-channel dependencies, channel-related features and responses can be more effectively adaptively recalibrated. Finally, the obtained channel information weights are weighted to the input of the channel fusion unit (CIF) through a simple matrix multiplication, resulting in the output data of the channel fusion unit (CIF).

[0068] Furthermore, the upsampling convolution unit UWC performs 2*2 deconvolution on the feature information of the upper layer to obtain a feature map with higher resolution. At the same time, the number of channels of the feature information is halved through two convolution layers to reduce information redundancy. The aggregation unit CAT simply connects the channel information obtained by the channel fusion unit CIF and the spatial information obtained by the spatial fusion unit SIF. In this embodiment, the two types of information are simply superimposed.

[0069] In step S30, when calculating the difference between the segmentation result and the label information of the brain CT sample image, it is necessary to select an appropriate loss function. Among them, in medical image segmentation tasks, the Dice loss function is the most widely used loss function, which is used to measure the difference between the predicted result and the gold standard. This method directly optimizes the evaluation criteria and can achieve higher accuracy. However, in the problem of brain ischemic area segmentation, the ischemic area often occupies a smaller part of the entire image, which also causes an extreme imbalance in the data categories. In order to weaken this imbalance, the recall rate of pixel classification is improved at the expense of a certain degree of accuracy. Therefore, in this embodiment 1, the Tversky loss function is preferably used as the loss function for network training. The model is trained using preprocessed brain ischemic CT image data to achieve the optimal convergence state, thereby obtaining a segmentation model. Among them, the specific calculation process of the loss function and the update process of the model parameters are technologies well known to those skilled in the art and will not be described in detail here.

[0070] In the training method of the brain CT image segmentation model of the present embodiment 1, the spatial fusion unit is used to extract the rich spatial information in the encoding stage, so that the decoding layer can also use the rich spatial information in the shallow layer, thereby improving the accuracy of the segmentation. At the same time, the channel fusion unit is used to establish a dynamic non-dependent relationship between channels, which simplifies the learning process and significantly enhances the characterization ability of the model. In addition, a non-local attention module (Non-local) is introduced into the model to increase the extraction of features using global context information. When calculating the response of a certain position, the non-local attention module considers the weighting of the features of all channel positions and spatial positions. In this way, the detection of cerebral ischemic areas is improved and false positives are suppressed. Therefore, the training method provided in the present embodiment 1 greatly improves the segmentation and detection performance and efficiency of the segmentation model for large-area cerebral ischemia in clinical practice.

[0071] Example 2 also provides a method for segmenting brain CT images, which includes the following steps: step S100, obtaining a brain CT image to be detected; step S200, inputting the brain CT image into a brain CT image segmentation model trained according to the training method of Example 1, and the segmentation model outputs a detection result.

[0072] Furthermore, in the actual diagnosis process, it also includes the ischemic area quantification and prediction steps and the result visualization step. Specifically, the brain CT image is input into the trained segmentation model, and the ischemic probability of each pixel in the slice will be obtained. Then the pixels with a probability greater than 0.5 are regarded as ischemic pixels, and the other pixels are regarded as background pixels, so as to obtain the segmentation map of the ischemic area. Secondly, the segmentation map of all pixels is used to calculate the 3D connected components in the segmentation map and remove small connected components to reduce the impact of false positives. Then the pixel sum of all remaining connected components is calculated, and the ischemic volume is obtained based on the actual voxel value. Finally, we calculate the ischemic volume exceeding 71cm 3 Patients with acute ischemic stroke are considered LHI patients. The trained cerebral ischemia segmentation model and detection algorithm can be packaged and deployed on common Windows platforms. For brain CT images showing ischemia, ischemic lesions can be accurately labeled and the presence of large-scale cerebral ischemia can be effectively diagnosed. If no ischemic lesions are detected, the patient does not have ischemic cerebral infarction. This effectively achieves automatic segmentation and detection of hyperacute cerebral ischemia, reducing the time required for manual observation, deliberation, and judgment of large-scale cerebral infarction. It can serve as a computer-assisted tool to provide objective evidence for medical research such as stroke.

[0073] In order to verify the segmentation performance and detection shape of the segmentation model trained by the training method of the first embodiment, experimental verification was performed on the collected hyperacute HLI dataset and hyperacute ischemic stroke segmentation dataset. The specific algorithm is implemented in Python language based on the Keras framework and Tensorflow backend, and is trained using 4 24G TITAN RTXGPUs. 85% of the training data is randomly divided into training set data, and the remaining data is divided into validation set data. The network was trained for 50 epochs using the training set data, and batch_size was set to 16. At the same time, the segmentation network was trained using the adaptive moment estimation (Adam) optimizer, where beta_1 was set to 0.9, beta_2 was set to 0.999, and epsilon was set to 10 -8 , the initial learning rate is set to 10 -4 , and a polynomial decay of 0.9 power is performed in each round.

[0074]

[0075] Table 1. Segmentation performance of different methods on the hyperacute cerebral ischemia dataset

[0076]

[0077] Table 2. Detection performance of different methods on the hyperacute HLI dataset

[0078] Therefore, different methods were used to evaluate segmentation on the hyperacute cerebral ischemia dataset, assuming consistent training parameters. The DSC (Dice similarity coefficient) and IOU (Intersection over union) coefficient results are shown in Table 1. The results of different methods on the hyperacute HLI dataset are shown in Table 2. These two experiments effectively validate the superiority of the segmentation model obtained in Example 1 for automatic segmentation and detection of large-area cerebral ischemia during the hyperacute phase.

[0079] like Figure 6As shown, this third embodiment discloses a training device for a brain CT image segmentation model, the training device comprising a first acquisition unit 300, a second acquisition unit 400, and a training unit 500. The first acquisition unit 300 is used to obtain an input feature map obtained by the encoding network to be trained based on a brain CT sample image with label information, wherein the input feature map includes several feature maps with different numbers of channels; the second acquisition unit 400 is used to obtain a segmentation result obtained by the decoding network to be trained based on the input feature map, wherein the decoding network performs spatial information fusion processing and channel information fusion processing on the feature maps with different numbers of channels, and aggregates the obtained spatial fusion features and channel fusion features to obtain a segmentation result; the training unit 500 is used to calculate the difference between the segmentation result and the label information of the brain CT sample image, and update the model parameters of the encoding network and the decoding network based on the difference.

[0080] This fourth embodiment further discloses a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the training method for the brain CT image segmentation model of the first embodiment or the brain CT image segmentation method of the second embodiment.

[0081] This fifth embodiment also discloses a computer device, at the hardware level, such as Figure 7 As shown, the computer device includes a processor 12, an internal bus 13, a network interface 14, and a computer-readable storage medium 11. The processor 12 reads the corresponding computer program from the computer-readable storage medium and then runs it, forming a request processing device at the logical level. Of course, in addition to software implementation, one or more embodiments of this specification do not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc., that is, the execution subject of the following processing flow is not limited to each logical unit, but can also be hardware or logic devices. The computer-readable storage medium 11 stores a computer program, and when the computer program is executed by the processor, it implements the training method of the brain CT image segmentation model of Example 1 or the brain CT image segmentation method of Example 2.

[0082] Computer-readable storage media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer-readable storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage, quantum memory, graphene-based storage media or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device.

[0083] The above describes in detail the specific implementation methods of the present invention. Although some embodiments have been shown and described, those skilled in the art should understand that these embodiments can be modified and improved without departing from the principles and spirit of the present invention, the scope of which is defined by the claims and their equivalents. These modifications and improvements should also be within the scope of protection of the present invention.

Claims

1. A method for training a segmentation model for a brain CT image, characterized in that: The segmentation model to be trained includes an encoding network and a decoding network, and the training method includes: Obtaining an input feature map obtained by the encoding network to be trained based on a brain CT sample image with labeled information, wherein the input feature map includes several feature maps with different numbers of channels, including: performing convolution processing on the brain CT sample image with labeled information to obtain a bottom-level feature map; performing convolution and pooling processing on the bottom-level feature map in sequence for several times to obtain several intermediate-layer feature maps with increasing numbers of channels, wherein the input of the first convolution and pooling processing is the bottom-level feature map, and after each convolution and pooling processing, an intermediate-layer feature map is output and used as the input of the next convolution and pooling processing; performing non-local attention processing on the intermediate-layer feature map output after the last convolution and pooling processing to obtain a high-level feature map, wherein the bottom-level feature map, other intermediate-layer feature maps except the intermediate-layer feature map output after the last convolution and pooling processing, and the high-level feature map constitute the input feature map; Obtaining a segmentation result obtained by a decoding network to be trained according to the input feature map, wherein the decoding network performs spatial information fusion processing and channel information fusion processing on feature maps with different numbers of channels, and aggregates the obtained spatial fusion features and channel fusion features to obtain a segmentation result; Calculating the difference between the segmentation result and the label information of the brain CT sample image, and updating the model parameters of the encoding network and the decoding network according to the difference; The decoding network includes a plurality of fusion modules cascaded in sequence from high level to low level, each of the fusion modules includes a spatial fusion unit, a channel fusion unit, an aggregation unit and an upsampling convolution unit, the spatial fusion unit is used to perform weighted processing on the spatial information of the feature map, the channel fusion unit is used to perform weighted processing on the channel information of the feature map, the aggregation unit is used to aggregate the output data of the spatial fusion unit and the channel fusion unit, the upsampling convolution unit is used to perform upsampling, deconvolution and convolution processing on the data output by the aggregation unit of the fusion module of the previous level, and use the obtained data as input data of the channel fusion unit, wherein the input data of the upsampling convolution unit of the highest level fusion module is the high-level feature map, the input data of the spatial fusion unit of each fusion module is the other feature maps in the input feature map except the high-level feature map, and the number of channels of the input data of the spatial fusion unit decreases with the level, and the output data of the aggregation unit of the lowest level fusion module is the segmentation result; The method for the spatial fusion unit to perform weighted processing on the spatial information of the feature map includes: respectively calculating the average value and the maximum value of the spatial information set of each pixel of the input feature map to obtain an average value feature map and a maximum value feature map; respectively performing convolution processing on the average value feature map and the maximum value feature map, and activating them through a PReLu function to obtain a spatial information weight; and performing a matrix multiplication operation on the spatial information weight and the input feature map to obtain output data of the spatial fusion unit; The method for the channel fusion unit to perform weighted processing on the channel information of the feature map includes: performing global maximization processing and global average pooling processing on the input feature map respectively, and performing matrix addition operation on the results obtained by the global maximization processing and the global average pooling processing to obtain the channel information weight; performing sigmoid transformation on the input feature map according to the channel information weight to obtain the output data of the channel fusion unit.

2. The method for training a brain CT image segmentation model according to claim 1, characterized in that: The convolution pooling processing method includes: Perform two convolutions and one maximum pooling on the input to obtain the features to be output; The number of channels of the feature to be output is doubled to obtain an intermediate layer feature map.

3. A method for segmenting a brain CT image, characterized in that: The segmentation method comprises: Acquire a brain CT image to be tested; The brain CT image is input into a segmentation model of the brain CT image obtained by training using the training method according to any one of claims 1 to 2, and the segmentation model outputs a detection result.

4. A training device for a brain CT image segmentation model, characterized in that: The training device comprises: A first acquisition unit is used to obtain an input feature map obtained by the encoding network to be trained according to the brain CT sample image with label information, wherein the input feature map includes several feature maps with different numbers of channels, including: performing convolution processing on the brain CT sample image with label information to obtain a bottom-level feature map; performing convolution and pooling processing on the bottom-level feature map in sequence for several times to obtain a plurality of intermediate-layer feature maps with increasing numbers of channels, wherein the input of the first convolution and pooling processing is the bottom-level feature map, and after each convolution and pooling processing, an intermediate-layer feature map is output and used as the input of the next convolution and pooling processing; performing non-local attention processing on the intermediate-layer feature map output after the last convolution and pooling processing to obtain a high-level feature map, wherein the bottom-level feature map, other intermediate-layer feature maps except the intermediate-layer feature map output after the last convolution and pooling processing, and the high-level feature map constitute the input feature map; A second acquisition unit is used to acquire a segmentation result obtained by a decoding network to be trained according to the input feature map, wherein the decoding network performs spatial information fusion processing and channel information fusion processing on feature maps with different numbers of channels, and aggregates the obtained spatial fusion features and channel fusion features to obtain a segmentation result; A training unit, configured to calculate a difference between the segmentation result and the label information of the brain CT sample image, and update model parameters of the encoding network and the decoding network according to the difference; The decoding network includes a plurality of fusion modules cascaded in sequence from high level to low level, each of the fusion modules includes a spatial fusion unit, a channel fusion unit, an aggregation unit and an upsampling convolution unit, the spatial fusion unit is used to perform weighted processing on the spatial information of the feature map, the channel fusion unit is used to perform weighted processing on the channel information of the feature map, the aggregation unit is used to aggregate the output data of the spatial fusion unit and the channel fusion unit, the upsampling convolution unit is used to perform upsampling, deconvolution and convolution processing on the data output by the aggregation unit of the fusion module of the previous level, and use the obtained data as input data of the channel fusion unit, wherein the input data of the upsampling convolution unit of the highest level fusion module is the high-level feature map, the input data of the spatial fusion unit of each fusion module is the other feature maps in the input feature map except the high-level feature map, and the number of channels of the input data of the spatial fusion unit decreases with the level, and the output data of the aggregation unit of the lowest level fusion module is the segmentation result; The method for the spatial fusion unit to perform weighted processing on the spatial information of the feature map includes: respectively calculating the average value and the maximum value of the spatial information set of each pixel of the input feature map to obtain an average value feature map and a maximum value feature map; respectively performing convolution processing on the average value feature map and the maximum value feature map, and activating them through a PReLu function to obtain a spatial information weight; and performing a matrix multiplication operation on the spatial information weight and the input feature map to obtain output data of the spatial fusion unit; The method for the channel fusion unit to perform weighted processing on the channel information of the feature map includes: performing global maximization processing and global average pooling processing on the input feature map respectively, and performing matrix addition operation on the results obtained by the global maximization processing and the global average pooling processing to obtain the channel information weight; performing sigmoid transformation on the input feature map according to the channel information weight to obtain the output data of the channel fusion unit.

5. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 2 is implemented.

6. A computer device, characterized in that: The computer device comprises a computer-readable storage medium, a processor and a computer program stored in the computer-readable storage medium, and the computer program implements the method according to any one of claims 1 to 2 when executed by the processor.

Citation Information

Patent Citations

  • Image semantic segmentation method and device based on codec

    CN111292330A

  • Medical image automatic segmentation method based on multi-path attention fusion

    CN111681252A