Method, system, device, processor and computer-readable storage medium for implementing multi-parameter MRI image lesion segmentation
By using a deep neural network model in prostate cancer foci segmentation combined with multi-parameter MRI data, cascade pyramid convolution and dual input channel attention module, the lack of performance of single sequence MRI in prostate cancer segmentation is solved, and efficient multi-scale foci segmentation and diagnostic support is achieved.
Patent Information
- Application Number
- CN202111326387.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-10
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2041-11-10
AI Technical Summary
When using a single sequence of MRI, existing prostate cancer lesion segmentation methods are prone to ignore different forms of mutual information, resulting in poor model segmentation performance, especially in small-target segmentation difficulties.
Using a deep neural network model method, combining multiple sequence data of multi-parameter nuclear magnetic resonance image (mpMRI) (such as ADC, T2W, DWI), a variety of MRI modal data are fused through the cascaded pyramid convolution module and the dual-input channel attention module to enhance feature extraction and segmentation performance.
Automatic segmentation of multi-scale prostate cancer lesions is achieved, segmentation performance is improved, noise interference is reduced, model classification ability of small target areas is enhanced, and the clinical diagnosis and treatment of prostate diseases is supported.
Smart Images

Figure CN114022462B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of automatic segmentation of medical images, and in particular to the field of semantic segmentation in image processing, and specifically refers to a method, system, device, processor and computer-readable storage medium thereof for implementing multi-parameter lesion segmentation of magnetic resonance imaging based on a deep neural network model. Background Art
[0002] Prostate cancer (PCa) is the second most deadly disease in men after lung cancer. If prostate cancer can be detected and treated as early as possible, the survival rate of patients can be effectively improved. Multi-parametric MRI (mpMRI) is an advanced prostate imaging method that combines conventional prostate MRI sequences with one or more functional imaging techniques. It is considered to be the best imaging technology for the clinical diagnosis of prostate cancer. However, the clinical diagnosis of prostate cancer based on mpMRI requires the professional knowledge of radiologists as a basis, and the judgments of different physicians will have certain deviations.
[0003] Medical image segmentation is a hot topic in the field of medical image analysis, and many scholars have proposed many different segmentation algorithms to address different challenges. Early work on PCa detection and segmentation focused on manual feature selection methods, which used predefined image features to build feature empirical models to achieve PCa lesion segmentation; then deep learning methods were widely used in the field of medical image segmentation, but there are few methods that use CNN to segment PCa lesions from prostate mpMRI. Among the existing prostate cancer lesion segmentation methods, there is a PCa detection method based on T2W images. However, using a single MRI sequence may ignore different forms of mutual information, thereby preventing the model from achieving better segmentation performance; there are also multi-channel codec networks based on mpMRI designed to achieve PCa detection and classification, but there are still problems such as network parameter redundancy and difficulty in segmenting small targets. Summary of the invention
[0004] The purpose of the present invention is to overcome the shortcomings of the above-mentioned prior art and provide a method, system, device, processor and computer-readable storage medium thereof for realizing multi-parameter MRI image lesion segmentation based on a deep neural network model with diverse detection dimensions and a wide range of applications.
[0005] In order to achieve the above objectives, the method, system, device, processor and computer-readable storage medium for implementing multi-parameter MRI image lesion segmentation based on a deep neural network model of the present invention are as follows:
[0006] The method for implementing multi-parameter MRI image lesion segmentation based on a deep neural network model is mainly characterized in that the method comprises the following steps:
[0007] (1) Input any combination of samples of imaging sequences including prostate multi-parameter MRI images, such as ADC, T2W, and DWI, and perform rigid matching operation;
[0008] (2) extracting the region of interest from the processed image and sending it to the prostate cancer lesion segmentation network for feature processing through the encoder;
[0009] (3) The encoder outputs the feature map and inputs it into the cross-connection layer of the cascaded pyramid convolution module for convolution and feature map sampling processing;
[0010] (4) After the decoder performs feature map upsampling, the features output by the cross-connection layer are transmitted to the dual-input channel attention module for feature fusion processing;
[0011] (5) Training the prostate cancer lesion segmentation network of the prostate multi-parameter magnetic resonance imaging to obtain lesion segmentation results.
[0012] Preferably, the step (2) is specifically as follows:
[0013] A preset number of convolution modules in a pre-trained ResNeXt network is used to retain the feature maps of each downsampling layer in each convolution module through an encoder to obtain the number of channels of the corresponding feature map.
[0014] Preferably, the preset number of convolution modules are set as the first five convolution modules in the ResNeXt network, wherein the first convolution module uses a convolution kernel size of 7×7, and the remaining four convolution modules use convolution kernels of sizes 3×3 and 1×1, respectively.
[0015] Preferably, the number of channels of the feature maps obtained by five downsamplings increases successively, and the size of each feature map decreases successively, which are 1 / 2, 1 / 4, 1 / 8, 1 / 16 and 1 / 32 of the original image respectively.
[0016] Preferably, the step (3) specifically comprises the following steps:
[0017] The encoder described in (3.1) outputs feature maps corresponding to four cascaded pyramid convolution modules from the first layer to the fourth layer, wherein the output of each large kernel convolution is fused pixel by pixel with the feature map after 1×1 convolution of the original input feature map as the input of the next convolution;
[0018] The cascaded pyramid convolution module described in (3.2) uses convolution decomposition to decompose a large kernel convolution into a dual-branch structure, where one branch is composed of x×1 and 1×y in series, and the convolution order of the other branch is 1×y and x×1. The outputs of the two branches are added element by element to obtain the final output;
[0019] (3.3) The results of multiple large kernel convolutions are concatenated on the channel to retain the feature information of small target objects to the greatest extent;
[0020] (3.4) According to the difference in the number and size of convolution kernels used by the four groups of cascaded pyramids corresponding to the output of the first four layers of the encoder, the sizes of feature maps of different sizes of the encoder are adapted.
[0021] Preferably, the step (4) specifically comprises the following steps:
[0022] The first channel of the dual-input channel attention module described in (4.1) inputs the first feature map through the cascaded pyramid convolution module in the cross-connection path The second channel is upsampled through the decoding layer and input into the second feature map
[0023] (4.2) Concatenate the first feature map and the second feature map in the channel dimension to obtain Among them, c, h, w and d are the number of channels, height, width and depth of the feature map respectively, C(X 1 ,X 2 ) is the concatenated feature map;
[0024] The dual-input channel attention module described in (4.3) performs a global average pooling operation after fusing the first feature map and the second feature map of the input in the channel dimension, and obtains the global information feature vector according to the following formula:
[0025]
[0026] Among them, i, j, k, and f represent the height, width, depth, and channel of the feature map respectively.
[0027] (4.4) The feature vector dimension is reduced to the number of channels c through 1×1 convolution, and the feature vector is normalized using the Sigmoid activation function. The channel attention vector CA is obtained according to the following formula:
[0028] CA=σ(W×v f + b);
[0029] Among them, W and b are the convolution kernel parameters, v f is the feature graph.
[0030] The attention vector described in (4.5) is multiplied with the output feature map of the cascade pyramid convolution module in the channel dimension to enhance the discriminability of the shallow features of the network, and then the deep features are connected to the output end through residual connection to become the output of the dual-input channel attention module, which is specifically implemented by the following formula:
[0031]
[0032] in, represents the multiplication of channel dimensions, and O is the output feature map of the attention module.
[0033] Preferably, the decoder is specifically composed of four dual-input channel attention modules connected in series. The decoder upsamples the feature vector output by each layer through a bilinear interpolation operation, and fuses it with the feature vector output by the cascaded pyramid convolution module of the previous layer, thereby gradually restoring the feature map to the original input size, and the output layer uses a Softmax function to output the probability of the category to which the pixels of each feature map belong.
[0034] More preferably, the step (5) is to obtain the lesion segmentation result through a loss function, specifically: wherein the total loss function is:
[0035] L′ total =L′ bce +L′ dice ;
[0036] Among them, L′ bce is the pixel-level binary cross entropy loss, L′ dice is the Dice loss, and the two losses are:
[0037]
[0038] Among them, x i,j is the probability of the predicted category, y i,j For real marking.
[0039] The system for implementing multi-parameter MRI image lesion segmentation based on a deep neural network model using the above method has the following main features: the system comprises:
[0040] The target extraction processing module is used to extract the target of the region of interest from the input imaging sequences ADC, T2W, and DWI of the prostate multi-parameter nuclear magnetic resonance image;
[0041] A size unification processing module, connected to the target extraction processing module, is used to perform a rigid registration operation on the extracted target images and unify the sizes of the target images to the same size;
[0042] A lesion segmentation neural network processing module is connected to the size unification processing module and is used to input the target image after size unification processing into the prostate cancer lesion segmentation network of the prostate multi-parameter nuclear magnetic resonance image to perform convolution and feature map sampling processing through an encoder;
[0043] A cascade pyramid convolution module, connected to the lesion segmentation neural network processing module, is used to receive the feature map output by the encoder and input it to the cross-connection layer of the cascade pyramid convolution module for group convolution and residual splicing to retain the feature information of the corresponding target object; and
[0044] The dual-input channel attention module is connected to the cascade pyramid convolution module and is used to fuse the output feature map of the cascade pyramid convolution module in the cross-connection path and the output feature map after feature upsampling of the decoding layer in the channel dimension and perform a global average pooling operation to obtain a feature vector of the corresponding channel.
[0045] The device for implementing multi-parameter MRI image lesion segmentation based on a deep neural network model has the following main features: the device comprises:
[0046] a processor configured to execute computer-executable instructions;
[0047] A memory stores one or more computer executable instructions, which, when executed by the processor, implement the various steps of the method for implementing multi-parameter magnetic resonance image lesion segmentation based on a deep neural network model as described above.
[0048] The processor for implementing multi-parameter MRI image lesion segmentation based on a deep neural network model has the main feature that the processor is configured to execute computer executable instructions. When the computer executable instructions are executed by the processor, the various steps of the method for implementing multi-parameter MRI image lesion segmentation based on a deep neural network model are implemented.
[0049] The main feature of this computer-readable storage medium is that a computer program is stored thereon, and the computer program can be executed by a processor to implement the various steps of the above-mentioned method for multi-parameter magnetic resonance image lesion segmentation based on a deep neural network model.
[0050] The method, system, device, processor and computer-readable storage medium for implementing multi-parameter nuclear magnetic resonance image lesion segmentation based on the deep neural network model of the present invention are adopted, and a multi-parameter sequence prostate cancer lesion segmentation method and framework based on a deep neural network with a codec structure is provided. The segmentation method integrates multiple MRI modality data, adopts modules such as cascaded pyramid convolution and channel attention, fully integrates deep and shallow layer feature information, reduces noise interference, and segments multi-scale prostate cancer lesion targets, providing strong support for the clinical diagnosis and treatment of prostate diseases, reducing the time spent on diagnosis, and providing more effective information processing methods and means for the screening, detection and diagnosis of prostate cancer. Automatic segmentation of MRI prostate cancer lesions can be achieved.
[0051] At the same time, considering that a single MRI sequence may ignore different forms of mutual information, thus hindering the model from achieving better segmentation performance, we use three sequence data for channel merging: T2W, ADC and DWI. The use of ADC and DWI sequences can supplement the lesion feature information and greatly improve the segmentation performance. In view of the large differences in the shapes and sizes of PCa lesions in different cases and the large number of small target areas, a cascaded pyramid convolution module is designed. The cascaded pyramid convolution module can capture the detailed spatial positioning information carried in the feature map generated by the encoder at multiple scales, and fuse local and global information at multiple scales to reduce the loss of spatial positioning information and improve the model's ability to classify pixels. At the same time, the use of the cascaded pyramid convolution module can also improve the under-segmentation of the model. In order to enhance the model's attention to the target area, a dual-input channel attention module is designed, which uses the semantic information of the deep features of the network to guide the shallow output to obtain features with higher discrimination ability and strengthen the feature extraction capabilities of each stage of the network. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 The present invention is a flowchart of a method for implementing multi-parameter MRI image lesion segmentation based on a deep neural network model.
[0053] Figure 2 A schematic diagram of extracting regions of interest and unifying sizes of the multi-parameter MRI image lesion segmentation method based on a deep neural network of the present invention.
[0054] Figure 3 It is the cascaded pyramid convolution module of the present invention.
[0055] Figure 4 It is a schematic diagram of splitting the large kernel convolution of the cascaded pyramid convolution module of the present invention.
[0056] Figure 5 Schematic diagram of the structure of the dual-input channel attention module of the present invention.
[0057] Figure 6 The figure is a schematic diagram of the segmentation results of the multi-parameter MRI image lesion segmentation method implemented by the present invention. DETAILED DESCRIPTION
[0058] In order to more clearly describe the technical content of the present invention, further description is given below in conjunction with specific embodiments.
[0059] Before describing in detail the embodiments according to the present invention, it should be noted that, hereinafter, relational terms such as first and second are used only to distinguish one entity or action from another entity or action, and do not necessarily require or imply any actual such relationship or order between such entities or actions. The terms "comprise", "include" or any other variants are intended to cover non-exclusive inclusion, whereby a process, method, article or device including a series of elements includes not only these elements, but also other elements not explicitly listed, or elements inherent to such process, method, article or device.
[0060] See also Figure 1 As shown, the method for implementing multi-parameter MRI image lesion segmentation based on a deep neural network model comprises the following steps:
[0061] (1) Input any combination of samples of ADC, T2W, and DWI of the imaging sequence of prostate multi-parameter magnetic resonance image (mpMRI) for rigid matching operation and unify the size to the same size, such as Figure 2 As shown;
[0062] The rigid matching operation used above is a conventional image processing method known to ordinary technicians in this technical field, which belongs to the common knowledge in this field and will not be explained separately here.
[0063] (2) extracting the region of interest from the processed image and sending it to the prostate cancer lesion segmentation network for feature processing through the encoder;
[0064] In practical applications, the encoder in the designed mpMRI prostate cancer (PCa) lesion segmentation network can be any commonly used deep neural network structure, taking the pre-trained ResNeXt101 as an example;
[0065] (3) The encoder outputs the feature map and inputs it into the cross-connection layer of the cascaded pyramid convolution module for convolution and feature map sampling processing;
[0066] (4) After the decoder performs feature map upsampling, the features output by the cross-connection layer are transmitted to the dual-input channel attention module for feature fusion processing;
[0067] (5) Training the prostate cancer lesion segmentation network of the prostate multi-parameter magnetic resonance imaging to obtain lesion segmentation results.
[0068] As a preferred embodiment of the present invention, the step (2) is specifically as follows:
[0069] A preset number of convolution modules in a pre-trained ResNeXt network is used to retain the feature maps of each downsampling layer in each convolution module through an encoder to obtain the number of channels of the corresponding feature map.
[0070] As a preferred embodiment of the present invention, the preset number of convolution modules is set as the first five convolution modules in the ResNeXt network, wherein the first convolution module uses a convolution kernel size of 7×7, and the remaining four convolution modules use convolution kernels of sizes 3×3 and 1×1, respectively.
[0071] As a preferred embodiment of the present invention, the number of channels of the feature map obtained by five downsamplings increases successively, and the size of each feature map decreases successively, which are 1 / 2, 1 / 4, 1 / 8, 1 / 16 and 1 / 32 of the original image respectively.
[0072] In actual application, the first five convolution blocks in the pre-trained ResNeXt network are used. Each convolution module uses grouped convolution and residual connection, which can improve the network accuracy without increasing (or even reducing) the complexity of the model. The convolution kernel size used in the first convolution module is 7×7, and the remaining four convolution modules use convolution kernels of sizes 3×3 and 1×1 respectively.
[0073] Through the pre-trained encoder, the feature map of each downsampling layer is retained. The number of channels of the five feature maps obtained by downsampling is increased successively, and the size of the feature maps is decreased successively, which are 1 / 2, 1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the original image respectively.
[0074] See also Figure 3 As shown, as a preferred embodiment of the present invention, the step (3) specifically includes the following steps:
[0075] The encoder described in (3.1) outputs feature maps corresponding to four cascaded pyramid convolution modules from the first layer to the fourth layer, wherein the output of each large kernel convolution is fused pixel by pixel with the feature map after 1×1 convolution of the original input feature map as the input of the next convolution;
[0076] The cascaded pyramid convolution module described in (3.2) uses convolution decomposition to decompose a large kernel convolution into a dual-branch structure, where one branch is composed of x×1 and 1×y in series, and the convolution order of the other branch is 1×y and x×1. The outputs of the two branches are added element by element to obtain the final output;
[0077] (3.3) The results of multiple large kernel convolutions are concatenated on the channel to retain the feature information of small target objects to the greatest extent;
[0078] (3.4) According to the difference in the number and size of convolution kernels used by the four groups of cascaded pyramids corresponding to the output of the first four layers of the encoder, the sizes of feature maps of different sizes of the encoder are adapted.
[0079] In actual applications, the encoder outputs from the first to the fourth layer correspond to four cascaded pyramid convolution modules. The output of each large kernel convolution is fused pixel by pixel with the feature map after the 1×1 convolution of the original input feature map, and then used as the input of the next convolution. The sizes of the large kernel convolutions we use include 15×15, 9×9, and 5×5.
[0080] See also Figure 4 As shown in the figure, large kernel convolution is used in the cascaded pyramid convolution module. To reduce the computational complexity, convolution decomposition is used to decompose a large kernel convolution into a dual-branch structure, where one branch is composed of x×1 and 1×y in series, and the convolution order of the other branch is 1×y and x×1. The outputs of the two branches are added element by element to get the final output.
[0081] The input of the dual-input channel attention module consists of two parts. One part is the output feature map of the cascaded pyramid convolution module in the spanning connection path. The other part is the feature map after upsampling of the corresponding decoding layer Where c, h, w and d are the number of channels, height, width and depth of the feature map respectively. The two input feature maps are concatenated in the channel dimension to obtain
[0082] See also Figure 5 As shown, as a preferred embodiment of the present invention, the step (4) specifically includes the following steps:
[0083] The first channel of the dual-input channel attention module described in (4.1) inputs the first feature map through the cascaded pyramid convolution module in the cross-connection path The second channel is upsampled through the decoding layer and input into the second feature map
[0084] (4.2) Concatenate the first feature map and the second feature map in the channel dimension to obtain Among them, c, h, w and d are the number of channels, height, width and depth of the feature map respectively, C(X 1 ,X 2 ) is the concatenated feature map;
[0085] The dual-input channel attention module described in (4.3) performs a global average pooling operation after fusing the first feature map and the second feature map of the input in the channel dimension, and obtains the global information feature vector according to the following formula:
[0086]
[0087] Among them, i, j, k, and f represent height, width, depth, and channel respectively.
[0088] (4.4) The feature vector dimension is reduced to the number of channels c through 1×1 convolution, and the feature vector is normalized using the Sigmoid activation function. The channel attention vector CA is obtained according to the following formula:
[0089] CA=σ(W×v f + b);
[0090] Among them, W and b are the convolution kernel parameters, v f is the feature graph.
[0091] The value of each element of the obtained attention vector is between 0 and 1: CA∈[0,1], and the sum is 1, that is, |CA|=1.
[0092] The attention vector described in (4.5) is multiplied with the output feature map of the cascade pyramid convolution module in the channel dimension to enhance the discriminability of the shallow features of the network, and then the deep features are connected to the output end through residual connection to become the output of the dual-input channel attention module, which is specifically implemented by the following formula:
[0093]
[0094] in, represents the multiplication of channel dimensions, and O is the output feature map of the attention module.
[0095] The dual-input channel attention module uses the feature maps of the shallow and deep parts of the network to further extract important content from the feature maps generated by the encoder, so that the model pays more attention to the target area and has a higher ability to distinguish different categories of areas.
[0096] As a preferred embodiment of the present invention, the decoder is specifically composed of four dual-input channel attention modules connected in series. The decoder upsamples the feature vector output by each layer through a bilinear interpolation operation, and fuses it with the feature vector output by the cascaded pyramid convolution module of the previous layer, thereby gradually restoring the feature map to the original input size, and the output layer uses a Softmax function to output the probability of the category to which the pixels of each feature map belong.
[0097] As a preferred embodiment of the present invention, the step (5) is to obtain the lesion segmentation result through a loss function, specifically: wherein the total loss function is:
[0098] L′ total =L′ bce +L′ dice ;
[0099] Among them, L′ bce is the pixel-level binary cross entropy loss, L′ dice is the Dice loss, and the two losses are:
[0100]
[0101] Among them, x i,j is the probability of the predicted category, y i,j For real marking.
[0102] In a preferred embodiment, the loss function in the prostate cancer segmentation network training process described in step (5) specifically includes the following steps:
[0103] Prediction probability plot of the model output True segmentation label map Since the target area of the PCa segmentation task has only one category, and the model input is a two-dimensional image, the PCa lesion segmentation training data is an image cropped according to the prostate ROI, and the problem of category imbalance is improved. The total loss function is:
[0104] L′ total =L′ bce +L′ dice
[0105] Among them, L′ bce is the pixel-level binary cross entropy loss, L′ dice is the Dice loss, and the two losses are:
[0106]
[0107] After obtaining the loss function of supervised information, the back propagation algorithm and the ADAM optimization algorithm with parameters β1=0.9, β2=0.999 are used to minimize the loss function of supervised information to train the target segmentation model in step (5).
[0108] The model obtained in the above manner is a multi-parameter MRI image lesion segmentation model based on a deep neural network. When using the trained model, the prostate MR image to be segmented is input into the deep neural network to obtain a segmentation result map. The segmentation results of this method in our data set are shown in the following table:
[0109] DSC (%) ABD(mm) RVD (%) Method of the present invention 82.11±0.95 3.64±0.91 -8.66±3.77
[0110] The segmentation diagram is as follows Figure 6 As shown in the figure, the experiment on the dataset used 5-fold cross validation to calculate the mean and standard deviation of each evaluation indicator. The evaluation indicators used included: Dice Similarity Coefficient (DSC), Average Boundary Distance (ABD) and Relative Volume Difference (RVD).
[0111] Dice similarity coefficient (DSC) is the most important indicator for evaluating segmentation results in the field of medical image segmentation. It is used to evaluate the similarity between the segmentation result and the true segmentation label. The calculation formula is:
[0112]
[0113] Among them, X and Y represent the model output segmentation map and the true segmentation label respectively. The value range of DSC is between 0 and 1. The larger the DSC value, the closer the prediction result is to the true label.
[0114] The average boundary distance (ABD) is used to calculate the average distance between the predicted segmentation result boundary and the true segmentation label boundary, which can reflect the accuracy of the segmentation result edge. The calculation formula is:
[0115]
[0116] Among them, X s and Y s They represent the set of edge points of the predicted result and the true segmentation label map, respectively. d(·,·) is the Euclidean distance between the two points. The Euclidean distance between two points in n-dimensional space can be expressed as:
[0117]
[0118] The calculation method of ABD can be summarized as follows: for each point in a given edge point set, calculate the minimum Euclidean distance with another edge point set and average all the results.
[0119] The relative volume difference can reflect the under-segmentation or over-segmentation state of the model, which is determined by the ratio of the voxel volume of the predicted segmentation result to the voxel volume of the true segmentation label. The calculation formula is:
[0120]
[0121] It can be seen that a negative RVD value indicates that the model prediction result is under-segmentation, and a positive RVD value indicates that the prediction result is over-segmentation.
[0122] The system for implementing multi-parameter MRI image lesion segmentation based on a deep neural network model using the above method comprises:
[0123] The target extraction processing module is used to extract the target of the region of interest from the input imaging sequences ADC, T2W, and DWI of the prostate multi-parameter nuclear magnetic resonance image;
[0124] A size unification processing module, connected to the target extraction processing module, is used to perform a rigid registration operation on the extracted target images and unify the sizes of the target images to the same size;
[0125] A lesion segmentation neural network processing module is connected to the size unification processing module and is used to input the target image after size unification processing into the prostate cancer lesion segmentation network of the prostate multi-parameter nuclear magnetic resonance image to perform convolution and feature map sampling processing through an encoder;
[0126] A cascade pyramid convolution module, connected to the lesion segmentation neural network processing module, is used to receive the feature map output by the encoder and input it to the cross-connection layer of the cascade pyramid convolution module for group convolution and residual splicing to retain the feature information of the corresponding target object; and
[0127] The dual-input channel attention module is connected to the cascade pyramid convolution module and is used to fuse the output feature map of the cascade pyramid convolution module in the cross-connection path and the output feature map after feature upsampling of the decoding layer in the channel dimension and perform a global average pooling operation to obtain a feature vector of the corresponding channel.
[0128] The device for implementing multi-parameter MRI image lesion segmentation based on a deep neural network model comprises:
[0129] a processor configured to execute computer-executable instructions;
[0130] A memory stores one or more computer executable instructions, which, when executed by the processor, implement the various steps of the method for implementing multi-parameter magnetic resonance image lesion segmentation based on a deep neural network model as described above.
[0131] The processor implements multi-parameter MRI image lesion segmentation based on a deep neural network model, wherein the processor is configured to execute computer executable instructions. When the computer executable instructions are executed by the processor, the various steps of the method for implementing multi-parameter MRI image lesion segmentation based on a deep neural network model are implemented.
[0132] The computer-readable storage medium stores a computer program thereon, and the computer program can be executed by a processor to implement the various steps of the method for implementing multi-parameter magnetic resonance image lesion segmentation based on a deep neural network model as described above.
[0133] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code that includes one or more executable instructions for implementing the steps of a specific logical function or process, and the scope of the preferred embodiments of the present invention includes alternative implementations in which functions may not be performed in the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present invention belong.
[0134] It should be understood that each part of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution device.
[0135] A person of ordinary skill in the art may understand that all or part of the steps of the method for implementing the above-mentioned embodiment may be completed by instructing the relevant hardware through a program, and the program may be stored in a computer-readable storage medium, which, when executed, includes one of the steps of the method embodiment or a combination thereof.
[0136] The above integrated modules can be implemented in the form of hardware or software functional modules. If the integrated modules are implemented in the form of software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium.
[0137] The storage medium mentioned above can be a read-only memory, a magnetic disk or an optical disk, etc.
[0138] In the description of this specification, the description with reference to the terms "an embodiment", "some embodiments", "example", "specific example", or "embodiment" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.
[0139] Although the embodiments of the present invention have been shown and described above, it is to be understood that the above embodiments are exemplary and are not to be construed as limitations of the present invention. A person skilled in the art may change, modify, replace and vary the above embodiments within the scope of the present invention.
[0140] The method, system, device, processor and computer-readable storage medium for implementing multi-parameter nuclear magnetic resonance image lesion segmentation based on the deep neural network model of the present invention are adopted, and a multi-parameter sequence prostate cancer lesion segmentation method and framework based on a deep neural network with a codec structure is provided. The segmentation method integrates multiple MRI modality data, adopts modules such as cascaded pyramid convolution and channel attention, fully integrates deep and shallow layer feature information, reduces noise interference, and segments multi-scale prostate cancer lesion targets, providing strong support for the clinical diagnosis and treatment of prostate diseases, reducing the time spent on diagnosis, and providing more effective information processing methods and means for the screening, detection and diagnosis of prostate cancer. Automatic segmentation of MRI prostate cancer lesions can be achieved.
[0141] At the same time, considering that a single MRI sequence may ignore different forms of mutual information, thus hindering the model from achieving better segmentation performance, we use three sequence data for channel merging: T2W, ADC and DWI. The use of ADC and DWI sequences can supplement the lesion feature information and greatly improve the segmentation performance. In view of the large differences in the shapes and sizes of PCa lesions in different cases and the large number of small target areas, a cascaded pyramid convolution module is designed. The cascaded pyramid convolution module can capture the detailed spatial positioning information carried in the feature map generated by the encoder at multiple scales, and fuse local and global information at multiple scales to reduce the loss of spatial positioning information and improve the model's ability to classify pixels. At the same time, the use of the cascaded pyramid convolution module can also improve the under-segmentation of the model. In order to enhance the model's attention to the target area, a dual-input channel attention module is designed, which uses the semantic information of the deep features of the network to guide the shallow output to obtain features with higher discrimination ability and strengthen the feature extraction capabilities of each stage of the network.
[0142] In this specification, the present invention has been described with reference to specific embodiments thereof. However, it is apparent that various modifications and variations may be made without departing from the spirit and scope of the present invention. Therefore, the specification and drawings should be regarded as illustrative rather than restrictive.
Claims
1. A method for implementing multi-parameter MRI image lesion segmentation based on a deep neural network model, characterized in that: The method comprises the following steps: (1) Input any combination of samples of imaging sequences including prostate multi-parameter MRI images, such as ADC, T2W, and DWI, and perform rigid matching operation; (2) extracting the region of interest from the processed image and sending it to the prostate cancer lesion segmentation network for feature processing through the encoder; (3) The encoder outputs the feature map and inputs it into the cross-connection layer of the cascaded pyramid convolution module for convolution and feature map sampling processing; (4) After the decoder performs feature map upsampling, the features output by the cross-connection layer are transmitted to the dual-input channel attention module for feature fusion processing; (5) training the prostate cancer lesion segmentation network of the prostate multi-parameter magnetic resonance imaging to obtain lesion segmentation results; The step (2) is specifically as follows: A preset number of convolution modules in a pre-trained ResNeXt network is used to retain the feature maps of each downsampling layer in each convolution module through an encoder to obtain the number of channels of the corresponding feature map; The preset number of convolution modules is set as the first five convolution modules in the ResNeXt network, wherein the convolution kernel size used by the first convolution module is 7×7, and the remaining four convolution modules use convolution kernel sizes of 3×3 and 1×1 respectively; The number of channels of the feature maps obtained by five downsamplings increases successively, and the size of each feature map decreases successively, which are 1 / 2, 1 / 4, 1 / 8, 1 / 16 and 1 / 32 of the original image respectively; The step (3) specifically comprises the following steps: The encoder described in (3.1) outputs feature maps corresponding to four cascaded pyramid convolution modules from the first layer to the fourth layer, wherein the output of each large kernel convolution is fused pixel by pixel with the feature map after 1×1 convolution of the original input feature map as the input of the next convolution; The cascaded pyramid convolution module described in (3.2) uses convolution decomposition to decompose a large kernel convolution into a dual-branch structure, where one branch is composed of x×1 and 1×y in series, and the convolution order of the other branch is 1×y and x×1. The outputs of the two branches are added element by element to obtain the final output; (3.3) The results of multiple large kernel convolutions are concatenated on the channel to retain the feature information of small target objects to the greatest extent; (3.4) According to the difference in the number and size of convolution kernels used by the four groups of cascaded pyramids corresponding to the output of the first four layers of the encoder, the sizes of feature maps of different sizes of the encoder are adapted.
2. The method for implementing multi-parameter MRI lesion segmentation based on a deep neural network model according to claim 1, characterized in that: The step (4) specifically comprises the following steps: The first channel of the dual-input channel attention module described in (4.1) inputs the first feature map through the cascaded pyramid convolution module in the cross-connection path The second channel is upsampled through the decoding layer and input into the second feature map (4.2) Concatenate the first feature map and the second feature map in the channel dimension to obtain Among them, c, h, w and d are the number of channels, height, width and depth of the feature map respectively, C(X 1 ,X 2 ) is the concatenated feature map; The dual-input channel attention module described in (4.3) performs a global average pooling operation after fusing the first feature map and the second feature map of the input in the channel dimension, and obtains the global information feature vector according to the following formula: Among them, i, j, k, and f represent the height, width, depth, and number of channels of the feature map respectively; (4.4) The feature vector dimension is reduced to the number of channels c through 1×1 convolution, and the feature vector is normalized using the Sigmoid activation function. The channel attention vector CA is obtained according to the following formula: CA=σ(W×v f +b) Among them, W and b are the convolution kernel parameters, v f is the feature map; The attention vector described in (4.5) is multiplied with the output feature map of the cascade pyramid convolution module in the channel dimension to enhance the discriminability of the shallow features of the network, and then the deep features are connected to the output end through residual connection to become the output of the dual-input channel attention module, which is specifically implemented by the following formula: in, represents the multiplication of channel dimensions, and O is the output feature map of the attention module.
3. The method for implementing multi-parameter MRI lesion segmentation based on a deep neural network model according to claim 2, characterized in that: The decoder is specifically composed of four dual-input channel attention modules connected in series. The decoder upsamples the feature vector output by each layer through a bilinear interpolation operation, and fuses it with the feature vector output by the cascaded pyramid convolution module of the previous layer, thereby gradually restoring the feature map to the original input size, and the output layer uses the Softmax function to output the probability of the category to which the pixels of each feature map belong.
4. The method for implementing multi-parameter MRI lesion segmentation based on a deep neural network model according to claim 3, characterized in that: The step (5) is to obtain the lesion segmentation result through the loss function, specifically: Among them, the total loss function is: L ′ total =L ′ bce +L ′ dice ; Among them, L′ bce is the pixel-level binary cross entropy loss, L′ dice is the Dice loss, and the two losses are: Among them, x i,j is the probability of the predicted category, y i,j For real marking.
5. A system for implementing multi-parameter MRI image lesion segmentation based on a deep neural network model using the method of claim 4, characterized in that: The system comprises: The target extraction processing module is used to extract the target of the region of interest from the input imaging sequences ADC, T2W, and DWI of the prostate multi-parameter nuclear magnetic resonance image; A size unification processing module, connected to the target extraction processing module, is used to perform a rigid registration operation on the extracted target images and unify the sizes of the target images to the same size; A lesion segmentation neural network processing module is connected to the size unification processing module and is used to input the target image after size unification processing into the prostate cancer lesion segmentation network of the prostate multi-parameter nuclear magnetic resonance image to perform convolution and feature map sampling processing through an encoder; A cascade pyramid convolution module, connected to the lesion segmentation neural network processing module, is used to receive the feature map output by the encoder and input it to the cross-connection layer of the cascade pyramid convolution module for group convolution and residual splicing to retain the feature information of the corresponding target object; and The dual-input channel attention module is connected to the cascade pyramid convolution module and is used to fuse the output feature map of the cascade pyramid convolution module in the cross-connection path and the output feature map after feature upsampling of the decoding layer in the channel dimension and perform a global average pooling operation to obtain a feature vector of the corresponding channel.
6. A device for implementing multi-parameter MRI image lesion segmentation based on a deep neural network model, characterized in that: The device comprises: a processor configured to execute computer executable instructions; A memory storing one or more computer executable instructions, wherein when the computer executable instructions are executed by the processor, the steps of the method for implementing multi-parameter magnetic resonance image lesion segmentation based on a deep neural network model as described in any one of claims 1 to 4 are implemented.
7. A processor for implementing multi-parameter MRI image lesion segmentation based on a deep neural network model, characterized in that: The processor is configured to execute computer executable instructions. When the computer executable instructions are executed by the processor, the various steps of the method for implementing multi-parameter magnetic resonance image lesion segmentation based on a deep neural network model as described in any one of claims 1 to 4 are implemented.
8. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and the computer program can be executed by a processor to implement the various steps of the method for implementing multi-parameter magnetic resonance image lesion segmentation based on a deep neural network model as described in any one of claims 1 to 4.
Citation Information
Patent Citations
Brain tumor segmentation network and segmentation method based on U-Net network
CN111192245A
Gland cell image segmentation method and device based on edge sensing network
CN113034505A