A colorectal cancer h&e staining pathological image semantic segmentation method and system
By constructing a multi-scale, multi-level semantic feature extraction and symmetric spiral backbone feature fusion network, and combining it with weighted model training, accurate segmentation of colorectal cancer pathological images was achieved, solving the problems of high segmentation difficulty and low accuracy in existing technologies, and improving the segmentation accuracy of colorectal cancer pathological images.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- KUNMING UNIV OF SCI & TECH
- Filing Date
- 2023-09-14
- Publication Date
- 2026-06-12
AI Technical Summary
Existing technologies face challenges in segmenting colorectal cancer pathological images due to their high segmentation difficulty and low segmentation accuracy, making it impossible to accurately identify multiple categories of colorectal cancer pathological images.
A semantic segmentation method for H&E-stained pathological images of colorectal cancer is adopted. Through image acquisition, dataset segmentation, construction of basic semantic segmentation network, multi-scale and multi-level semantic feature extraction, training of symmetric spiral backbone feature fusion network and weight model, the method can accurately segment the colorectal cancer region, normal mucosa region and stromal region.
This technology achieves precise masked segmentation of colorectal cancer pathological images, improving segmentation accuracy and robustness, and solving the problems of high segmentation difficulty and low accuracy in existing technologies.
Smart Images

Figure CN117095173B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer vision technology for H&E staining medical pathological images of rectal cancer, specifically to a semantic segmentation method and system for H&E staining pathological images of colorectal cancer. Background Technology
[0002] Segmentation of histopathological images of colorectal cancer is a crucial step in computer-aided diagnostic systems. According to relevant studies, the 5-year survival rate for colorectal cancer patients who are diagnosed and treated promptly can reach as high as 90%. In contrast, if diagnosis is delayed and cancer cells spread, the 5-year survival rate drops to as low as 14%. Therefore, timely, objective, and accurate detection of colorectal cancer is one of the important means to improve patient survival rates. Traditional medical image segmentation methods include thresholding, edge detection, region growing, and feature-based segmentation methods. While these methods have some applications in medical image segmentation, they typically require manual selection of thresholds or parameters, leading to subjective and dependent results that are easily influenced by the operator's subjective judgment. Deep learning-based medical segmentation methods automatically learn feature representations of images through multi-layer neural networks, avoiding the tedious process of manually designing features. The network can learn more discriminative features from large amounts of data, improving the accuracy and robustness of segmentation.
[0003] Due to the complex tissue structure and diverse pathological changes in colorectal cancer pathological images, current visual segmentation algorithms face technical challenges such as high segmentation difficulty and low segmentation accuracy when processing colorectal cancer pathological images, resulting in the inability to accurately identify multiple categories of colorectal cancer pathological images. Summary of the Invention
[0004] This invention provides a semantic segmentation method and system for H&E-stained pathological images of colorectal cancer, which solves the technical problems of high segmentation difficulty and low segmentation accuracy in the prior art due to the complex tissue structure and diverse pathological changes of colorectal cancer pathological images.
[0005] According to a first aspect of the present invention, a semantic segmentation method for H&E-stained pathological images of colorectal cancer is provided, comprising: acquiring images of a target pathological region using an image acquisition device to obtain a pathological image dataset of the target pathological region; dividing the pathological image dataset into segments according to the size of the pathological images to determine a training dataset, a validation dataset, and a test dataset; configuring experimental parameters and constructing a basic semantic segmentation network based on the training dataset, the validation dataset, and the test dataset; extracting features from the training dataset using a grouped feature sequence extraction module based on convolutional kernels to obtain multi-scale, multi-level semantic features; and constructing a symmetrical feature segmentation network based on the grouped feature sequence extraction module using multi-branch and spiral arrangements. A spiral backbone feature fusion network parses the depth and width information of the training dataset, validation dataset, and test dataset based on the multi-scale, multi-level semantic features to obtain a corrected image data test set. Based on the corrected image data test set, the class probabilities are obtained through the symmetric spiral feature fusion network for fine segmentation. The symmetric spiral feature fusion network is then embedded into the basic semantic segmentation network to obtain an updated image semantic segmentation network. Based on the updated image semantic segmentation network, a weighted model for segmenting colorectal cancer regions, normal mucosal regions, and stromal regions is trained using the image data training set. This weighted model is then used to segment images within the target pathological regions in the corrected image data test set.
[0006] According to a second aspect of the present invention, a semantic segmentation system for H&E-stained pathological images of colorectal cancer is provided, comprising: a pathological image acquisition module, wherein the pathological image acquisition module is used to acquire images of a target pathological region through an image acquisition device to acquire a pathological image dataset of the target pathological region; a segmentation module, wherein the segmentation module is used to segment the pathological image dataset according to the pathological image size, determine a training dataset, a validation dataset, and a test dataset, and configure experimental parameters and build a basic semantic segmentation network according to the training dataset, the validation dataset, and the test dataset; a multi-scale feature extraction module, wherein the multi-scale feature extraction module is used to extract features from the training dataset based on a grouped feature sequence extraction module constructed based on convolutional kernels to obtain multi-scale, multi-level semantic features; and a feature fusion network construction module, wherein the feature fusion network construction module is used to construct a feature fusion network based on the grouped feature sequence extraction module. A symmetrical spiral backbone feature fusion network is constructed using multiple branching and spiral arrangements. Based on the multi-scale, multi-level semantic features, the training, validation, and test datasets are analyzed for depth and width information to obtain a corrected image data test set. An updated image semantic segmentation network acquisition module is then used to obtain class probabilities for fine segmentation based on the corrected image data test set through the symmetrical spiral feature fusion network. This symmetrical spiral feature fusion network is then embedded into the basic semantic segmentation network to obtain an updated image semantic segmentation network. Finally, an image segmentation module is used to train a weighted model for segmenting colorectal cancer regions, normal mucosal regions, and stromal regions based on the updated image semantic segmentation network and the image data training set. This weighted model is then used to segment images within the target pathological regions in the corrected image data test set.
[0007] The beneficial effects that can be achieved by adopting one or more technical solutions according to the present invention are as follows:
[0008] Images of the target pathological region are acquired using image acquisition equipment to obtain a pathological image dataset. The dataset is then segmented according to the image size to determine training, validation, and test datasets. Based on these datasets, experimental parameters are configured, and a basic semantic segmentation network is built. A grouped feature sequence extraction module, constructed using convolutional kernels, extracts features from the training dataset, acquiring multi-scale, multi-level semantic features. A symmetrical spiral backbone feature fusion network is then constructed using multi-branch and spiral arrangements based on these grouped feature sequence extraction modules. Finally, depth and complexity analysis is performed on the training, validation, and test datasets based on the multi-scale, multi-level semantic features. Width information is analyzed to obtain a corrected image data test set. Based on the corrected image data test set, a symmetric spiral feature fusion network is used to obtain class probabilities for fine segmentation. The symmetric spiral feature fusion network is embedded into a basic semantic segmentation network to obtain an updated image semantic segmentation network. Based on the updated image semantic segmentation network, a weight model for segmenting colorectal cancer regions, normal mucosal regions, and stroma regions is trained according to the image data training set. The weight model is then used to segment images within the target pathological region in the corrected image data test set. This solves the technical problem of high difficulty and low accuracy in class segmentation of colorectal cancer H&E stained pathological images in existing visual segmentation algorithms, and achieves accurate mask segmentation of colorectal cancer. Attached Figure Description
[0009] To more clearly illustrate the technical solutions in this invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. The accompanying drawings, which constitute a part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0010] Figure 1 This is a flowchart illustrating a semantic segmentation method for H&E-stained pathological images of colorectal cancer provided in an embodiment of the present invention.
[0011] Figure 2 This is a schematic diagram of a grouping feature sequence extraction module in semantic segmentation of H&E-stained pathological images of colorectal cancer according to an embodiment of the present invention;
[0012] Figure 3 This is a schematic diagram of the deep feature extraction layer in a semantic segmentation method for H&E-stained pathological images of colorectal cancer according to an embodiment of the present invention;
[0013] Figure 4This is a schematic diagram of an encoder and decoder in semantic segmentation of H&E-stained pathological images of colorectal cancer according to an embodiment of the present invention;
[0014] Figure 5 This is a schematic diagram of the structure of a semantic segmentation system for H&E-stained pathological images of colorectal cancer provided in an embodiment of the present invention.
[0015] Figure labeling: Pathological image acquisition module 11, segmentation module 12, multi-scale feature extraction module 13, feature fusion network construction module 14, updated image semantic segmentation network acquisition module 15, image segmentation module 16. Detailed Implementation
[0016] To make the objectives, technical solutions, and advantages of the present invention more apparent, exemplary embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of the present invention, and not all embodiments of the present invention. It should be understood that the present invention is not limited to the exemplary embodiments described herein.
[0017] The terminology used in this specification is for describing embodiments and not for limiting the invention. As used in this specification, the singular terms “a,” “an,” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. When used in this specification, the terms “comprising” and / or “including” specify the presence of a step, operation, element, and / or component, but do not preclude the presence or addition of one or more other steps, operations, elements, components, and / or groups thereof.
[0018] Unless otherwise defined, all terms (including technical and scientific terms) used in this specification shall have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. Terms, as defined in common dictionaries, should not be interpreted in an idealized or overly formal sense unless expressly defined herein. Throughout this specification, the same reference numerals denote the same elements.
[0019] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for display, data used for analysis, etc.) involved in this invention are all information and data authorized by the user or fully authorized by all parties.
[0020] Example 1
[0021] This application provides a semantic segmentation method for H&E-stained pathological images of colorectal cancer, which is described below. Figure 1 , Figure 2 , Figure 3 and Figure 4 The method, as described above, includes:
[0022] The target pathological region is captured by an image acquisition device to obtain a pathological image dataset of the target pathological region;
[0023] After a possible diagnosis of colorectal cancer, doctors will perform a biopsy or surgical resection to obtain tissue samples. These samples will be submitted to pathologists for examination. Pathologists will perform fixation, sectioning, and H&E staining to observe the morphology and structure of the tissue cells under a microscope. H&E staining is a commonly used pathological staining technique that uses heme and eosin staining agents, resulting in purple cell nuclei and pink cytoplasm. In modern medical practice, H&E-stained colorectal cancer pathology samples are usually digitally scanned. The image acquisition equipment is a digital scanning device that converts the tissue sections into high-resolution digital images. These digital images can be collected and processed to form an image dataset of the target pathological area, providing a data basis for subsequent cutting.
[0024] In a preferred embodiment, it further includes:
[0025] Based on the image acquisition boundary of the image acquisition device, a pixel range value is determined; within the pixel range value, the lesion area, normal mucosal area, and matrix target area are labeled at the pixel level to obtain multiple labeled areas; the multiple labeled areas are marked with different pixels to obtain multiple labeled regions; the multiple labeled regions are integrated to obtain the pathological image dataset.
[0026] Based on the image acquisition from the image acquisition device, the pixel range values of the target regions of colorectal cancer lesions, normal mucosa, and stroma are determined. Within the pixel range values, a professional annotation tool is used to perform pixel-level annotation on the lesion region, normal mucosa region, and stroma target region to obtain multiple annotation regions. These multiple annotation regions are then labeled with different pixels to obtain multiple labeled regions. All the labels corresponding to these multiple labeled regions are integrated on a mask to obtain the pathological image dataset within the target pathological region, providing a data foundation for subsequent image segmentation.
[0027] The pathological image dataset is segmented according to the size of the pathological images to determine the training dataset, validation dataset, and test dataset. Experimental parameters are configured and a basic semantic segmentation network is built based on the training dataset, the validation dataset, and the test dataset.
[0028] In a preferred embodiment, it further includes:
[0029] The pathological image dataset is segmented based on the size of the pathological images to obtain segmented digital pathological image information. The segmented digital pathological image information is then sequentially traversed and identified to obtain segmented pathological image recognition information. Based on the segmented pathological image recognition information, lesion regions, normal mucosal regions, and matrix target regions are determined. Based on the lesion regions, normal mucosal regions, and matrix target regions, the pathological image dataset is divided into an image data training set, an image data validation set, and an image data test set according to a predetermined ratio. Foreground target instances are obtained based on the target pathological regions and analyzed to obtain the structural and color features of the foreground target instances. Based on the structural and color features, experimental parameters are configured and a basic semantic segmentation network is built.
[0030] The pathological image dataset is segmented based on the size of the pathological images to obtain segmented digital pathological image information. This segmented digital pathological image information is then sequentially traversed and identified to obtain segmented pathological image recognition information. Based on this information, lesion regions, normal mucosa regions, and stromal target regions containing colorectal cancer lesions, normal mucosa regions, and stromal target regions are determined. Based on these lesion regions, normal mucosa regions, and stromal target regions, the pathological image dataset is divided into a training set, a validation set, and a test set according to a predetermined ratio (e.g., 8:8:2).
[0031] Based on the image data training set, image data validation set, and image data test set, a network training strategy suitable for the H&E staining pathology dataset of colorectal cancer was determined. Foreground target instances in the target pathological region were analyzed. H&E staining pathology images are high-resolution, complex background data. Aggregated colorectal cancer lesions and normal mucosal edges are extremely similar, while dispersed colorectal cancer cells are contained within the stroma and have similar colors. Therefore, it is necessary to obtain the structural and color features of the foreground target instances. Based on these structural and color features, experimental parameters were configured, and a basic semantic segmentation network was built. Targeting the computer configuration and the structural and color features of foreground target instances in the H&E staining pathology images of colorectal cancer, an experimental platform was built and experimental parameters were configured for the basic semantic segmentation network, determining a network training strategy suitable for the colorectal cancer semantic segmentation dataset.
[0032] For example, the input pathological images were uniformly 448×448 pixels. During training, the deep learning framework was PyTorch 1.6.0, Python version 3.6, CUDA 11.2, and cuDNN 8.0.5 were used to accelerate model training. TensorRT 7.1.3.4 was used to accelerate model testing. LovaszSoftmax, an IOU-based loss function, was used as the training loss function. The values of all hyperparameters are shown in Table 1.
[0033] Table 1
[0034]
[0035] During training, model weights are saved every 4500 steps, for a total of 20 models. The last five trained models are then compared to examine their accuracy and inference speed. This comparison selects the best-performing base semantic segmentation network. The selected best model is then validated on a validation set to ultimately evaluate the performance of the base semantic segmentation network.
[0036] A grouped feature sequence extraction module based on convolutional kernels is used to extract features from the training dataset and obtain multi-scale and multi-level semantic features.
[0037] In a preferred embodiment, it further includes:
[0038] The grouped feature sequence extraction module extracts pathological image contour edge detail information from the training dataset to obtain a feature extraction map. Based on the feature extraction map, upsampling and downsampling are performed to obtain the multi-scale multi-level semantic features. The upsampling result is used as the input feature of the next information node. The input feature is downsampled. The multi-scale multi-level semantic features are obtained by concatenating the grouped feature sequence extraction module. The upsampled multi-scale semantic information is used as the prior feature information of the downsampled multi-scale semantic information.
[0039] The grouped feature sequence extraction module extracts the contour edge details of the pathological image to obtain a feature extraction map. Based on this feature extraction map, upsampling and downsampling are performed to obtain the multi-scale, multi-level semantic features. Specifically, the feature map obtained by the grouped feature sequence extraction module is upsampled and used as the input feature for the next information node. The received information is concatenated by the grouped feature sequence extraction module. The upsampled multi-scale semantic information is used as the prior feature information for the downsampled multi-scale semantic information. A multi-scale generation control module is determined based on the optimal segmentation effect to obtain the multi-scale, multi-level semantic features. Multi-scale, multi-level semantic features are obtained by replacing the traditional two convolutional layers with convolutional kernels of different sizes for upsampling and downsampling. Different dilation rates are used for these convolutional kernels to expand the receptive field. A 1*1 convolutional layer is introduced as a residual structure to supplement the large receptive field feature map information through sparse sampling. A max-pooling layer is used to eliminate non-maximum values in the feature information, thereby reducing the computational complexity of the upper layers and accelerating the convergence speed of the basic semantic segmentation network by removing useless information from the semantic features.
[0040] Furthermore, the calculation formula output by the calculation group feature sequence extraction module is as follows:
[0041]
[0042] Where X i_j-1 X i-1_j Input features, X, to the grouped feature sequence extraction module i_j The output features of the grouped feature sequence extraction module are s and d, which are the stride and dilation rate of the dilated convolution, n*n represents the size of the convolution kernel, MaxPool(·) represents the maximum pooling with a kernel size of 2, Conv(·) represents the convolution operation, and Cat(·) represents the concatenation operation.
[0043] Based on the grouped feature sequence extraction module, a symmetrical spiral backbone feature fusion network is constructed through multi-branch and spiral arrangements. The training dataset, validation dataset, and test dataset are analyzed for depth and width information according to the multi-scale and multi-level semantic features to obtain the corrected image data test set.
[0044] In a preferred embodiment, it further includes:
[0045] The grouped feature sequence extraction module is serially deepened in the same dimension to obtain a multi-dimensional network. A pooling layer is used to perform downsampling operations to expand the network depth of each dimension and obtain network depth expansion features. Width downsampling is performed on the network of each dimension to obtain multiple branch structures, and features are fused with the network depth expansion features to obtain the symmetric spiral backbone feature fusion network.
[0046] In a preferred embodiment, it further includes:
[0047] The encoder backbone network is constructed based on the multi-scale, multi-level semantic features, and the construction formula is as follows:
[0048]
[0049] Where: function θ[·] represents the feature extraction block implemented by convolution, batch normalization and ReLU activation function, C[·] represents concatenation, D[·] represents downsampling operation, n represents the number of layers after downsampling operation, m represents the number of feature extraction blocks in the feature extraction layer, n=5, m=5.
[0050] In a preferred embodiment, it further includes:
[0051] Based on the multi-scale, multi-level semantic feature extraction, feature layer information of the encoder and decoder is constructed, and feature extraction blocks are fused to obtain feature extraction block aggregation information. The feature extraction block aggregation information is used to generate parameter weights to regulate information features at each scale. The feature extraction blocks are connected in a predetermined serial-parallel manner using a symmetrical structure in the depth and width directions to form an information decoder. The information decoder is used to parse the depth and width information of the training dataset, validation dataset, and test dataset to obtain the corrected image data test set.
[0052] The grouped feature sequence extraction module is concatenated in the same dimension to deepen the network. Downsampling is performed using a 2×2 max-pooling layer to expand the network depth by 5 times in each dimension, obtaining network depth-expanded features. Width downsampling divides the overall network into five branches, expanding the network width. Feature fusion is then performed using the network depth-expanded features, maintaining the same scale for feature extraction blocks within the same branch and different scales between different branches. The pathological images are then convolved at the same scale as each branch, and the resulting feature layers are added to the corresponding branch structures to enhance the extraction capability of each feature branch for different features. The training, validation, and test datasets, derived from multi-scale, multi-level semantic features, are used to construct the encoder and decoder. Feature extraction blocks are then fused. The feature extraction module aggregates information to generate parameter weights to regulate information features at each scale. A symmetrical structure is used to assemble the feature extraction blocks into an information decoder through a specific serial-parallel connection in both depth and width directions, resulting in the symmetrical spiral backbone feature fusion network.
[0053] exist Figure 4 In the encoder planar structure diagram, from the x-axis of the network... 0 Feature extraction block to x 4-0 The direction of the curved arrow indicates the main forward direction of the network. The network depth is increased along the deep feature extraction layer. The structure diagram of the deep feature extraction layer is shown below. Figure 3 As shown, depth downsampling connects the depth feature extraction layers of different dimensions. Width downsampling divides the overall network into five branches, expanding the network width and performing depth-wise feature fusion. The feature extraction blocks of each branch have the same scale, while the scales of different branches differ. This perfectly matches the five convolutional operations of different scales in the depth feature extraction layers, thus obtaining feature information at different scales. kThis involves performing convolution operations on the original image at the same scale as each branch, and then supplementing the resulting feature layers into the corresponding branch structures to enhance the ability of each feature branch to extract different features. Simultaneously, this low-level semantic information can be directly fused with high-level semantic information through wide downsampling, constructing short-range dependencies in the model and obtaining better global context feature extraction capabilities. After final width downsampling, five feature maps of the same resolution are obtained. These five feature maps at different scales are fused using two 3×3 convolutional kernels to obtain the symmetric spiral backbone feature fusion network. Through the combination of feature extraction blocks, depth and width are perfectly combined, achieving this dual expansion of depth and width. Furthermore, the depth and width of the network can be simultaneously changed by altering the number of feature extraction blocks in the feature extraction layer.
[0054] The encoder backbone network is constructed based on the multi-scale, multi-level semantic features, and the construction formula is as follows:
[0055]
[0056] Where: the function θ[·] represents the feature extraction block implemented by convolution, batch normalization, and ReLU activation function, C[·] represents concatenation, and D[ · ] represents the downsampling operation, n represents the number of layers after the downsampling operation, and m represents the number of feature extraction blocks in the feature extraction layer, n=5, m=5.
[0057] Based on the multi-scale, multi-level semantic feature extraction, feature layer information of the encoder and decoder is constructed, and feature extraction blocks are fused to obtain feature extraction block aggregation information. Parameter weights are generated using the feature extraction block aggregation information to regulate information features at each scale. Using a symmetrical structure, the feature extraction blocks are connected in a predetermined serial-parallel manner along the depth and width directions to form an information decoder. The information decoder is used to parse the depth and width information of the training dataset, validation dataset, and test dataset to obtain the corrected image data test set.
[0058] Based on the corrected image data test set, the category probabilities are obtained through the symmetric spiral feature fusion network for fine segmentation. The symmetric spiral feature fusion network is then embedded into the basic semantic segmentation network to obtain the updated image semantic segmentation network.
[0059] In a preferred embodiment, it further includes:
[0060] Based on the characteristics of pathological images, each segmented image in the training dataset, validation dataset, and test dataset is labeled with a category to obtain label definition information. The semantic information of the encoder constructed from the multi-scale and multi-level semantic features is passed through an average pooling layer and a fully connected layer to obtain the category probability of each segmented image. When defining the category label for each segmented image, the cross-entropy loss function and the training loss function are used as a joint loss function.
[0061] In a preferred embodiment, it further includes:
[0062] Formula for constructing category label definitions:
[0063]
[0064] n represents the number of categories in each partitioned image;
[0065] The joint loss function formula is as follows:
[0066] L = l L +λl CE Where λ is the trade-off coefficient between the training loss function and the cross-entropy loss, l CE Let l be the cross-entropy loss function. L This is the training loss function.
[0067] Based on the corrected image data test set, the category probabilities are obtained through the symmetric spiral feature fusion network for fine segmentation. The symmetric spiral feature fusion network is then embedded into the basic semantic segmentation network to obtain the updated image semantic segmentation network.
[0068] Based on the characteristics of pathological images, each segmented image in the training dataset, validation dataset, and test dataset is assigned a category label, resulting in label definition information. The category label definition formula is as follows:
[0069]
[0070] In the formula, n represents the number of categories in each colorectal cancer image.
[0071] When an image contains only one category, matrix, its category label is STR. When multiple categories exist, the category pixels other than STR are accumulated on the segmentation mask. The accumulated values of each category pixel are compared, and the segmentation label corresponding to the largest value is taken as the classification label (images generally do not contain three or more categories of tissue).
[0072] The richest semantic information of the encoder, constructed from the multi-scale, multi-level semantic features, is then passed through an average pooling layer and a fully connected layer to obtain the class probability of each segmented image. Specifically, when defining the class label for each segmented image, the cross-entropy loss function and the training loss function are used as a joint loss function. That is, when calculating the classification task, the cross-entropy loss function is used. CE When calculating the loss for multi-class segmentation, a loss function based on IOU, lovaszSoftmax, is used as the training loss function. The joint loss function is formed by combining these two loss functions.
[0073] L = l L +λl CE
[0074] In the formula: where λ is the trade-off between lovaszSoftmax loss and cross-entropy loss, and is set to 10 in all experiments.
[0075] Based on the updated image semantic segmentation network, a weighted model for segmenting colorectal cancer regions, normal mucosal regions, and stromal regions is trained according to the image data training set. The weighted model is then used to segment images within the target pathological region in the corrected image data test set.
[0076] The image data training set was used to train an updated image semantic segmentation network, generating a weighted model suitable for segmenting four semantic classes: H&E-stained colorectal cancer cells, normal mucosa, stroma, and background. Then, the effectiveness of this weighted model was validated on the image data validation set. Four metrics—mean Intersection Over Union (mIOU), mean Dice coefficient (mDice), mean Accuracy (mAcc), and mean Precision (mPer)—were used to comprehensively evaluate the segmentation capability of H&E-stained colorectal cancer images. The weighted model for colorectal cancer lesion segmentation was applied to the target pathological region in the image data validation set to segment image regions containing colorectal cancer cells, normal mucosa, and stroma. The segmentation results show that the weighted model for colorectal cancer lesion segmentation can be better applied to the target region in the image data validation set, thus achieving more accurate multi-class pathological image segmentation.
[0077] Example 2
[0078] Based on the same inventive concept as the semantic segmentation method for H&E staining pathological images of colorectal cancer in the foregoing embodiments, such as Figure 5 As shown, this application also provides a semantic segmentation system for H&E-stained pathological images of colorectal cancer, the system comprising:
[0079] The pathological image acquisition module 11 is used to acquire images of the target pathological region through an image acquisition device and acquire a pathological image dataset of the target pathological region.
[0080] The segmentation module 12 is used to segment the pathological image dataset according to the size of the pathological image, determine the training dataset, the validation dataset and the test dataset, and configure experimental parameters and build a basic semantic segmentation network according to the training dataset, the validation dataset and the test dataset.
[0081] Multi-scale feature extraction module 13 is used to extract features from the training dataset based on a grouped feature sequence extraction module constructed by convolutional kernels, thereby obtaining multi-scale and multi-level semantic features.
[0082] The feature fusion network construction module 14 is used to construct a symmetric spiral backbone feature fusion network based on the grouped feature sequence extraction module through multi-branch arrangement and spiral arrangement, and to analyze the depth and width information of the training dataset, validation dataset and test dataset according to the multi-scale and multi-level semantic features to obtain the corrected image data test set.
[0083] The updated image semantic segmentation network acquisition module 15 is used to obtain class probabilities for fine segmentation based on the corrected image data test set through the symmetric spiral feature fusion network, and to embed the symmetric spiral feature fusion network into the basic semantic segmentation network to obtain the updated image semantic segmentation network.
[0084] Image segmentation module 16 is used to train a weight model for segmenting colorectal cancer region, normal mucosa region and stroma region based on the updated image semantic segmentation network and the image data training set, and to segment the image within the target pathological region in the corrected image data test set using the weight model.
[0085] Furthermore, the pathological image acquisition module 11 is also used for:
[0086] Based on the image acquisition boundary of the image acquisition device, determine the pixel range values;
[0087] Within the specified pixel range, pixel-level annotations are performed on the lesion area, normal mucosal area, and target stromal area to obtain multiple annotated areas;
[0088] The multiple labeled regions are marked using different pixels to obtain multiple labeled regions;
[0089] The multiple marked regions are integrated to obtain the pathological image dataset.
[0090] Furthermore, the cutting and dividing module 12 is also used for:
[0091] The pathological image dataset is segmented and divided based on the size of the pathological images to obtain segmented digital pathological image information;
[0092] The segmented digital pathological image information is sequentially traversed and identified to obtain the segmented pathological image recognition information;
[0093] Based on the identification information of the cut pathological images, the lesion area, normal mucosal area, and target stromal area are determined;
[0094] Based on the lesion area, normal mucosal area, and matrix target area, the pathological image dataset is divided into an image data training set, an image data validation set, and an image data test set according to a predetermined ratio;
[0095] Based on the target pathological region, foreground target instances are obtained and analyzed to obtain the structural and color features of the foreground target instances;
[0096] Based on the structural and color features, experimental parameters were configured and a basic semantic segmentation network was built.
[0097] Furthermore, the multi-scale feature extraction module 13 is also used for:
[0098] The grouped feature sequence extraction module extracts the contour edge detail information of the pathological images in the training dataset to obtain a feature extraction map.
[0099] The multi-scale, multi-level semantic features are obtained by upsampling and downsampling based on the feature extraction map. The upsampling result is used as the input feature of the next information node. The input feature is downsampled. The multi-scale, multi-level semantic features are obtained by concatenating the grouped feature sequence extraction module. The upsampled multi-scale semantic information is used as the prior feature information of the downsampled multi-scale semantic information.
[0100] Furthermore, the feature fusion network construction module 14 is also used for:
[0101] The grouped feature sequence extraction module is connected in the same dimension to deepen the network, resulting in a multi-dimensional network. A pooling layer is used to perform downsampling operations, and the network depth of each dimension is expanded to obtain network depth-expanded features.
[0102] Width downsampling is performed on the network in each dimension to obtain multiple branch structures, and these are then fused with the network depth augmentation features to obtain the symmetric spiral backbone feature fusion network.
[0103] Furthermore, the feature fusion network construction module 14 is also used for:
[0104] The encoder backbone network is constructed based on the multi-scale, multi-level semantic features, and the construction formula is as follows:
[0105]
[0106] Where: function θ[·] represents the feature extraction block implemented by convolution, batch normalization and ReLU activation function, C[·] represents concatenation, D[·] represents downsampling operation, n represents the number of layers after downsampling operation, m represents the number of feature extraction blocks in the feature extraction layer, n=5, m=5.
[0107] Furthermore, the feature fusion network construction module 14 is also used for:
[0108] Based on the multi-scale, multi-level semantic feature extraction, feature layer information of the encoder and decoder is constructed, and feature extraction blocks are fused to obtain feature extraction block aggregation information.
[0109] The feature extraction block aggregates information to generate parameter weights, which are used to regulate information features at various scales.
[0110] The feature extraction blocks are combined into an information decoder using a symmetrical structure in both the depth and width directions through a predetermined serial-parallel connection.
[0111] The information decoder is used to parse the depth and width information of the training dataset, validation dataset, and test dataset to obtain the corrected image data test set.
[0112] Furthermore, the updated image semantic segmentation network acquisition module 15 is also used for:
[0113] Based on the characteristics of pathological images, each segmented image in the training dataset, validation dataset, and test dataset is labeled with a category to obtain label definition information;
[0114] The semantic information of the encoder constructed from the multi-scale and multi-level semantic features is passed through an average pooling layer and a fully connected layer to obtain the class probability of each image.
[0115] Specifically, when defining category labels for each segmented image, the cross-entropy loss function and the training loss function are used as a joint loss function.
[0116] Furthermore, the updated image semantic segmentation network acquisition module 15 is also used for:
[0117] Formula for constructing category label definitions:
[0118]
[0119] n represents the number of categories in each partitioned image;
[0120] The joint loss function formula is as follows:
[0121] L = l L +λl CE Where λ is the trade-off coefficient between the training loss function and the cross-entropy loss, l CE Let l be the cross-entropy loss function. L This is the training loss function.
[0122] The specific example of the semantic segmentation method for colorectal cancer H&E-stained pathological images in the aforementioned Embodiment 1 is also applicable to the semantic segmentation system for colorectal cancer H&E-stained pathological images in this embodiment. Through the foregoing detailed description of the semantic segmentation method for colorectal cancer H&E-stained pathological images, those skilled in the art can clearly understand the semantic segmentation system for colorectal cancer H&E-stained pathological images in this embodiment. Therefore, for the sake of brevity, it will not be described in detail here.
[0123] It should be understood that various forms of the process shown above can be used, with steps rearranged, added, or deleted, as long as the desired result of the technical solution disclosed in this invention can be achieved, and this is not limited herein.
[0124] Note that the above description is merely a preferred embodiment of the present invention and the technical principles employed. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions can be made without departing from the scope of protection of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments, and may include many other equivalent embodiments without departing from the concept of the present invention, the scope of which is determined by the scope of the appended claims.
Claims
1. A method for semantic segmentation of colorectal cancer H&E staining pathology images, characterized in that, The method includes: Images of the target pathological region are acquired using an image acquisition device to obtain a pathological image dataset of the target pathological region; The pathological image dataset is segmented according to the size of the pathological images to determine the training dataset, validation dataset, and test dataset. Experimental parameters are configured and a basic semantic segmentation network is built based on the training dataset, the validation dataset, and the test dataset. A grouped feature sequence extraction module based on convolutional kernels is used to extract features from the training dataset and obtain multi-scale and multi-level semantic features. Based on the grouped feature sequence extraction module, a symmetrical spiral backbone feature fusion network is constructed through multi-branch and spiral arrangements. The training dataset, validation dataset, and test dataset are analyzed for depth and width information according to the multi-scale and multi-level semantic features to obtain the corrected image data test set. Based on the corrected image data test set, the category probabilities are obtained through the symmetric spiral backbone feature fusion network for fine segmentation. The symmetric spiral backbone feature fusion network is then embedded into the basic semantic segmentation network to obtain the updated image semantic segmentation network. Based on the updated image semantic segmentation network, a weight model for segmenting colorectal cancer region, normal mucosa region and stroma region is trained according to the image data training set, and the image within the target pathological region in the corrected image data test set is segmented using the weight model. The grouped feature sequence extraction module constructs a symmetric spiral backbone feature fusion network through multi-branch and spiral arrangements, including: The grouped feature sequence extraction module is connected in the same dimension to deepen the network, resulting in a multi-dimensional network. A pooling layer is used to perform downsampling operations, and the network depth of each dimension is expanded to obtain network depth-expanded features. The width of the network in each dimension is downsampled to obtain multiple branch structures, and these are then fused with the network depth augmentation features to obtain the symmetric spiral backbone feature fusion network. The grouped feature sequence extraction module constructs a symmetric spiral backbone feature fusion network through multi-branch and spiral arrangements, including: The encoder backbone network is constructed based on the multi-scale, multi-level semantic features, and the construction formula is as follows: ; Where: function This represents a feature extraction block implemented using convolution, batch normalization, and ReLU activation functions. Indicates splicing, This indicates a downsampling operation, where n represents the number of layers after the downsampling operation, and m represents the number of feature extraction blocks in the feature extraction layer. n=5, m=5.
2. The method as described in claim 1, characterized in that, The acquisition of the pathological image dataset of the target pathological region includes: Based on the image acquisition boundary of the image acquisition device, determine the pixel range value; Within the specified pixel range, pixel-level annotations are performed on the lesion area, normal mucosal area, and target stromal area to obtain multiple annotated areas; The multiple labeled regions are marked using different pixels to obtain multiple labeled regions; The multiple marked regions are integrated to obtain the pathological image dataset.
3. The method as described in claim 1, characterized in that, The step of segmenting the pathological image dataset according to the size of the pathological images to determine the training dataset, validation dataset, and test dataset, and configuring experimental parameters and building a basic semantic segmentation network based on the training dataset, the validation dataset, and the test dataset, includes: The pathological image dataset is segmented and divided based on the size of the pathological images to obtain segmented digital pathological image information; The segmented digital pathological image information is sequentially traversed and identified to obtain the segmented pathological image recognition information; Based on the identification information of the cut pathological images, the lesion area, normal mucosal area, and target stromal area are determined; Based on the lesion area, normal mucosal area, and matrix target area, the pathological image dataset is divided into an image data training set, an image data validation set, and an image data test set according to a predetermined ratio; Based on the target pathological region, foreground target instances are obtained and analyzed to obtain the structural and color features of the foreground target instances; Based on the structural and color features, experimental parameters were configured and a basic semantic segmentation network was built.
4. The method as described in claim 1, characterized in that, The kernel-based grouped feature sequence extraction module extracts features from the training dataset to obtain multi-scale, multi-level semantic features, including: The grouped feature sequence extraction module extracts the contour edge detail information of the pathological images in the training dataset to obtain a feature extraction map. The multi-scale, multi-level semantic features are obtained by upsampling and downsampling based on the feature extraction map. The upsampling result is used as the input feature of the next information node. The input feature is downsampled. The multi-scale, multi-level semantic features are obtained by concatenating the grouped feature sequence extraction module. The upsampled multi-scale semantic information is used as the prior feature information of the downsampled multi-scale semantic information.
5. The method as described in claim 1, characterized in that, The step involves parsing the depth and width information of the training dataset, validation dataset, and test dataset based on the multi-scale, multi-level semantic features to obtain a corrected image data test set, including: Based on the multi-scale, multi-level semantic feature extraction, feature layer information of the encoder and decoder is constructed, and feature extraction blocks are fused to obtain feature extraction block aggregation information. The feature extraction block aggregates information to generate parameter weights, which are used to regulate information features at various scales. The feature extraction blocks are combined into an information decoder using a symmetrical structure in both the depth and width directions through a predetermined serial-parallel connection. The information decoder is used to parse the depth and width information of the training dataset, validation dataset, and test dataset to obtain the corrected image data test set.
6. The method as described in claim 1, characterized in that, The process of obtaining class probabilities through the symmetric spiral backbone feature fusion network for fine segmentation includes: Based on the characteristics of pathological images, each segmented image in the training dataset, validation dataset, and test dataset is labeled with a category to obtain label definition information; The semantic information of the encoder constructed from the multi-scale and multi-level semantic features is passed through an average pooling layer and a fully connected layer to obtain the class probability of each image. Specifically, when defining category labels for each segmented image, the cross-entropy loss function and the training loss function are used as a joint loss function.
7. The method as described in claim 6, characterized in that, The step of defining category labels for each segmented image in the training dataset, validation dataset, and test dataset based on pathological image characteristics to obtain label definition information also includes: Formula for constructing category label definitions: ; n represents the number of categories in each partitioned image; The formula for the joint loss function is as follows: Where λ is the trade-off coefficient between the training loss function and the cross-entropy loss. Let cross-entropy be the loss function. This is the training loss function.
8. A semantic segmentation method and system for H&E-stained pathological images of colorectal cancer, characterized in that, The system includes: A pathological image acquisition module is used to acquire images of a target pathological region through an image acquisition device, and to acquire a pathological image dataset of the target pathological region. The segmentation module is used to segment the pathological image dataset according to the size of the pathological image, determine the training dataset, the validation dataset and the test dataset, and configure experimental parameters and build a basic semantic segmentation network based on the training dataset, the validation dataset and the test dataset. A multi-scale feature extraction module is used to extract features from the training dataset based on a grouped feature sequence extraction module constructed from convolutional kernels, thereby obtaining multi-scale, multi-level semantic features. The feature fusion network construction module is used to construct a symmetric spiral backbone feature fusion network based on the grouped feature sequence extraction module through multi-branch and spiral arrangements. The module analyzes the depth and width information of the training dataset, validation dataset, and test dataset according to the multi-scale and multi-level semantic features to obtain the corrected image data test set. The updated image semantic segmentation network acquisition module is used to obtain class probabilities for fine segmentation based on the corrected image data test set through the symmetric spiral backbone feature fusion network, and to embed the symmetric spiral backbone feature fusion network into the basic semantic segmentation network to obtain the updated image semantic segmentation network. An image segmentation module is used to train a weighted model for segmenting colorectal cancer regions, normal mucosal regions, and stroma regions based on the updated image semantic segmentation network and the image data training set, and to segment images within the target pathological region in the corrected image data test set using the weighted model. The grouped feature sequence extraction module is connected in the same dimension to deepen the network, resulting in a multi-dimensional network. A pooling layer is used to perform downsampling operations, and the network depth of each dimension is expanded to obtain network depth-expanded features. The width of the network in each dimension is downsampled to obtain multiple branch structures, and these are then fused with the network depth augmentation features to obtain the symmetric spiral backbone feature fusion network. The encoder backbone network is constructed based on the multi-scale, multi-level semantic features, and the construction formula is as follows: ; Where: function This represents a feature extraction block implemented using convolution, batch normalization, and ReLU activation functions. Indicates splicing, This indicates a downsampling operation, where n represents the number of layers after the downsampling operation, and m represents the number of feature extraction blocks in the feature extraction layer. n=5, m=5.