Information meat segmentation method and equipment based on regularization and multi-model fusion
By combining PVTv2 and CNNs to build a model based on regularization and multi-model fusion, the accuracy of fuzzy boundaries and small-volume polyps in polyp segmentation is solved, and the credible measurement of segmentation results is achieved, which improves the accuracy and credibility of polyp segmentation.
Patent Information
- Application Number
- CN202510483224.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-07-29
AI Technical Summary
The existing polyp segmentation methods are inaccurate when facing fuzzy boundaries and small volume polyps, and lack credible measurements of segmentation results, resulting in a high risk of misdiagnosis.
Using a method based on regularization and multi-model fusion, the model is constructed using the improved pyramid vision transformer PVTv2 and convolutional neural network CNNs. Combining evidence regularization terms and multi-model fusion rules, segmentation accuracy is measured through Dice Score and mIoU indicators, local channel attention mechanism and evidence regularization terms are introduced, and the model's learning ability in the zero-evidence area is improved.
It improves the accuracy and credibility of polyp segmentation, especially in the case of blurred boundaries and small polyps, which can more accurately locate polyps, provide credibility measurements for segmentation results, and reduce the risk of misdiagnosis.
Smart Images

Figure CN120388176A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical image segmentation, and particularly relates to a polyp segmentation method based on regularization and multi-model fusion and a computer device. Background Art
[0002] Early detection and resection of polyps are of crucial significance for the prevention and treatment of colorectal cancer. Research shows that early detection of colorectal cancer can increase the 5-year survival rate from 18% in the worst case to 88.5%. With the progress of medical imaging technology, especially endoscopic technology, doctors can clearly observe the interior of the colorectum through colonoscopy for polyp detection. Colonoscopy is considered the gold standard for detecting and removing colorectal lesions. It can provide the appearance characteristics and location information of polyps. Regular screening can detect and remove polyps before they turn into cancer, thus preventing colorectal cancer. However, manual identification and segmentation of polyps are not only time-consuming and laborious but also vulnerable to the influence of doctors' experience and fatigue, resulting in the risk of missed diagnosis or misdiagnosis. Therefore, researchers and engineers have been seeking to automate this process using artificial intelligence and machine learning technologies.
[0003] Polyp segmentation is a basic task in the field of medical image analysis, aiming to accurately locate polyps at an early stage and is of great significance for the clinical prevention of colorectal cancer. Polyp segmentation has gone through two development stages. In the early stage, it mainly relied on manual feature extraction, such as texture, geometric features, simple linear iterative clustering superpixels, etc. However, the results of these methods have low quality, poor generalization ability, are difficult to capture global context information, and are not robust enough for complex scenarios. Later, with the development of deep learning, medical image segmentation using convolutional neural networks (CNNs) has greatly promoted quantitative pathological evaluation, diagnostic support systems, and tumor analysis.
[0004] However, standard AI models face an open-set recognition problem in practical applications. AI models are trained with limited data, but in practical applications, the model may encounter some difficult samples. At this time, the standard AI model may give a wrong and overly confident conclusion. This may lead to misdiagnosis in clinical applications. Therefore, it is very necessary to develop a polyp segmentation method that can achieve credibility measurement. At the same time, existing deep learning-based polyp segmentation methods lack credibility measurement for polyp segmentation results. Especially when facing difficult samples, their segmentation results are unreliable and cannot meet the reliability requirements of clinical work for polyp segmentation prediction results. Summary of the Invention
[0005] The main objective of the present invention is to provide an information-based polyp segmentation algorithm based on regularization and multi-model fusion, aiming to solve the problems of inaccurate segmentation of polyps with fuzzy boundaries and small volumes in current polyp segmentation methods and the lack of measurement of the credibility of segmentation results.
[0006] The present invention is implemented as follows. An information-based polyp segmentation method based on regularization and multi-model fusion includes the following steps: S1. Obtain a polyp segmentation image dataset; S2. Build Model One by adding an evidence regularization term to the improved Pyramid Vision Transformer PVTv2, and build Model Two by adding an evidence regularization term to the Convolutional Neural Network CNNs; S3. Use the polyp segmentation dataset in step S1 to train Model One and Model Two respectively to obtain two sets of evidence and Dirichlet parameters; S4. Based on the two sets of evidence and Dirichlet parameters obtained in S3, obtain the final evidence and Dirichlet parameters through the multi-model fusion rule; S5. For the evidence and Dirichlet parameters obtained in S4, calculate the prediction for polyps and the estimation of uncertainty according to the uncertainty estimation module and the class probability calculation formula; Step S6: Use the Dice Score and mIoU metrics to measure the segmentation accuracy.
[0007] Furthermore, divide the publicly available polyp segmentation dataset in step S1 into a training set and a test set according to a certain ratio.
[0008] Furthermore, step S2 includes the following steps: Step S2.1: Build an encoder module; use PVTv2 (hereinafter collectively referred to as Model One) and Res2Net (hereinafter collectively referred to as Model Two) as encoders respectively for multi-level mapping; for the model using PVTv2, the data generates features in four stages through the encoder, denoted by ; where contains detailed texture information of the target, contains high-dimensional semantic information. Model Two uses an improved Residual Network Res2Net as the encoder to extract features at five levels through convolutional layers , and performs parallel connection on the high-level features as the final output; Step S2.2: Construct the decoder module of Model 1. Model 1 adopts a cascaded attention decoding module, which consists of three parts: a convolutional module (UpConv) for upsampling, a local channel feature enhancement module (LCFE) for feature enhancement, and a convolutional attention module (CAM) for enhancing the robustness of the feature map. By processing the features of the four stages obtained in S2.1, latent information is gradually mined and background information is suppressed. Specifically, in the order of process the features. For , use convolutional kernels for processing, and then enhance the robustness of the feature map through the CAM module. The result is input into the UpConv module to restore the feature map to the original size; for the feature , construct the LCFE module, use the LCFE to extract features to obtain local information, and concatenate the result of inputting the output of the previous CAM module into the UpConv module with the extracted features. The output is sent into the CAM module and restored to high resolution through the UpConv; finally, the final results of each feature are aggregated and predicted to obtain the final segmentation map.
[0009] Step S2.3: Construct the decoder module of Model 2. Model 2 adopts a reverse attention mechanism module (RA) and constructs a feature enhancement module LCFE to decode the high-level feature . Specifically, decode in the order of ; process the previous-level feature through the RA module to improve the decoder's perception ability of details; extract local information from the output result and the output of the previous layer through the LCFE module to improve the decoding quality. In this way, different segmentation maps will be output for the feature , and this process is called deep reliable supervision.
[0010] Step S2.4: Add an evidence regularization term: For each category, the amount of evidence output by an evidence model before passing through the activation function is non-positive. Therefore, using an exponential activation function can generate a larger gradient update. At the same time, for evidence amounts less than zero, even a small change will cause a huge change in the output, resulting in a smaller gradient update compared to common softplus and ReLu. Therefore, the exponential activation function is selected. After the model uses the exponential activation function, a custom evidence regularization term is used to weaken the impact of the zero-evidence region on the overall performance. The specific formula is as follows:
[0011] The parameter represents the uncertainty value, which is obtained by dividing the number of non-real category nodes by the total number of nodes. It determines the relative importance of the correct evidence regularization term and is regarded as a constant during the model update process. Indicates the true class prediction evidence. This means that only the evidence related to the true type is considered, and for non-true classes, their corresponding gradients are zero. This represents that the regularization term only updates the target nodes.
[0012] Furthermore, steps S2.2 and S2.3 include the following steps: Construct the feature enhancement module LCFE: Since the boundary between the polyp and the surrounding mucosa is relatively blurred, in order to enable more accurate boundary localization for polyp segmentation, we invented the feature enhancement module LCFE. LCFE adaptively emphasizes important features and enhances the key information between adjacent channels through the attention gate AG and the local channel attention mechanism. Specifically, first define the gating coefficient and the calculation process of the AG operation:
[0013]
[0014] Indicates the gating coefficient, and represent the function ReLU and the Sigmoid activation function, , and represent 1×1 convolution, is the batch normalization operation; and represent the skip connection feature and the upsampling feature respectively; Using AG can gradually suppress irrelevant background regions based on the spatial information of the image using the grid technique.
[0015] Secondly, different feature channels usually represent different semantic information. In order to achieve a powerful representation through good integration, we use a local channel attention mechanism to explore cross-channel interactions while mining the key clues between channels. Specifically, perform element-wise multiplication on the feature of AG and the upsampling feature and use the skip connection as well as
[0016] is a 3×3 convolution, is the element-wise multiplication, is the element-wise addition; Then, use the channel-level global average pooling GAP to aggregate the convolutional features, and input the result into a one-dimensional convolution and the Sigmoid function to obtain the channel attention; Finally, multiply the input feature by the channel attention, and pass through Channel reduction is performed by convolution to obtain the final output:
[0017] Among them, is a 1×1 convolution, represents a 1D convolution with a convolution kernel of k, represents the Sigmoid function. By multiplying the channel attention and the input features, the key channels can be highlighted, redundant channels or noise can be suppressed, thereby enhancing the semantics.
[0018] The size of the convolution kernel K is calculated as follows:
[0019] Among them, represents the nearest odd number, and C represents the number of channels.
[0020] Furthermore, step S2.2 includes the following steps: Use the convolutional attention module to refine the feature map. The CAM module consists of channel attention, spatial attention, and convolutional blocks, and is represented as follows:
[0021] The input data is denoted as x, the spatial attention mechanism is marked as SA, and the channel attention mechanism is denoted as CA. The role of the spatial attention mechanism is to identify and enhance the regions that need attention in the feature map; while the channel attention mechanism is responsible for determining which feature maps are important. In addition, there is a part called the convolutional block (ConvBlock), which is used to further strengthen the features generated by using the spatial and channel attention mechanisms. Specifically, a convolutional block consists of two 3×3 convolutional layers, and each convolutional layer is followed by a batch normalization layer and a ReLU activation function layer. The following is the structural description of the convolutional block:
[0022] The ReLu activation layer is denoted as The batch normalization operation is denoted as , and is used to represent the convolutional layer. After passing through the CAM, pixel grouping can be performed to suppress background information.
[0023] Furthermore, step S3 includes the following steps: Define the loss function L for model training: Define the modified cross-entropy function , and it is defined as:
[0024]
[0025] Use and to represent the label value and predicted probability of the m-th pixel of the n-th class. Use C to represent the number of classes, and represent the class assignment probability on the simplex as and represent the digital matrix function as , is the concentration parameter of the m-th sample of the polynomial beta function, is the m-dimensional unit simplex. Setting the cross-entropy loss function in this way can relate the Dirichlet distribution and the belief distribution, and obtain the probabilities of different classes and the uncertainties of different pixels based on the evidence collected from the backbone.
[0026] Next, introduce the KL divergence to ensure that the evidence generated by the wrong label is close to 0. The calculation process is as follows:
[0027] The gamma function is represented as , represents the calibration parameter of the Dirichlet distribution to ensure that the evidence of the true class is not 0.
[0028] Introduce the Dice loss function to improve the segmentation performance. The calculation formula is:
[0029] Since a new evidence regularization term is introduced in step S2.4 to avoid the zero-evidence region during the training process, a new evidence regularization term is added to the loss function , and its calculation formula is as follows:
[0030] Therefore, the final total loss function L of the model is defined as: . Training with the loss function defined in this way can maximize the correct evidence, minimize the wrong evidence, and avoid the zero-evidence region.
[0031] Calculate the evidence and Dirichlet parameters for Model 1 and Model 2: After the decoders of Model 1 and Model 2, use the exponential activation function to output the evidence of the segmentation result
[0032] Next, obtain the Dirichlet parameters , so Model 1 and Model 2 obtain the Dirichlet parameters respectively.
[0033] Further, step S4 includes the following steps: According to Dempster-Shafer evidence theory, we can combine evidence from different sources, taking into account all available evidence, to obtain a belief mass. Therefore, we fuse Model One and Model Two to obtain a belief mass. This belief mass is calculated by the following formula: Obtained, where , ; for , use the formula , for u, use the formula , where represents the measure of the conflicting part between two mass sets. The obtained combined mass can ensure that when both models have high uncertainty, the final prediction has low confidence, and when the models have low uncertainty simultaneously, the final prediction may have high confidence.
[0034] After obtaining the combined mass M, calculate the evidence and Dirichlet parameters after fusing Model One and Model Two according to the formula:
[0035]
[0036] Therefore, the fused multi-model fusion evidence e and the fused Dirichlet parameters are obtained .
[0037] Further, step S5 includes the following steps: Through a credible segmentation framework defined by subjective logic, derive the probabilities and uncertainties of different class segmentation problems based on the evidence; through , where , Define the coordinate probability of the nth class and the overall uncertainty value. Use the fused evidence obtained in step S4 . The belief mass and uncertainty of the (i, j)th pixel are expressed as:
[0038]
[0039] The Dirichlet intensity is calculated by the formula , C represents the total number of classes; this represents that the more evidence the pixel obtains for the nth class, the higher its probability.
[0040] On the other hand, the present invention also discloses a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to execute the steps of the above method.
[0041] On yet another aspect, the present invention also discloses a computer device comprising a memory and a processor, the memory storing a computer program, which, when executed by the processor, causes the processor to execute the steps of the above method.
[0042] Compared with the prior art, the beneficial effects of the present invention are as follows: (1) The present invention uses a local channel attention mechanism to fuse cascaded features, and gradually restores the details of the edge on the premise of accurately positioning polyps, improving the polyp positioning ability affected by the diversity of polyp shape, size, color and texture and the blurred boundary in polyp segmentation.
[0043] (2) The present invention constructs a credible polyp segmentation model based on subjective logic evidence theory, derives the probability and uncertainty of the polyp segmentation problem, and measures the credibility of the segmentation result.
[0044] (3) The present invention provides an information-based polyp segmentation method based on regularization and multi-model fusion, and improves the learning ability of the model on zero-evidence training samples by adding an evidence regularization term to the training of the model.
[0045] (4) By performing credible fusion on the models based on PVTv2 and CNNs encoders, comprehensively considering the probability distribution and uncertainty of the output, and integrating the above different models, the accuracy and credibility of polyp segmentation are further improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] The specification drawings forming a part of this application are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings: Figure 1 is a schematic diagram of the steps of an embodiment of the present invention; Figure 2 is a structural diagram of an information-based polyp segmentation model of an embodiment of the present invention; Figure 3 is a structural diagram of a boundary-guided feature unit provided by an embodiment of the present invention; Figure 4 is a structural diagram of Model 2 provided by an embodiment of the present invention; Figure 5 is a structural diagram of an uncertainty estimation unit provided by an embodiment of the present invention; Figure 6It is a schematic diagram of the segmentation result and uncertainty visualization result obtained by the information-based polyp segmentation network based on regularization and multi-model fusion. Detailed implementation manner
[0047] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0048] Specifically, as Figure 1 shown, the embodiments of the present invention provide a technical solution: An information-based polyp segmentation method based on regularization and multi-model fusion, comprising the following steps: S1. Obtain a polyp segmentation image data set; S2. Build Model 1 by adding an evidence regularization term based on the improved Pyramid Vision Transformer PVTv2, and build Model 2 by adding an evidence regularization term based on the Convolutional Neural Network CNNs; S3. Use the polyp segmentation data set in step S1 to train Model 1 and Model 2 respectively to obtain two sets of evidence and Dirichlet parameters; S4. Based on the two sets of evidence and Dirichlet parameters obtained in S3, obtain the final evidence and Dirichlet parameters through the multi-model fusion rule; S5. For the evidence and Dirichlet parameters obtained in S4, calculate the prediction for the polyp and the estimation of uncertainty according to the uncertainty estimation module and the class probability calculation formula; S6. Use the Dice Score and mIoU metrics to measure the segmentation accuracy.
[0049] The following is a specific description. Step S1 includes the following steps: The present invention uses five publicly available datasets, and 1,450 images from the datasets CVC-ClinicDB and Kvasir are used as training materials during training. Among them, the data in CVC-ClinicDB (also known as CVC-612) comes from 612 pictures in 25 colonoscopy videos, and the image size is 384×288. Among them, 550 are used for training and 112 are used for testing. The Kvasir dataset contains 1,000 polyp images and corresponding annotations, 900 are used for training, and 100 are used for testing. To test the generalization ability of the model, three brand-new datasets are used as the test sets, namely: CVC-300, which comes from the EndoScene test dataset. The EndoScene dataset includes 912 images of 44 colonoscopy sequences of 36 patients. We use CVC-300 containing 609 pictures as a test set. CVC-ColonDB also comes from EndoScene and contains 380 images extracted from 15 colonoscopy sequences. The ETIS dataset contains 196 images in 34 colonoscopy videos, with an image size of 1,225×966, which is the dataset with the largest images, and the polyps in the images are the smallest and very difficult to distinguish, making it the test set with the greatest challenge.
[0050] As Figure 2 shown, the constructed credible polyp segmentation model framework based on regularization and multi-model fusion in step S2 includes an encoder, a decoder, uncertainty estimation, and credible fusion, and includes the following steps: Step S2.1: Construct an encoder module. PVTv2 (hereinafter collectively referred to as Model 1) and Res2Net (hereinafter collectively referred to as Model 2) are respectively used as encoders for multi-level mapping. For the model using PVTv2, the data generates four stages of features through the encoder, represented by X1~X4. X1 contains detailed texture information of the target, and X2~X4 contain more advanced semantic information. For the model using Res2Net, the data generates five stages of features through the encoder, represented by , and aggregates advanced features through parallel connection .
[0051] Step S2.2: Construct a decoder module. Model 1 adopts a cascaded attention decoding module, which consists of three parts: Attention gate module for cascaded feature fusion, edge-guided feature module for enhancing boundary representation, and convolutional attention module for robust enhanced feature maps. By processing the features of the four stages obtained in S2.1, latent information is gradually mined and background information is suppressed. Model 2 uses a reverse attention mechanism and a feature enhancement module based on local channel attention mechanism (LCFE) for the three high-level features obtained in S2.1 for decoding.
[0052] Step S2.3: Construct an uncertainty module. The output of step S2.2 is obtained through the operation of a non-negative activation function to obtain an evidence output, and the probabilities and uncertainties of different categories are constructed through the Dirichlet distribution.
[0053] Step S2.2 is used to construct a decoder module, and the specific steps are as follows: In Model 1, first, the high-level feature X4 obtained in step S2.1 is input into the convolutional attention module (CAM). CAM consists of a channel attention, a spatial attention, and a convolutional block. The specific expression is as follows:
[0054] where x represents the input data, SA represents spatial attention, and CA represents channel attention. The output obtained through the convolutional attention module (CAM) is sampled 32 times on one path to obtain a final result. On the other path, after passing through the upsampling layer, it is fused with the feature X3 in the feature attention gate (AG) module, and then the edge information is mined through the edge-guided feature module (EFM). Finally, the convolutional attention module (CAM) is used for layering, and the above operations are repeated in the following two layers. The structure of the edge-guided feature module (EFM) used in Model 1 is as Figure 3 shown. First, given the input feature and the high-level feature obtained through upsampling . First, perform element-wise multiplication on them, append skip connections, and perform 3×3 convolution to obtain the initial fused feature , which is expressed as:
[0055] where is a 3×3 convolution, represents element-wise multiplication, represents element-wise addition. Next, in order to strengthen the feature representation, we use local attention to explore the key feature channels. Specifically, the convolutional features are fused through channel-level global average pooling (GAP), and then the corresponding channel attention is obtained through one-dimensional convolution and the Sigmoid function. Finally, the channel attention and the input feature Multiply them, and then reduce the number of channels through convolution to obtain the output , and the specific calculation process is as follows:
[0056] Among them, is a 1×1 convolution, represents the 1×1 convolution of the convolution kernel k, represents the Sigmoid function. The size k of the convolution kernel can be calculated by the formula:
[0057] Among them, represents the odd number closest to *, and C represents the number of channels of.
[0058] In Model 2, first, the high-level features are initially decoded in the pixel decoder, and then on one path, the output is put into the reverse attention mechanism module to highlight the saliency of the target area and input into the local attention-based feature enhancement module (LCEF). On the other path, the result is directly input into the local attention module. On the LCEF module, the attention gate (AG) and the local attention mechanism adaptively highlight the significant features in the task and mine the key information between adjacent channels for enhancement. AG uses grid attention technology to gradually suppress the irrelevant background areas, and the gating coefficient and AG operations are as follows:
[0059]
[0060] Among them, represents the gating coefficient, and respectively represent the ReLU and Sigmoid activation functions, , and represent 1×1 convolutions, is the batch normalization operation. and respectively represent the skip connection feature and the reverse attention feature.
[0061] Then, the same edge-guided feature module (EFM) as in Model 1 is used. And after obtaining the result, the deep supervision mechanism is used to ensure that the feature hierarchies of the model at different scales are effectively optimized, and the above steps are repeated to obtain a final result.
[0062] As Figure 5 shown, step S2.3 constructs an uncertainty module, and the specific steps are as follows: The output results of the encoder and decoder are passed through softplus to obtain the evidence , where > 0, and H and W respectively represent the height and width of the input data. Next, subjective logic associates the evidence with the Dirichlet distribution through the Dirichlet parameter . Finally, the more evidence of the nth class the (i, j) pixel obtains, the higher its probability; on the contrary, the higher the uncertainty
[0063] Specifically, step S3 includes the following steps: Adopt a combination of cross-entropy loss function, KL divergence loss function, Dice loss function, and regularization term for the information-meat segmentation architecture model based on regularization and multi-model fusion constructed in step S2; The definition of the cross-entropy loss function is as follows:
[0064] where and respectively represent the label value and predicted probability of the mth sample of the nth class
[0065] Then, within the framework of evidence theory, SL associates the Dirichlet distribution with the belief distribution, and based on the evidence collected from the backbone, obtains the probabilities of different classes and the uncertainties of different isotopes. The above formula can be further improved as follows:
[0066] where represents the digital matrix function is the class assignment probability on a simplex is the concentration parameter of the mth sample of the polynomial beta function is the m-dimensional unit simplex
[0067] The definition of the KL divergence loss function is as follows:
[0068] where is the gamma function represents the calibration parameter of the Dirichlet distribution, which is used to ensure that the true class evidence is not misinterpreted as 0
[0069] The definition of the Dice loss function is as follows:
[0070] The present invention uses a new evidence regularization term to reduce the impact of the zero-evidence region on the performance of the evidence model. Its calculation is as follows:
[0071] Among them, represents the uncertainty value output by the evidence model, represents the prediction evidence of the true class, and the regularization term determines the relative importance of correct evidence regularization compared to evidence loss and false evidence regularization, and is regarded as a constant during the parameter update process.
[0072] Therefore, the overall loss function of the network is defined as follows:
[0073] This loss function can maximize the correct evidence, minimize the false evidence, and avoid the zero-evidence region during training.
[0074] Step S4 uses the following steps: The polyp segmentation network in this example is implemented based on the Pytorch framework and accelerated using an NVIDIA GEFORCE RTX 3090 graphics card. For the backbone network, the weights pre-trained on ImageNet are used. The AdamW optimizer is used, and the learning rates and weight decays of Model 1 and Model 2 are set to 1e-5 and 1e-4 respectively, the batch size is set to 16, and the number of iterations is 100 and 60 respectively. We resize the images to 352×352 and use a multi-scale training strategy {0.75, 1.0, 1.25}.
[0075] Step S5 uses the following steps: First, calculate the combined mass; it is composed of the mass and fused together. By recombining the compatible parts of the given two sets of beliefs and ignoring the mutually exclusive parts, the combined belief can be obtained. The specific calculation rules are as follows: First, calculate the combined mass M. It is based on the mass and fused together, recombining the compatible parts of the given two sets of beliefs and ignoring the mutually exclusive parts. The specific calculation rules are as follows:
[0076]
[0077]
[0078] Among them, is a measure of the conflicting part in the two masses, is used for normalization.
[0079] Then, according to the following formula, the multi-model fusion evidence e and the corresponding parameters of the fused Dirichlet distribution are obtained. are as follows:
[0080]
[0081] Finally, the final probabilities of each class and the overall uncertainty are obtained.
[0082] Furthermore, step S6 includes the following steps: In this example, two widely used evaluation metrics for image segmentation, Dice Score and average IoU, are adopted, and the specific calculation rules are as follows:
[0083]
[0084]
[0085] Among them, A represents the polyp lesion area marked by the doctor, and B represents the polyp lesion area segmented by this implementation method. lesion area, represents the IoU of the i-th test image.
[0086] Table 1 Quantitative comparison results of polyp segmentation networks (the best results are in bold)
[0087] The polyp segmentation network constructed by the present invention is compared with other 5 representative polyp segmentation networks (U-Net, U-Net++, ParNet, SANet, PVT-CASACDE) in a comparative experiment. The results of the quantitative comparison experiment are shown in Table 1. It can be seen from Table 1 that most of the results of the network of the present invention are better than the classical medical image segmentation line and the other 4 advanced polyp segmentation methods. The qualitative comparison experiment results of the polyp segmentation network constructed by the present invention and the other 5 representative polyp segmentation networks are as Figure 6 shown. It can be found that the polyp segmentation method constructed by the present invention has better effects, especially for small polyps with low boundary contrast, and the uncertainty of the segmentation results can be visualized.
[0088] On the one hand, the present invention also discloses a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to execute the steps of the above method.
[0089] In another aspect, the present invention also discloses a computer device, including a memory and a processor. The memory stores a computer program. When the computer program is executed by the processor, the processor is caused to execute the steps of the above method.
[0090] In yet another embodiment provided by the present application, there is also provided a computer program product containing instructions. When it runs on a computer, the computer is caused to execute any one of the above-described mobile source emission prediction methods based on temporal feature migration.
[0091] It can be understood that the system, device, and storage medium provided by the embodiments of the present invention correspond to the method provided by the embodiments of the present invention. For the explanations, examples, and beneficial effects of related content, reference can be made to the corresponding parts in the above method.
[0092] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that includes one or more integrated available media. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk (SSD)).
[0093] It should be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising said element.
[0094] Each embodiment in this specification is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and for the relevant parts, reference can be made to the partial description of the method embodiment.
[0095] The embodiments of the present invention provide a method for information-based meat segmentation based on regularization and multi-model fusion. There are many methods and ways to specifically implement this technical solution. The above is only the specific implementation method of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention. Each component not clearly defined in the embodiments of the present invention can be implemented using existing technologies.
Claims
1. An information-based meat segmentation method based on regularization and multi-model fusion, characterized in that It includes the following steps: S1. Obtain a polyp segmentation image dataset; S2. Based on the improved Pyramid Vision Transformer PVTv2, add an evidence regularization term to construct Model 1, and based on Convolutional Neural Networks CNNs, add an evidence regularization term to construct Model 2; S3. Use the polyp segmentation dataset in step S1 to train Model 1 and Model 2 respectively, obtaining two sets of evidence and Dirichlet parameters; S4. Based on the two sets of evidence and Dirichlet parameters obtained in S3, obtain the final evidence and Dirichlet parameters through a multi-model fusion rule; S5. For the evidence and Dirichlet parameters obtained in S4, calculate the prediction for polyps and the estimation of uncertainty according to the uncertainty estimation module and the class probability calculation formula.
2. The method for information-based meat segmentation based on regularization and multi-model fusion according to claim 1, characterized in that It also includes S6. Use the Dice Score and mIoU metrics to measure the segmentation accuracy.
3. The information-based meat segmentation method based on regularization and multi-model fusion according to claim 1, characterized in that Step S2 includes the following steps: Step S2.1: Construct an encoder module; Model 1 uses Pyramid Vision Transformer v2 (PVTv2) as the encoder, and the input data passes through the encoder to generate features in four stages , where contains detailed texture information of the target, and the features of the remaining three stages contain high-dimensional semantic information; Model 2 uses the improved residual network Res2Net as the encoder to extract five levels of features through convolutional layers , and performs parallel connection on the high-level features as the final output; Step S2.2: Construct the decoder module of Model 1; Model 1 adopts a cascaded attention decoding module, and the decoding module includes an upsampling convolution module UpConv, a feature enhancement module LCFE based on local channel attention, and a convolutional attention module CAM, which processes features in the order of respectively; For the feature , it is processed using a 1×1 convolutional kernel, and the result is fed into the constructed convolutional attention module CAM to enhance the robustness of the feature map. The enhanced result is output and restored to the original image size using upsampling; For the feature , a feature enhancement module LCFE is constructed, and local information is extracted from the original feature by using the feature enhancement module LCFE to obtain . The output of the convolutional attention module CAM in the previous layer is upsampled by the upsampling convolutional module UpConv and then cascaded with the extracted feature; the cascaded result is sent to the convolutional attention module CAM to enhance the information of the feature map, and then restored to the high resolution through upsampling. Finally, the output result of the convolutional attention module CAM is sent to the prediction head, and the predictions are aggregated to generate the final segmentation map. Step S2.3: Construct the decoder model of Model 2; use the Reverse Attention Mechanism Module RA and construct the Local Feature Enhancement Module LCFE to decode the high-level features ; specifically, decode in the order of ; use the Reverse Attention Mechanism Module RA to improve the decoder's perception ability of details for the features and the features of the previous level; then extract local information and improve the decoding quality for the results and the results output by the features of the previous layer in the Local Feature Enhancement Module LCFE Step S2.4: Add an evidence regularization term; after using the exponential activation function, use a new type of evidence regularization term to reduce the impact of the zero-evidence region on the performance of the evidence model, and its calculation is as follows: Among them, , K is the node of the non-real category, S is the total number of nodes, and its value is the uncertainty value output by the evidence model. represents the predicted evidence of the real category, and the regularization term determines the relative importance of the correct evidence regularization term compared to the evidence loss and the false evidence regularization, and is regarded as a constant during the model parameter update process; only the relevant evidence of the real category nodes is considered, and the amount of evidence of the counterexamples is not included. For the nodes of the non-real category k gt, their corresponding gradients are zero, which means that the regularization term will not impose any influence or update on these non-target category nodes.
4. The information-based meat segmentation method based on regularization and multi-model fusion according to claim 3, characterized in that Construct a feature enhancement module LCFE, including: Adaptively highlight the key features in the task through the Attention Gate AG and the local channel attention mechanism, and mine the key information between adjacent channels for enhancement; the specific process is as follows: First, define the gating coefficient and the calculation of the Attention Gate AG operation: Among them, represents the gating coefficient, and represent the ReLU and Sigmoid activation functions, , and represent 1×1 convolution, is the batch normalization operation; and represent the skip connection feature and the upsampling feature respectively; Next, perform element-wise multiplication between the features of the attention gate AG and the upsampled features and append a skip connection as well as a 3×3 convolution to obtain the initial fusion feature : , Among them, is a 3×3 convolution, is an element-wise multiplication, is an element-wise addition; Then, use a channel-level global average pooling GAP to aggregate the convolutional features, and then obtain the corresponding channel attention through a one-dimensional convolution and the Sigmoid function; Finally, multiply the channel attention with the input features and reduce the number of channels through 1×1 convolution to obtain the final output , which is calculated as follows: Among them, is a 1×1 convolution, represents a 1D convolution with a convolution kernel of k, represents the Sigmoid function; The calculation of the convolution kernel K size is as follows: Among them, represents the nearest odd number, and C represents the number of channels.
5. The information-based meat segmentation method based on regularization and multi-model fusion according to claim 3, characterized in that Construct a Convolutional Attention Module CAM, including: This module consists of a channel attention, a spatial attention, and a convolutional block, and is represented by the following formula: Among them, x represents the input data, SA represents the spatial attention, and CA represents the channel attention; the spatial attention determines the positions to be focused on in the feature map, and then enhances these features; the channel attention determines the feature maps to be focused on; ConvBlock represents the convolutional block, and the convolutional block is used to further enhance the features generated by the CA and SA operations; the convolutional block consists of two 3×3 convolutional layers, and each layer is followed by a batch normalization layer and a ReLU activation layer. The convolutional block is represented by the following formula: Among them, is a ReLU activation layer, is a batch normalization operation, is a 3×3 convolutional layer.
6. The method for information-based meat segmentation based on regularization and multi-model fusion according to claim 3, wherein Step S3 includes the following steps: Total loss function of the model is defined as the modified cross-entropy function, i.e., the KL divergence function, i.e., the Dice loss function and the sum of the evidence regularization terms defined in step S2.4; training with this single loss function can maximize the correct evidence, minimize the incorrect evidence, and avoid zero-evidence regions during training; among them, the improved version of the cross-entropy loss function used is modified to: Among them, and represent the label value and predicted probability of the m-th pixel of the n-th class respectively. C represents the number of classes. is the class assignment probability on a simplex. represents the digital matrix function. is the concentration parameter of the m-th sample. is the polynomial beta function of is the m-dimensional unit simplex; To improve the Dice, a segmentation performance metric in the semantic segmentation task, introduce the Dice loss function: To ensure that incorrect labels produce less evidence, or close to 0, introduce the KL divergence: Among them, is the gamma function, represents the correction parameter of the Dirichlet distribution, which is used to ensure that the true class evidence is not misrecognized as 0.
7. The method for information-based meat segmentation based on regularization and multi-model fusion according to claim 6, characterized in that, Calculate the evidence and Dirichlet parameters for the defined Model 1 and Model 2, and their definitions are: Connect an exponential activation function after the decoder to output the evidence of the segmentation result Then, the evidence and the Dirichlet distribution are related through subjective logic to obtain the Dirichlet parameters ; For Model 1, since the evidence it obtained is , the Dirichlet parameter is obtained; for Model 2, since the evidence it obtained is , the Dirichlet parameter is obtained.
8. The method for information-based meat segmentation based on regularization and multi-model fusion according to claim 1, characterized in that Step S4 includes the following steps: Combine the evidence from Model 1 and Model 2 according to Dempster - Shafer evidence to obtain a belief mass; according to the formula obtained, where , ; for , use the formula , for u, use the formula , where the is a measure of the conflicting part between the two mass sets; After obtaining the joint quality M, calculate the evidence and Dirichlet parameters after the fusion of Model 1 and Model 2 respectively according to the following formula: According to the above rules, the multi-model fusion evidence e and the corresponding parameters of the fused Dirichlet distribution are finally obtained. 9. The method for information-based meat segmentation based on regularization and multi-model fusion according to claim 1, wherein, Step S5 includes the following steps: Define a credible segmentation framework according to subjective logic, and derive the probabilities and uncertainties of different types of segmentation problems based on evidence; define , where , represent the probability and the overall uncertainty value at the coordinates of the nth class respectively; then, according to the fused evidence obtained in step S4, where , and H and W are the width and height of the input data; finally, the belief mass and uncertainty of the (i,j) pixel are expressed as: Among them, represents the Dirichlet intensity, and C represents the total number of categories; this means that the more evidence of the nth category a pixel obtains, the higher its probability is. On the contrary, the greater its uncertainty is.
10. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the computer program is executed by the processor, the processor is caused to perform the steps of the method according to any one of claims 1 to 9.
Citation Information
Cited By
Trusted medical image segmentation method based on non-independent evidence fusion and uncertainty decoupling
CN121962619A
A credible medical image segmentation method decoupling non-independent evidence fusion from uncertainty
CN121962619B