An image segmentation method based on the U-EM network

By adopting U-EM network in image segmentation, combining EM routing algorithm and convolution capsule layer, the problem of lack of spatial information in features in deep learning methods is solved, and more accurate and effective image segmentation is achieved.

CN114511491BActive Publication Date: 2025-06-20BEIJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111465652.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-03
Publication Date
2025-06-20
Estimated Expiration
2041-12-03

AI Technical Summary

Technical Problem

The image segmentation method based on deep learning lacks spatial information in features due to convolution and pooling operations, making it difficult to achieve accurate image segmentation in complex scenarios.

Method used

Using the image segmentation method based on U-EM network, the PrimaryCaps layer and the convolution capsule layer are combined with the EM routing algorithm to iteratively update the Pose matrix and the Activation matrix, and use spatial features to perform image segmentation.

Benefits of technology

It improves the accuracy and effect of image segmentation, can better utilize the spatial characteristics of the image, and enhances the segmentation ability in complex scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114511491B_ABST
    Figure CN114511491B_ABST
Patent Text Reader

Abstract

The present invention discloses an image segmentation method implemented based on the U-EM network, which combines the advantages of the Unet and the EM capsule network, and proposes a new image segmentation model. It not only calculates by combining the image features of each layer, but also can utilize the spatial features for the final result output. When predicting the output result, it simultaneously considers the high-level features of the image and the spatial features of the image, can better utilize the image spatial features, and improve the effect and accuracy of the final image segmentation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly to an image segmentation method implemented based on a U-EM network. Background Art

[0002] Image segmentation is a technology and process of dividing an image into several specific regions with unique properties and extracting the target of interest. It is a key step from image processing to image analysis. With the rapid development of computer technology, computer vision has gradually formed its own scientific system. As an important branch in the field of image processing, image segmentation plays an increasingly important role. Image segmentation refers to dividing an image into non-overlapping and meaningful sub-regions. Pixel points in the same region have a certain correlation, and pixel points in different regions have a certain difference, that is, a process of assigning the same label to pixels with the same properties in the picture.

[0003] The earliest image segmentation methods were applied in the field of medical image processing to perform medical analysis and diagnosis after segmenting specific targets in the images. Since the medical image scene is simple and the background and the target are clearly distinguishable, most of the methods in this field are rough pixel-level segmentations through simple threshold-based methods. With the complication of the segmentation scene, the requirements for segmentation technology have become more and more strict, and segmentation methods based on edges, regions, clustering, graph theory, and specific theories have emerged one after another, and the segmentation effect has thus been improved. Especially after introducing deep learning into the field of image processing, more accurate semantic information has been given to the segmented regions, and breakthrough progress has been made in the image segmentation problem. For example, the emergence of deep learning-based image segmentation methods such as the full convolutional network (FCN), pyramid scene parsing network (PSPNet), DeepLab, and Mask RCNN has continuously improved the segmentation accuracy and made the segmentation process more intelligent.

[0004] However, since the convolutional operations in deep learning-based image segmentation methods are locally connected and parameter-shared, unlike a graph structure, the mutual correlation and mutual positional relationship between different features are not considered, and there is actually no available spatial information in the feature maps obtained by convolution. In addition, due to pooling operations in deep learning-based image segmentation methods, serious information loss occurs. For example, in max pooling, only the most active neurons are retained and passed to the next layer, resulting in the loss of valuable spatial information. Therefore, using a capsule network based on spatial feature encoding (EM capsule network) can well solve this problem.

[0005] The structure of the EM capsule network is as Figure 1As shown, ReLU Conv1 is a conventional convolutional layer with a 5×5 convolutional kernel, a stride of 2, 32 (A = 32) output channels (feature maps), and the ReLU activation function. In the PrimaryCaps layer, we use a 1×1 convolutional kernel to convert 32 channels into 32 (B = 32) primary capsules, each capsule containing a 4×4 matrix and an activation value. Here, a conventional convolutional layer is still used to implement PrimaryCaps, and every 4×4+1 neurons are combined into a capsule. After PrimaryCaps is a ×× convolutional capsule layer ××ConvCaps1 with a 3×3 (K = 3) convolutional kernel and a stride of 2. ConvCaps1 uses capsules as both input and output. ConvCaps1 is similar to a conventional convolutional layer, except that it uses EM routing to calculate the capsule output. Then the capsule output of ConvCaps1 is fed into ConvCaps2. ConvCaps2 is another convolutional capsule layer with a stride of 1. The output capsules of ConvCaps2 are connected to ClassCapsules through a 1×1 convolutional kernel, and each classification is represented by a capsule (in MNIST, there are 10 classes, E = 10). EM routing is used to calculate the pose matrices and output activation values of ConvCaps1, ConvCaps2, and Class Capsules. In a CNN, we slide the same filter in the spatial dimension to calculate the same feature map. The position is not considered when detecting the same feature. Similarly, in EM routing, we share the same transformation matrix W in the spatial dimension ij to calculate the votes. However, the technology for image segmentation does not utilize the spatial features of the features after feature extraction, which may lead to image segmentation failure or degraded performance in some scenarios (such as when objects overlap or are too close to each other). To address this shortcoming, the present invention will utilize the property that the EM capsule network can well utilize spatial features to improve the above deficiencies and achieve image segmentation. Summary of the Invention

[0006] In view of the above problems, the present invention proposes an image segmentation method based on a U-EM network.

[0007] To achieve the above object, the present invention provides the following technical solutions:

[0008] An image segmentation method based on a U-EM network, comprising the following steps:

[0009] Step 1: Input a single-channel image of 512×512, and generate a feature map of 512×512×16 through convolution with 16 3×3 convolutional kernels

[0010] Step 2: Input the obtained feature map into the Primary Capsule layer to generate a Pose matrix of 512×512×16×(4×4) and an Activation matrix of 512×512×16×1;

[0011] Step 3: Pool the obtained Pose matrix and Activation matrix through a 2×2 maxpooling layer to generate a matrix of 256×256×16×(4×4 + 1);

[0012] Step 4: Pass the pooled matrix through the convolutional capsule layer and use the EM routing algorithm to iterate three times to update the parameters of the Pose matrix and Activation matrix. The number of channels is twice the number of input channels, obtaining a matrix of 256×256×64×(4×4 + 1);

[0013] Step 5: Repeat Steps 3 and 4 twice. For the last convolutional capsule to prepare for the next step, use the same number of channels as the input matrix to obtain a 128×128×64×(4×4 + 1) matrix;

[0014] Step 6: Upsample the obtained matrix to restore its size, and merge the obtained matrix with the feature matrix of the same layer to obtain a matrix of 128×128×128×(4×4 + 1);

[0015] Step 7: Pass the merged matrix through the convolutional capsule layer and use the EM routing algorithm to iterate three times to update the parameters of the Pose matrix and Activation matrix. The number of channels is one-fourth of the number of input channels, obtaining a matrix of 128×128×32×(4×4 + 1);

[0016] Step 8: Repeat Steps 6 and 7 twice. When passing through the convolutional capsule layer for the last time, the number of generated channels is one-half of the input, obtaining a 512×512×16×(4×4 + 1) matrix;

[0017] Step 9: Pass through a fully connected layer to achieve pixel-level prediction and restore the original image.

[0018] Furthermore, in Step 4, the EM routing is used to iteratively calculate the pose matrix and the activation values of the output capsules. The EM method alternately calls Step E and Step M to fit the data points to a mixture Gaussian model. Among them, Step E determines the probability R assigned to each data point by the parent capsule ij , and Step M recalculates the values of the Gaussian model based on R ij . Repeat the iteration 3 times. The final a j is the output of the parent capsule, and the 16 μ of the final Gaussian model will form the 4×4 pose matrix of the parent capsule.

[0019] Furthermore, the EM routing algorithm in step four is as follows:

[0020] procedure EM ROUTING(a, V)

[0021]

[0022] for t iterations do

[0023]

[0024]

[0025] return a, M

[0026] where a and V are the activation value and the vote of the sub-capsule respectively, and R ij is initialized with a uniform distribution to assign probabilities, that is, at the beginning, each sub-capsule has the same association with any parent capsule.

[0027] Furthermore, call step M to calculate the updated Gaussian model (μ, σ) and the parent activation a j , based on a, V, and the current R ij , and call step E to recalculate the assignment probability R j based on the new Gaussian model and the new a ij .

[0028] Furthermore, the algorithm of step M is as follows:

[0029] procedure M-STEP(a, R, V, j)

[0030]

[0031]

[0032]

[0033]

[0034] a j ← sigmoid(λ(β a - ∑ h cost h ))

[0035] In step M, based on the activation a i of the sub-capsule, the current r ij and the vote V, calculate μ and σ, and step M recalculates the cost and activation a j of the parent capsule, β v and βα Train separately, and the reciprocal λ of the temperature parameter increases by 1 after each routing iteration.

[0036] Furthermore, the algorithm of step E is as follows:

[0037] procedure E-STEP(μ,σ,a,V,i)

[0038]

[0039]

[0040] In step E, based on the new μ, σ and a j Recalculate the assignment probability R ij , which increases if the vote is closer to the μ assignment of the updated Gaussian model.

[0041] Furthermore, use the propagation loss as the loss function for backpropagation to train W, β v and β α , and the loss of the i-th class is defined as:

[0042] L i =(max(0,m-(a t -a i ))) 2

[0043] a t is the activation value of the target class, a i is the activation value of class i, and the total loss is:

[0044]

[0045] Compared with the prior art, the beneficial effects of the present invention are:

[0046] The image segmentation method based on the U-EM network proposed by the present invention combines the advantages of the Unet and the EM capsule network. It is a new image segmentation model that not only calculates by combining the image features of each layer, but also can use the spatial features for the final result output. When predicting the output result, it takes into account both the high-level features of the image and the spatial features of the image, can better utilize the image spatial features, and improve the effect and accuracy of the final image segmentation. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments recorded in the present invention. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings.

[0048] Figure 1 It is the network structure diagram of the EM capsule.

[0049] Figure 2 It is the network structure diagram of the U-EM capsule provided by the embodiment of the present invention.

[0050] Figure 3 It is the pooling operation diagram provided by the embodiment of the present invention.

[0051] Figure 4 It is the convolution schematic diagram provided by the embodiment of the present invention.

[0052] Figure 5 It is the transposed convolution schematic diagram provided by the embodiment of the present invention. Detailed implementation manners

[0053] To better understand the technical solution, the method of the present invention will be described in detail below with reference to the accompanying drawings.

[0054] The image segmentation method based on the U-EM network proposed by the present invention, the network structure diagram of the U-EM capsule is as Figure 2 shown, and includes the following steps:

[0055] Step 1: Input a single-channel image of 512×512, and generate a feature map of 512×512×16 through convolution with 16 3×3 convolution kernels;

[0056] Step 2: Input the obtained feature map into the PrimaryCapsule layer to generate a Pose matrix of 512×512×16×(4×4) and an Activation matrix of 512×512×16×1;

[0057] Step 3: Pool the obtained Pose matrix and Activation matrix through a 2×2 maxpooling layer to generate a matrix of 256×256×16×(4×4 + 1);

[0058] Step 4: Pass the pooled matrix through the convolutional capsule layer, and use the EM routing algorithm to iterate three times to update the parameters of the Pose matrix and the Activation matrix. The number of channels is twice the number of input channels, and a matrix of 256×256×64×(4×4 + 1) is obtained;

[0059] Step 5: Repeat Step 3 and Step 4 twice. For the last convolutional capsule to prepare for the next step, the same number of channels as the input matrix will be used to obtain a 128×128×64×(4×4 + 1) matrix;

[0060] Step 6: Upsample the obtained matrix to restore its size, and merge the obtained matrix with the feature matrix of the same layer to obtain a matrix of 128×128×128×(4×4 + 1);

[0061] Step 7: Pass the merged matrix through the convolutional capsule layer, use the EM routing algorithm to iterate three times, update the parameters of the Pose matrix and the Activation matrix, with the number of channels being one-fourth of the input number of channels, to obtain a matrix of 128×128×32×(4×4 + 1);

[0062] Step 8: Repeat Step 6 and Step 7 twice. When passing through the convolutional capsule layer for the last time, the number of generated channels is one-half of the input, to obtain a matrix of 512×512×16×(4×4 + 1);

[0063] Step 9: Pass through a fully connected layer to achieve pixel-level prediction and restore the original image.

[0064] 1. EM Routing Algorithm

[0065] Use the EM routing to iteratively calculate the pose matrix and the activation values of the output capsules. The EM method alternately calls Step E and Step M to fit the data points to a mixture Gaussian model. Step E determines the probability R assigned to each data point by the parent capsule ij . Step M recalculates the values of the Gaussian model based on R ij . We repeat the iteration 3 times. The final a j is the output of the parent capsule. The 16 μs of the final Gaussian model will form the 4×4 pose matrix of the parent capsule.

[0066] The EM routing algorithm is as follows:

[0067] procedure EM ROUTING(a, V)

[0068]

[0069] for t iterations do

[0070]

[0071]

[0072] return a, M

[0073] The above a and V are the activation values and the votes of the child capsules respectively. We initialize the assignment probability R with a uniform distribution ij , that is, at the beginning, the child capsules have the same association with any parent capsule. We call Step M to calculate the updated Gaussian model (μ, σ) and the parent activation a j, based on a, V, and the current R ij . Then we call step E based on the new Gaussian model and the new a j Recalculate the assignment probability R ij .

[0074] Details of step M:

[0075] procedure M-STEP(a, R, V, j)

[0076]

[0077]

[0078]

[0079]

[0080] a j ← sigmoid(λ(β a - ∑ h cost h ))

[0081] In step M, we calculate μ and σ based on the activation a of the sub-capsule i , the current r ij and the vote V. Step M also recalculates the cost and activation a of the parent capsule j . β v and β α will be trained separately. In our implementation, λ (the reciprocal of the temperature parameter) increases by 1 after each routing iteration.

[0082] Details of step E:

[0083] procedure E-STEP(μ, σ, a, V, i)

[0084]

[0085]

[0086] In step E, we recalculate the assignment probability R based on the new μ, σ, and a j If the vote is closer to μ of the updated Gaussian model, the assignment increases. ij .

[0087] 2. Loss function

[0088] Matrix capsules require a loss function to train W, β v and β α。We choose spread loss as the main loss function for backpropagation. The loss for the i-th class (not the true label t) is defined as:

[0089] L i =(max(0,m-(a t -a i ))) 2

[0090] at is the activation value of the target class (true label), and a i is the activation value of class i. The total loss is:

[0091]

[0092] 3. Pooling Layer

[0093] After obtaining the feature map through the convolutional layer, the next step is to use these features for integration and classification. Theoretically, all features extracted by convolution can be used as the input to a classifier (such as a softmax classifier), but this will face a huge amount of computation. For example, for a 300×300 input image (assuming only one channel), after convolution with 100 3×3 convolutional kernels, the size of the resulting feature matrix is (300 - 3 + 1)×(300 - 3 + 1) = 88804. It is obviously not good to input all this data into the classifier at once.

[0094] At this time, we will use the pooling layer to reduce the dimension of the obtained feature map. Suppose we get a local feature, which is a magnified local view of an image with a large resolution. Then we can approximately view the pixels (feature values) around some pixels, statistically summarize the feature values at a certain position and its adjacent positions in the plane, and use the summarized result as the value of this position in the plane.

[0095] Common pooling methods include Max Pooling and Average Pooling. We use the pooling function to further process the feature map results obtained from the convolution operation. Pooling will statistically summarize the feature values at a certain position and its adjacent positions in the plane, and use the summarized result as the value of this position in the plane. Max Pooling will calculate the maximum value within the position and its adjacent matrix region and use this maximum value as the value of this position. Average Pooling will calculate the average value within the position and its adjacent matrix region and use this value as the value of this position. Using pooling will not change the depth of the data matrix, but only reduce the height and width, achieving the purpose of dimensionality reduction. The max pooling operation will select the maximum value within the stride region to represent the pooling result of this region. The pooling operation is asFigure 3 as shown

[0096] 4. Upsampling and Transposed Convolution

[0097] In the field of deep learning applied to computer vision, after the input image extracts features through a convolutional neural network (CNN), the output size often becomes smaller. Sometimes, we need to restore the image to its original size for further calculations (e.g., semantic segmentation of images). This operation of enlarging the image size to achieve the mapping of the image from a small resolution to a large resolution is called upsampling (Upsample). There are three common methods of upsampling: bilinear interpolation, transposed convolution, and unpooling. Here, we only discuss transposed convolution. The so-called transposed convolution, also known as the inverse convolution, is not a complete inverse process of the forward convolution. In one sentence, transposed convolution is a special forward convolution. First, the size of the input image is enlarged by padding with zeros according to a certain ratio, then the convolution kernel is rotated, and then the forward convolution is performed. The schematic diagrams of convolution and transposed convolution are as shown in Figure 4 and Figure 5 as shown. Among them, Figure 4 in [diagram 1], the convolution kernel size is 3, the stride is 1, and there is no padding. Figure 5 in [diagram 2], the convolution kernel size is 3, the stride is 1, and the padding is 2.

[0098] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features, but these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An image segmentation method implemented based on the U-EM network, characterized in that, It includes the following steps: Step 1: Input a single-channel image of 512×512, and generate a feature map of 512×512×16 through convolution with 16 3×3 convolutional kernels; Step 2: Input the obtained feature map into the Primary Capsule layer to generate a Pose matrix of 512×512×16×(4×4) and an Activation matrix of 512×512×16×1; Step 3: Pool the obtained Pose matrix and Activation matrix through a 2×2 maxpooling layer to generate a matrix of 256×256×16×(4×4 + 1); Step 4: Input the pooled matrix into the convolutional capsule layer, and use the EM routing algorithm to iterate three times to update the parameters of the Pose matrix and Activation matrix. The number of channels is twice the number of input channels, and a matrix of 256×256×64×(4×4 + 1) is obtained; Step 5: Repeat Steps 3 and 4 twice. For the last convolutional capsule to prepare for the next step, the same number of channels as the input matrix is used, and a matrix of 128×128×64×(4×4 + 1) is obtained; Step 6: Upsample the obtained matrix to restore its size, and merge the obtained matrix with the feature matrix of the same layer to obtain a matrix of 128×128×128×(4×4 + 1); Step 7: Input the merged matrix into the convolutional capsule layer, and use the EM routing algorithm to iterate three times to update the parameters of the Pose matrix and Activation matrix. The number of channels is one-fourth of the number of input channels, and a matrix of 128×128×32×(4×4 + 1) is obtained; Step 8: Repeat Steps 6 and 7 twice. When passing through the convolutional capsule layer for the last time, the number of generated channels is one-half of the input, and a matrix of 512×512×16×(4×4 + 1) is obtained; Step 9: Pass through a fully connected layer to achieve pixel-level prediction and restore the original image.

2. The image segmentation method implemented based on the U-EM network according to claim 1, characterized in that, Step 4: Use the EM routing to iteratively calculate the pose matrix and the activation values of the output capsules. The EM method alternately calls Step E and Step M to fit the data points to the Gaussian mixture model. Among them, Step E determines the probability R assigned to each data point of the parent capsule ij , and Step M recalculates the values of the Gaussian model based on R ij . Repeat the iteration 3 times. The final aj is the output of the parent capsule, and the 16 μs of the final Gaussian model will form the 4×4 pose matrix of the parent capsule 3. The image segmentation method implemented based on the U-EM network according to claim 2, characterized in that, The EM routing algorithm in Step 4 is as follows: procedure EM ROUTING(a,V) for t iterations do return a,M where a and V are the activation value and the vote of the sub-capsule, respectively, and R ij is initialized with a uniform distribution to assign probabilities, i.e., at the beginning, a sub-capsule has the same association with any parent capsule.

4. The image segmentation method implemented based on the U-EM network according to claim 3, characterized in that, Call step M to calculate the updated Gaussian model (μ, σ) and the parent activation a j , based on a, V, and the current R ij , call step E to recalculate the assignment probability R based on the new Gaussian model and the new a j ij .​ 5. The image segmentation method implemented based on the U-EM network according to claim 4, characterized in that, The algorithm for Step M is as follows: procedure M-STEP(a,R,V,j) a j ← sigmoid(λ(β a - ∑ h cost h )) In step M, based on the activation a of the child capsules i , the current r ij and the vote V, μ and σ are calculated, and step M recalculates the cost and activation a of the parent capsule j , β v and β α are trained separately, and the reciprocal λ of the temperature parameter increases by 1 after each routing iteration.

6. The image segmentation method implemented based on the U-EM network according to claim 5, characterized in that, The algorithm for Step E is as follows: procedure E-STEP(μ,σ,a,V,i) In step E, based on the new μ, σ, and a j Recalculate the assignment probability R ij , which increases if the vote is closer to the μ assignment of the updated Gaussian model.

7. The image segmentation method implemented based on the U-EM network according to claim 6, wherein, Use the propagation loss as the loss function for backpropagation to train \(W\) and \(\beta\). v and \(\beta\) α The loss for the \(i\)-th class is defined as: L i = (max(0, m - (a t - a i ))) 2 a t is the activation value of the target class, a i is the activation value of class i, and the total loss is:

Citation Information

Patent Citations

  • A novel biomedical image automatic segmentation method based on a U-net network structure

    CN109191476A

  • Gesture image segmentation and recognition method based on improved capsule network and algorithm

    CN110032925A