Image segmentation method based on longitudinally stacked re-parameterized structure model
By adopting a reparameterized model with multiple 3×3 convolutional longitudinal superposition structures in the image segmentation model, the shortcomings of the existing models in memory usage and feature fusion are solved, and higher segmentation accuracy and inference speed are achieved.
Patent Information
- Application Number
- CN202510288092.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-03-11
AI Technical Summary
The existing image segmentation model takes up a large memory during training, and the feature fusion method between branches is relatively simple, making it difficult to fully tap the complementary information between features, resulting in poor segmentation effect on edge details and small target areas.
A reparameterization model based on longitudinal superposition structure is adopted, and the original 3×3 convolution module is replaced with a multi-3×3 convolution longitudinal superposition structure. In the training stage, features are extracted by upsampling and multi-layer 3×3 convolution layers, and in the inference stage, they are reparameterized into a single 3×3 convolution to accelerate inference.
Through hierarchical feature extraction, multi-scale information is captured, segmentation accuracy for small targets and weak edge areas is improved, and the inference speed is significantly improved, which is suitable for practical application deployment.
Smart Images

Figure CN120182601A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision, and particularly relates to an image segmentation method based on a reparameterization model with a vertical stacking structure of (multiple 3x3 convolutions). Background Art
[0002] Image segmentation is a technology and process of dividing an image into several specific regions with unique properties and extracting the target of interest. It classifies pixels belonging to the same class into one class, thereby understanding the scene and content of the image. It has applications in many fields, such as identifying roads, vehicles, and pedestrians in autonomous driving; identifying organs and tissues in CT scans and MRI images in medical image analysis; identifying terrain, vegetation, and buildings in satellite and aerial images in remote sensing image processing; identifying abnormal behaviors and target tracking in video surveillance in security monitoring, etc.
[0003] With the help of an image segmentation model, pixels or regions in an image can be classified according to their features, and the image can be segmented into multiple meaningful regions or objects to better understand and analyze the image content. Image segmentation models play an important role in computer vision and image processing.
[0004] Structure reparameterization refers to the process of converting a model of one structure form into another structure form by performing specific transformations and parameter rearrangements on the network structure during neural network training and inference, so as to achieve purposes such as improving model performance, reducing computational complexity, or increasing inference efficiency. Its core idea is to represent a complex network structure in different forms during the training stage and the inference stage: during the training stage, a more complex and flexible structure is adopted, which is beneficial for the model to better learn features and capture complex patterns in the data; while during the inference stage, it is reparameterized into a simpler and more efficient structure to accelerate the inference speed and reduce memory occupancy, etc. The common method is to change the topology of the network. For example, some branch structures are multi-branch connection methods during training, and are reparameterized into a single-path structure during inference. For example, the residual connection in ResNet has a main path and a residual path during training, and during inference, the parameters of the residual connection can be merged into the main path to form a more concise linear structure, improving the inference speed.
[0005] Currently, the mainstream methods are all horizontally extended (for example, RepVGG [1] extends the feature expression ability during training through a multi-branch structure, and DBB [2] enhances the model capacity through diverse branch combinations). The main problems of such methods are as follows: (1) The multi-branch structure leads to a large memory footprint during training, restricting the expansion scale of the model; (2) The feature fusion method between branches is relatively simple, making it difficult to fully exploit the complementary information between features. When reflected in image segmentation applications, the negative impacts on the image segmentation effect are mainly: (1) Segmentation discontinuities are likely to occur in key areas such as edge details; (2) The segmentation accuracy for small targets or weak edge areas is insufficient. Summary of the Invention
[0006] To solve the above problems, the present invention proposes an image segmentation method based on a vertically stacked reparameterized structure model. When performing image segmentation, this method uses a vertically extended structure to reparameterize the model. Specifically, the steps of this method include: 1) Acquire and process the image; 2) Use the image to be processed as the input of the image segmentation model, and the output of the image segmentation model is the segmented image.
[0007] The image segmentation model is an image segmentation model reparameterized based on a vertically stacked structure; in the network structure of the image segmentation model, all 3×3 convolution modules are replaced with a multi-3×3 convolution vertically stacked structure.
[0008] (1) During the training stage of the image segmentation model, the multi-3×3 convolution vertically stacked structure replaces the 3×3 convolution in the image segmentation model. Let the data dimensions of the input be H×W×Cin, which are the height, width, and number of input channels of the original image or image features respectively, and the data dimensions of the processing result output be H×W×Cout, which are the height, width, and number of output channels of the output image or image features respectively. Then, in the multi-3×3 convolution vertically stacked structure, the operation process for the input data is as follows:
[0009] S1. Obtain 2H×2W×Cin data through upsampling;
[0010] S2. After passing through one or more 3×3 convolution layers, the number of channels changes to obtain 2H×2W×Cin;
[0011] S3. Use a 3×3 convolution with a convolution stride of 2 to change the resolution and number of channels to H×W×Cout and output.
[0012] (2) During the inference stage, the multi-3×3 convolution vertically stacked structure is changed to a single 3×3 convolution to perform inference on the input data.
[0013] The operation process in the training stage of the image segmentation model further includes S4, enabling skip connections: if Cout = Cin, the processing result of step S3) is directly added to the input and used as the final output.
[0014] In the training stage of the image segmentation model, the upsampling in step S1 can be implemented by the bilinear interpolation method.
[0015] The image segmentation model is a Unet network model or a SegNet network model; it is also applicable to DeepLab series models, including DeepLabV3, DeepLabV3+, PSPNet and other models.
[0016] The method for reparameterizing the structure of stacking multiple 3x3 convolutions longitudinally in the present invention can be applied to various existing segmentation models to replace the basic 3×3 convolution structure. In the training stage, the present invention uses an upsampling layer and a structure of stacking multiple 3×3 convolutions longitudinally to replace the original simple convolution, which is beneficial for the model to better learn features and capture complex patterns in the data; in the inference stage, it is reparameterized into a single 3×3 convolution to accelerate the inference speed and reduce the occupation of computing resources.
[0017] Compared with horizontal expansion, the advantages of the longitudinal expansion of the present invention are mainly: (1) Through hierarchical feature extraction, it can better capture multi-scale information; (2) The degree of parameter sharing is high, and the video memory occupation is less.
[0018] Reflected in the application of image segmentation, the positive effects on the image segmentation effect are mainly: (1) The segmentation accuracy of small targets and weak edge regions is improved; (2) While ensuring the accuracy, the inference speed is significantly improved, which is beneficial for actual application deployment. Description of the Drawings
[0019] Figure 1(a) shows the structure of the model training stage of the present invention;
[0020] Figure 1(b) shows the structure of the model inference stage of the present invention;
[0021] Figure 2(a) shows a typical Unet network structure;
[0022] Figure 2(b) shows the Unet network structure based on the model of the embodiment of the present invention. Detailed Embodiments
[0023] The present invention will be further described below in conjunction with the drawings and specific embodiments.
[0024] An image segmentation method based on a vertically stacked reparameterized structure model. When performing image segmentation, this method uses a vertically extended structure reparameterized model. Specifically, the steps of this method are as follows: 1) Acquire the processed image; 2) Use the image to be processed as the input of the image segmentation model, and the output of the image segmentation model is the segmented image; the image segmentation model is an image segmentation model reparameterized based on a vertically stacked structure; in the network structure of the image segmentation model, all 3×3 convolution modules are replaced with a multi-3×3 convolution vertically stacked structure.
[0025] The design model of the multi-3×3 convolution vertically stacked structure is shown in Figures 1(a) and 1(b).
[0026] I. Figure 1(a) shows the multi-3×3 convolution vertically stacked structure used in the model training stage to replace the 3×3 convolution in various existing image segmentation models. The input data dimensions H×W×Cin are the height, width, and number of input channels of the original image or image features, respectively.
[0027] The image resolution is increased through the upsampling module (which can be implemented using the bilinear interpolation method) to obtain 2H×2W×Cin data;
[0028] After passing through one or more 3×3 convolutional layers, the number of channels changes to obtain 2H×2W×Cin; finally, a 3×3 convolution with a convolution stride of 2 is used to change the resolution and number of channels to H×W×Cout. The black dashed arrow is a skip connection, which is an optional item to directly add the processing result to the input for output (it is required that Cout = Cin). This multi-3×3 convolution vertically stacked structure has more parameters and flexible configuration, which is beneficial for the model to better learn features and capture complex patterns in the data.
[0029] II. In the inference stage, the entire structure can be reparameterized into a single 3x3 convolution in Figure 1(b) to accelerate the inference speed, reduce memory occupancy, and maintain the computational complexity of the original model.
[0030] The specific implementation process of the module reparameterization method using PyTorch is as follows:
[0031] 1. Combine the weights and biases of the 3×3 convolution conv1 and the 3x3 convolution conv2 to obtain w0 and b0.
[0032] First, perform a permute(1,0,2,3) operation on the weight w1 of the 3×3 convolution conv1 to adjust its dimension order.
[0033] Then, perform a flip([2,3]) operation on the processed w1 to flip its last two dimensions.
[0034] Use the processed w1 as the convolutional kernel to perform a convolution operation on the weights w2 of the 3×3 convolution conv2 to obtain the weights w0.
[0035] Multiply the bias b1 of the 3×3 convolution conv1 by w2 and sum over the [1, 2, 3] dimensions, then add b2 as the bias b0.
[0036] 2. Perform matrix multiplication and reshaping on the combined weights w0 to achieve downsampling.
[0037] First, perform a reshaping operation on w0, reshaping it from its original shape to the shape of (w_cout, w_cin, matrix_shape[0]), where matrix_shape[0] is the number of rows of the transformation matrix (i.e., 25). This reshaping operation is to make the shape of w0 match the dimensions of the transformation matrix for matrix multiplication. (Reference)
[0038] Then, use the torch.matmul function provided by pytorch to perform matrix multiplication on the reshaped w0 and the transformation matrix, and reshape the result back to the shape of (w_cout, w_cin, 3, 3) to obtain the new w0.
[0039] 3. Combine the skip connections (if the skip connections are selected to be enabled)
[0040] Generate an identity mapping convolutional kernel weight w3 with 1 only at the center position of the convolutional kernel and 0 elsewhere, and a convolutional kernel bias b3 of all 0s. Use the method in step 1 to combine them with w0 and b0 to obtain the new w0 and b0.
[0041] 4. Output w0 and b0 as the weight and bias parameters for combining the 3×3 convolution in the inference stage.
[0042] Take the Unet in Figures 2(a) and 2(b) as an example. Figure 2(a) is a typical Unet, where the 3×3 convolution that can be replaced by the multi-3×3 convolution longitudinal stacking structure proposed in the present invention is marked. Figure 2(b) is the Unet replaced by the multi-3×3 convolution longitudinal stacking structure proposed in the present invention. It can be seen that the method of the present invention is simple and easy to implement, does not bring major changes to the original model, is especially convenient in programming implementation, and avoids potential hidden dangers caused by large-scale modification of the model.
[0043] Specific implementation process of the segmentation method:
[0044] 1. Prepare the training dataset X = {x m , z m} m m-1 The test dataset Y = {y n , z n}n n-1 , where x and y are images, and z is the segmentation label.
[0045] 2. Construct a Unet network and replace all 3×3 convolution modules in it with the multi-3x3 convolution vertically stacked structure proposed by the present invention.
[0046] 3. Use the training dataset to train the network to obtain the trained model M.
[0047] 4. Use the structural reparameterization method to reparameterize the multi-3×3 convolution vertically stacked structure in the trained M to obtain the reparameterized model Mc, which has the same structure as the original Unet model.
[0048] 5. Use Mc to test the test dataset to obtain the image segmentation result.
[0049] Appendix: Transformation matrix (25×9)
[0050]
[0051] References:
[0052] [1] Ding X, Zhang X, Ma N, et al. Repvgg: Making vgg-style convnets great again. Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2021:13733-13742.
[0053] [2] Ding X, Zhang X, Han J, et al. Diverse branch block: Building a convolution as an inception-like unit. Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2021:10886-10895.
Claims
1. An image segmentation method based on a vertically stacked reparameterized structural model. When performing image segmentation, the method adopts a vertically expanded structural reparameterized model. Specifically, the steps of the method include: First, the image is processed at the acquisition site; then, the image to be processed is used as the input of the image segmentation model, and the output of the image segmentation model is the segmented image; The image segmentation model is characterized in that the image segmentation model is a re-parameterized image segmentation model based on a longitudinal stacking structure; in the network structure of the image segmentation model, all 3×3 convolution modules are replaced by a multi-3×3 convolution longitudinal stacking structure; (I) In the training stage of the image segmentation model, the 3×3 convolution in the image segmentation model is replaced by a multi-3×3 convolution vertical stacking structure; assuming that the input data dimensions H×W×Cin are the height, width and number of input channels of the original image or image feature, and the data dimensions H×W×Cout of the processing result output are the height, width and number of output channels of the output image or image feature, then in the multi-3×3 convolution vertical stacking structure, the operation process of the input data is: S1, after upsampling, obtains 2H×2W×Cin data; S2, after one or more 3×3 convolutional layers, the number of channels changes to 2H×2W×Cin; S3, use 3×3 convolution with a convolution step size of 2 to change the resolution and number of channels to H×W×Cout and output; (ii) In the inference stage, the multiple 3×3 convolution vertical stacking structure is changed to a single 3×3 convolution to infer the input data.
2. The image segmentation method according to claim 1, characterized in that The operation process in the training stage of the image segmentation model also includes S4, enabling skip connection: if Cout=Cin, the processing result of step S3) is directly added to the input as the final output.
3. The image segmentation method according to claim 1, characterized in that the model The implementation method of reparameterization is: 1) Merge the weights and biases of 3×3 convolution conv1 and 3×3 convolution conv2 respectively to obtain weight w0 and bias b0. The steps include: 1.1) Perform permute(1,0,2,3) operation on the weight w1 of 3×3 convolution conv1 to adjust the dimension order of w1; perform flip([2,3]) operation on the processed w1 to flip the last two dimensions of w1; 1.2) Use w1 processed in step 1.1) as the convolution kernel to perform a convolution operation on the weight w2 of the 3×3 convolution conv2 to obtain the weight w0; 1.3) Multiply the bias b1 of the 3×3 convolution conv1 by w2 and sum them in the [1,2,3]th dimension, and add the bias b2 of the 3×3 convolution conv2 as the bias b0; 2) The combined weight w0 is matrix multiplied and reshaped to achieve downsampling. The steps include: 2.1) Reshape w0 from its original shape to the shape of (w_cout, w_cin, matrix_shape[0]); matrix_shape[0] is the number of rows of the transformation matrix; make the shape of w0 match the dimension of the transformation matrix for matrix multiplication; 2.2) Use pytorch's torch.matmul function to perform matrix multiplication on w0 obtained in step 2.1) and the transformation matrix, and reshape the result into the shape of (w_cout, w_cin, 3, 3) to obtain the new w0; Finally, w0 and b0 are used as the weight and bias parameters of the 3×3 convolution in the inference stage, respectively.
4. The image segmentation method according to claim 2, characterized in that the model The implementation method of reparameterization is: 1) Merge the weights and biases of 3×3 convolution conv1 and 3×3 convolution conv2 respectively to obtain weight w0 and bias b0. The steps include: 1.1) Perform permute(1,0,2,3) operation on the weight w1 of 3×3 convolution conv1 to adjust the dimension order of w1; perform flip([2,3]) operation on the processed w1 to flip the last two dimensions of w1; 1.2) Use w1 processed in step 1.1) as the convolution kernel to perform a convolution operation on the weight w2 of the 3×3 convolution conv2 to obtain the weight w0; 1.3) Multiply the bias b1 of the 3×3 convolution conv1 by w2 and sum them in the [1,2,3]th dimension, and add the bias b2 of the 3×3 convolution conv2 as the bias b0; 2) The combined weight w0 is matrix multiplied and reshaped to achieve downsampling. The steps include: 2.1) Reshape w0 from its original shape to the shape of (w_cout, w_cin, matrix_shape[0]); matrix_shape[0] is the number of rows of the transformation matrix; make the shape of w0 match the dimension of the transformation matrix for matrix multiplication; 2.2) Perform matrix multiplication on w0 obtained in step 2.1) and the transformation matrix, and reshape the result into the shape of (w_cout,w_cin,3,3) to obtain the new w0; 3) Merge jump connection: Generate an identical mapping convolution kernel weight w3 with only the center position of the convolution kernel being 1 and the rest being all 0, and a convolution kernel bias b3 with all 0; Use the method in step 1) to merge w3 and b3 with w0 and b0 respectively to obtain new w0 and b0; Finally, the new w0 and b0 are used as the weight and bias parameters of the merged 3×3 convolution in the inference stage, respectively.
5. The image segmentation method according to claim 1 or 2, characterized in that In the training phase of the image segmentation model, the upsampling in step S1 is implemented using a bilinear interpolation method.
6. The image segmentation method according to claim 1 or 2, characterized in that The image segmentation model is a Unet network model or a SegNet network model.
7. The image segmentation method according to claim 1 or 2, characterized in that The image segmentation model is a DeepLabV3, DeepLabV3+ or PSPNet network model.
Citation Information
Patent Citations
Method for segmenting vasa sanguinea retinae image
CN106651846A
Image segmentation method based on convolutional network
CN111612008A
Medical image segmentation method, system and equipment based on convolution re-parameterization
CN115829986A
Polyp segmentation algorithm based on re-parameterization and convolution block attention
CN117994273A
Edge calculation-oriented reparametric neural network architecture search method
US20230076457A1