Image fine structure intelligent detection algorithm based on double-branch encoder
By employing a dual-branch encoder-based intelligent image fine structure detection algorithm, combined with a lightweight pre-trained model and a structure-aware visual state space module, the problems of weak global modeling capability and high computational overhead in existing technologies are solved, achieving efficient and accurate fine structure segmentation.
Patent Information
- Application Number
- CN202511590490.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-03
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2045-11-03
AI Technical Summary
In existing technologies, convolutional neural networks have weak global modeling capabilities, visual Transformers have high computational overhead, and basic visual models have insufficient generalization ability on specific tasks, making it difficult to efficiently handle image segmentation tasks with fine structures.
We employ an intelligent image fine structure detection algorithm based on a dual-branch encoder, combined with a lightweight pre-trained model and a structure-aware visual state space module. Through a two-level cross-modal fusion mechanism, we improve the model's generalization ability and computational efficiency in scenarios with few samples.
It achieves efficient capture of long-distance dependencies and complex irregular shapes of fine structures with low computational cost, improving the accuracy and boundary clarity of segmentation results, and is suitable for practical application deployment.
Smart Images

Figure CN121074418A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer vision and digital image processing, and particularly relates to an image fine structure intelligent detection algorithm based on a double-branch encoder. BACKGROUND
[0002] In the field of deep learning-based image segmentation, the current mainstream technical solutions are mainly based on two types of architectures, convolutional neural networks and visual Transformers.
[0003] Traditional convolutional neural networks rely on local convolution kernel sliding window operations to extract features, and their inherent inductive bias enables them to efficiently capture local features; however, in order to obtain global information, convolutional neural networks must stack multiple layers of convolution or use large kernel convolution, which is inefficient and difficult to directly establish dependencies between distant pixels. For fine structure segmentation tasks with long and irregular structures and large spans, it is difficult to maintain their continuity and integrity.
[0004] Visual Transformers use self-attention mechanisms that can calculate the correlation weights between any two image blocks in a sequence, and have strong global modeling capabilities by nature; however, the computational complexity of their self-attention mechanism is proportional to the square of the input sequence length, resulting in huge computational and memory overhead when processing high-resolution images, making it difficult to be deployed in real time on edge devices with limited computing resources.
[0005] In recent years, models such as EdgeSAM have emerged, which obtain a lightweight image encoder based on convolutional networks by knowledge distillation of the SAM model based on ViT, a visual base model.
[0006] Although the EdgeSAM encoder can output rich general visual prior features, when it is directly used for specific tasks (such as fine structure segmentation), an effective mechanism is still needed to combine its prior knowledge with task-specific features, otherwise the model's generalization ability and accuracy will still be insufficient when the size of the specific dataset is small, therefore, the present application proposes an image fine structure intelligent detection algorithm based on a double-branch encoder to solve the problems existing in the prior art. SUMMARY
[0007] In view of the problems of weak global modeling capability of the CNN in the prior art, large computing overhead of the ViT, and insufficient generalization capability of the visual basic model on specific tasks and insufficient utilization of prior knowledge, the purpose of the present application is to provide an image fine structure intelligent detection algorithm based on a double-branch encoder, which fully utilizes rich general visual prior features provided by a visual basic model, enhances the generalization capability of the model in a few-sample scene, constructs a fine structure extraction branch through a specially designed structure perception visual state space module, efficiently captures long-distance dependence and complex irregular morphology of the fine structure at a lower computing cost, and effectively fuses general prior features and task-specific features through a two-level cross-modal fusion mechanism, improves the representation capability of the model for fine structure features, and has a low computing complexity while maintaining high precision, and is suitable for practical application deployment.
[0008] To achieve the purpose of the present application, the present application realizes the following technical scheme: an image fine structure intelligent detection algorithm based on a double-branch encoder, comprising the following steps: Step one, construction of an image segmentation system, constructing an image segmentation system comprising a feature prior branch based on a lightweight pre-trained model, a fine structure extraction branch based on a structure perception visual state space module, a cross-modal fusion module, and a multi-scale feature complementary mapping decoder; Step two, feature extraction, inputting an original image into the feature prior branch and the fine structure extraction branch of the image segmentation system respectively, outputting, by the feature prior branch, hierarchical features of four scales F s1 , F s2 , F s3 , F s4 and a F final feature with rich semantics in the last layer, and outputting, by the fine structure extraction branch, five different levels of features; Step three, preliminary fusion, then outputting each level of feature by the fine structure extraction branch, inputting the features of the same level of the feature prior branch into the cross-modal fusion module for first-stage fusion together, and obtaining fused features F fusion1 , F fusion2 , F fusion3 , and F fusion4 ; Step four, secondary fusion, inputting the preliminarily fused features F fusion1 , F fusion2 , and F fusion3 into the corresponding subsequent levels of the fine structure extraction branch in sequence for further processing, and then inputting the features F fusion1 , F fusion2 , F fusion3 , and F fusion4 into the cross-modal fusion module for second-stage fusion together with the final semantic feature F finalThe second stage fusion is performed through the cross-modal fusion module to obtain final encoder features F1, F2, F3 and F4; Step five, decoding segmentation, inputting the final encoder features F1, F2, F3 and F4 into the multi-scale feature complementary mapping decoder to obtain a final binary segmentation map.
[0009] Further improvement lies in that the training strategy of the lightweight pre-training model in the step one adopts a fine-tuning method of only updating the parameters of the normalization layer.
[0010] Further improvement lies in that the fine structure extraction branch in the step two first divides the input image into image patches and adds position encoding, and then sends the image patches into the SAVSS module for processing, wherein the SAVSS module is composed of a gating bottleneck convolution layer, a two-dimensional selective scanning module with a structure perception scanning strategy, a pixel attention guiding fusion module and a residual connection.
[0011] Further improvement lies in that the specific process of the gating bottleneck convolution layer is as follows: S1, for the input feature x, after the bottleneck convolution, group normalization and ReLU activation function, g1(x) and x1 are obtained, which are represented by the following formula g1(x) = ReLU(GN(BottConv(x))) x1 = ReLU(GN(BottConv(g1(x)))) In the formula, BottConv represents bottleneck convolution, and GN represents group normalization. S2, g2(x) is obtained from another path through bottleneck convolution, group normalization and ReLUU activation function, which is represented by g2(x) = ReLU(GN(BottConv(x))); S3, then g2(x) and x1 are combined through Hadamard product to obtain the gating feature map m(x) = g2(x)⊙x1; S4, then the gating feature map m(x) is further refined to obtain x2 = ReLU(GN(BottConv(m(x)))); S5, finally, x2 and the original input x are connected through residual connection to obtain the final output y = x2 + x.
[0012] Further improvement lies in that the scanning strategy includes horizontal alternate scanning, vertical alternate scanning, main diagonal scanning and secondary diagonal scanning, the horizontal alternate scanning is to start from the last row and scan from left to right and from right to left alternately, the vertical alternate scanning is to start from the first column and scan from top to bottom and from bottom to top alternately, the main diagonal scanning is to scan in the diagonal direction from top left to bottom right, and the secondary diagonal scanning is to scan in the diagonal direction from top right to bottom left.
[0013] A further improvement is that the pixel attention-guided fusion module is used to fuse the initial sequence features x1 and the scanned sequence features x2, and the fusion process is expressed by the following formula. f1(x) =(Norm(BottConv(x1)) f2(x) = Norm(BottConv((x2))) σ = Sigmoid(Norm(BottConv( f2(x)⊙f1(x)))) y = (1-σ) ⊙ f1(x) + σ ⊙ f2(x) Where ⊙ represents the Hadamard product, Norm represents batch normalization, and BottConv represents the bottleneck convolution.
[0014] A further improvement is made in step four, where the cross-modal fusion module first uses a 1×1 point convolution and size scaling operation to convert the feature prior branch's features F before fusion. s Features F of the fine structure extraction branch m Alignment is performed in both channel and spatial dimensions, and the aligned features F s and F m The features are concatenated along the channel dimension, and then subjected to 1×1 point convolution, group normalization, and ReLU activation to obtain the intermediate feature F. mid =ReLU(GN(PConv(Concat(F s ,F m )))), for F mid Two independent 1×1 point convolutions and group normalization operations are performed respectively. Convolutional modulation is then used to generate two attention weight maps, A1=GN(PConv(Fmid)) and A2=GN(PConv(Fmid)). Finally, the features F are fused. fusion F is obtained by weighted summation and concatenation with residuals. fusion = A1⊙F s +A2⊙F m +F m .
[0015] A further improvement is made in step five, where the multi-scale feature complementarity mapping decoder consists of a linear layer, a deconvolutional upsampling layer, and a feature complementarity mapping module. Specifically, the input multi-scale features F1, F2, F3, and F4 are first adjusted for the number of channels using a linear layer, and then the deconvolutional upsampling layer is used to upsample each feature map to a uniform size, as expressed by the following formula. Wherein Upsample represents a dynamic convolution layer, and MLP represents a linear layer; the features after upsampling are spliced in the channel dimension, and then fed into a feature complementary mapping module, to enhance the guidance of semantics to space through channel attention and spatial attention mechanisms, and to strengthen the complement of space to semantics through a spatial weight; finally, a final segmentation result image is output through a convolution layer, represented as .
[0016] The application has the advantages that: the application introduces an EdgeSAM-based feature prior branch, so that the algorithm has excellent performance and generalization ability, and can more accurately process diversified complex scenes. The innovative structure perception visual state space module in the branch can better maintain the continuity and integrity of fine structures. The two-level cross-modal fusion module can adaptively and efficiently fuse general prior features and task-specific features, avoids information conflict or loss caused by simple splicing or addition, fully utilizes the synergistic advantages of the double-branch architecture, and improves the quality and efficiency of feature fusion. The pixel attention guided fusion module and the multi-scale feature complementary mapping decoder jointly act on the structure perception visual state space module (SAVSS), can strengthen the perception and reconstruction ability of fine structure edges and other detailed information, and make the segmentation result more accurate and the boundary clearer. BRIEF DESCRIPTION OF DRAWINGS
[0017] Figure 1 It is a general architecture diagram of the application.
[0018] Figure 2 It is a structure perception visual state space module (SAVSS) architecture diagram.
[0019] Figure 3 It is a structure diagram of the gating bottleneck convolution layer. DETAILED DESCRIPTION
[0020] In order to deepen the understanding of the application, the application will be further described in combination with the embodiments below, and the embodiments are only used to explain the application and do not constitute a limitation on the protection scope of the application.
[0021] According to Figure 1 , Figure 2 and Figure 3 , the application provides an image fine structure intelligent detection algorithm based on a double-branch encoder, including the following steps: Step one, construction of an image segmentation system, constructing an image segmentation system containing a feature prior branch based on a lightweight pre-training model, a fine structure extraction branch based on a structure perception visual state space module (SAVSS), a cross-modal fusion module and a multi-scale feature complementary mapping decoder, the general architecture is shown in the accompanying drawingsFigure 1 as shown in the accompanying drawings; The training strategy of the lightweight pre-training model adopts a fine-tuning method of updating only the normalization layer parameters to better adapt to the downstream fine structure segmentation task and efficiently utilize the pre-training knowledge, which significantly reduces the training parameter quantity and reduces the overfitting risk.
[0022] Step two, feature extraction, the original image is input into the feature prior branch and the fine structure extraction branch of the image segmentation system respectively, the feature prior branch outputs hierarchical features F s1 , F s2 , F s3 , F s4 of four scales and the F final feature of the last layer has rich semantics. The fine structure extraction branch first divides the input image into image patches and adds position encoding, and then sends it into the SAVSS module for processing, wherein the SAVSS module is composed of a gating bottleneck convolution layer, a two-dimensional selective scanning module with a structure-aware scanning strategy, a pixel attention guiding fusion module and a residual connection, as shown in the accompanying drawings. Figure 2 The structure of the gating bottleneck convolution layer is shown in the accompanying drawings. Figure 3 The specific process is as follows: S1, for the input feature x, after bottleneck convolution (BottConv), group normalization (GN) and ReLU activation function, g1(x) and x1 are obtained, which are represented by the following formula g1(x) = ReLU(GN(BottConv(x))) x1 = ReLU(GN(BottConv(g1(x)))) In the formula, BottConv represents bottleneck convolution, and GN represents group normalization. S2, g2(x) is obtained by another path through bottleneck convolution, group normalization and ReLUU activation function, which is represented by g2(x) = ReLU(GN(BottConv(x))); S3, then g2(x) and x1 are combined by Hadamard product to obtain the gating feature map m(x) = g2(x)⊙x1; S4, then further refine the gating feature map m(x) to obtain x2 = ReLU(GN(BottConv(m(x)))); S5, finally, X2 and the original input x are connected by residual connection to obtain the final output y = x2 + x.
[0023] This layer is used for preliminary feature transformation and gating adjustment to enhance the feature expression ability.
[0024] Two-dimensional selective scanning module: As different scanning strategies affect the sensitivity of the model to the direction, in order to make the two-dimensional selective scanning module better perceive irregular fine structures, a structure perception scanning strategy more in line with the extension characteristics of fine structures is designed, including horizontal alternate scanning, vertical alternate scanning, main diagonal scanning and secondary diagonal scanning; Horizontal alternate scanning: starting from the last row, scanning from left to right and from right to left alternately; Vertical alternate scanning: starting from the first column, scanning from top to bottom and from bottom to top alternately; Main diagonal scanning: scanning in the diagonal direction from top left to bottom right; Secondary diagonal scanning: scanning in the diagonal direction from top right to bottom left.
[0025] Each path expands the two-dimensional image features into a one-dimensional sequence along a certain direction, which is input into the core SSM equation for processing: ①, Discrete parameter calculation
[0026] In the formula, G and H are learnable parameter matrices, which control the dynamics of hidden states and the influence of inputs respectively; Δ is a learnable time step parameter, which is used to control the granularity of discretization; and are the discretized state transition matrix and input influence matrix.
[0027] ②, Hidden state update Hidden state y k The update formula at time step k is
[0028] In the formula, y k -1 represents the hidden state of the previous time step, r k is the input of the current time step.
[0029] ③, Output calculation The output c k of the current time step is determined by the hidden state and the input, c k =E yk +T rk , where E and T are learnable output projection matrices.
[0030] After processing, the output sequences of the four paths are reorganized into feature maps and added to obtain the output of this module. This strategy enables the model to fully perceive the context information of fine structures in different directions.
[0031] Pixel attention oriented fusion module: used to fuse the initial sequence feature x1 (the output of the gated bottleneck convolution layer) and the scanned sequence feature x2 (the output of the two-dimensional selective scanning module), and the fusion process is represented by the following formula f1(x) =(Norm(BottConv(x1)) f2(x)= Norm(BottConv((x2))) σ = Sigmoid(Norm(BottConv( f2(x)⊙f1(x)))) y = (1-σ) ⊙ f1(x) + σ ⊙ f2(x) Where is the Hadamard product, Norm is the batch normalization, and BottConv is the bottleneck convolution.
[0032] Step three, preliminary fusion, and then output each level feature by the microstructure extraction branch, which is input into the cross-modal fusion module together with the feature of the same level of the feature prior branch to obtain the fused feature F fusion1 , F fusion2 , F fusion3 and F fusion4 ; Step four, secondary fusion, sequentially input the preliminary fused features F fusion1 , F fusion2 and F fusion3 into the subsequent corresponding levels of the microstructure extraction branch for further processing, and then fuse the features F fusion1 , F fusion2 , F fusion3 and F fusion4 with the final semantic feature Ffinal of the feature prior branch through the cross-modal fusion module to obtain the final encoder features F1, F2, F3 and F4; In the cross-modal fusion module, due to the differences in channel number and resolution of the features of the two branches, the cross-modal fusion module first aligns the features F s of the feature prior branch and the features F m of the microstructure extraction branch in channel and spatial dimensions before fusion, and then splices the aligned features F s and F m in the channel dimension, and finally obtains the intermediate feature F mid =ReLU(GN(PConv(Concat(F s , F m), two independent 1x1 point convolutions and group normalization operations are performed on Fmid respectively, two attention weight maps A1=GN(PConv(F mid ) and A2=GN(PConv(Fmid)) are generated by using convolution modulation, and finally fused features F fusion are obtained by weighted summation and residual connection. fusion = A1⊙F s + A2⊙F m + F m .
[0033] Step five, decoding segmentation, inputting the final encoder features F1, F2, F3 and F4 into a multi-scale feature complementary mapping decoder to obtain a final binary segmentation map.
[0034] The multi-scale feature complementary mapping decoder is composed of a linear layer, a deconvolution upsampling layer and a feature complementary mapping module. Specifically, the input multi-scale features F1, F2, F3 and F4 are first adjusted in channel number using a linear layer, and then each feature map is upsampled to a unified size using a deconvolution upsampling layer, which is represented by the following formula
[0035] In the formula, UpSample represents a dynamic convolution layer, and MLP represents a linear layer; the upsampled features are then concatenated in the channel dimension, and then input into the feature complementary mapping module, which enhances the guidance of semantics by channel weight and the complement of semantics by spatial weight through channel attention and spatial attention mechanisms; finally, the final segmentation result map is output through a convolution layer, represented as .
[0036] The basic principles, main features and advantages of the present application are shown and described above. It should be understood by those skilled in the art that the present application is not limited by the above examples, and the above examples and descriptions in the specification are only to illustrate the principles of the present application. Without departing from the spirit and scope of the present application, various changes and improvements can be made to the present application, and these changes and improvements all fall within the scope of the claimed present application. The scope of protection of the present application is defined by the appended claims and their equivalents.
Claims
1. A dual-branch encoder-based image microstructure intelligent detection algorithm, characterized in that, The method comprises the following steps: Step one, construction of an image segmentation system, which comprises a feature prior branch based on a lightweight pre-training model, a fine structure extraction branch based on a structure perception visual state space module, a cross-modal fusion module, and a multi-scale feature complementary mapping decoder; Step two, feature extraction, the original image is input into the feature prior branch and the fine structure extraction branch of the image segmentation system respectively, the feature prior branch outputs F s1 , F s2 , F s3 , F s4 four scale hierarchical features and the last layer has rich semantic F final features of five different levels of features; Step three, preliminary fusion, then extract branch output each level feature from the microstructure, with the same level feature of the feature prior branch input into the cross-modal fusion module for the first stage fusion, get the fused feature F fusion1 , F fusion2 , F fusion3 and F fusion4 ; Step four, secondary fusion, the features F fusion1 , F fusion2 and F fusion3 are sequentially input into the subsequent corresponding levels of the fine structure extraction branch for continuous processing, and then the features F fusion1 , F fusion2 , F fusion3 and F fusion4 are respectively combined with the final semantic features F final of the feature prior branch to obtain the final encoder features F1, F2, F3 and F4 through the second stage fusion of the cross-modal fusion module. Step five, decoding segmentation, in which the final encoder features F1, F2, F3, and F4 are input into the multi-scale feature complementary mapping decoder to obtain a final binary segmentation map.
2. The image fine structure intelligent detection algorithm based on a double-branch encoder according to claim 1, characterized in that: The training strategy of the lightweight pre-training model in step one adopts a fine-tuning method of updating only the parameters of the normalization layer.
3. The image fine structure intelligent detection algorithm based on a double-branch encoder according to claim 1, characterized in that: The fine structure extraction branch in step two first divides the input image into image patches and adds position encoding, and then sends the image patches into the SAVSS module for processing, wherein the SAVSS module comprises a gated bottleneck convolution layer, a two-dimensional selective scanning module with a structure perception scanning strategy, a pixel attention guiding fusion module, and a residual connection.
4. The image fine structure intelligent detection algorithm based on a double-branch encoder according to claim 3, characterized in that: The specific process of the gated bottleneck convolution layer is as follows: S1, for the input feature x, after bottleneck convolution, group normalization, and ReLU activation function, g1(x) and x1 are obtained, which are represented by the following formula g1(x) = ReLU(GN(BottConv(x))) x1 = ReLU(GN(BottConv(g1(x)))) where BottConv represents bottleneck convolution, and GN represents group normalization; S2, g2(x) is obtained from another path by bottleneck convolution, group normalization, and ReLUU activation function, which is represented by g2(x) = ReLU(GN(BottConv(x))); S3, then g2(x) and x1 are combined by Hadamard product to obtain the gated feature map m(x) = g2(x)⊙x1; S4, then x2 = ReLU(GN(BottConv(m(x)))) is obtained by further refining the gated feature map m(x); S5, finally, x2 and the original input x are connected by residual connection to obtain the final output y = x2 + x.
5. The image fine structure intelligent detection algorithm based on a double-branch encoder according to claim 3, characterized in that: The scanning strategy comprises horizontal alternate scanning, vertical alternate scanning, main diagonal scanning, and secondary diagonal scanning, wherein the horizontal alternate scanning starts from the last row and scans from left to right and from right to left alternately; the vertical alternate scanning starts from the first column and scans from top to bottom and from bottom to top alternately; the main diagonal scanning is performed in the diagonal direction from the top left to the bottom right; and the secondary diagonal scanning is performed in the diagonal direction from the top right to the bottom left.
6. The image microstructure intelligent detection algorithm based on a double-branch encoder according to claim 3, characterized in that: The pixel attention guiding fusion module is used to fuse the initial sequence feature x1 and the scanned sequence feature x2, and the fusion process is represented by the following formula f1(x) = Norm(BottConv(x1)) f2(x) = Norm(BottConv((x2)) σ = Sigmoid(Norm(BottConv(f2(x)⊙f1(x)))) y = (1-σ)⊙f1(x) + σ⊙f2(x) where⊙ is Hadamard product, Norm is batch normalization, and BottConv is bottleneck convolution.
7. The image microstructure intelligent detection algorithm based on a double-branch encoder according to claim 1, characterized in that: The cross-modal fusion module in step four first adopts a 1x1 point convolution and a size scaling operation to fuse the features F s of the feature prior branch before fusion m Aligning in the channel and spatial dimensions, the aligned features F s and F m are spliced in the channel dimension. The spliced features are subjected to a 1x1 point convolution, group normalization, and a ReLU activation function to obtain intermediate features F mid =ReLU(GN(PConv(Concat(F s ,F m )))). mid Two independent 1x1 point convolution and group normalization operations are respectively performed on F fusion to generate two attention weight maps A1=GN(PConv(Fmid)) and A2=GN(PConv(Fmid)) using convolution modulation. The final fused features F fusion are obtained through weighted summation and residual connection s . m +F m .
8. The image microstructure intelligent detection algorithm based on a double-branch encoder according to claim 1, characterized in that: The multi-scale feature complementary mapping decoder in the fifth step is composed of a linear layer, a deconvolution upsampling layer and a feature complementary mapping module. Specifically, the input multi-scale features F1, F2, F3 and F4 are first adjusted in channel number using a linear layer, and then each feature map is upsampled to a uniform size using a deconvolution upsampling layer, represented by the following formula In the formula, UpSample represents a dynamic convolution layer, and MLP represents a linear layer. The upsampled features are then spliced in the channel dimension, and then input into the feature complementary mapping module. Through the channel attention and spatial attention mechanisms, the channel weight is used to enhance the guidance of semantics to space, and the spatial weight is used to strengthen the complement of space to semantics. Finally, the final segmentation result image is output through a convolution layer, represented by .
Citation Information
Patent Citations
Remote sensing image semantic segmentation method based on double-branch multi-scale fusion network
CN119579891A
Medical image segmentation method and device based on spatial perception and frequency domain information
CN120298441A
Classroom behavior identification method based on improved YOLOv12 model
CN120452061A
Remote sensing image segmentation method and system based on convolution-state space fusion and position trigger
CN120635462A
Image segmentation method and device, equipment and medium
CN120726065A
Cited By
Remote sensing image segmentation method based on multi-scale gating bottleneck convolution scanning
CN121640277A
A remote sensing image segmentation method based on multi-scale gated bottleneck convolution scanning
CN121640277B