A real-time semantic segmentation method based on space-frequency cooperative feature enhancement network
Patent Information
- Application Number
- CN202611243821.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-17
- Publication Date
- 2026-09-25
AI Technical Summary
[0004]本发明的目的在于克服现有实时语义分割方法在复杂背景与尺度变化场景下易出现边界模糊、细小目标漏分等问题,同时避免引入过高的推理计算开销,从而在保证工程部署效率的前提下提升语义分割结果的边界精度与区域一致性
[0042]本发明的有益效果在于:通过在主干网络关键连接处引入多尺度特征提取模块,增强不同尺度目标与纹理信息的表征能力,降低下采样带来的细节衰减;通过卷积线性融合注意力模块将卷积注意力的局部结构建模优势与线性注意力的全局语义建模能力进行互补融合,缓解边界模糊、细长结构断裂以及长距离依赖不足问题;通过在训练阶段引入空频协同增强模块将频域低频与高频信息分别用于区域一致性与边界细节的引导学习,提高语义区域内部一致性与边界定位精度,同时空频协同增强模块仅用于训练,以确保推理效率与工程部署的可行性。
Smart Images

Figure CN122821140A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of computer vision and real-time semantic segmentation, specifically to a real-time semantic segmentation method based on a space-frequency collaborative feature enhancement network. Background Technology
[0002] In the field of autonomous driving, real-time semantic segmentation technology plays a crucial role. This technology determines the semantic category of each pixel in the input image, providing the autonomous driving system with a precise outline of the target region, thereby enhancing its environmental perception capabilities. However, in practical applications such as autonomous driving, existing real-time semantic segmentation methods still face challenges due to factors such as large target scale variations, complex backgrounds, blurred boundaries, and occlusion. These challenges include insufficient segmentation accuracy, unclear boundaries, and the tendency to miss small targets.
[0003] In existing technologies, semantic segmentation methods based on convolutional neural networks rely on convolutional operators to extract local features in the spatial domain. While offering strong advantages in local inductive bias and computational efficiency, their limited receptive field makes it difficult to effectively model long-distance dependencies on high-resolution feature maps, easily leading to inconsistencies within semantic regions, confusion between class boundaries, and breaks in elongated structures. To enhance global modeling capabilities, some methods introduce self-attention or Transformer structures to establish relationships between distant pixels through global association. However, the computational complexity of standard self-attention typically increases quadratically with the number of tokens, significantly increasing computational and memory overhead at higher input resolutions. Furthermore, pure attention structures often lack the local inductive bias of convolution; if the fusion method is inappropriate, problems such as insufficient local details and unstable boundary localization may still occur. Therefore, how to maintain inference efficiency while simultaneously ensuring local detail modeling and global semantic consistency is a crucial issue in the design of semantic segmentation networks. Summary of the Invention
[0004] The purpose of this invention is to overcome the problems of blurred boundaries and missed small targets that existing real-time semantic segmentation methods are prone to in complex backgrounds and scale-changing scenarios, while avoiding the introduction of excessive inference computation overhead, thereby improving the boundary accuracy and regional consistency of semantic segmentation results while ensuring engineering deployment efficiency.
[0005] The semantic segmentation network adopts an encoder-decoder architecture, where the encoder serves as the backbone network and is responsible for extracting multi-scale features; the decoder is used to fuse multi-scale features and output pixel-level class prediction maps.
[0006] To achieve the above objectives, the present invention adopts the following technical solution:
[0007] A real-time semantic segmentation method based on a space-frequency collaborative feature enhancement network includes the following steps:
[0008] Step 1: Perform standardization preprocessing on the acquired raw image to obtain the input image for the network;
[0009] Step 2: The input image is fed into the backbone of the semantic segmentation network for feature extraction. First, the low-level information of the image is extracted through the convolutional embedding module and downsampled to obtain the initial feature F0. Then, the residual convolutional block extracts the same-scale features and outputs the feature F1. After nonlinear activation of the feature F1, channel expansion and spatial downsampling are completed. Finally, the residual convolutional block extracts local features and outputs the feature. ;
[0010] Step 3: Add features After nonlinear activation, spatial downsampling and channel dimensionality expansion are performed. Then, the input is fed into a stacked convolutional linear fusion attention module to output features. ;
[0011] Step 4: Introduce a multi-scale feature extraction module to extract features. The output feature Y after being enhanced by the multi-scale feature extraction module;
[0012] Step 5: Downsample feature Y while expanding its channel dimension, then feed it into a stacked convolutional linear fusion attention module to output the feature. Finally, the decoder outputs a pixel-level category prediction map;
[0013] Step 6: Introduce the space-frequency co-enhancement module as an auxiliary training branch during the network training phase. The space-frequency co-enhancement module only participates in the forward and backward propagation during the training phase and does not participate in the computation during the inference phase.
[0014] Furthermore, the backbone network of the semantic segmentation network includes four feature extraction stages: stage1, stage2, stage3, and stage4. Stage1 and stage2 use stacked residual convolutional blocks to perform preliminary feature extraction and downsampling operations; stage3 and stage4 use stacked convolutional linear fusion attention modules to perform feature processing operations.
[0015] Furthermore, the internal structure of the convolutional linear fusion attention module adopts a parallel structure of convolutional attention branches and linear attention branches. For the input feature X, the specific operation of the convolutional attention branch is as follows:
[0016] (1)
[0017] Among them, W h and W w These correspond to fixed 7×1 and 1×7 convolution mappings, respectively. This indicates that a softmax normalization operation is performed in the spatial dimension to generate attention weights;
[0018] The linear attention branch reduces complexity by employing a linear attention approximation. The query Q, key K, and value V are generated by 1×1 convolutions and feature mapping is used. The specific steps are as follows:
[0019] (2)
[0020] in, , To avoid division by zero for minimal constants, E is a local position enhancement term. Local enhancement is achieved using 5×5 depthwise separable convolutions and position embedding is achieved using 3×3 depthwise separable convolutions.
[0021] Furthermore, the enhancement process of the input features by the multi-scale feature extraction module is as follows;
[0022] (3)
[0023] Where X represents the input feature. and represents a 1×1 convolution used for channel remapping; DWConv represents a depthwise separable convolution. This indicates channel-dimensional splicing.
[0024] Furthermore, the specific process of step 6 is as follows:
[0025] Step 6.1: Perform a frequency domain transformation on the input image X according to equation (4) to obtain the spectrum:
[0026] (4)
[0027] Wherein, RFFT2 represents the two-dimensional real-number fast Fourier transform;
[0028] Step 6.2: Construct the low-pass mask M according to equations (5) and (6):
[0029] (5)
[0030] (6)
[0031] in, For element-wise multiplication, the low-frequency region in M is 1, and the rest is 0 or a smooth transition. IRFFT2 represents the two-dimensional real number inverse fast Fourier transform.
[0032] Step 6.3: Equations (7) and (8) generate guiding gating weights based on low / high frequency components:
[0033] (7)
[0034] (8)
[0035] Among them, X low X represents the consistency information between the outline and the region. high Represents boundary, texture, and detail information. The function is a Sigmoid function, with the gate weights restricted to [0,1].
[0036] Step 6.4: Equation (9) modulates the spatial features:
[0037] (9)
[0038] Among them, G low and G high The gating weights are represented by α and β, which are coefficients used to balance high-frequency boundary enhancement and low-frequency consistency enhancement.
[0039] Furthermore, the loss function during the training phase consists of two parts: the backbone output CWD Loss and the auxiliary prediction Pa output from the space-frequency co-enhancement module. The CWD Loss formula is as follows:
[0040] (10) (11)
[0041] Where c=1,2...C represents the channel index, i=1,2...H·W represents the spatial position, and X T and X S The feature maps of the Transformer branch and the CNN branch are transformed by the feature activation function. This transforms the feature maps of CNN and Transformer into a unified channel probability distribution, eliminating the difference between the two in terms of feature scale.
[0042] The beneficial effects of this invention are as follows: By introducing a multi-scale feature extraction module at the key connection of the backbone network, the representation ability of targets and texture information at different scales is enhanced, and the detail attenuation caused by downsampling is reduced; by using a convolutional linear fusion attention module, the local structure modeling advantage of convolutional attention and the global semantic modeling ability of linear attention are complemented and fused, alleviating the problems of boundary ambiguity, slender structure breakage, and insufficient long-distance dependence; by introducing a space-frequency collaborative enhancement module during the training phase, low-frequency and high-frequency information in the frequency domain are used for guiding the learning of regional consistency and boundary details, respectively, thereby improving the consistency within the semantic region and the accuracy of boundary localization. At the same time, the space-frequency collaborative enhancement module is only used for training to ensure inference efficiency and the feasibility of engineering deployment. Attached Figure Description
[0043] Figure 1 This is a structural diagram of the semantic segmentation network and the space-frequency collaborative enhancement module (SPCE) of this invention;
[0044] Figure 2 This is a structural diagram of the Multi-Scale Feature Extraction Module (MSFE) of this invention;
[0045] Figure 3 This is a block structure diagram of the Convolutional Linear Fusion Attention Module (CLAF) of the present invention;
[0046] Figure 4 These are comparison diagrams showing the effects of examples of the present invention. Detailed Implementation
[0047] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. It should be understood that the specific embodiments described herein are only for explaining the present invention and are not intended to limit the present invention.
[0048] A real-time semantic segmentation method based on a space-frequency collaborative feature enhancement network includes the design of a multi-scale feature extraction module, a convolutional linear fusion attention module, and a space-frequency collaborative enhancement module, specifically divided into the following parts:
[0049] This embodiment is an encoder-decoder semantic segmentation network. The encoder is the backbone network used to extract multi-scale features; the decoder is used to fuse multi-scale features and output a pixel-level class prediction map. The semantic segmentation network is as follows: Figure 1 As shown, SAFNet (Spatial-Frequency Co-enhancement Network) is primarily an encoder-decoder structure. The input image first enters the Stem, then passes through two shallow Conv Blocks (stages 1 and 2) to extract local texture and low / mid-level spatial features. These features are then fed into the MSFE for multi-scale enhancement, followed by two CLAF Blocks (stages 3 and 4) for local-global feature modeling, and finally fed into the Decoder to output the segmentation result.
[0050] The working process of a semantic segmentation network is as follows:
[0051] Step 1: Obtain the input image to be segmented The input image is obtained by normalizing, scaling, and data augmentation. The input image I is then fed into the backbone of the semantic segmentation network for feature extraction. The backbone network consists of four feature extraction stages: stage1, stage2, stage3, and stage4. Specifically, stage1 and stage2 use stacked residual convolutional blocks to perform preliminary feature extraction and downsampling operations; stage3 and stage4 use stacked convolutional linear fusion attention modules to perform feature processing operations.
[0052] Step 2: Input image I is fed into the initial convolutional embedding module to extract low-level information and downsample the input resolution from (H0, W0) to (H0 / 4, W0 / 4) to obtain the initial feature F0; then, F0 is input to form stage 1 by stacking residual convolutional blocks, and the output feature is... Next, F1 undergoes nonlinear activation processing and is fed into stage 2, where the number of channels is expanded from C to 2C while downsampling is performed. Then, local features are extracted through residual convolution blocks, finally yielding the output features of stage 2. ;
[0053] Step 3: Input feature map F2 into stage 3. First, apply non-linear activation to F2 and then send it into stage 3. While downsampling, expand the number of channels from 2C to 4C. Then, it enters a linear fusion attention module consisting of two stacked convolutions, and finally obtains the output features of stage 3. The convolutional linear fusion attention module employs a parallel structure of convolutional attention branches and linear attention branches. For the input feature X, the specific operation of the convolutional attention branch is as follows:
[0054] (1)
[0055] Among them, W h and W w These correspond to fixed 7×1 and 1×7 convolution mappings, respectively. This indicates that a softmax normalization operation is performed in the spatial dimension to generate attention weights;
[0056] The linear attention branch reduces complexity by employing a linear attention approximation. The query Q, key K, and value V are generated by 1×1 convolutions and feature mapping is used. The specific steps are as follows:
[0057] (2)
[0058] in, Its function is to map features to a non-negative region, thereby improving the numerical stability of linear attention. To avoid division by zero for minimal constants, E is a local position enhancement term. Local enhancement is achieved using 5×5 depthwise separable convolutions and position embedding is achieved using 3×3 depthwise separable convolutions.
[0059] Step 4: A multi-scale feature extraction module is introduced between stage 3 and stage 4. The feature Y enhanced by the multi-scale feature extraction module is input into stage 4 for further deep semantic extraction. The enhancement process of the input feature X by the multi-scale feature extraction module is performed according to equation (3).
[0060] (3)
[0061] in, and represents a 1×1 convolution used for channel remapping; DWConv represents a depthwise separable convolution. Indicates channel-dimensional splicing;
[0062] Multi-scale feature extraction module, such as Figure 2 As shown, MSFE can be divided into two parts. The first part is separation: the input features are first converted into sequence form by Conv2Seq, then processed by Linear Projection, and then split into three groups of features at different scales: 0.5x, 1x, and 2x. The purpose of this is to allow the same feature to cover small-scale, medium-scale, and large-scale contexts simultaneously. The second part is MultiDWConv: each scale branch is processed by DWConv, then concatenated by Concat, and finally added to the input residual by Add & Dropout. The box in the lower left corner represents the recovery process: Concat -> Channel Restore -> Seq2Conv, which restores the multi-scale sequence features into convolutional feature maps. Finally, the MSFE output is the enhanced multi-scale spatial features.
[0063] Step 5: The features Y input to stage 4 are downsampled while the number of channels is expanded from 4C to 8C. They then enter two stacked convolutional linear fusion attention modules to obtain the output features of stage 4. Finally, the decoder outputs a pixel-level category prediction map;
[0064] Convolutional linear fusion attention module, such as Figure 3As shown, CLAF has two parallel branches. The upper branch is the convolutional attention branch: after the input X passes through a Norm, features are extracted along the height and width directions to obtain Xh and Xw, which are then added together to obtain Xconv. This branch focuses on capturing local boundaries, directional structures, and spatial details. The lower branch is the linear attention branch: after posing the input features, Q, K, and V are generated, and ELU+1 is used to ensure the stability of linear attention computation. Then, Q@context and the normalization factor Z are calculated, and combined with LePE local position enhancement, and then Xlinear is obtained through 1x1 Conv Projection. This branch is mainly responsible for long-distance dependencies and global semantic consistency. Finally, Xconv and Xlinear are fused in Fusion, and then passed through MLP to output the final enhanced features. The core logic of CLAF is: the convolutional branch preserves details, the linear attention branch supplements the global context, and the fusion improves both boundary representation and semantic consistency.
[0065] Step Six: Introduce a space-frequency cooperative enhancement module into the network as an auxiliary training branch. This module only participates in the forward and backward propagation during the training phase and does not participate in computation during the inference phase, thus not introducing additional overhead. The specific operation is as follows:
[0066] Step 6.1: Perform a frequency domain transformation on the input image X according to equation (4) to obtain the spectrum:
[0067] (4)
[0068] Wherein, RFFT2 represents the two-dimensional real-number fast Fourier transform;
[0069] Step 6.2: Construct the low-pass mask M according to equations (5) and (6). Specific operations are as follows:
[0070] (5)
[0071] (6)
[0072] in, For element-wise multiplication, the low-frequency region in M is 1, and the rest is 0 or a smooth transition. IRFFT2 represents the two-dimensional real number inverse fast Fourier transform.
[0073] Step 6.3: Equations (7) and (8) generate guiding gating weights based on low / high frequency components:
[0074] (7)
[0075] (8)
[0076] Among them, X lowX represents the consistency information between the outline and the region. high Represents boundary, texture, and detail information. The function is a Sigmoid function, with the gate weights restricted to [0,1].
[0077] Step 6.4: Equation (9) modulates the spatial features:
[0078] (9)
[0079] Among them, G low and G high The gating weights are α and β, which are coefficients used to balance high-frequency boundary enhancement and low-frequency consistency enhancement.
[0080] Space-frequency cooperative enhancement module, such as Figure 1 As shown: Figure 1 The Transformer Block -> Decoder -> Head in the upper left corner is an auxiliary semantic branch during the training phase, used to provide semantic alignment loss, and does not participate in inference. The SPCE in the middle dashed box is the frequency domain auxiliary enhancement module during training. It is derived from the middle features of the main branch and generates auxiliary supervision signals through frequency domain decomposition and gated modulation. During inference, SPCE and the auxiliary branch above will be removed. Therefore, only the main path Stem -> Conv Block -> MSFE -> CLAF -> Decoder is retained during deployment.
[0081] SPCE is structured as follows: Input -> RFFT2 -> Low Freq / High Freq -> Gatelow / Gatehigh -> Collaborative Modulation -> Projection. The low-frequency branch primarily constrains region consistency, while the high-frequency branch primarily enhances boundaries and details. The output enters the training head and participates in frequency loss.
[0082] Step 7: Explain the composition of the loss function during the training phase, which consists of two parts: the backbone output CWD Loss and the auxiliary prediction Pa output from the space-frequency co-enhancement module. The CWD Loss formula is as follows:
[0083] (10) (11)
[0084] Where c=1,2...C represents the channel index, i=1,2...H·W represents the spatial position, and X T and X SThe feature maps of the Transformer branch and the CNN branch are transformed by the feature activation function. This transforms the feature maps of CNN and Transformer into a unified channel probability distribution, eliminating the difference between the two in terms of feature scale.
[0085] Example:
[0086] This invention presents a real-time semantic segmentation method based on a space-frequency collaborative feature enhancement network. It adopts an encoder-decoder overall architecture, with the network backbone consisting of a Stem module, a multi-level residual convolution module, an MSFE multi-scale feature enhancement module, a CLAF convolution-linear attention fusion module, and a decoder. During the training phase, an SPCE space-frequency collaborative enhancement auxiliary branch is added for feature supervision and optimization. This auxiliary branch is removed during the inference phase, resulting in no additional computational overhead and enabling high-precision, high-real-time pixel-level semantic segmentation.
[0087] 1. Configuration parameters
[0088] Training hardware used an RTX 3090 GPU, and edge deployment hardware used a Jetson Orin NX. The optimizer was AdamW, with an initial learning rate of 4×10, weight decay of 0.0125, and a multinomial decay strategy with a decay coefficient of 0.9. The datasets used were the Cityscapes dataset for autonomous driving scenarios and the ADE20K dataset for general scenarios. The network in this invention includes two versions: a lightweight version of SAFNet-S (5.0M parameters) and a standard version of SAFNet-B (18.7M parameters). The optimal hyperparameter configuration for the SPCE module was: high-frequency enhancement coefficient α=0.3, low-frequency consistency coefficient β=0.2; loss balance coefficient η=0.5, and linear attention anti-zero constant ε=10.
[0089] 2. Overall Network Implementation Process
[0090] Step 1: Perform data preprocessing on the input image, including normalization, scaling, random flipping, and scale perturbation, to obtain a standardized network input image. Step 2: The input image is downsampled fourfold by the Stem module to extract initial low-level features. Shallow texture and mid-level spatial features are then extracted using Stage 1 and Stage 2 residual convolution modules. Step 3: The MSFE module is used to achieve parallel enhancement of 0.5×, 1×, and 2× multi-scale features, enriching multi-scale contextual information. Step 4: The CLAF module, stacked in Stage 3 and Stage 4, captures local edge details through convolutional attention branches and models global long-distance dependencies through linear attention branches, achieving local and global feature fusion optimization. Step 5: During training, the SPCE spatial-frequency collaborative enhancement branch is introduced to decompose intermediate features into high- and low-frequency features using Fourier transform. Feature modulation is achieved through dual-gated weights, and training is jointly supervised by semantic alignment loss, frequency domain alignment loss, and segmentation loss. Step 6: The decoder fuses shallow detail features and high-level semantic features, outputting a pixel-level semantic segmentation prediction map. Step 7: During model inference, all auxiliary training branches are removed, retaining only the backbone network for efficient inference.
[0091] 3. Experimental Comparison Results and Analysis
[0092] Table 1. Performance comparison of the present invention and existing real-time segmentation algorithms on the Cityscapes dataset.
[0093]
[0094] As shown in Table 1, the comparative results show that on the Cityscapes dataset, a core dataset for autonomous driving, SAFNet-B, as presented in this invention, exhibits the best overall performance compared to existing mainstream real-time segmentation algorithms. Compared to the high-precision model PIDNet-M, it reduces the number of parameters by 45.6%, increases the inference frame rate by 38.7%, and achieves higher segmentation accuracy. Compared to SCTNet-B with the same number of parameters, it improves segmentation accuracy by 0.2%. Compared to classic real-time models such as DDRNet and RTFormer, it simultaneously achieves higher segmentation accuracy and inference speed. The lightweight version, SAFNet-S, has only 5.0M parameters and can achieve an ultra-high frame rate inference of 115.1 FPS, balancing lightweight design and segmentation accuracy, making it suitable for high-speed real-time segmentation scenarios.
[0095] Table 2. Performance comparison of the present invention and existing algorithms on the ADE20K dataset.
[0096]
[0097] As shown in Table 2, the comparison results show that on the general scenario ADE20K dataset, the SAFNet-B segmentation accuracy of this invention reaches 44.8% mIoU, which is significantly better than all the mainstream real-time models compared. It is 2.7% higher than RTFormer-B and 1.8% higher than SCTNet-B, and the inference frame rate can reach 114.7 FPS. This verifies that the invention has excellent cross-scene generalization ability and can be adapted to autonomous driving and general scene segmentation tasks at the same time. It solves the technical defects of traditional real-time models, such as poor generalization and low segmentation accuracy in complex scenes.
[0098] Table 3. Performance Comparison of Jetson Orin NX Edge Device Deployment
[0099]
[0100] As shown in Table 3, the comparative results show that in industrial deployment scenarios using embedded edge hardware such as Jetson Orin NX and TensorRT acceleration, the SAFNet-S inference frame rate of this invention reaches 31.6 FPS, significantly higher than mainstream lightweight models such as SeaFormer, PIDNet-S, and DDRNet-23-S, while maintaining a high segmentation accuracy of 78.3%. This demonstrates that the auxiliary training architecture of this invention, which has no inference overhead, is extremely suitable for low-computing-power edge devices, solving the engineering pain points of insufficient real-time performance and severe accuracy degradation after deployment of existing lightweight models, and possessing extremely high industrial application value.
[0101] Table 4 Ablation Experiment Results of Each Module of the Invention
[0102]
[0103] As shown in Table 4, the comparative results analysis shows that the ablation experiments verified the necessity and synergistic gains of the three core modules of this invention. Adding the MSFE multi-scale enhancement module alone improves multi-scale feature representation capabilities, increasing accuracy by 0.3%. After superimposing the CLAF convolutional-linear attention fusion module, it balances local details and global semantics, further improving accuracy by 1.2%. Finally, after introducing the SPCE space-frequency collaborative auxiliary branch, the accuracy reaches the optimal 44.8%, with no decrease in inference frame rate. This proves that the SPCE module only optimizes the training process without increasing inference overhead; the modules are not simply superimposed but can form a synergistic enhancement effect, demonstrating significant technical advantages. Table 5 shows the SPCE module hyperparameter optimization experiment.
[0104]
[0105] As shown in Table 5, the comparative results show that the optimal ratio α=0.3 and β=0.2 was selected through multiple sets of hyperparameter experiments. Enabling only single low-frequency or high-frequency enhancement, or excessively enhancing high- and low-frequency features, all lead to a decrease in segmentation performance. This parameter combination can accurately balance high-frequency edge detail enhancement and low-frequency region consistency optimization, achieving the optimal synergistic effect of spatial frequency features. This provides sufficient experimental evidence for the optimal customized parameter configuration of this invention.
[0106] Table 6 Comparison Experiments of Different Auxiliary Enhancement Methods
[0107]
[0108] As shown in Table 6, the comparative results show that the SPCE (Space-Frequency Enhancement) scheme proposed in this invention achieves the highest segmentation accuracy compared to mainstream auxiliary training schemes such as traditional boundary enhancement, contextual detail enhancement, and prototype feature enhancement. This demonstrates that the optimization approach of combining frequency domain features with spatial domain features possesses significant innovation and technical advantages compared to traditional single-dimensional feature enhancement methods, simultaneously addressing the dual technical problems of discontinuous region segmentation and blurred target edges.
[0109] Combined with the visual illustrations, Figure 4 As shown, a qualitative comparative analysis of the segmentation performance of the present invention (SAFNet) and existing technologies is conducted:
[0110] First, in terms of large-scale scene region segmentation, existing real-time segmentation algorithms are prone to problems such as fragmented region segmentation, semantic confusion, and large-area holes. This invention relies on the low-frequency feature constraint effect of the SPCE spatial-frequency collaborative enhancement mechanism, which can effectively ensure the integrity and smooth consistency of segmentation of large-area scene regions such as roads, vegetation, and buildings. The region prediction results are coherent and free of fragmented missegmentation, and the global semantic discrimination accuracy is significantly better than existing comparative algorithms.
[0111] Secondly, in the segmentation of small and fine-edged targets, traditional algorithms suffer from defects such as blurred edges, missing contours, missed segmentation, and undersegmentation when dealing with small and narrow targets like pedestrians, traffic signs, utility poles, and road edges. This invention, through high-frequency feature enhancement and the CLAF local-global fusion mechanism, accurately strengthens the detailed features of target edges and refines the semantic contours of small targets. This enables complete, clear, and accurate segmentation of small targets, effectively solving the problem of low edge fitting accuracy in existing technologies.
[0112] Third, regarding anti-interference and generalization performance in complex scenes, existing algorithms are prone to misclassification and boundary drift in complex interference scenarios such as changes in illumination, scene occlusion, and blurred distant views. This invention combines multi-scale feature enhancement with bidirectional optimization in the frequency and spatial domains, enabling it to adaptively adapt to scene features of different scales and resolutions, resulting in stronger anti-interference capabilities and superior stability and robustness of the segmentation results.
[0113] The visualization results and the aforementioned quantitative experimental data corroborate each other. This invention, through collaborative optimization of various modules, achieves a triple performance improvement in global region integrity, local edge accuracy, and robustness in complex scenes without sacrificing inference real-time performance. From both visual qualitative and quantitative data perspectives, the innovation and practicality of this invention's technical solution are verified. This invention, through the collaborative design of three core modules—MSFE, CLAF, and SPCE—innovatively introduces a space-frequency collaborative auxiliary training mechanism, significantly improving the global region consistency and local edge detail accuracy of semantic segmentation without increasing model inference computational overhead. This invention effectively solves the technical pain points of existing real-time semantic segmentation methods, such as the difficulty in balancing accuracy and speed, missed segmentation of small targets, blurred edges, weak generalization ability, and poor real-time edge deployment. The model can be efficiently adapted to industrial application scenarios such as in-vehicle autonomous driving perception and embedded mobile real-time scene parsing, demonstrating significant technological advancement and practical value.
Claims
1. A real-time semantic segmentation method based on a space-frequency collaborative feature enhancement network, characterized in that, Includes the following steps: Step 1: Perform standardization preprocessing on the acquired raw image to obtain the input image for the network; Step 2: The input image is fed into the backbone of the semantic segmentation network for feature extraction. First, the low-level information of the image is extracted through the convolutional embedding module and downsampled to obtain the initial feature F0. Then, the residual convolutional block extracts the same-scale features and outputs the feature F1. After nonlinear activation of the feature F1, channel expansion and spatial downsampling are completed. Finally, the residual convolutional block extracts local features and outputs the feature. ; Step 3: Add features After nonlinear activation, spatial downsampling and channel dimensionality expansion are performed. Then, the input is fed into a stacked convolutional linear fusion attention module to output features. ; Step 4: Introduce a multi-scale feature extraction module to extract features. The output feature Y after being enhanced by the multi-scale feature extraction module; Step 5: Downsample feature Y while expanding its channel dimension, then feed it into a stacked convolutional linear fusion attention module to output the feature. Finally, the decoder outputs a pixel-level category prediction map; Step 6: Introduce the space-frequency co-enhancement module as an auxiliary training branch during the network training phase. The space-frequency co-enhancement module only participates in the forward and backward propagation during the training phase and does not participate in the computation during the inference phase.
2. The real-time semantic segmentation method based on a space-frequency collaborative feature enhancement network according to claim 1, characterized in that, The backbone network of the semantic segmentation network includes four feature extraction stages: stage1, stage2, stage3, and stage4. Stage1 and stage2 use stacked residual convolutional blocks to perform preliminary feature extraction and downsampling operations; stage3 and stage4 use stacked convolutional linear fusion attention modules to perform feature processing operations.
3. The real-time semantic segmentation method based on a space-frequency cooperative feature enhancement network according to claim 1, characterized in that, The internal structure of the convolutional linear fusion attention module adopts a parallel structure of convolutional attention branches and linear attention branches. For input feature X, the specific operation of the convolutional attention branch is as follows: (1); Among them, W h and W w These correspond to fixed 7×1 and 1×7 convolution mappings, respectively. This indicates that a softmax normalization operation is performed in the spatial dimension to generate attention weights; The linear attention branch reduces complexity by employing a linear attention approximation. The query Q, key K, and value V are generated by 1×1 convolutions and feature mapping is used. The specific steps are as follows: (2); in, , This represents the activation function. To avoid division by zero for a minimal constant, E is a local position enhancement term. Local enhancement is achieved using a 5×5 depthwise separable convolution and position embedding is achieved using a 3×3 depthwise separable convolution.
4. The real-time semantic segmentation method based on a space-frequency collaborative feature enhancement network according to claim 1, characterized in that, The enhancement process of the input features by the multi-scale feature extraction module is as follows; (3); Where X represents the input feature. and represents a 1×1 convolution used for channel remapping; DWConv represents a depthwise separable convolution. This indicates channel dimension splicing.
5. The real-time semantic segmentation method based on a space-frequency collaborative feature enhancement network according to claim 1, characterized in that, The specific process of step 6 is as follows: Step 6.1: Perform a frequency domain transformation on the input image X according to equation (4) to obtain the spectrum: (4); Wherein, RFFT2 represents the two-dimensional real-number fast Fourier transform; Step 6.2: Construct the low-pass mask M according to equations (5) and (6): (5) ; (6) ; in, For element-wise multiplication, the low-frequency region in M is 1, and the rest is 0 or a smooth transition. IRFFT2 represents the two-dimensional real number inverse fast Fourier transform. Step 6.3: Equations (7) and (8) generate guiding gating weights based on low / high frequency components: (7) ; (8) ; Among them, X low X represents the consistency information between the outline and the region. high Represents boundary, texture, and detail information. The function is a Sigmoid function, with the gate weights restricted to [0,1]. Step 6.4: Equation (9) modulates the spatial features: (9); Among them, G low and G high The gating weights are represented by α and β, which are coefficients used to balance high-frequency boundary enhancement and low-frequency consistency enhancement.
6. The real-time semantic segmentation method based on a space-frequency collaborative feature enhancement network according to claim 1, characterized in that, The training phase loss function consists of two parts: the backbone output CWD Loss and the auxiliary prediction Pa output from the space-frequency co-enhancement module. The CWD Loss formula is as follows: (10); (11); Where c=1,2...C represents the channel index, i=1,2...H·W represents the spatial position, and X T and X S The feature maps of the Transformer branch and the CNN branch are transformed by the feature activation function. This transforms the feature maps of CNN and Transformer into a unified channel probability distribution, eliminating the difference between the two in terms of feature scale.