Multi-scale real-time semantic segmentation method and system based on three-branch structure
By adopting a multi-scale method of three-branch structure in real-time semantic segmentation, the problems of small receptive fields and insufficient feature extraction are solved, more efficient feature extraction and semantic information fusion are achieved, and the accuracy and efficiency of real-time semantic segmentation are improved.
Patent Information
- Application Number
- CN202510101542.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art has problems such as small receptive fields and insufficient feature extraction in real-time semantic segmentation, which limits the accuracy of real-time segmentation.
A multi-scale real-time semantic segmentation method based on a three-branch structure is adopted, and the semantic branches, detail branches and boundary branches are processed in parallel, and multi-scale semantic information is fused to improve feature extraction capabilities.
It improves the accuracy of semantic segmentation, realizes efficient feature extraction and semantic information fusion in real-time environment, and balances prediction efficiency and accuracy.
Smart Images

Figure CN119992093A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of real-time semantic segmentation of images, and in particular to a multi-scale real-time semantic segmentation method and system based on a three-branch structure. Background Art
[0002] The statements in this section merely mention background art related to the present invention and do not necessarily constitute prior art.
[0003] Image semantic segmentation technology based on convolutional neural networks (CNNs) is widely used in computer vision tasks, such as autonomous driving and robot vision. Convolutional neural networks are one of the key technologies to achieve intelligence, safety and efficiency, and are crucial for real-time environmental perception and decision-making.
[0004] In real-time semantic segmentation, multi-scale learning strategies are widely used to balance inference speed and accuracy. However, these multi-scale methods do not consider the impact of CNN receptive field on feature discriminability. Current semantic segmentation models, such as DeepLabv3, introduce a pyramid pooling module to complete the fusion of multi-scale features. DeepLabv3+ aggregates contextual information at multiple scales based on a pyramid structure; LRDNet combines decomposed convolution and deep convolution to improve the efficiency of feature extraction. They all have problems such as small receptive field and insufficient feature extraction, which limits the accuracy of real-time segmentation.
[0005] There are also some methods that propose to use multi-scale information to improve the network structure, such as the two-branch structure proposed by BiseNet and the use of an asymmetric encoder-decoder structure to optimize the inference speed and segmentation ability, but there is still room for improvement in feature extraction. Although these methods can extract certain multi-scale information, due to the small receptive field of the backbone CNN network and insufficient feature extraction, the accuracy is still limited. Summary of the invention
[0006] In order to address the deficiencies in the prior art, the present invention provides a multi-scale real-time semantic segmentation method, system, electronic device, computer-readable storage medium and computer program product based on a three-branch structure, constructs a multi-scale real-time semantic segmentation network based on a three-branch structure, and improves the accuracy of semantic segmentation.
[0007] In a first aspect, the present invention provides a multi-scale real-time semantic segmentation method based on a three-branch structure;
[0008] A multi-scale real-time semantic segmentation method based on a three-branch structure, comprising:
[0009] Obtain the image to be segmented and perform feature extraction to generate a comprehensive feature representation;
[0010] Processing the semantic branch, detail branch and boundary branch set in parallel with the comprehensive feature representation input, and fusing the processing results to obtain a real-time semantic segmentation image;
[0011] The detail branch fuses the multi-scale semantic information output by the semantic branch based on pixel-attention, and the boundary branch predicts the boundary area based on the multi-scale semantic information output by the semantic branch.
[0012] In some implementations, feature extraction of the image to be segmented is specifically performed by inputting the image to be segmented into a simple inverted residual module, squeezing the receiving field with a single-branch structure and performing point-by-point convolution operations to construct a comprehensive feature representation.
[0013] In some implementations, inputting the comprehensive feature representation into a semantic branch for processing includes:
[0014] The comprehensive feature representation is processed sequentially by a plurality of extended residual modules to generate an intermediate feature representation;
[0015] The intermediate feature representation is processed by dilated convolutions with different dilation rates, and the processing results are interleaved to obtain multi-scale semantic information.
[0016] In some implementations, processing the comprehensive feature representation into a detail branch includes:
[0017] Extracting features from the comprehensive feature representation to obtain detail features; calculating pixel-attention based on the detail features and multi-scale semantic information;
[0018] Guided by pixel-attention, detail features and multi-scale semantic information are fused to generate detail branch feature maps.
[0019] In some implementations, the comprehensive feature representation is fused with the output features of the semantic branch to obtain a boundary branch feature map.
[0020] In some embodiments, the semantic branch includes an extended residual module and a multi-scale semantic clustering pyramid module connected in sequence, and the number of the extended residual modules is multiple.
[0021] In a second aspect, the present invention provides a multi-scale real-time semantic segmentation system based on a three-branch structure;
[0022] A multi-scale real-time semantic segmentation system based on a three-branch structure, comprising:
[0023] The feature extraction module is configured to: obtain the image to be segmented and perform feature extraction to generate a comprehensive feature representation;
[0024] A real-time semantic segmentation module is configured to: process the semantic branch, detail branch and boundary branch set in parallel for the comprehensive feature representation input, and fuse the processing results to obtain a real-time semantic segmentation image;
[0025] The detail branch fuses the multi-scale semantic information output by the semantic branch based on pixel-attention, and the boundary branch predicts the boundary area based on the multi-scale semantic information output by the semantic branch.
[0026] In a third aspect, the present invention provides an electronic device;
[0027] An electronic device comprises a memory, a processor and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the above-mentioned multi-scale real-time semantic segmentation method based on a three-branch structure.
[0028] In a fourth aspect, the present invention provides a computer-readable storage medium;
[0029] A computer-readable storage medium stores a computer program / instruction, which, when executed by a processor, implements the steps of the above-mentioned multi-scale real-time semantic segmentation method based on a three-branch structure.
[0030] In a fifth aspect, the present invention provides a computer program product;
[0031] A computer program product includes a computer program / instruction, which, when executed by a processor, implements the steps of the above-mentioned multi-scale real-time semantic segmentation method based on a three-branch structure.
[0032] Compared with the prior art, the present invention has the following beneficial effects:
[0033] 1. The technical solution provided by the present invention constructs a multi-scale real-time semantic segmentation network based on a three-branch structure, thereby improving the feature extraction capability during real-time semantic segmentation; effectively integrates multi-scale local semantics through semantic branches, so that the network has rich semantic information; selectively integrates information from semantic branches through detail branches, and directly adds features provided by semantic branches point by point through boundary branches to predict boundary areas; a relatively good balance is achieved in terms of prediction efficiency and prediction accuracy, further promoting the practical application of semantic segmentation.
[0034] 2. The technical solution provided by the present invention quickly downsamples and expands the receptive field through a simple inverted residual module to maintain a high feature extraction efficiency, builds a stronger and more comprehensive feature representation through an extended residual module, and further enhances detail information and enriches semantic information through a multi-scale semantic aggregation pyramid module to more efficiently extract multi-scale features. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] The accompanying drawings in the specification, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.
[0036] Figure 1 A schematic diagram of a network architecture of a multi-scale real-time semantic segmentation network based on a three-branch structure provided in an embodiment of the present invention;
[0037] Figure 2 A schematic diagram of a network architecture of a simple inverted residual module provided in an embodiment of the present invention;
[0038] Figure 3 A schematic diagram of a network architecture of an extended residual module provided in an embodiment of the present invention;
[0039] Figure 4 A schematic diagram of the network architecture of a multi-scale semantic clustering pyramid module provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0040] It should be noted that the following detailed descriptions are exemplary and are intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used in the present invention have the same meanings as those commonly understood by those skilled in the art to which the present invention belongs.
[0041] It should be noted that the terms used herein are only for describing specific embodiments, and are not intended to limit exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that the terms "include" and "have" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0042] In the absence of conflict, the embodiments of the present invention and the features of the embodiments may be combined with each other.
[0043] Embodiment 1
[0044] The existing image semantic segmentation is limited by the small receptive field and insufficient feature extraction capability, and the accuracy of real-time segmentation needs to be improved; therefore, the present invention provides a multi-scale real-time semantic segmentation method based on a three-branch structure, constructs a multi-scale real-time semantic segmentation network based on the three-branch structure, and achieves the accuracy of real-time semantic segmentation.
[0045] Next, combine Figure 1-Figure 4 , a multi-scale real-time semantic segmentation method based on a three-branch structure disclosed in this embodiment is described in detail. The multi-scale real-time semantic segmentation method based on a three-branch structure comprises the following steps:
[0046] S1. Obtain the image to be segmented and perform feature extraction to generate a comprehensive feature representation.
[0047] Furthermore, the image to be segmented is processed by a plurality of simple inverse residual modules connected in sequence, and a final output is obtained as a comprehensive feature representation.
[0048] Since the feature map of the image to be segmented is small and contains weak semantic information, less information can be collected. Therefore, in this step, one-step feature extraction is more efficient than two-step feature extraction. In this embodiment, a simple inverted residual module is designed, and a single-branch structure is used to squeeze the receiving field to avoid invalid calculations.
[0049] Specifically, the simple inverted residual module includes a 3x3 convolutional layer, batch normalization, relu activation function and a 1x1 convolutional layer connected in sequence. The input of the 3x3 convolutional layer and the output of the 1x1 convolutional layer are added in the channel dimension; it is expressed as:
[0050] Qut SIR =F C×H×W +f 1×1 ReLu(BN(f 3×3 F C×H×W ));
[0051] In the formula, F C×H×W represents the input feature map. For the first simple inverted residual module, it is the image to be segmented. For the second simple inverted residual module, it is the output feature map of the first simple inverted residual module. 3×3 represents 3×3 convolution, f 1×1 represents 1×1 convolution, BN represents batch normalization, and ReLu represents the relu activation function.
[0052] Based on this, the simple inverted residual module described in this embodiment first expands the number of channels of the input features to three times the original through ordinary 3x3 convolution, batch normalization (BN) layer and ReLU layer. The 3x3 convolution of the initial feature extraction plays an important role in activating and simplifying the regional features; then, point-by-point convolution is used to reduce the channel dimension; finally, it is added to the input feature map to build a stronger and more comprehensive feature representation.
[0053] S2. The comprehensive feature representation is input into the semantic branch, detail branch and boundary branch set in parallel for processing, and the processing results are fused to obtain a real-time semantic segmentation image.
[0054] Furthermore, the semantic branch includes three extended residual modules and one multi-scale semantic aggregation pyramid module connected in sequence, the detail branch includes two pixel-attention guided fusion modules connected in sequence, and the boundary branch includes two point-by-point summation modules connected in sequence; the output of the first extended residual module is respectively input into the first pixel-attention guided fusion module and the first point-by-point summation module, the output of the second extended residual module is respectively input into the second pixel-attention guided fusion module and the second point-by-point summation module, the output of the second pixel-attention guided fusion module, the second point-by-point summation module and the multi-scale semantic aggregation pyramid module are respectively input into the boundary attention guided fusion module, and after being processed by the boundary attention guided fusion module, a real-time semantic segmentation image is output.
[0055] Specifically, the extended residual module includes a 3×3 convolutional layer connected in sequence, a channel separation layer, three hole convolutional layers set in parallel, a splicing layer, a 1×1 convolutional layer and a fusion layer, and the input feature map is input into the 3×3 convolutional layer and the fusion layer respectively; the multi-scale semantic aggregation pyramid module includes a 3×3 convolutional layer connected in sequence, five 3×3 hole convolutional layers set in parallel, a splicing layer and a point-by-point addition layer, and a 1×1 convolutional layer set in parallel with the 3×3 convolutional layer. Point-by-point addition layers are set between the first, second, fourth and fifth 3×3 hole convolutional layers and the splicing layers, and the output of the third 3×3 hole convolutional layer is input into four point-by-point addition layers respectively, and the output of the 1×1 convolutional layer is input into the point-by-point addition layer.
[0056] As an implementation method, S2 includes:
[0057] S201, input the comprehensive feature representation into the semantic branch for processing to obtain multi-scale semantic information.
[0058] Furthermore, first, the overall feature representation is processed in sequence through three extended residual modules, and the intermediate feature representation integrating multi-scale local semantics is output by the third extended residual module; then the intermediate feature representation is input into the multi-scale semantic aggregation pyramid module, and processed by dilated convolutions with different dilation rates, and the processing results are interlaced to generate multi-scale semantic information.
[0059] Among them, the extended residual module first performs a 3×3 convolution on the input feature map, and then divides the convolution result into three concise feature maps with different area forms according to the number of channels. The three concise feature maps are input into three 3×3 dilated convolutions with different dilation rates for processing, and the processing results are spliced according to the channel dimension. The spliced results are processed by 1×1 convolution to obtain the output feature map.
[0060] Specifically, the data processing flow of the extended residual module is expressed as:
[0061] utdWR (X) = F C×H×W +f 1×1 (Concat(d 1 ,d 2 ,…,d n ));
[0062]
[0063] In the formula, F C×H×W represents the input feature map. For the first extended residual module, the input feature map is a comprehensive feature representation. For the second extended residual module, the input feature map is the output of the first extended residual module. For the third extended residual module, the input feature map is the output of the second extended residual module. 3×3 represents 3×3 convolution, f 1×1 represents 1×1 convolution, represents a 3×3 atrous convolutional layer, Concat is a cascade for achieving feature fusion, BN represents batch normalization, and ReLu represents the relu activation function.
[0064] Here, the dilated convolution rates are 1, 3, and 5 respectively.
[0065] In order to make full use of all feature maps of each region size and extract multi-scale information more effectively, in this embodiment, the extended residual module adopts a three-branch structure; at the same time, the extended residual module is used to downsample the feature map of the semantic branch to 1 / 64 to extract more effective information and obtain a larger receptive field; secondly, the appropriate number of modules is adjusted to achieve the best trade-off between accuracy and efficiency.
[0066] The data processing flow of the multi-scale semantic clustering pyramid module is expressed as:
[0067]
[0068] Where n = 2, ..., 5, F C×H×W represents the input feature map, i.e., the output of the third extended residual module; f 3×3 represents 3×3 convolution, f 1×1 represents 1×1 convolution, represents the nth 3×3 deep convolutional layer, Concat is the cascade for achieving feature fusion, BN represents batch normalization, and ReLu represents the relu activation function.
[0069] The features have passed through the extended residual module, which effectively integrates multi-scale local semantics and further enhances the detail information, making the network rich in semantic information. The multi-scale semantic aggregation pyramid module is used to reconstruct multi-scale contextual information, which enhances the feature representation ability of semantic information.
[0070] S202: Input the comprehensive feature representation into the detail branch for processing, combine the multi-scale semantic information, and generate a detail branch feature map.
[0071] Specifically, the overall feature representation and the output of the first extended residual module are processed by the first pixel-attention-guided fusion module to obtain an intermediate detail feature map; the intermediate detail feature map and the output of the second extended residual module are processed by the second pixel-attention-guided fusion module to obtain a detail branch feature map.
[0072] For example, the vectors of corresponding pixels in the feature maps of the detail and semantic branches are represented as and Then the output of the Sigmoid function in the pixel-attention-guided fusion module can be expressed as:
[0073]
[0074] In the formula, f p represents the convolution operation applied to the detail branch,,f i Represents the convolution operation applied to the semantic branch.
[0075] The output of the pixel-attention-guided fusion module is expressed as:
[0076]
[0077] The detail branch selectively fuses the information in the semantic branch using a pixel-attention-guided fusion module, where σ represents the probability that two pixels belong to the same object. Considering the semantic richness and accuracy of the semantic branch, a higher σ value indicates a better understanding of the semantic branch. On the contrary, when σ is low, the The degree of dependence will decrease.
[0078] S203, processing the overall feature representation input boundary branch, fusing the output features of the semantic branch, and obtaining a boundary branch feature map.
[0079] The semantics within each object pixel remain consistent and only diverge at the boundaries between adjacent objects, resulting in non-zero semantic differences at these boundaries. In this embodiment, the boundary branch is attached to the two-branch network structure of high-frequency semantic information. For the boundary branch, a shallower network is used to extract high-frequency features, and the features provided by the semantic branch are directly fused in a point-by-point addition manner to predict the boundary area.
[0080] Specifically, the overall feature representation is directly fused with the output of the first extended residual module through point-by-point addition to obtain an intermediate boundary feature map; the intermediate boundary feature map is directly fused with the output of the second extended residual module through point-by-point addition to obtain a boundary branch feature map.
[0081] S204, the final outputs of the semantic branch, the boundary branch and the detail branch are fused through the boundary attention guided fusion module to obtain a real-time semantic segmentation image.
[0082] Considering the boundary features extracted by the boundary branch, the three branches are integrated; the boundary attention is used to guide the fusion module, which guides the fusion of detail and semantic branches based on the boundary branch; considering that the detail branch preserves spatial details more effectively, its information in the boundary area is prioritized, and the contextual features of the semantic branch are used to supplement other areas to obtain a more accurate feature map.
[0083] For example, and Represent the vectors of corresponding pixels in the feature maps of the semantic branch, detail branch, and boundary branch respectively. The data processing flow of the boundary attention guided fusion module is expressed as:
[0084]
[0085] Where f represents the combination of convolution operation, batch normalization and ReLU activation. When Sigmoid>0.5, the model tends to rely more on complex details, otherwise, contextual information is preferred.
[0086] In summary, this embodiment constructs a multi-scale real-time semantic segmentation network through a simple inverted residual module, a semantic branch, a detail branch, a boundary branch and a boundary attention guided fusion module.
[0087] Furthermore, before executing S1, it also includes: training a multi-scale real-time semantic segmentation network.
[0088] As an implementation method, the specific process of training a multi-scale real-time semantic segmentation network is as follows:
[0089] Step 1: Pre-train the multi-scale real-time semantic segmentation network.
[0090] Specifically, the multi-scale real-time semantic segmentation network is pre-trained on ImageNet using the cityscapes on the CamVid dataset. The training process involves 90 epochs in total, and the initial learning rate is set to 0.1, which is reduced by a factor of 10 at the 30th and 60th epochs. To augment the data, the images are randomly cropped to 224×224 and flipped horizontally.
[0091] Step 2: Build a training set.
[0092] First, we selected the publicly available CamVid and Cityscapes datasets. The images in these datasets are all from driving scenes in different cities and are accompanied by high-quality pixel-level annotation information. We screened out the categories that meet the actual application needs from the annotation information, and at the same time eliminated the inapplicable categories and set them as ignored categories.
[0093] Then, the RGB images are normalized to eliminate the negative impact that abnormal sample data may bring. Secondly, in order to improve the segmentation accuracy and generalization ability of the model, a variety of data enhancement methods are used, including random cropping and scaling, random horizontal flipping, random adjustment of brightness, saturation, contrast, and application of Gaussian blur.
[0094] Step 3: Train the multi-scale real-time semantic segmentation network using the training set.
[0095] Specifically, the polygon-based strategy is used to adjust the learning rate. The number of training epochs, initial learning rate, weight decay, crop size, and batch size for Cityscape and CamVid can be summarized as [484, 1e -2 ,5e -4 ,1024×1024,12] and [200,1e -3 ,5e -4 ,960×720,12]. Subsequently, the pre-trained model is fine-tuned once the learning rate drops to 5e -4 The training process is terminated below to reduce the risk of overfitting.
[0096] Furthermore, a semantic header is placed at the output of the first pixel-attention guided fusion module to generate an additional semantic loss l 0 , so as to better optimize the entire network. Using weighted binary cross entropy loss l 1 Instead of dice loss, it handles the imbalance problem of boundary detection, because for small objects, coarse boundaries are more likely to highlight the boundary area and enhance features. 2 and l 3 represents CE loss; uses boundary-aware cross entropy loss l 3 , and uses the output of the boundary head to synchronize the semantic segmentation and boundary detection tasks, thereby enhancing the function of the attention-guided fusion module; the calculation of BAS-Loss can be expressed as:
[0097]
[0098] Among them, t is the predefined threshold, b i 、s i,c , They are respectively the boundary head output of class c, the segmentation truth value and prediction result of the i-th pixel of class c.
[0099] Therefore, the final loss of the multi-scale real-time semantic segmentation network is:
[0100] Loss = λ 0 l 0 +λ 1 l 1 +λ 2 l 2 +λ 3 l 3
[0101] Based on empirical evidence, the training loss parameter of the network is determined as: 0 =0.4,λ 1 =20,λ 2 =1,λ 3 =1 and t=0.8.
[0102] Step 4: Verify.
[0103] During the validation process, the model was initially validated using a validation set from cityscapes. The inference speed was evaluated on a standardized platform equipped with an RTX3080 GPU, PyTorch 1.8 version, CUDA 11.2, cuDNN 8.0, and Windows Conda environment. In order to evaluate the performance and efficiency of the proposed network (single-scale test), the average cross-linking (mIoU) and frames per second (FPS) were used to evaluate the segmentation accuracy and latency.
[0104] The model can achieve 80.5% mIoU and 50.4FPS on the Cityscapes dataset, which is 1.5% better than the 78.0% mIoU of the classic real-time semantic segmentation network SegNext-T-Seg75. In addition, the model can achieve 86.1% mIoU and 165.2FPS on the CamVid dataset, which is 0.6% more accurate and 5.3fps faster than the most advanced model LCFNet.
[0105] Embodiment 2
[0106] This embodiment discloses a multi-scale real-time semantic segmentation system based on a three-branch structure, including:
[0107] The feature extraction module is configured to: obtain the image to be segmented and perform feature extraction to generate a comprehensive feature representation;
[0108] A real-time semantic segmentation module is configured to: process the semantic branch, detail branch and boundary branch set in parallel for the comprehensive feature representation input, and fuse the processing results to obtain a real-time semantic segmentation image;
[0109] The detail branch fuses the multi-scale semantic information output by the semantic branch based on pixel-attention, and the boundary branch predicts the boundary area based on the multi-scale semantic information output by the semantic branch.
[0110] It should be noted that the feature extraction module and the real-time semantic segmentation module correspond to the steps in Embodiment 1, and the examples and application scenarios implemented by the modules and the corresponding steps are the same, but are not limited to the contents disclosed in Embodiment 1. It should be noted that the modules as part of the system can be executed in a computer system such as a set of computer executable instructions.
[0111] Embodiment 3
[0112] Embodiment 3 of the present invention provides an electronic device, including a memory and a processor, and computer instructions stored in the memory and executed on the processor. When the computer instructions are executed by the processor, the steps of the above-mentioned multi-scale real-time semantic segmentation method based on the three-branch structure are completed.
[0113] Embodiment 4
[0114] Embodiment 4 of the present invention provides a computer-readable storage medium for storing computer instructions. When the computer instructions are executed by a processor, the steps of the above-mentioned multi-scale real-time semantic segmentation method based on a three-branch structure are completed.
[0115] Embodiment 5
[0116] Embodiment 5 of the present invention provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of the above-mentioned multi-scale real-time semantic segmentation method based on a three-branch structure.
[0117] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0118] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0119] These computer program instructions can also be loaded onto a computer or other programmable data processing device, and a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0120] The description of each embodiment in the above embodiments has different emphases. For parts not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0121] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A multi-scale real-time semantic segmentation method based on a three-branch structure, characterized in that: include: Obtain the image to be segmented and perform feature extraction to generate a comprehensive feature representation; Processing the semantic branch, detail branch and boundary branch set in parallel with the comprehensive feature representation input, and fusing the processing results to obtain a real-time semantic segmentation image; The detail branch fuses the multi-scale semantic information output by the semantic branch based on pixel-attention, and the boundary branch predicts the boundary area based on the multi-scale semantic information output by the semantic branch.
2. The multi-scale real-time semantic segmentation method based on the three-branch structure according to claim 1, characterized in that: The feature extraction of the image to be segmented is specifically as follows: the image to be segmented is input into a simple inverted residual module, the receiving field is squeezed with a single branch structure, and a point-by-point convolution operation is performed to construct a comprehensive feature representation.
3. The multi-scale real-time semantic segmentation method based on the three-branch structure according to claim 1, characterized in that: Inputting the comprehensive feature representation into the semantic branch for processing includes: The comprehensive feature representation is processed sequentially by a plurality of extended residual modules to generate an intermediate feature representation; The intermediate feature representation is processed by dilated convolutions with different dilation rates, and the processing results are interleaved to obtain multi-scale semantic information.
4. The multi-scale real-time semantic segmentation method based on the three-branch structure according to claim 1, characterized in that: Processing the comprehensive feature representation input into the detail branch includes: Extracting features from the comprehensive feature representation to obtain detail features; calculating pixel-attention based on the detail features and multi-scale semantic information; Guided by pixel-attention, detail features and multi-scale semantic information are fused to generate detail branch feature maps.
5. The multi-scale real-time semantic segmentation method based on a three-branch structure as claimed in claim 1, characterized in that: The processing of the comprehensive feature representation input to the boundary branch is specifically as follows: the comprehensive feature representation is fused with the output features of the semantic branch to obtain a boundary branch feature map.
6. The multi-scale real-time semantic segmentation method based on a three-branch structure as claimed in claim 1, characterized in that: The semantic branch includes an extended residual module and a multi-scale semantic clustering pyramid module connected in sequence, and the number of the extended residual modules is multiple.
7. A multi-scale real-time semantic segmentation system based on a three-branch structure, characterized by: include: The feature extraction module is configured to: obtain the image to be segmented and perform feature extraction to generate a comprehensive feature representation; A real-time semantic segmentation module is configured to: process the semantic branch, detail branch and boundary branch set in parallel for the comprehensive feature representation input, and fuse the processing results to obtain a real-time semantic segmentation image; The detail branch fuses the multi-scale semantic information output by the semantic branch based on pixel-attention, and the boundary branch predicts the boundary area based on the multi-scale semantic information output by the semantic branch.
8. An electronic device comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the multi-scale real-time semantic segmentation method based on a three-branch structure according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program / instruction stored thereon, characterized in that: When the computer program / instruction is executed by a processor, the steps of the multi-scale real-time semantic segmentation method based on a three-branch structure described in any one of claims 1 to 6 are implemented.
10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instruction is executed by a processor, the steps of the multi-scale real-time semantic segmentation method based on a three-branch structure described in any one of claims 1 to 6 are implemented.