Real-time semantic segmentation method and system based on context calibration and enhancement

By adopting context calibration and enhancement methods in real-time semantic segmentation technology, the problems of context mismatch and high computational complexity are solved, and efficient and highly accurate semantic segmentation is achieved, which is suitable for applications in a variety of intelligent fields.

CN119992192APending Publication Date: 2025-05-13SHANDONG NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510079006.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

When existing real-time semantic segmentation technology deals with the problems of context mismatch and high computational complexity, it is difficult to achieve efficient and highly accurate semantic segmentation in real-time scenarios.

Method used

The real-time semantic segmentation method based on context calibration and enhancement is adopted to calibrate the context feature information of the input features through the context calibration module, and the feature alignment module is used to match and fusion the context feature information to generate a semantic segmentation feature map.

Benefits of technology

It realizes efficient scene perception tasks, reduces the computational complexity, improves the accuracy and efficiency of semantic segmentation, and is suitable for fields such as autonomous driving and augmented reality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992192A_ABST
    Figure CN119992192A_ABST
Patent Text Reader

Abstract

The invention discloses a real-time semantic segmentation method and system based on context calibration and enhancement, and belongs to the technical field of image real-time semantic segmentation. Comprises: acquiring a to-be-processed image; processing the to-be-processed image through a trained real-time semantic segmentation network, calibrating context feature information of input features by using a context calibration module, establishing effective matching between pixel points and the context feature information, and generating a semantic segmentation feature map; wherein when the real-time semantic segmentation network is trained, context feature information extraction is carried out by utilizing a training branch, and matching and fusion of the context feature information are carried out through a feature alignment module. The prediction efficiency and the prediction accuracy are well balanced, and the practical application of semantic segmentation is further promoted; the problem that existing real-time semantic segmentation is high in calculation cost and not suitable for practical application is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of real-time semantic segmentation of images, and in particular to a real-time semantic segmentation method and system based on context calibration and enhancement. Background Art

[0002] The statements in this section merely mention background art related to the present invention and do not necessarily constitute prior art.

[0003] Semantic segmentation is a key task in visual scene analysis, which aims to assign a specific category label to each pixel in a given image. In the past few decades, many semantic segmentation methods based on deep convolutional neural networks have been proposed and achieved excellent performance; this success has led to the widespread application of semantic segmentation in intelligent fields such as autonomous driving, medical imaging diagnosis, robotic surgery, and remote sensing imaging.

[0004] Context modeling has been proven to be an effective method to improve semantic segmentation performance and is widely used in real-time semantic segmentation to balance inference speed and accuracy. For example, PSPNet captures context information at different scales through the pyramid pooling module, which improves the accuracy of semantic segmentation. DeepLabv3+ aggregates context information at multiple scales based on the pyramid structure. However, they do not solve the problem of context mismatch, and the high computational complexity hinders their widespread application in real-time scenarios.

[0005] Existing decoder-based context modeling methods often lack the flexibility to adapt to various inputs and ignore the intrinsic changes in context requirements between them, resulting in a mismatch between pixels and their contexts. Although the spatial attention mechanisms (especially self-attention) adopted by models such as DANet and SFNet can effectively capture meaningful features of each pixel while suppressing irrelevant information, their large computational cost makes them unsuitable for real-time applications. Summary of the invention

[0006] In order to address the deficiencies in the prior art, the present invention provides a real-time semantic segmentation method, system, electronic device, computer-readable storage medium and computer program product based on context calibration and enhancement, which has high prediction accuracy and small parameter amount and is easy to apply.

[0007] In a first aspect, the present invention provides a real-time semantic segmentation method based on context calibration and enhancement;

[0008] A real-time semantic segmentation method based on context calibration and enhancement, comprising:

[0009] Get the image to be processed;

[0010] The image to be processed is processed by a trained real-time semantic segmentation network, the context feature information of the input feature is calibrated by a context calibration module, an effective match is established between the pixel points and the context feature information, and a semantic segmentation feature map is generated;

[0011] When training the real-time semantic segmentation network, the training branch is used to extract context feature information, and the feature alignment module is used to match and fuse the context feature information.

[0012] In some embodiments, the real-time semantic segmentation network includes a semantic segmentation branch and a training branch arranged in parallel;

[0013] The semantic segmentation branch includes a plurality of preliminary feature extraction modules, a plurality of semantic gap reduction modules and a context calibration module connected in sequence, and the training branch includes a plurality of Segformer modules and a feature alignment module, wherein the feature alignment module is arranged between the Segformer module and the semantic gap reduction module.

[0014] In some embodiments, the semantic gap reduction module includes an attention unit, a batch normalization layer, and a feed-forward network unit connected in sequence.

[0015] In some embodiments, processing the image to be processed by using a trained real-time semantic segmentation network includes:

[0016] Performing preliminary feature extraction on the extraction to be processed to generate an initial feature map;

[0017] The initial feature map is converted into pixel blocks with learnable kernels of different sizes, and pixel-by-pixel convolution operations are performed. The batch normalized residuals are added to determine the intermediate feature map.

[0018] The multi-scale context feature information of the intermediate feature map is extracted and a pyramid context is formed, and the pixel-context affinity is calculated in combination with the softmax layer to generate a semantic segmentation feature map.

[0019] In some implementations, the calibrating the context feature information of the input feature using the context calibration module to establish an effective match between the pixel point and the context feature information includes:

[0020] The input features are processed through multiple cascaded adaptive pooling layers to generate multi-scale pooling results and tile connections to form a pyramid context;

[0021] The reduced dimensionality features are reconstructed, and the pixel-context affinity is calculated using the softmax layer combined with the pyramid context. The pixel-context affinity is reshaped and de-redundant and added to generate a semantic segmentation feature map.

[0022] In some embodiments, the matching and fusion of context feature information through the feature alignment module is specifically: with the goal of minimizing the alignment loss of semantic content, the output feature maps of the real-time semantic segmentation branch and the training branch in the real-time semantic segmentation network are reshaped and normalized respectively.

[0023] In a second aspect, the present invention provides a real-time semantic segmentation system based on context calibration and enhancement;

[0024] A real-time semantic segmentation system based on context calibration and enhancement, comprising:

[0025] The acquisition module is configured to: acquire the image to be processed;

[0026] The real-time semantic segmentation module is configured to: process the image to be processed by a trained real-time semantic segmentation network, calibrate the context feature information of the input feature by using the context calibration module, establish an effective match between the pixel points and the context feature information, and generate a semantic segmentation feature map;

[0027] When training the real-time semantic segmentation network, the training branch is used to extract context feature information, and the feature alignment module is used to match and fuse the context feature information.

[0028] In a third aspect, the present invention provides an electronic device;

[0029] An electronic device comprises a memory, a processor and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the above-mentioned real-time semantic segmentation method based on context calibration and enhancement.

[0030] In a fourth aspect, the present invention provides a computer-readable storage medium;

[0031] A computer-readable storage medium stores a computer program / instruction thereon, which, when executed by a processor, implements the steps of the above-mentioned real-time semantic segmentation method based on context calibration and enhancement.

[0032] In a fifth aspect, the present invention provides a computer program product;

[0033] A computer program product includes a computer program / instruction, which, when executed by a processor, implements the steps of the above-mentioned real-time semantic segmentation method based on context calibration and enhancement.

[0034] Compared with the prior art, the present invention has the following beneficial effects:

[0035] 1. The technical solution provided by the present invention designs a real-time semantic segmentation network based on context calibration and enhancement. Its backbone network has the advantages of high prediction accuracy, small number of parameters, fast and lightweight, and has achieved a good balance between prediction efficiency and prediction accuracy; it can realize efficient scene perception tasks and further promote its application in many fields such as autonomous driving, augmented reality, and video surveillance. ;

[0036] 2. The technical solution provided by the present invention designs a preliminary feature extraction module using attention units, feedforward network units and batch normalization, and only uses convolution operations to capture the ability of remote context, thereby minimizing the semantic difference between the output features of the semantic segmentation branch and the training branch, and facilitating the semantic segmentation branch to better learn the training branch during training.

[0037] 3. The technical solution provided by the present invention constructs a context calibration module to customize the context for each pixel and capture the most informative context for accurate classification. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] The accompanying drawings in the specification, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.

[0039] Figure 1 A schematic diagram of the network architecture of a real-time semantic segmentation network provided by an embodiment of the present invention;

[0040] Figure 2 A schematic diagram of the network architecture of a preliminary feature extraction module provided in an embodiment of the present invention;

[0041] Figure 3 A schematic diagram of a network architecture of a context calibration module provided in an embodiment of the present invention;

[0042] Figure 4 A schematic diagram of the network architecture of a feature alignment module provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0043] It should be noted that the following detailed descriptions are exemplary and are intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used in the present invention have the same meanings as those commonly understood by those skilled in the art to which the present invention belongs.

[0044] It should be noted that the terms used herein are only for describing specific embodiments, and are not intended to limit exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that the terms "include" and "have" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0045] In the absence of conflict, the embodiments of the present invention and the features of the embodiments may be combined with each other.

[0046] Embodiment 1

[0047] Existing semantic segmentation ignores the intrinsic changes of contextual features between input features and has a high computational cost; therefore, the present invention provides a real-time semantic segmentation method based on context calibration and enhancement, which uses a cascaded pyramid pooling module to efficiently capture nested contexts and aggregates the private context of each pixel based on pixel-context similarity to achieve context feature calibration.

[0048] Next, combine Figure 1-Figure 4 , a real-time semantic segmentation method based on context calibration and enhancement disclosed in this embodiment is described in detail. The real-time semantic segmentation method based on context calibration and enhancement includes:

[0049] S1. Obtain an image to be processed.

[0050] S2. Process the image to be processed through the trained real-time semantic segmentation network to generate a semantic segmentation feature map.

[0051] Furthermore, the real-time semantic segmentation network includes a semantic segmentation branch and a training branch set in parallel. The semantic segmentation branch includes two preliminary feature extraction modules, two semantic gap reduction modules and one context calibration module connected in sequence. The training branch includes a feature alignment module and five Segformer modules. The output of the semantic gap reduction module and the output features of the third and fourth Segformer modules are input into the feature alignment module for processing.

[0052] In this embodiment, the preliminary feature extraction module is a ResNet module; the semantic gap reduction module includes an attention unit, a batch normalization layer, a feedforward network unit and a batch normalization layer connected in sequence, the input feature and the output residual of the first batch normalization layer are added, and the output obtained by the residual addition is then added to the residual of the second batch normalization layer; the feedforward network unit includes a 3×3 regular convolution layer and a 3×3 dilated convolution layer with a hole rate of 3 connected in sequence; the context calibration module includes a 1×1 convolution layer, a pyramid pooling block set in parallel, a reshaping layer and a standard convolution layer, a softmax layer, and a reshaping layer connected in sequence, the output of the first reshaping layer is multiplied by the output of the pyramid pooling block and then input into the softmax layer, the output of the softmax layer is multiplied by the output of the pyramid pooling block and then input into the second reshaping layer, and the output of the second reshaping layer is fused with the output of the second standard convolution layer to generate a semantic segmentation feature map.

[0053] The feature alignment module includes two feature alignment branches with the same structure and arranged in parallel. The feature alignment branch includes two reshaping layers, a Softmax layer and a reshaping layer connected in sequence. The output of the first reshaping layer is multiplied by the output of the second reshaping layer and then input into the Softmax layer. The output of the Softmax layer is multiplied by the output of the first reshaping layer and then input into the third reshaping layer. Subsequently, the input is added to the output of the third reshaping layer, and finally the added results are processed by the loss function respectively.

[0054] In the semantic segmentation branch, the input image first uses the ResNet module to extract preliminary features. Then, the semantic gap between the preliminary features and the output features of the training branch is reduced by the semantic gap reduction module. The context information is then calibrated using the context calibration module to establish an effective match between the pixel points and the pool-based context to enhance its recognition ability.

[0055] Furthermore, during the training process, the input image is input into the semantic segmentation branch and the training branch for processing respectively, and rich contextual information is extracted through a training branch that only trains; finally, through the feature alignment module and the loss function, the deep matching and fusion of the contextual feature information of the semantic segmentation branch and the training branch are achieved, so that the training branch can guide the training of the semantic segmentation branch and improve the training accuracy of the semantic segmentation branch.

[0056] The training branch is composed of multiple Segformer modules, which have high training accuracy but a large number of parameters. Therefore, in this embodiment, during the training process, the semantic gap reduction module and the feature alignment module are used to reduce the gap between the output features of the semantic segmentation branch and the training branch, and the lightweight semantic segmentation branch is trained under the guidance of the training branch, while ensuring the segmentation accuracy and efficiency of the semantic segmentation branch.

[0057] As an implementation mode, S2 specifically includes:

[0058] S201. Process the image to be processed in sequence through two ResNet modules to obtain an initial feature map.

[0059] S201, sequentially processing the initial feature map through two semantic gap reduction modules to generate an intermediate feature map.

[0060] As different types of networks, there are obvious differences in the feature representations extracted by the semantic segmentation branch and the training branch. Directly aligning features between the semantic segmentation branch and the training branch will increase the difficulty of the learning process, resulting in limited performance improvement. Therefore, in this embodiment, a semantic gap reduction module is used to learn to extract high-quality context feature information from the training branch. The semantic gap reduction module can be expressed as:

[0061] f = Norm(x + Attention(x));

[0062] y = Norm(f + FFN(f));

[0063] Among them, Attention represents the attention unit, FFN represents the feedforward network unit, Norm represents batch normalization, and x, f, and y represent input, hidden features, and output, respectively.

[0064] Next, taking the processing of the initial feature map by the first semantic gap reduction module as an example, the specific data processing flow of the semantic gap reduction module is further explained:

[0065] First, the initial feature map is processed by the attention unit, in which convolution is used to convert the initial feature map into pixel blocks with learnable kernels of different sizes, that is, the convolution kernel of the convolution operation in this unit can change with the change of the input features; then it is convolved with the initial feature map pixel by pixel to obtain the attention feature map; it is expressed as:

[0066]

[0067] In the formula, X, K, K T denote the initial feature map, learnable query and key respectively, X∈R n×C×H×W ,K∈R C×N×k×k , C, H, W represent the channel, height and width of the feature map respectively, N represents the number of learnable parameters, k represents the kernel size of the learnable parameters; θ represents group double normalization, Represents a convolution operation.

[0068] Here, the input size n×C×H×W is adjusted to the size of the convolution kernel C×N×k×k through the convolution operation. The convolution kernel changes with the input features, so it is a learnable kernel.

[0069] Then, the attention feature map is batch normalized and added to the residual of the primary feature map. The residual added features enter the feedforward network unit, where the features first pass through a 3×3 regular convolution, and then pass through a 3×3 dilated convolution with a dilation rate of 3; the feedforward network unit can be expressed as:

[0070] Qut FFN =Gelu(Norm(F C×H×W )f 3×3 )df 3×3 ;

[0071] Where Norm represents batch normalization, F C×H×W represents the input feature map, f 3×3 represents 3×3 convolution, df 3×3 represents a 3×3 atrous convolutional layer, and Gelu represents the Gelu activation function.

[0072] Finally, the output of the feedforward network unit is batch normalized and added to the residual of the attention feature map to obtain the output feature.

[0073] S203: calibrate the context feature information of the intermediate feature map using the context calibration module, establish an effective match between the pixel points and the context feature information, and generate a semantic segmentation feature map.

[0074] Context can provide rich information for scene classification and help correct unintentional misclassification. However, previous methods usually assume that context plays an equal role in the classification of each pixel, which inevitably leads to the problem of context mismatch. Therefore, in this embodiment, a context calibration module is constructed to customize and enhance the semantic context of a single pixel.

[0075] Specifically, the data processing flow of the context calibration module is as follows:

[0076] (1) For the intermediate feature map X∈R C×H×W Apply a 1×1 convolutional layer to produce dimensionally reduced features Q∈R c×H×W , minimizing unnecessary calculations.

[0077] (2) Use the pyramid pooling block to obtain multi-scale context feature information. Specifically, five cascaded adaptive pooling layers are used to generate the pooling result, which is then flattened and connected to form a pyramid multi-scale context Z∈R c×M Here, the pyramid pooling block is represented as:

[0078] PPM(X)=Concat(Reshape(AvgPooli(X))),i=1,2,3,4,5;

[0079] In the formula, AvgPool represents the average pooling operation, Reshape represents the reshaping operation, and Concat represents the concatenation operation of the channel dimension.

[0080] At the same time, the dimension reduction features are reconstructed and transposed to R N×c , and use the above context representation to perform matrix multiplication, combined with a softmax layer to calculate pixel-context affinity, the calibrated semantic context feature information E∈R C×N Reshaped back to R C×H×W . The tanh function is applied to refine the context, remove redundancy and emphasize useful information, and the elements are added together to produce the final output semantic segmentation feature map Y∈R C×H×W .

[0081] By default, the pooling layer sizes are set to [1,2,3,4,6] in sequence, processing this context to produce two context representations with different channel counts.

[0082] Exemplarily, the specific formula is as follows:

[0083] X=F C×H×W f 3×3 ;

[0084]

[0085] In the formula, F C×H×W represents the intermediate feature map, f 3×3 represents 3×3 convolution, PPM represents pyramid pooling module, Softmax represents Softmax activation function, That is the obtained pixel-context affinity, Reshape represents the reshaping activation function, and tanh represents the tanh function operation.

[0086] As an implementation method, before S2, it also includes: training a real-time semantic segmentation network; the specific process is as follows:

[0087] Step 1: Select the publicly available ADE20K, COCO-Stuff-10K, and Cityscapes datasets, and then perform preprocessing and data augmentation operations on these images. The images in these datasets are from different scenes in different cities and are accompanied by high-quality pixel-level annotation information. From the annotation information, we filter out categories that meet actual application requirements, and at the same time remove inapplicable categories and set them as ignored categories to construct a training set.

[0088] Here, data augmentation operations include random cropping and scaling, random horizontal flipping, and Gaussian blurring.

[0089] Step 2: Pre-train the real-time semantic segmentation network and fine-tune it on the semantic segmentation dataset.

[0090] Specifically, before fine-tuning the models, they are pre-trained on ImageNet. During the pre-training phase, the MMClassification development code base is used, and the training configuration of swing-transformer is adhered to on the ImageNet-1K dataset.

[0091] Step 3: Train the real-time semantic segmentation network using the training set.

[0092] Here, it should be noted that a simple and effective feature alignment module for the training process is proposed.

[0093] Directly aligning the features output by the semantic segmentation branch and the training branch will damage the network's supervision of the true labels during the training process. At the same time, although the features of the semantic gap reduction module reduce the semantic gap, the feature dimension sizes are still different. The features of both sides can be aligned more effectively through the feature alignment module.

[0094] Therefore, during the training process, first, the feature maps output by the semantic segmentation branch are processed by a 3×3 convolutional layer to generate feature maps with dimensions of C×H×W; then, these feature maps are reshaped into a matrix of size C×N, where N=H×W; next, the transpose of the matrix B is calculated, denoted as B T (shape is N×C) and multiply it by the matrix D (shape is C×N); apply the softmax function to the resulting matrix to get the weights S (shape is N×N).

[0095] Then, the original feature map is reshaped into a matrix of size C×N again. The transposed feature B T Multiply by matrix D to get the softmax weight matrix S. The resulting feature map E (of shape C×N) is multiplied by the transpose of S (of shape N×N), scaled by parameter α. After reshaping the product back to dimension C×H×W, it is element-wise summed with the original input feature map A to produce the final output F (of shape C×H×W). The training branch follows the same process, except that the weight S is of shape C×C.

[0096] This module can not only capture the spatial relationship between different locations in the feature map, but also compress and deform the global information to obtain a broader global view, thereby enriching the contextual information and improving the segmentation accuracy of similar objects. Finally, different features are downsampled or upsampled for projection alignment to achieve deep matching and fusion of contextual feature information.

[0097] Exemplarily, the module can be expressed as:

[0098]

[0099]

[0100] In the formula, X is the input feature map, Softmax is the Softmax activation function, Reshape is the reshape activation function, and Y CNN and Y TF They are the feature outputs from the image segmentation branch and the feature outputs from the training branch, respectively.

[0101] The network can be decomposed into a real-time semantic segmentation branch, a training branch, and a fusion segmentation stage. First, the given image is extracted with preliminary features through the Resnet module in the real-time semantic segmentation branch; then the semantic gap reduction module is used to reduce the semantic gap between the output features of the ResNet module and the output features of the training branch. Subsequently, the context information in the training branch can be better learned, and the context calibration module is used to calibrate the context information, establish an effective match between the pixel points and the pool-based context, and enhance its recognition ability.

[0102] Transformer has begun to emerge in the field of graphics, but its computational consumption is very high. The encoder and decoder are redesigned in the Segformer module, and an efficient network structure with powerful feature extraction capabilities is proposed; at the same time, the given image is also processed by using only the trained SegFormer, which not only ensures effective feature extraction but also avoids excessive computational redundancy. Finally, the feature map after the two branches enters the fusion segmentation stage, which includes a feature alignment module and a loss function. The feature alignment module receives feature inputs from different positions and modes, establishes a model of interdependence between context features, and realizes deep matching and fusion of context feature information.

[0103] At the same time, in order to better synchronize context information, an alignment loss that emphasizes semantic content rather than spatial details is needed. In this embodiment, CWD Loss is used as the alignment loss, which is superior to other loss functions. In short, CWD loss can be described as:

[0104]

[0105] Among them, c = 1, 2, ..., C represents the channel, i = 1, 2, ..., H × W represents the spatial position, and Represent the feature maps of the semantic segmentation branch and the training branch, respectively. The function converts feature activations into channel-wise probability distributions. is a hyperparameter called temperature. We carefully tune the hyperparameter throughout the network until we achieve the best trade-off between accuracy and efficiency.

[0106] For example, for the Cityscape dataset, the AdamW optimizer was used to train all models with an initial learning rate of 0.0004 and a weight decay factor of 0.0125; a polynomial learning rate schedule with a power of 0.9 was used to gradually reduce the learning rate. In addition, various data augmentation techniques such as random cropping, scaling, and horizontal flipping were also combined. For ADE20K, all models were trained with a batch size of 32 for a total of 160,000 iterations. The only difference in the training settings compared to Cityscape is the crop size, which is set to 512×512 and the initial learning rate is 0.0005. The rest of the training configuration remains the same as that of Cityscape. For the training of the COCO dataset, the AdamW optimizer was used with a weight decay value of 0.00006 and an initial learning rate set to 0.01. Our data augmentation strategy includes random cropping to a size of 640×640 and random scaling in the range of 0.5 to 2.0. The rest of the training details are the same as those used in the Cityscape dataset configuration.

[0107] During inference, the model was initially trained using the training and validation sets from Cityscape. The inference speed was evaluated on a standardized platform equipped with an RTX 4090 GPU, PyTorch version 1.11, CUDA12.4, cuDNN 8.9, and Linux Conda environment. In order to evaluate the performance and efficiency of the proposed network (single-scale test), the segmentation accuracy and latency were evaluated using mean cross-linking (mIoU) and frames per second (FPS); the model can achieve 79.8% mIoU and 120.4FPS on the Cityscapes dataset. Compared with the 78.9% mIoU of the state-of-the-art real-time semantic segmentation network LCFNet-slim, there is a 0.9% performance advantage. In addition, the model can achieve 178.1FPS and 42.3% mIoU on the ADE20K dataset, and 35.6% mIoU at a speed of 174.5FPS on the COCO-Stuff-10K dataset.

[0108] Embodiment 2

[0109] This embodiment discloses a real-time semantic segmentation system based on context calibration and enhancement, including:

[0110] The acquisition module is configured to: acquire the image to be processed;

[0111] The real-time semantic segmentation module is configured to: process the image to be processed by a trained real-time semantic segmentation network, calibrate the context feature information of the input feature by using the context calibration module, establish an effective match between the pixel points and the context feature information, and generate a semantic segmentation feature map;

[0112] When training the real-time semantic segmentation network, the training branch is used to extract context feature information, and the feature alignment module is used to match and fuse the context feature information.

[0113] It should be noted that the acquisition module and the real-time semantic segmentation module correspond to the steps in Embodiment 1, and the examples and application scenarios implemented by the modules and the corresponding steps are the same, but are not limited to the contents disclosed in Embodiment 1. It should be noted that the modules as part of the system can be executed in a computer system such as a set of computer executable instructions.

[0114] Embodiment 3

[0115] Embodiment 3 of the present invention provides an electronic device, including a memory and a processor, and computer instructions stored in the memory and executed on the processor. When the computer instructions are executed by the processor, the steps of the above-mentioned real-time semantic segmentation method based on context calibration and enhancement are completed.

[0116] Embodiment 4

[0117] Embodiment 4 of the present invention provides a computer-readable storage medium for storing computer instructions. When the computer instructions are executed by a processor, the steps of the above-mentioned real-time semantic segmentation method based on context calibration and enhancement are completed.

[0118] Embodiment 5

[0119] Embodiment 5 of the present invention provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of the above-mentioned real-time semantic segmentation method based on context calibration and enhancement.

[0120] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1A device that provides the functions specified in a block or multiple blocks.

[0121] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0122] These computer program instructions can also be loaded onto a computer or other programmable data processing device, and a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0123] The description of each embodiment in the above embodiments has different emphases. For parts not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0124] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A real-time semantic segmentation method based on context calibration and enhancement, characterized in that: include: Get the image to be processed; The image to be processed is processed by a trained real-time semantic segmentation network, the context feature information of the input feature is calibrated by a context calibration module, an effective match is established between the pixel points and the context feature information, and a semantic segmentation feature map is generated; When training the real-time semantic segmentation network, the training branch is used to extract context feature information, and the feature alignment module is used to match and fuse the context feature information.

2. The real-time semantic segmentation method based on context calibration and enhancement as claimed in claim 1, characterized in that: The real-time semantic segmentation network includes a semantic segmentation branch and a training branch arranged in parallel; The semantic segmentation branch includes a plurality of preliminary feature extraction modules, a plurality of semantic gap reduction modules and a context calibration module connected in sequence, and the training branch includes a plurality of Segformer modules and a feature alignment module, wherein the feature alignment module is arranged between the Segformer module and the semantic gap reduction module.

3. The real-time semantic segmentation method based on context calibration and enhancement as claimed in claim 2, characterized in that: The semantic gap reduction module includes an attention unit, a batch normalization layer, and a feed-forward network unit connected in sequence.

4. The real-time semantic segmentation method based on context calibration and enhancement as claimed in claim 1, characterized in that: The processing of the image to be processed by the trained real-time semantic segmentation network includes: Performing preliminary feature extraction on the extraction to be processed to generate an initial feature map; The initial feature map is converted into pixel blocks with learnable kernels of different sizes, and pixel-by-pixel convolution operations are performed. The batch normalized residuals are added to determine the intermediate feature map. The multi-scale context feature information of the intermediate feature map is extracted and a pyramid context is formed, and the pixel-context affinity is calculated in combination with the softmax layer to generate a semantic segmentation feature map.

5. The real-time semantic segmentation method based on context calibration and enhancement as claimed in claim 1, characterized in that: The process of calibrating the context feature information of the input feature using the context calibration module to establish an effective match between the pixel point and the context feature information includes: The input features are processed through multiple cascaded adaptive pooling layers to generate multi-scale pooling results and tile connections to form a pyramid context; The reduced dimensionality features are reconstructed, and the pixel-context affinity is calculated using the softmax layer combined with the pyramid context. The pixel-context affinity is reshaped and de-redundant and added to generate a semantic segmentation feature map.

6. The real-time semantic segmentation method based on context calibration and enhancement as claimed in claim 1, characterized in that: The matching and fusion of context feature information through the feature alignment module is specifically as follows: with the goal of minimizing the alignment loss of semantic content, the output feature maps of the real-time semantic segmentation branch and the training branch in the real-time semantic segmentation network are reshaped and normalized respectively.

7. A real-time semantic segmentation system based on contextual calibration and enhancement, characterized in that: include: The acquisition module is configured to: acquire the image to be processed; The real-time semantic segmentation module is configured to: process the image to be processed by a trained real-time semantic segmentation network, calibrate the context feature information of the input feature by using the context calibration module, establish an effective match between the pixel points and the context feature information, and generate a semantic segmentation feature map; When training the real-time semantic segmentation network, the training branch is used to extract context feature information, and the feature alignment module is used to match and fuse the context feature information.

8. An electronic device comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the real-time semantic segmentation method based on context calibration and enhancement as described in any one of claims 1-6.

9. A computer-readable storage medium having a computer program / instruction stored thereon, characterized in that: When the computer program / instruction is executed by a processor, the steps of the real-time semantic segmentation method based on context calibration and enhancement described in any one of claims 1 to 6 are implemented.

10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instruction is executed by a processor, the steps of the real-time semantic segmentation method based on context calibration and enhancement described in any one of claims 1 to 6 are implemented.