Image semantic segmentation method, device and medium based on multi-scale feature fusion

By introducing a multi-scale feature fusion module and a spatial detail enhancement module in the image semantic segmentation network, the problem of insufficient utilization of multi-scale features in the prior art is solved, and the accuracy and accuracy of image segmentation are improved.

CN120182612BActive Publication Date: 2025-08-12EAST CHINA JIAOTONG UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510669016.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-08-12
Estimated Expiration
2045-05-23

AI Technical Summary

Technical Problem

The prior art image segmentation method lacks sufficient semantic information in complex backgrounds and cannot effectively process multi-scale features, resulting in low segmentation accuracy.

Method used

A multi-scale feature fusion module is introduced in the image semantic segmentation network. Through the improved SegFormer network, the different scale features are fused at the bridge between the encoder and the decoder, combined with the spatial detail enhancement module, the multi-scale feature map is extracted and fused to output more accurate segmentation results.

Benefits of technology

The accuracy of image segmentation is improved, especially in complex contexts, and the segmentation accuracy of the same type of target when the scale changes greatly in close and long-range areas is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182612B_ABST
    Figure CN120182612B_ABST
Patent Text Reader

Abstract

The present application discloses a method, device and medium for image semantic segmentation based on multi-scale feature fusion, which relates to the field of image processing technology. The method includes obtaining an image to be segmented; inputting the image to be segmented into a trained image semantic segmentation network to obtain an image semantic segmentation result; the image semantic segmentation network adopts an improved SegFormer network, and introduces a multi-scale feature fusion module at the bridge between the encoder and the decoder; the multi-scale feature fusion module is used to fuse different scale features according to the semantic information output by the encoder, and output multiple fused feature maps of different scales; the decoder is used to perform feature fusion on multiple fused feature maps of different scales to obtain the image semantic segmentation result. The present application introduces a multi-scale feature fusion module into the image semantic segmentation network, which can help the segmentation network better understand and process image information at different scales, thereby improving the accuracy of the segmentation result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular to a method, device and medium for image semantic segmentation based on multi-scale feature fusion. Background Art

[0002] Image segmentation is a crucial component of many visual understanding systems. It involves dividing an image (or video frame) into multiple segments or objects. Image segmentation plays a central role in a wide range of applications, including medical image analysis, autonomous vehicles, video surveillance, and augmented reality.

[0003] Semantic segmentation, a core task in computer vision, aims to assign a semantic label to each pixel in an image, thereby achieving pixel-level classification. This task has broad applications in diverse fields, such as autonomous driving, medical image analysis, and remote sensing image processing. Traditional image segmentation methods focus on simple pixel-based region delineation, often relying on techniques such as edge detection, region growing, or threshold segmentation. While these methods are effective in simple scenes, they perform less well in complex backgrounds and lack sufficient semantic information, making them incapable of meeting the sophisticated requirements of practical applications. With the widespread application of deep learning in computer vision, semantic segmentation technology has garnered unprecedented attention and development. In recent years, convolutional neural networks (CNNs) and their variants have become mainstream techniques in semantic segmentation and have achieved remarkable progress. However, as the complexity of the task increases, effectively extracting and utilizing multi-scale features in images remains a key challenge in improving deep learning semantic segmentation performance. Summary of the Invention

[0004] The purpose of this application is to provide an image semantic segmentation method, device and medium based on multi-scale feature fusion, which can improve the accuracy of image semantic segmentation.

[0005] To achieve the above objectives, this application provides the following solutions:

[0006] In a first aspect, the present application provides an image semantic segmentation method based on multi-scale feature fusion, comprising:

[0007] Obtain the image to be segmented;

[0008] The image to be segmented is input into a trained image semantic segmentation network to obtain an image semantic segmentation result; the image semantic segmentation network adopts an improved SegFormer network; in the improved SegFormer network, a multi-scale feature fusion module is introduced at the bridge of the encoder and the decoder; the multi-scale feature fusion module is used to fuse features of different scales according to the semantic information output by the encoder, and output multiple fused feature maps of different scales; the decoder is used to perform feature fusion on multiple fused feature maps of different scales to obtain the image semantic segmentation result.

[0009] In a second aspect, the present application provides an image semantic segmentation device based on multi-scale feature fusion, comprising:

[0010] An image acquisition module, used to acquire the image to be segmented;

[0011] An image semantic segmentation module is used to input the image to be segmented into a trained image semantic segmentation network to obtain an image semantic segmentation result; the image semantic segmentation network adopts an improved SegFormer network; in the improved SegFormer network, a multi-scale feature fusion module is introduced at the bridge between the encoder and the decoder; the multi-scale feature fusion module is used to fuse features of different scales according to the semantic information output by the encoder, and output multiple fused feature maps of different scales; the decoder is used to perform feature fusion on multiple fused feature maps of different scales to obtain the image semantic segmentation result.

[0012] In a third aspect, the present application provides a computer device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-mentioned image semantic segmentation method based on multi-scale feature fusion.

[0013] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-mentioned image semantic segmentation method based on multi-scale feature fusion.

[0014] According to the specific embodiments provided in this application, this application discloses the following technical effects:

[0015] The present application provides an image semantic segmentation method, device and medium based on multi-scale feature fusion, which obtains an image to be segmented; inputs the image to be segmented into a trained image semantic segmentation network to obtain an image semantic segmentation result; the image semantic segmentation network adopts an improved SegFormer network; in the improved SegFormer network, a multi-scale feature fusion module is introduced at the bridge between the encoder and the decoder; the multi-scale feature fusion module is used to fuse different scale features according to the semantic information output by the encoder, and output multiple fused feature maps of different scales; the decoder is used to perform feature fusion on multiple fused feature maps of different scales to obtain the image semantic segmentation result. The present application introduces a multi-scale feature fusion module into the image semantic segmentation network, which can help the segmentation network better understand and process image information at different scales, thereby improving the accuracy of the segmentation results. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0017] Figure 1 This is an application environment diagram of an image semantic segmentation method based on multi-scale feature fusion in one embodiment of the present application;

[0018] Figure 2 A flowchart of an image semantic segmentation method based on multi-scale feature fusion provided in one embodiment of the present application;

[0019] Figure 3 A schematic diagram of the structure of an image semantic segmentation network provided in one embodiment of the present application;

[0020] Figure 4 A schematic diagram of the structure of a multi-scale feature fusion module provided in one embodiment of the present application;

[0021] Figure 5 A schematic diagram of the structure of a spatial detail enhancement module provided in one embodiment of the present application;

[0022] Figure 6 A schematic diagram of the structure of a decoder provided in one embodiment of the present application;

[0023] Figure 7 A schematic diagram of the visualization results of segmentation of the same data set using different methods provided in an embodiment of the present application;

[0024] Figure 8A schematic diagram of the functional modules of an image semantic segmentation device based on multi-scale feature fusion provided in one embodiment of the present application;

[0025] Figure 9 A schematic diagram of the structure of a computer device provided in one embodiment of the present application. DETAILED DESCRIPTION

[0026] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0027] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0028] The image semantic segmentation method based on multi-scale feature fusion provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown, the terminal communicates with the server via a network. The data storage system can store data that the server needs to process. The data storage system can be set up separately, integrated on the server, or placed on the cloud or other servers. The terminal can send the image to be segmented to the server. After the server receives the image to be segmented, the server inputs the image to be segmented into a trained image semantic segmentation network to obtain the image semantic segmentation result. The image semantic segmentation network uses an improved SegFormer network. In the improved SegFormer network, a multi-scale feature fusion module is introduced at the bridge between the encoder and decoder. The multi-scale feature fusion module is used to fuse features of different scales based on the semantic information output by the encoder, outputting multiple fused feature maps of different scales. The decoder is used to perform feature fusion on the multiple fused feature maps of different scales to obtain the image semantic segmentation result. The server can feed back the obtained video tags for the video to the terminal. In addition, in some embodiments, the image semantic segmentation method based on multi-scale feature fusion can also be implemented independently by the server or the terminal. For example, the terminal can directly perform image semantic segmentation based on multi-scale feature fusion on the image to be segmented, or the server can obtain the image to be segmented from the data storage system and perform image semantic segmentation based on multi-scale feature fusion.

[0029] The terminals may be, but are not limited to, various desktop computers, laptops, smart phones, tablet computers, IoT devices, and portable wearable devices. The server may be implemented as an independent server or a server cluster consisting of multiple servers, or as a cloud server.

[0030] In an exemplary embodiment, Figure 2 As shown, a method for image semantic segmentation based on multi-scale feature fusion is provided. The method is executed by a computer device, specifically a computer device such as a terminal or a server, or a terminal and a server. In the embodiment of the present application, the method is applied to Figure 1 The server in is used as an example for description, including the following steps 101 to 102.

[0031] Step 101: Obtain an image to be segmented.

[0032] Step 102: input the image to be segmented into the trained image semantic segmentation network to obtain the image semantic segmentation result; the image semantic segmentation network adopts the improved SegFormer network; Figure 3 As shown in the figure, in the improved SegFormer network, a multi-scale feature fusion module is introduced at the bridge of the encoder and the decoder; the multi-scale feature fusion module is used to fuse features of different scales according to the semantic information output by the encoder, and output multiple fused feature maps of different scales; the decoder is used to perform feature fusion on multiple fused feature maps of different scales to obtain the image semantic segmentation result.

[0033] By implementing the above steps 101 to 102, the present application introduces a multi-scale feature fusion module into the image semantic segmentation network. Multi-scale feature fusion, as an effective strategy, can help the image semantic segmentation network better understand and process image information at different scales, thereby improving the accuracy of the segmentation results and improving the defect of the existing technology that the scale of the same type of target in the near and distant views varies greatly, resulting in low segmentation accuracy.

[0034] In another exemplary embodiment of the present application, the multi-scale feature fusion module includes fusing features of different scales through convolution group downsampling and DySample upsampling, and outputting multiple feature maps of different scales in combination with the pyramid segmentation attention mechanism. Figure 4 As shown, the multi-scale feature fusion module includes: N multi-scale feature fusion groups; N is the number of encoding layers of the encoder.

[0035] The nth multi-scale feature fusion group includes the first convolution unit (corresponding to Figure 4 Conv1 in), n-1 downsampling units (corresponding to Figure 4Conv2 in), Nn upsampling units (corresponding to Figure 4 DySample and Conv3 in ), the first addition unit, the pyramid segmentation attention mechanism unit (corresponding to Figure 4 PSA in); n=1, 2, ..., N.

[0036] The input of the first convolution unit of the nth multi-scale feature fusion group is the output of the nth encoding layer.

[0037] The inputs of the n-1 downsampling units of the n-th multi-scale feature fusion group are the outputs of the 1st to n-1th encoding layers respectively.

[0038] The inputs of the Nn upsampling units of the nth multi-scale feature fusion group are the outputs of the n+1th coding layer to the Nth coding layer respectively.

[0039] The input of the first addition unit of the nth multi-scale feature fusion group is the output of the first convolution unit, n-1 downsampling units, and Nn upsampling units of the nth multi-scale feature fusion group.

[0040] The input of the pyramid segmentation attention mechanism unit of the nth multi-scale feature fusion group is the output of the first addition unit of the nth multi-scale feature fusion group.

[0041] The outputs of the pyramid segmentation attention mechanism units of the nth multi-scale feature fusion group are respectively the inputs of each decoding layer of the decoder.

[0042] Figure 4 An example of a multi-scale feature fusion module is shown, including four multi-scale feature fusion groups. Figure 4 The expression of the multi-scale feature fusion module is shown in the following formula:

[0043]

[0044] in, Represents the feature map output by the encoder at the nth stage (nth encoding layer), n=1, 2, 3, 4; Represents the fused feature map output by the n-th multi-scale feature fusion group; Conv 3×3 It is represented as a convolution with a kernel size of 3×3 and a stride of 1; PSA represents the pyramid segmentation attention mechanism. The process of upsampling operation at each magnification can be expressed as follows:

[0045]

[0046] in, Indicates an upsampling operation with a magnification of 2; Indicates an upsampling operation with a magnification of 4; Indicates an upsampling operation with a magnification of 8; It represents a DySample upsampling operation with a magnification of 2; It represents a DySample upsampling operation with a magnification of 4; Conv 3×3 It is represented as a convolution with a kernel size of 3×3 and a stride of 1. The process of downsampling operations at each reduction ratio can be expressed as follows:

[0047]

[0048] in, Indicates a downsampling operation with a reduction ratio of 2; Indicates a downsampling operation with a reduction ratio of 4; Indicates a downsampling operation with a reduction ratio of 8; Conv 3×3 It is represented as a convolution with a kernel size of 3×3 and a stride of 2; Conv 5×5 It is represented as a convolution with a kernel size of 5×5 and a stride of 4.

[0049] In another exemplary embodiment of the present application, semantic segmentation requires accurate boundary demarcation of each category. Shallow features have more detailed information but lack semantic information, resulting in inaccurate segmentation results. Deep features are rich in semantic information, but due to downsampling operations, the spatial resolution is reduced and the edges are blurred. Therefore, it is necessary to retain more details while maintaining semantic accuracy, so that the segmentation edges are more accurate. Based on this, Figure 3 As shown, the improved SegFormer network also includes a spatial detail enhancement branch, which extracts and outputs spatial detail features through an independent branch structure, including a spatial detail enhancement module (corresponding to Figure 3 The SDEM module in the ), a multiplication unit and a second addition unit; the spatial detail enhancement module includes a sequentially connected convolution combination and a spatial attention mechanism unit (for Figure 5 SAM in ); the convolution combination includes multiple second convolution units connected in sequence (for Figure 5 Conv7×7, Conv3×3, Conv1×1 in .

[0050] The input of the first second convolution unit is the image to be segmented; the output of the spatial attention mechanism unit is connected to the input of the multiplication unit; the input of the multiplication unit is also connected to the output of the last decoding layer of the decoder; the input of the second addition unit is connected to the output of the multiplication unit and the output of the last decoding layer of the decoder; the output of the second addition unit is the semantic segmentation result of the image.

[0051] The Spatial Detail Enhancement Module (SDEM) consists of four convolutional units and a spatial attention mechanism. Each convolutional unit is followed by a batch normalization layer and a ReLU activation function. The first two convolutional units (Conv7×7) have a kernel size of 7, the third (Conv3×3) has a kernel size of 3, and the fourth (Conv1×1) has a kernel size of 1. Furthermore, to preserve the relationships between pixels in the image, convolutions are configured to avoid downsampling and maintain a constant resolution. The working process of the Spatial Detail Enhancement Module (SDEM) can be expressed as follows:

[0052]

[0053] Where, F SDEM represents the output of the spatial detail enhancement module, SAM represents the spatial attention mechanism, Input Represents the input of the spatial detail enhancement module, that is, the image to be segmented.

[0054] After the spatial detail enhancement module, multiplication unit and second addition unit are introduced into the improved SegFormer network, the decoder outputs a fused semantic feature map, and the output of the second addition unit is the image semantic segmentation result.

[0055] The output of the SDEM module is multiplied pixel by pixel with the decoder fused features, and then a residual connection is performed. The process can be expressed as follows:

[0056]

[0057] Among them, F is the image semantic segmentation result; It represents the features after the decoder is fused layer by layer, that is, the fused semantic feature map; residual processing is used to fuse the boundary feature information (output of the SDEM module) with the feature semantic information (output of the decoder). The residual connection is to prevent the output of the SDEM module from having too much influence on the fused features output by the decoder, thereby affecting the performance of the segmentation model.

[0058] In another exemplary embodiment of the present application, Figure 3 As shown in Figure 1, the encoder (based on SegFormer) consists of multiple sequentially connected encoding layers, each of which uses a Transformer block. Each Transformer block contains a hybrid feedforward network and an efficient self-attention mechanism layer.

[0059] In another exemplary embodiment of the present application, Figure 6As shown in the figure, the decoder includes multiple decoding layers connected in sequence; a splicing layer is provided between two adjacent decoding layers; the input of each splicing layer also includes multiple fusion feature maps output by the multi-scale feature fusion module; the decoding layer uses a multi-layer perceptron (MLP). The first-stage jump connection of the encoder corresponds to the last stage of the decoder, and the second-stage jump connection corresponds to the penultimate stage. The encoder and each stage of the multi-scale feature fusion module have a one-to-one correspondence, so the output of the first stage of the multi-scale feature fusion module corresponds to the last stage of the decoder. Therefore, as shown in the figure, Figure 3 and Figure 6 As shown, the calculation process of the decoder can be expressed as follows:

[0060]

[0061] in, Represents the fused feature map output by the n-th multi-scale feature fusion group; Represents a splicing operation.

[0062] In another exemplary embodiment of the present application, the image to be segmented is input into a trained image semantic segmentation network to obtain an image semantic segmentation result, specifically including:

[0063] (1) The image to be segmented is input into an encoder, and each encoding layer of the encoder extracts semantic information of the image.

[0064] (2) The semantic information extracted from each coding layer is input into each multi-scale feature fusion group of the multi-scale feature fusion module to obtain multiple fusion feature maps of different scales.

[0065] (3) Input the fused feature maps of different scales into the decoder to obtain the fused semantic feature map.

[0066] (4) Inputting the image to be segmented into the spatial detail enhancement module to extract the spatial detail feature information of the image.

[0067] (5) Obtain the image semantic segmentation result based on the fusion of semantic feature map and spatial detail feature information.

[0068] In another exemplary embodiment of the present application, the image to be segmented is input into a trained image semantic segmentation network to obtain an image semantic segmentation result, specifically including:

[0069] (1) Obtaining a training set; the training set includes a number of image samples to be segmented and corresponding image semantic segmentation result samples.

[0070] The training set can be selected from the public dataset Cityscapes, which is divided into 34 categories of semantic segmentation targets. The model is trained using 19 of these categories. It contains 5,000 street view images with a resolution of 1024×2048, including 2,975 training sets, 500 validation sets, and 1,525 test sets. Pixel annotations are provided in the training set, validation set, and test set.

[0071] During the training process, the dataset is randomly flipped in the horizontal and vertical directions, and the images are randomly cropped or padded to avoid overfitting, so that the input image resolution is 1024×1024.

[0072] During training, different network hyperparameters were selected to achieve better performance. The Adam optimizer was used to optimize the network during training, with an initial learning rate of 6e-5. A warmup strategy was also used, gradually increasing the learning rate from the initial warmup value to the set learning rate over the first 1500 training iterations. The Poly strategy was then used to dynamically adjust the learning rate.

[0073] (2) Input the image samples to be segmented in the training set into the initial image semantic segmentation network to obtain the image semantic segmentation prediction samples.

[0074] (3) Calculate the loss error based on the image semantic segmentation prediction samples and the corresponding image semantic segmentation result samples.

[0075] (4) Adjust the network parameters of the initial image semantic segmentation network according to the loss error.

[0076] (5) Until the loss error converges or the maximum number of iterations is reached, the trained image semantic segmentation network is obtained.

[0077] (6) Inputting the image to be segmented into the trained image semantic segmentation network to obtain the image semantic segmentation result.

[0078] In another exemplary embodiment of the present application, after obtaining the trained image semantic segmentation network, it is further tested using a test set, and the model weight file with the best segmentation effect is saved.

[0079] During the evaluation phase, this application uses multiple indicators to evaluate the segmentation effect. The intersection over union (IoU), also known as the Jaccard index, is one of the most commonly used indicators in semantic segmentation. It is defined as the intersection area between the predicted segmentation map and the ground truth divided by the union area between the predicted segmentation map and the label. The mean intersection over union (mIoU) is the average of the intersection-to-union ratios between the predicted segmentation areas and the true segmentation areas of each category. It is a standard performance measurement indicator in the field of image semantic segmentation. The average accuracy (mAcc) is the average of the accuracy of the model in each category. The specific formulas for the above evaluation indicators are given below:

[0080]

[0081]

[0082]

[0083] Where A and B are the predicted segmentation areas and labels, C is the number of classification labels, C+1 is the total number of image label categories, TP j It is j The number of pixels whose class was correctly predicted, FN j Is the other category predicted as category j The number of pixels.

[0084] This application introduces a multi-scale feature fusion module and a spatial detail enhancement module. On the one hand, the multi-scale feature fusion module uses a top-down feature fusion path in conjunction with pyramid segmentation attention to extract feature information at different scales, resulting in a feature map (a fused feature map of different scales). On the other hand, the spatial detail enhancement module weights the image to be segmented by adding boundary-weighted information to the regions near the image boundary to obtain boundary feature information. Finally, a decoder is used to fuse the fused feature maps at different scales to produce a fused semantic feature map. Residual processing is then used to fuse the boundary feature information with the fused semantic feature map, resulting in a more accurate segmentation result.

[0085] Table 1 below shows a comparison of the segmentation effects of the present invention and a variety of existing methods on the same dataset. The comparison results show that the network improved by the present invention has achieved better results in terms of the evaluation indicators mIoU and mAcc compared to mainstream semantic segmentation algorithms such as UperNet, Swin-B, and DeepLab-V3+. Compared with the suboptimal algorithm DeepLab-V3+ in terms of the evaluation indicator mIoU, the mIoU of the present invention has increased by 0.13%, and the mAc of the present invention has increased by 1.32%; compared with the suboptimal algorithm ACFNet in terms of the evaluation indicator mAc, the mIoU of the present invention has increased by 0.75%, and the mAc of the present invention has increased by 1.01%, and the number of parameters of the network of the present invention is only 45.8% and 16.1% of the two. This shows that the improved network of the present invention performs well in balancing the performance of the model and the amount of floating-point operations. When applied to the segmentation of urban street environments, it can more accurately segment urban street environments.

[0086]

[0087] Figure 7 The segmentation results of the method proposed in this application are compared with those of the above-mentioned existing methods. In order to more intuitively demonstrate the segmentation effect of the method proposed in this application, Figure 7 The visualization results of segmentation of different methods on the same dataset are given.

[0088] The present application also provides an application scenario, which applies the above-mentioned image semantic segmentation method based on multi-scale feature fusion. Specifically: the image semantic segmentation method based on multi-scale feature fusion provided in this embodiment can be applied in an autonomous driving scenario. The scenario includes an image acquisition link, an image processing link, and an autonomous driving control link; the image acquisition link is used to acquire obstacle images of the driving lane; the image processing link is used to perform semantic segmentation on the obstacle image to extract obstacle information; the autonomous driving control link is used to automatically control the vehicle driving according to the extracted obstacle information. The image semantic segmentation method based on multi-scale feature fusion provided in this embodiment belongs to the image processing link.

[0089] Based on the same inventive concept, an embodiment of the present application also provides an image semantic segmentation device based on multi-scale feature fusion for implementing the above-mentioned image semantic segmentation method based on multi-scale feature fusion. The implementation solution provided by the device is similar to the implementation solution described in the above-mentioned method. Therefore, the specific limitations in one or more embodiments of the image semantic segmentation device based on multi-scale feature fusion provided below can be found in the above-mentioned limitations on the image semantic segmentation method based on multi-scale feature fusion, and will not be repeated here.

[0090] In an exemplary embodiment, Figure 8As shown, an image semantic segmentation device based on multi-scale feature fusion is provided, comprising:

[0091] The image acquisition module M1 is used to acquire the image to be segmented.

[0092] The image semantic segmentation module M2 is used to input the image to be segmented into the trained image semantic segmentation network to obtain the image semantic segmentation result; the image semantic segmentation network adopts an improved SegFormer network; in the improved SegFormer network, a multi-scale feature fusion module is introduced at the bridge between the encoder and the decoder; the multi-scale feature fusion module is used to fuse different scale features according to the semantic information output by the encoder, and output multiple fused feature maps of different scales; the decoder is used to perform feature fusion on multiple fused feature maps of different scales to obtain the image semantic segmentation result.

[0093] In an exemplary embodiment, a computer device is provided. The computer device may be a server or a terminal. The internal structure diagram thereof may be as follows: Figure 9 As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O) and a communication interface. The processor, memory and input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store image semantic segmentation data. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a method for image semantic segmentation based on multi-scale feature fusion is implemented.

[0094] Those skilled in the art will understand that Figure 9 The structure shown in the figure is merely a block diagram of a portion of the structure related to the present application solution and does not constitute a limitation on the computer device to which the present application solution is applied. A specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement. In an exemplary embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the above-mentioned method embodiments are implemented.

[0095] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0096] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0097] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned embodiments. In particular, any reference to memory, database, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM may be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).

[0098] The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processors involved in the various embodiments provided herein may include, but are not limited to, general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic units, data processing logic units based on quantum computing, and the like.

[0099] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0100] This document uses specific examples to illustrate the principles and implementation methods of this application. The description of the above examples is only intended to help understand the method and core concept of this application. At the same time, for those skilled in the art, based on the concept of this application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.

Claims

1. A method for image semantic segmentation based on multi-scale feature fusion, characterized in that: include: Obtain the image to be segmented; Inputting the image to be segmented into a trained image semantic segmentation network to obtain an image semantic segmentation result; the image semantic segmentation network adopts an improved SegForme network; in the improved SegForme network, a multi-scale feature fusion module is introduced at the bridge between the encoder and the decoder; The multi-scale feature fusion module is used to fuse features of different scales according to the semantic information output by the encoder and output multiple fused feature maps of different scales; The decoder is used to perform feature fusion on multiple fusion feature maps of different scales to obtain the image semantic segmentation result; The multi-scale feature fusion module includes: N multi-scale feature fusion groups; N is the number of encoding layers of the encoder; The nth multi-scale feature fusion group includes the first convolution unit, n-1 downsampling units, Nn upsampling units, the first addition unit, and the pyramid segmentation attention mechanism unit; n=1, 2, ..., N; The input of the first convolution unit of the nth multi-scale feature fusion group is the output of the nth encoding layer; The inputs of the n-1 downsampling units of the n-th multi-scale feature fusion group are the outputs of the 1st to n-1th encoding layers respectively; The inputs of the Nn upsampling units of the nth multi-scale feature fusion group are the outputs of the n+1th encoding layer to the Nth encoding layer respectively; The input of the first addition unit of the nth multi-scale feature fusion group is the output of the first convolution unit, n-1 downsampling units, and Nn upsampling units of the nth multi-scale feature fusion group; The input of the pyramid segmentation attention mechanism unit of the nth multi-scale feature fusion group is the output of the first addition unit of the nth multi-scale feature fusion group; The outputs of the pyramid segmentation attention mechanism units of the nth multi-scale feature fusion group are the inputs of each decoding layer of the decoder; The improved SegFormer network further includes a spatial detail enhancement module, a multiplication unit and a second addition unit; the spatial detail enhancement module includes a convolution combination and a spatial attention mechanism unit connected in sequence; the convolution combination includes a plurality of second convolution units connected in sequence; The input of the first second convolution unit is the image to be segmented; the output of the spatial attention mechanism unit is connected to the input of the multiplication unit; the input of the multiplication unit is also connected to the output of the last decoding layer of the decoder; the input of the second addition unit is connected to the output of the multiplication unit and the output of the last decoding layer of the decoder; the output of the second addition unit is the semantic segmentation result of the image; The decoder includes multiple decoding layers connected in sequence; a splicing layer is provided between two adjacent decoding layers; the input of each splicing layer also includes multiple fusion feature maps output by a multi-scale feature fusion module; the decoding layer adopts a multi-layer perceptron.

2. The image semantic segmentation method based on multi-scale feature fusion according to claim 1, characterized in that: The encoder includes a plurality of encoding layers connected in sequence; the encoding layers adopt Transformer blocks.

3. The image semantic segmentation method based on multi-scale feature fusion according to claim 1, characterized in that: Inputting the image to be segmented into the trained image semantic segmentation network to obtain the image semantic segmentation result, specifically including: Inputting the image to be segmented into an encoder, wherein each encoding layer of the encoder extracts semantic information of the image; The semantic information extracted from each coding layer is input into each multi-scale feature fusion group of the multi-scale feature fusion module to obtain multiple fusion feature maps of different scales; Input the fused feature maps of different scales into the decoder to obtain the fused semantic feature map; Inputting the image to be segmented into a spatial detail enhancement module to extract spatial detail feature information of the image; The image semantic segmentation result is obtained by fusing the semantic feature map and spatial detail feature information.

4. The image semantic segmentation method based on multi-scale feature fusion according to claim 1, characterized in that: Inputting the image to be segmented into the trained image semantic segmentation network to obtain the image semantic segmentation result, specifically including: Obtaining a training set; the training set includes a number of image samples to be segmented and corresponding image semantic segmentation result samples; Input the image samples to be segmented in the training set into the initial image semantic segmentation network to obtain image semantic segmentation prediction samples; Calculate the loss error based on the image semantic segmentation prediction sample and the corresponding image semantic segmentation result sample; Adjust the network parameters of the initial image semantic segmentation network according to the loss error; Until the loss error converges or the maximum number of iterations is reached, the trained image semantic segmentation network is obtained; The image to be segmented is input into the trained image semantic segmentation network to obtain the image semantic segmentation result.

5. An image semantic segmentation device based on multi-scale feature fusion, characterized in that: include: An image acquisition module, used to acquire the image to be segmented; An image semantic segmentation module is used to input the image to be segmented into a trained image semantic segmentation network to obtain an image semantic segmentation result; the image semantic segmentation network adopts an improved SegFormer network; in the improved SegFormer network, a multi-scale feature fusion module is introduced at the bridge between the encoder and the decoder; The multi-scale feature fusion module is used to fuse features of different scales according to the semantic information output by the encoder and output multiple fused feature maps of different scales; The decoder is used to perform feature fusion on multiple fusion feature maps of different scales to obtain the image semantic segmentation result; The multi-scale feature fusion module includes: N multi-scale feature fusion groups; N is the number of encoding layers of the encoder; The nth multi-scale feature fusion group includes the first convolution unit, n-1 downsampling units, Nn upsampling units, the first addition unit, and the pyramid segmentation attention mechanism unit; n=1, 2, ..., N; The input of the first convolution unit of the nth multi-scale feature fusion group is the output of the nth encoding layer; The inputs of the n-1 downsampling units of the n-th multi-scale feature fusion group are the outputs of the 1st to n-1th encoding layers respectively; The inputs of the Nn upsampling units of the nth multi-scale feature fusion group are the outputs of the n+1th encoding layer to the Nth encoding layer respectively; The input of the first addition unit of the nth multi-scale feature fusion group is the output of the first convolution unit, n-1 downsampling units, and Nn upsampling units of the nth multi-scale feature fusion group; The input of the pyramid segmentation attention mechanism unit of the nth multi-scale feature fusion group is the output of the first addition unit of the nth multi-scale feature fusion group; The outputs of the pyramid segmentation attention mechanism units of the nth multi-scale feature fusion group are the inputs of each decoding layer of the decoder; The improved SegFormer network further includes a spatial detail enhancement module, a multiplication unit and a second addition unit; the spatial detail enhancement module includes a convolution combination and a spatial attention mechanism unit connected in sequence; the convolution combination includes a plurality of second convolution units connected in sequence; The input of the first second convolution unit is the image to be segmented; the output of the spatial attention mechanism unit is connected to the input of the multiplication unit; the input of the multiplication unit is also connected to the output of the last decoding layer of the decoder; the input of the second addition unit is connected to the output of the multiplication unit and the output of the last decoding layer of the decoder; the output of the second addition unit is the semantic segmentation result of the image; The decoder includes multiple decoding layers connected in sequence; a splicing layer is provided between two adjacent decoding layers; the input of each splicing layer also includes multiple fusion feature maps output by a multi-scale feature fusion module; the decoding layer adopts a multi-layer perceptron.

6. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the image semantic segmentation method based on multi-scale feature fusion according to any one of claims 1 to 4.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the image semantic segmentation method based on multi-scale feature fusion according to any one of claims 1 to 4 is implemented.