Image semantic segmentation method, device and equipment introducing decoupling residual attention

By introducing an image semantic segmentation method with decoupled residual attention, combined with an encoder, decoder and self-attention module, the problem of focal boundary blur in endoscopic image segmentation is solved, and a higher precision focal region segmentation is achieved.

CN120236072APending Publication Date: 2025-07-01SOUTHWEAT UNIV OF SCI & TECH +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311869974.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-29
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

The existing medical endoscopic image segmentation method is difficult to achieve high-precision segmentation when facing complex lesion structures. The traditional method is insufficient in flexibility, while deep learning methods have problems of blur or loss in lesion boundary recognition.

Method used

The image semantic segmentation method that introduces decoupled residual attention is adopted. Through the combination of encoder and decoder, cross-stage cross-fusion and boundary supervision decoder are used, and the decoupled residual self-attention module is combined to extract and fuse features of different dimensions, enhance boundary information, and improve segmentation effect.

Benefits of technology

It realizes clearer lesion area segmentation in endoscopic images, improves segmentation accuracy and boundary recognition capabilities, and improves segmentation effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236072A_ABST
    Figure CN120236072A_ABST
Patent Text Reader

Abstract

The invention provides an image semantic segmentation method, device and equipment introducing decoupling residual attention, and can be applied to the technical field of deep learning and image segmentation. The method comprises the steps that image data of an endoscope image are input into an encoder to obtain N encoding feature maps, N is a positive integer and is larger than or equal to 3, the ith encoding feature map is obtained by inputting the (i-1) th encoding feature map into the encoder, and i is a positive integer and 1lt; i < = N; performing feature fusion on the first N-2 coded feature maps of the N coded feature maps and the (N-1) th coded feature map to obtain N-2 fused feature maps; and respectively inputting the N-2 fusion feature maps, the (N-1) th coding feature map and the Nth coding feature map into a decoder to obtain a semantic segmentation image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the fields of deep learning and image segmentation, and particularly to an image semantic segmentation method, apparatus, and device incorporating decoupled residual attention. Background Art

[0002] With the development of computer technology, in the medical field, computer technology can be used to process medical images to reduce the complexity of medical images and assist medical staff in analyzing medical images. Currently, the segmentation methods for endoscopic images can be broadly divided into two categories: traditional segmentation methods and deep learning segmentation methods.

[0003] Medical images include in-vivo endoscopic images captured using an endoscope. Due to the presence of a large amount of noise in endoscopic images, it is difficult to clearly identify the lesion area from the endoscopic images. Using traditional segmentation methods requires manual adjustment of parameters, which is difficult to adapt to complex lesion structures and is limited to specific image types, lacking flexibility. Different feature extraction methods need to be manually designed for different types of endoscopic images, and the computational performance cannot handle large-scale and high-resolution data. Summary of the Invention

[0004] In view of the above problems, the present disclosure provides an image semantic segmentation method, apparatus, device, medium, and program product incorporating decoupled residual attention.

[0005] According to a first aspect of the present disclosure, there is provided an image semantic segmentation method incorporating decoupled residual attention, including: inputting the image data of an endoscopic image into an encoder to obtain N encoded feature maps, where N is a positive integer, N≥3, and the i-th encoded feature map is obtained by inputting the (i - 1)-th encoded feature map into the encoder, i is a positive integer, 1 < i ≤ N; respectively performing feature fusion on the first N - 2 encoded feature maps among the N encoded feature maps and the (N - 1)-th encoded feature map to obtain N - 2 fused feature maps; inputting the N - 2 fused feature maps, the (N - 1)-th encoded feature map, and the N-th encoded feature map into a decoder respectively to obtain a semantic segmentation image.

[0006] According to an embodiment of the present disclosure, respectively performing feature fusion on the first N - 2 encoded feature maps among the N encoded feature maps and the (N - 1)-th encoded feature map to obtain N - 2 fused feature maps includes: performing a convolution operation on the j-th encoded feature map to obtain the j-th first dimension-reduced feature map, where j is a positive integer, j ≤ N - 2; performing bilinear interpolation upsampling on the (N - 1)-th encoded feature map using adjacent pixel weighting to obtain an upsampled feature map; performing a convolution operation on the upsampled feature map to obtain a second dimension-reduced feature map; concatenating the j-th first dimension-reduced feature map and the second dimension-reduced feature map in the channel dimension to obtain the j-th fused feature map.

[0007] According to an embodiment of the present disclosure, the image semantic segmentation method introducing decoupled residual attention further includes: inputting the j-th fused feature map into a decoder to obtain a first decoded feature corresponding to the j-th encoded feature map; inputting the N-th encoded feature map and the (N - 1)-th encoded feature map into the decoder to obtain a second decoded feature.

[0008] According to an embodiment of the present disclosure, the decoder includes a boundary supervision decoder; inputting the j-th fused feature map into the decoder to obtain a first decoded feature corresponding to the j-th encoded feature map includes: obtaining a gradient magnitude response according to the boundary supervision decoder; obtaining the first decoded feature based on the gradient magnitude response and the j-th fused feature map.

[0009] According to an embodiment of the present disclosure, inputting N - 2 fused feature maps, the (N - 1)-th encoded feature map, and the N-th encoded feature map into the decoder respectively to obtain a semantic segmentation image includes: hierarchically splicing the first decoded feature and the second decoded feature to obtain the semantic segmentation image.

[0010] According to an embodiment of the present disclosure, the image semantic segmentation method introducing decoupled residual attention further includes: inputting a shallow encoded feature map into a decoupled residual self-attention module to obtain a self-attention feature map; using the self-attention feature map to update the shallow encoded feature map.

[0011] According to an embodiment of the present disclosure, the N encoded feature maps include M shallow encoded feature maps, where M is a positive integer and M < N; inputting the shallow encoded feature map into the decoupled residual self-attention module to obtain a self-attention feature map includes: projecting the shallow encoded feature map through a projection mapping matrix to obtain a projection vector, where the projection vector includes a query vector, a key vector, and a value vector; determining a position vector of the projection vector, where the position vector includes a query position vector, a key position vector, and a value position vector; obtaining the self-attention feature map based on the projection vector and the position vector.

[0012] According to an embodiment of the present disclosure, obtaining the self-attention feature map based on the projection vector and the position vector includes: determining position correlation according to the query vector and the query position vector; determining key correlation according to the key vector and the key position vector; determining vector correlation according to the query vector and the key vector; determining self-attention weights according to the position correlation, the key correlation, and the vector correlation; determining the self-attention feature map according to the self-attention weights, the value vector, and the value position vector.

[0013] A second aspect of the present disclosure provides an image semantic segmentation device introducing decoupled residual attention, including:

[0014] A data input module for inputting the image data of an endoscopic image into an encoder to obtain N encoded feature maps, where N is a positive integer, N≥3, and the i-th encoded feature map is obtained by inputting the (i - 1)-th encoded feature map into the encoder, i is a positive integer, 1 < i ≤ N;

[0015] A feature fusion module for respectively performing feature fusion on the first N - 2 encoded feature maps and the (N - 1)-th encoded feature map among the N encoded feature maps to obtain N - 2 fused feature maps;

[0016] A feature decoding module for respectively inputting the N - 2 fused feature maps, the (N - 1)-th encoded feature map, and the N-th encoded feature map into a decoder to obtain a semantic segmentation image.

[0017] A third aspect of the present disclosure provides an electronic device, including: one or more processors; a memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors are caused to execute the above method.

[0018] A fourth aspect of the present disclosure further provides a computer-readable storage medium having executable instructions stored thereon, and when the instructions are executed by a processor, the processor is caused to execute the above method.

[0019] A fifth aspect of the present disclosure further provides a computer program product including a computer program, and when the computer program is executed by a processor, the above method is implemented.

[0020] According to the embodiments of the present disclosure, by inputting the image data of an endoscopic image into an encoder, multiple encoded feature maps are obtained to extract features of different dimensions. Cross-level cross-fusion is performed on the encoded feature maps, so as to fuse the feature maps of deep features and the feature maps of shallow features, and a fused feature map that takes into account both deep features and shallow details is obtained. The decoder is used to process the encoded feature maps and the fused feature maps, so as to obtain a semantic segmentation image. Since the semantic segmentation image can contain features of multiple different dimensions, the boundary of the semantic segmentation image can be ensured to be clear and the segmentation effect can be improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Through the following description of the embodiments of the present disclosure with reference to the drawings, the above content and other objects, features, and advantages of the present disclosure will become clearer. In the drawings:

[0022] Figure 1 Schematically shows an application scenario diagram of an image semantic segmentation method, device, device, medium, and program product introducing decoupled residual attention according to an embodiment of the present disclosure;

[0023] Figure 2Schematically shows a flowchart of an image semantic segmentation method introducing decoupled residual attention according to an embodiment of the present disclosure;

[0024] Figure 3 Schematically shows a schematic diagram of a cross-level cross-fusion module according to an embodiment of the present disclosure;

[0025] Figure 4 Schematically shows a configuration diagram of a Gabor filter according to an embodiment of the present invention;

[0026] Figure 5 Schematically shows a structural diagram of a boundary supervision decoding module according to an embodiment of the present disclosure;

[0027] Figure 6 Schematically shows a schematic diagram of a decoupled self-attention module according to an embodiment of the present disclosure;

[0028] Figure 7 Schematically shows a structural diagram of an image semantic segmentation model introducing decoupled residual attention according to an embodiment of the present disclosure;

[0029] Figure 8 Schematically shows an input endoscopic image according to an embodiment of the present disclosure;

[0030] Figure 9A Schematically shows a schematic diagram of the real part response of a Gabor filter according to an embodiment of the present disclosure;

[0031] Figure 9B Schematically shows a schematic diagram of the amplitude response of a Gabor filter according to an embodiment of the present disclosure;

[0032] Figure 10 Schematically shows a schematic diagram of a segmentation result using an image semantic segmentation model introducing decoupled residual attention according to an embodiment of the present disclosure;

[0033] Figure 11 Schematically shows a structural block diagram of an image semantic segmentation device introducing decoupled residual attention according to an embodiment of the present disclosure; and

[0034] Figure 12 Schematically shows a block diagram of an electronic device suitable for implementing an image semantic segmentation method introducing decoupled residual attention according to an embodiment of the present disclosure. Detailed implementation manners

[0035] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the following detailed description, for the sake of explanation, numerous specific details are set forth to provide a comprehensive understanding of the embodiments of the present disclosure. However, obviously, one or more embodiments can also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present disclosure.

[0036] The terms used herein are merely for describing specific embodiments and are not intended to limit the present disclosure. The terms "including", "comprising", etc. used herein indicate the presence of the described features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0037] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0038] In the case of using expressions such as "at least one of A, B, and C, etc.", generally, it should be interpreted according to the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include, but is not limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.).

[0039] In the technical solution of the present invention, the user information involved (including but not limited to user personal information, user image information, user device information, such as location information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) are all information and data authorized by the user or fully authorized by all parties. And the processing of relevant data, such as collection, storage, use, processing, transmission, provision, disclosure, and application, all comply with the relevant laws, regulations, and standards of relevant countries and regions, take necessary confidentiality measures, do not violate public order and good customs, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0040] In the process of implementing the present disclosure, it is found that although traditional segmentation methods have good interpretability and low computational resource consumption, they are difficult to achieve high-precision segmentation results when facing complex esophageal endoscopy scenarios. The semantic segmentation method based on deep learning learns and segments the lesion area through training data, and can capture the lesion features while considering the generalization of segmentation for various complex lesions.

[0041] However, there are also a large number of problems with deep learning-based semantic segmentation methods. For example, the local perception and weight sharing of Convolutional Neural Networks (CNNs) can lead to their focus only on local features of images, lacking the ability to completely segment lesion areas with a relatively large area occupancy. During the downsampling process of feature extraction, it will inevitably cause the blurring or loss of lesion boundary information, affecting the performance of the segmentation model for accurate segmentation of lesion areas and resulting in insufficient accuracy of the segmentation results.

[0042] Embodiments of the present disclosure provide an image semantic segmentation method introducing decoupled residual attention, including: inputting image data of an endoscopic image into an encoder to obtain N encoded feature maps, where N is a positive integer, N≥3, and the i-th encoded feature map is obtained by inputting the (i - 1)-th encoded feature map into the encoder, i is a positive integer, 1 < i ≤ N; respectively performing feature fusion on the first N - 2 encoded feature maps among the N encoded feature maps and the (N - 1)-th encoded feature map to obtain N - 2 fused feature maps; inputting the N - 2 fused feature maps, the (N - 1)-th encoded feature map, and the N-th encoded feature map into a decoder respectively to obtain a semantic segmentation image.

[0043] Figure 1 Schematically shows an application scenario diagram of an image semantic segmentation method, device, equipment, medium, and program product introducing decoupled residual attention according to an embodiment of the present disclosure.

[0044] As Figure 1 shown, the application scenario 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links among the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0045] Users can use at least one of the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only for example).

[0046] The first terminal device 101, the second terminal device 102, and the third terminal device 103 may be an endoscope system including a flexible fiber optic sensor, an endoscope camera, a display screen, an operation and control device, or various other electronic devices with a display screen and supporting web browsing, including but not limited to smartphones, tablet computers, laptop computers, desktop computers, etc. The first terminal device 101, the second terminal device 102, and the third terminal device 103 may interact with the network 104 through a database.

[0047] The server 105 may be a server providing various services, such as a background management server (only for example) that supports the websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103. The background management server may analyze and process data such as user requests received, and feedback the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.

[0048] It should be noted that the image semantic segmentation method introducing decoupled residual attention provided by the embodiments of the present disclosure can generally be executed by the server 105. Correspondingly, the image semantic segmentation device introducing decoupled residual attention provided by the embodiments of the present disclosure can generally be deployed in the server 105. The image semantic segmentation method introducing decoupled residual attention provided by the embodiments of the present disclosure can also be executed by a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105. Correspondingly, the image semantic segmentation device introducing decoupled residual attention provided by the embodiments of the present disclosure can also be deployed in a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105.

[0049] The deployment process of the image semantic segmentation method introducing decoupled residual attention provided by the embodiments of the present disclosure is as follows:

[0050] (1) Develop a pilot system (build a network and hardware infrastructure, install and configure relevant software). Use a data set to train the network to obtain a preliminary segmentation model, build an interface software with core functions, and build a simulated patient information data database, and preliminary function tests and verifications can be carried out.

[0051] (2) Perform function tests, performance tests, and load tests according to the design. After building the experimental system, verify whether the web page can normally call the network model, and verify whether the multi-user login system is stuck.

[0052] (3) After passing the test, start planning the prototype system. After the test function is normal, start modifying and improving the network model and conducting more experiments, improving the interface presentation effect, ensuring the fluency and convenience of operation, ensuring the usability function of the interface, and ensuring the integrity function of the platform.

[0053] (4) Complete the network construction, software and hardware installation and configuration of the prototype system. After the platform network is made, build a local area network to ensure that after the code runs, multiple PC terminals can access within the local area network of the entire server.

[0054] (5) Migrate the data from the existing application to the current solution. Migrate the data such as the database of the experimental system to the current server, and migrate the designed lesion segmentation model to this server, and conduct experiments on the current platform web page to verify whether the function can be realized and whether the realized effect meets the expectations.

[0055] (6) Set up the administrator account through the database to complete the deployment of the platform. According to the function settings of the platform, different hospitals can set different administrator accounts to manage all users within their own hospitals, set up the initial administrator account to ensure the normal login of the system and the registration application of new users.

[0056] It should be understood that Figure 1 the numbers of the terminal devices, networks and servers in

[0057] are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks and servers. Figure 1 Based on the scenario described below Figures 2 - 8 、 Figure 9A 、 Figure 9B 、 Figure 10 the image semantic segmentation method introducing decoupled residual attention of the public implementation example will be described in detail.

[0058] Figure 2 Schematically shows a flowchart of the image semantic segmentation method introducing decoupled residual attention according to an embodiment of the present disclosure.

[0059] As Figure 2 shown, the image semantic segmentation introducing decoupled residual attention of this embodiment includes operation S210 to operation S230.

[0060] In operation S210, input the image data of the endoscopic image into the encoder to obtain N encoded feature maps, where N is a positive integer, N≥3, and the i-th encoded feature map is obtained by inputting the (i - 1)-th encoded feature map into the encoder, i is a positive integer, 1 < i ≤ N.

[0061] In operation S220, the first N - 2 encoded feature maps among the N encoded feature maps are respectively subjected to feature fusion with the (N - 1)-th encoded feature map to obtain N - 2 fused feature maps.

[0062] In operation S230, the N - 2 fused feature maps, the (N - 1)-th encoded feature map, and the N-th encoded feature map are respectively input into a decoder to obtain a semantic segmentation image.

[0063] According to an embodiment of the present disclosure, an encoder may perform downsampling processing such as convolution on the features of input image data to obtain output encoded feature maps. Among them, the parameters of the encoder used in the process of obtaining each encoded feature map may be different, as shown in formula (1):

[0064]

[0065] where f i is the i-th encoded feature map extracted by the encoder, f0 ∈ R H×W×C may be the image data of an endoscopic image, R is an image set, and H, W, and C are respectively the height, width, and number of channels of the image data. is the convolution process of the i-th layer of the encoder.

[0066] According to an embodiment of the present disclosure, the first N - 2 encoded feature maps and the (N - 1)-th encoded feature map may be respectively subjected to cross-level cross fusion (CLF) to obtain fused feature maps that take into account both deep features and shallow details.

[0067] According to an embodiment of the present disclosure, shallow encoded feature maps include more details of an endoscopic image, and deep encoded feature maps include more semantic features of the endoscopic image. By performing CLF grouped fusion on the shallow encoded feature maps and deep encoded feature maps obtained in the encoder, comprehensive information containing both details and semantics can be provided to the decoder without increasing the computational amount of additional feature maps.

[0068] According to an embodiment of the present disclosure, a decoder may perform upsampling processing such as interpolation or transposed convolution on the input encoded feature maps and fused feature maps to obtain a semantic segmentation image.

[0069] According to an embodiment of the present disclosure, by inputting the image data of an endoscopic image into an encoder, multiple encoded feature maps are obtained, and features of different dimensions are extracted. Cross-level cross fusion is performed on the encoded feature maps, so that the feature maps of deep features and the feature maps of shallow features are fused to obtain fused feature maps that take into account both deep features and shallow details. A decoder is used to process the encoded feature maps and the fused feature maps, thereby obtaining a semantic segmentation image. Since the semantic segmentation image can contain features of multiple different dimensions, the boundary of the semantic segmentation image can be ensured to be clear, and the segmentation effect can be improved.

[0070] According to an embodiment of the present disclosure, the first N - 2 encoded feature maps among the N encoded feature maps are respectively subjected to feature fusion with the (N - 1)-th encoded feature map to obtain N - 2 fused feature maps, including: performing a convolution operation on the j-th encoded feature map to obtain the j-th first dimensionality reduction feature map, where j is a positive integer and j ≤ N - 2; performing upsampling on the (N - 1)-th encoded feature map by using a bilinear interpolation upsampling method with adjacent pixel weighting to obtain an upsampled feature map; performing a convolution operation on the upsampled feature map to obtain a second dimensionality reduction feature map; and concatenating the j-th first dimensionality reduction feature map and the second dimensionality reduction feature map in the channel dimension to obtain the j-th fused feature map.

[0071] According to an embodiment of the present disclosure, the receptive field of the encoded feature map increases as the number of network layers deepens, so that more semantic features of the endoscopic image can be extracted. However, at the same time, the boundary detail information will also be lost as the receptive field increases. The semantic features generated during the encoding process contain context and semantic relationships, which are helpful for the use of the segmentation image. However, since the deep feature map lacks detail information and local context information, it is not conducive to the decoder to supplement the detail information lost during the encoding process through skip connections.

[0072] According to an embodiment of the present disclosure, a 1×1 convolutional kernel can be used to perform a convolution operation on the j-th encoded feature map to obtain the j-th first dimensionality reduction feature map, where the number of channels of the j-th first dimensionality reduction feature map is half of the number of channels of the j-th encoded feature map. The (N - 1)-th encoded feature map is subjected to upsampling by using a bilinear interpolation upsampling method with adjacent pixel weighting to obtain an upsampled feature map, where the number of channels of the upsampled feature map is greater than that of the j-th first dimensionality reduction feature map. A 1×1 convolutional kernel can be used to perform a convolution operation on the upsampled feature map to obtain a second dimensionality reduction feature map, where the number of channels of the second dimensionality reduction feature map is half of the number of channels of the j-th encoded feature map. Since the number of channels of both the j-th first dimensionality reduction feature map and the second dimensionality reduction feature map is half of the number of channels of the j-th encoded feature map, the j-th first dimensionality reduction feature map and the second dimensionality reduction feature map are concatenated in the channel dimension, and the number of channels of the obtained j-th fused feature map is the same as that of the j-th encoded feature map.

[0073] According to an embodiment of the present disclosure, the j-th fused feature map f′ can be calculated by formula (2) j :

[0074] f′ j = CLF(f j , f N-1 ) (2)

[0075] where CLF(·) is the processing of the cross-level cross-fusion layer and can be expressed as formula (3):

[0076] f' j = Concat(f 1×1 (f j ), f 1×1 (f up (f N-1 ))) (3)

[0077] where Concat(·) is to concatenate matrices in the channel dimension (Concatenate, Concat), f 1×1 (·) is a 1×1 convolution, and f up (·) is an upsampling operation.

[0078] According to an embodiment of the present disclosure, by respectively performing feature fusion on the first N - 2 encoded feature maps and the (N - 1)-th encoded feature map among N encoded feature maps, a fused feature map that combines details and semantic features of different levels of the endoscopic image can be obtained.

[0079] Figure 3 Schematically shows a schematic diagram of a cross-level cross-fusion module according to an embodiment of the present disclosure.

[0080] As Figure 3 shown, this cross-level cross-fusion module performs cross-level cross-fusion on f1 to f3 and f4 respectively. First, f1 is dimension-reduced to half of the original number of channels of f1 through a 1×1 convolution, and then processed using batch normalization and the ReLU activation function to obtain the first dimension-reduced feature map. f4 is upsampled using the bilinear interpolation method with adjacent pixel weighting so that the number of channels is greater than that of the first dimension-reduced feature map to obtain an upsampled feature map. Then, the upsampled feature map is dimension-reduced to half of the original number of channels of f1 through a 1×1 convolution to obtain a second dimension-reduced feature map. The first dimension-reduced feature map and the second dimension-reduced feature map are concatenated in the channel dimension to obtain the first fused feature map f'1 with the same number of channels as the first encoded feature map. The same method can be used to process f2 and f3 to obtain the second fused feature map f'2 with the same number of channels as the second encoded feature map and the third fused feature map f'3 with the same number of channels as the third encoded feature map, completing the operation of the cross-level cross-fusion module.

[0081] According to an embodiment of the present disclosure, the image semantic segmentation method introducing decoupled residual attention further includes: inputting the j-th fused feature map into a decoder to obtain a first decoded feature corresponding to the j-th encoded feature map; inputting the N-th encoded feature map and the (N - 1)-th encoded feature map into the decoder to obtain a second decoded feature.

[0082] According to an embodiment of the present disclosure, by inputting the N-th encoded feature map and the (N - 1)-th encoded feature map into the decoder, after operations such as convolution and concatenation, the second decoded feature U can be obtainedN-2 , as shown in formula (4):

[0083] U N-2 = N conv (f N-1 , f N )(4)

[0084] where N conv (·) represents operations such as convolution, splicing, and upsampling in the decoder.

[0085] According to an embodiment of the present disclosure, the decoder includes a boundary supervision decoder; inputting the j-th fusion feature map into the decoder to obtain a first decoded feature corresponding to the j-th encoded feature map, including: obtaining a gradient magnitude response according to the boundary supervision decoder; and obtaining the first decoded feature based on the gradient magnitude response and the j-th fusion feature map.

[0086] According to an embodiment of the present disclosure, the decoder may include a Boundary Supervised Decoding (BSD) module, and the BSD can be used to enhance the boundary information of the fusion feature map output by the CLF. Inputting the j-th fusion feature map into the decoder, and using the BSD to process to obtain a first decoded feature corresponding to the j-th encoded feature map.

[0087] According to an embodiment of the present disclosure, the boundary supervision decoding module may include a Gabor filter. Inputting the fusion feature map output by the CLF into the Gabor filter, setting filters with different scales for the fusion feature map, calculating the L2 norm of the four-direction feature maps, and obtaining the gradient magnitude response after convolving the endoscopic image with the Gabor filter to represent the change in the gray value of the image pixels in the direction. Among them, the larger the gradient magnitude response, the greater the pixel gradient at that position, indicating that there are important features such as edges or brightness at that position. The magnitude response generates a spatial attention map through a non-linear activation function, which is multiplied by the j-th encoded feature map to obtain a first decoded feature with enhanced edge information.

[0088] According to an embodiment of the present disclosure, the two-dimensional Gabor filter may include a two-dimensional Gaussian function and a sine function. The response range of the filter to the image space can be controlled by adjusting the mean and standard deviation of the two-dimensional Gaussian, and features of different sizes and scales in the image can be captured. By adjusting the frequency and phase of the sine function to control the response range of the filter to various texture frequencies and different directions in the image, textures and edges in the image can be detected.

[0089] According to an embodiment of the present disclosure, the boundary supervision decoding module may use a Gabor filter to form a boundary enhancement operator, take the gradient magnitude response after convolving the endoscopic image with the filter as the extracted feature, enhance the response of the boundary and surrounding pixels through a non-linear transformation function, adaptively extract the boundary weight information of the feature map to generate a spatial attention map, then multiply it with the fused feature map output by the CLF, integrate the boundary information into the multi-level feature map, and finally splice it with the encoded feature maps at each resolution level in the upsampling branch through a skip connection.

[0090] According to an embodiment of the present disclosure, the Gabor transform has good spatial locality and direction selectivity, and can capture the spatial frequency and local structural features in multiple directions of an image. The sinusoidal harmonic modulated by the Gaussian envelope of the Gabor wavelet filter can be represented by formula (5):

[0091]

[0092] wherein, the two-dimensional Gaussian function is used to determine the size and shape of the Gabor wavelet filter in the spatial domain, x r and y r are the rotated coordinates of the point (x, y) in the image spatial domain at a given angle . When the standard deviations σ x and σ y are relatively large, the response of the filter in the spatial domain is smoother and wider, and can capture larger-scale image features and structures. When the standard deviations σ x and σ y are relatively small, the response of the filter is sharper and more localized, and can detect local details and textures of the image. ω x =|ω|cosθ, ω y =|ω|sinθ, ω x and ω y are the projections in the x and y directions at a given angular frequency respectively, and are used to represent the sensitivity to different spatial frequencies of the image. Image regions with higher spatial frequencies usually represent detailed structures such as edges or textures, and image regions with lower spatial frequencies represent larger-scale structures.

[0093] Figure 4 Schematically shows a Gabor filter configuration diagram according to an embodiment of the present invention.

[0094] As Figure 4As shown, using five scales (2, 7, 9, 13, 19) and 4 directions (0, π / 4, π / 2, 3π / 4), 20 different Gabor filters can be obtained. The center point of each figure is the point with the strongest response of the filter, that is, the mean point of the Gaussian function. The corresponding intensity of the filter is negatively correlated with the distance from the center point. Affected by the oscillating waveform of the complex function, the response in the vertical direction of the filter is an alternating strong and weak response.

[0095] According to an embodiment of the present disclosure, the image feature W(x, y) extracted by the Gabor filter can be determined by formula (6):

[0096]

[0097] where * is the convolution operation, I(x, y) is the fused feature map, is the filter.

[0098] Figure 5 Schematically shows the structural diagram of the boundary supervision decoding module according to an embodiment of the present disclosure.

[0099] As Figure 5 shown, f′ j is the fused feature map. Different scales of filters are set according to the feature maps at different levels in the fused feature map. The L2 norm of the fused feature map in four directions is calculated at a specific scale to obtain the gradient amplitude response of the endoscopic image and the Gabor filter. The amplitude response generates a spatial attention map through the Sigmoid function. The spatial attention map is multiplied by the fused feature map to obtain the output feature map with enhanced edge information, that is, the first decoded feature Oj, as shown in formula (7):

[0100]

[0101] where Sigmoid(·) is the Sigmoid activation function, L2(·) is the L2 norm, is θ i angle of the Gabor filter.

[0102] According to an embodiment of the present disclosure, the N - 2 fused feature maps, the (N - 1)th encoded feature map, and the Nth encoded feature map are respectively input into the decoder to obtain the semantic segmentation image, including: hierarchically splicing the first decoded feature and the second decoded feature to obtain the semantic segmentation image.

[0103] According to an embodiment of the present disclosure, the second decoded feature and k first decoded features are hierarchically spliced to obtain the (k - 1)th splicing result, where 2 ≤ k ≤ N - 3.

[0104] According to an embodiment of the present disclosure, the splicing result can be determined by formula (8):

[0105] U k = N conv (BSD(CLF(f k+1 ,f N-1 )),U k+1 ) = N conv (BSD(f′ k+1 ),U k+1 ) (8)

[0106] Wherein, BSD(·) are operations such as convolution, splicing, and upsampling in boundary supervised decoding.

[0107] According to an embodiment of the present disclosure, by performing boundary supervised decoding on the first decoded feature corresponding to the first encoded feature map and the first splicing result, a semantic segmentation image O can be obtained, as shown in formula (9):

[0108] O = N conv (BSD(f′1), f1) (9)

[0109] According to an embodiment of the present disclosure, by determining the first decoded feature and the second decoded feature through a decoder, and hierarchically splicing the first decoded feature with enhanced edge information and the second decoded feature, the context information and detail information lost during the downsampling process can be compensated, and at the same time, the decoder's ability to express boundary information can be strengthened, obtaining a more accurate semantic segmentation image.

[0110] According to an embodiment of the present disclosure, an image semantic segmentation method introducing decoupled residual attention further includes: inputting a shallow encoded feature map into a decoupled residual self-attention module to obtain a self-attention feature map; using the self-attention feature map to update the shallow encoded feature map.

[0111] According to an embodiment of the present disclosure, the shallow encoded feature map contains more detail information of the endoscopic image. The shallow encoded feature map can be determined according to the size of N. For example, in the case of N = 5, the first encoded feature map and the second encoded feature map can be used as the shallow encoded feature map. Inputting the shallow encoded feature map into the decoupled residual self-attention module, the global and local information of the feature map fused through residual connection is obtained, that is, the self-attention feature map. The decoupled residual self-attention module includes a decoupled self-attention (DSA) module and a residual connection. The nth self-attention feature map F can be obtained through formula (10) n :

[0112] F n = σ(DSA W (DSA H (f n )) + f n) (10)

[0113] Among them, σ(·) is the ReLU activation function, and DSA W (·) represents the calculation process in the width direction of DSA, and DSA H (·) represents the calculation process in the height direction of DSA.

[0114] According to an embodiment of the present disclosure, after determining the nth self-attention feature map F n , F can be used n to update f n , so that the updated f n includes the global and local information of the encoded feature map.

[0115] According to an embodiment of the present disclosure, the N encoded feature maps include M shallow encoded feature maps, where M is a positive integer and M < N; the shallow encoded feature maps are input into the decoupled residual self-attention module to obtain self-attention feature maps, including: projecting the shallow encoded feature maps through a projection mapping matrix to obtain projection vectors, where the projection vectors include query vectors, key vectors, and value vectors; determining the position vectors of the projection vectors, where the position vectors include query position vectors, key position vectors, and value position vectors; and obtaining self-attention feature maps based on the projection vectors and the position vectors.

[0116] According to an embodiment of the present disclosure, the query vector Q, the key vector K, and the value vector V can be determined by the shallow encoded feature Figure X and the projection mapping matrix W QKV as shown in formula (11):

[0117] XW QKV =[Q, K, V] (11)

[0118] According to an embodiment of the present disclosure, the position vector can be used to represent the position information of the current feature map in all feature maps. The relative position parameters need to be initialized before network training. Appropriate initialization parameters can help the model converge to a better solution faster, reduce training time, and avoid problems such as gradient disappearance or explosion. By using a random sampling initialization method with a normal distribution having a mean of 0 and a standard deviation of 1, the parameters are within a small but appropriate range, which helps to learn and update the position information during the subsequent model training process and finally effectively represent the position information in the input endoscopic image.

[0119] According to an embodiment of the present disclosure, obtaining a self-attention feature map based on a projection vector and a position vector includes: determining query correlation according to a query vector and a query position vector; determining keyword correlation according to a keyword vector and a keyword position vector; determining vector correlation according to the query vector and the keyword vector; determining a self-attention weight according to the query correlation, the keyword correlation, and the vector correlation; and determining a self-attention feature map according to the self-attention weight, a value vector, and a value position vector.

[0120] According to an embodiment of the present disclosure, the query correlation can be obtained by performing matrix multiplication on a query vector Q and a query position vector Q em The keyword correlation can be obtained by performing matrix multiplication on a keyword vector K and a keyword position vector K em The vector correlation can be obtained by performing matrix multiplication on the query vector Q and the keyword vector K. As shown in formulas (12) to (14):

[0121] qr = Q·Q em (12)

[0122] kr = K·K em (13)

[0123] qk = Q·K (14)

[0124] According to an embodiment of the present disclosure, the query correlation, the keyword correlation, and the vector correlation are used to obtain the self-attention weight W of the relative position information through batch normalization and an activation function s , as shown in formula (15):

[0125] W s = Softmax(BN(Concat(qr, kr, qk))) (15)

[0126] where Softmax(·) is the softmax activation function and BN(·) is the batch normalization (Batch Normalization, BN) process.

[0127] According to an embodiment of the present disclosure, using the self-attention weight W s to weight the value vector and the value position vector can obtain the self-attention feature map f Figure X corresponding to the shallow encoding feature x , as shown in formula (16):

[0128] f x = Concat(W s ×V, W s ×V em ) (16)

[0129] Among them, V em is the value position vector.

[0130] According to an embodiment of the present disclosure, the decoupled self-attention decouples the two-dimensional self-attention into two one-dimensional self-attentions, which are calculated successively in the height and width directions of the feature map. While capturing the global information of the feature map, the computational complexity of the affinity in the self-attention is reduced. In order to make up for the image patch position information lost due to serialization in the self-attention, relative position encoding is embedded in the decoupled attention, which is beneficial for the model to learn the relative position relationship between different dimensions in the feature map.

[0131] Figure 6 Schematically shows a schematic diagram of a decoupled self-attention module according to an embodiment of the present disclosure.

[0132] As Figure 6 shown, by encoding the feature map f n the query vector Q, the key vector K, and the value vector V can be determined. According to the query vector Q, the key vector K, and the query position vector Q em and the key position vector K em determined by training, the query correlation qr, the key correlation kr, and the vector correlation qk can be determined. By performing channel dimension concatenation, batch normalization, and processing using a non-linear activation function on qr, kr, and qk, the self-attention weight W s is determined. Multiply the self-attention weight W s separately with the value vector V and the value position vector V em by matrix multiplication, and perform channel dimension concatenation to obtain the self-attention feature map f′ n .

[0133] Figure 7 Schematically shows a structural diagram of an image semantic segmentation model introducing decoupled residual attention according to an embodiment of the present disclosure.

[0134] As Figure 7As shown, the model structure includes an encoder and a decoder. Among them, the encoder includes a convolution operation and a decoupled residual self-attention module. The image data of the endoscopic image is input into the encoder, and multiple encoded feature maps can be obtained through multiple convolutions. Among them, each encoded feature map is obtained by performing a convolution operation on the previous encoded feature map. The shallow encoded feature maps in the encoded feature maps are input into the decoupled residual self-attention module. After being processed by convolution (Conv), normalization, and ReLU activation function, and after passing through DSA_height (computation in the decoupled self-attention height direction) and DSA_width (computation in the decoupled self-attention width direction), convolution, normalization, and ReLU activation function processing are performed again to obtain the output feature map of the decoupled self-attention. The output feature map is added to the encoded feature map to obtain the self-attention feature map corresponding to the encoded feature map, and the corresponding encoded feature map is updated using the self-attention feature map. The decoder includes a boundary supervision decoding module, which performs operations such as convolution, splicing, and upsampling on the two encoded feature maps with the largest depth to obtain a depth feature map. The fused feature map obtained through CLF is input into the boundary supervision decoding module and is convolved, spliced, upsampled, etc. with the depth feature map of the previous layer to obtain the depth feature map of the next layer. According to the depth feature map of the last layer, the boundary segmentation result of the endoscopic image, that is, the segmentation prediction image, can be determined.

[0135] Figure 8 Schematically shows an input endoscopic image according to an embodiment of the present disclosure.

[0136] As Figure 8 shown, the input endoscopic image can be a medical image of an esophageal endoscope.

[0137] Figure 9A Schematically shows a schematic diagram of the real part response of a Gabor filter according to an embodiment of the present disclosure.

[0138] Figure 9B Schematically shows a schematic diagram of the amplitude response of a Gabor filter according to an embodiment of the present disclosure.

[0139] According to Figure 9A and Figure 9B , when using this Gabor filter to process the endoscopic image as Figure 8 shown, the most obvious value ranges of the real part response and the amplitude response are when σ takes 7, 13, and 19. Therefore, calculate the L2 norm in four directions at the scales where σ takes 7, 13, and 19, and calculate the first decoded feature, which can ensure that the first decoded feature contains enhanced edge features.

[0140] According to an embodiment of the present disclosure, compared with existing image segmentation algorithms, the image semantic segmentation method introducing decoupled residual attention has better segmentation effect because it considers the image details and semantic features of endoscopic images. Table 1 shows the comparison results between the image semantic segmentation method (BCS-SegNet) introducing decoupled residual attention of the present disclosure and existing image segmentation algorithms.

[0141] Table 1

[0142] Year Network Type Method mIoU (%) Dice (%) FLOPs Number of Parameters 2015 CNN U-Net 80.09 87.19 226.15G 24.89M 2017 CNN DeepLabV3+ 79.80 88.53 264.60G 70.07M 2020 CNN <![CDATA[U 2 Net]]> 74.43 83.31 150.61G 43.99M 2020 Transformer Axial-DeepLab 81.80 89.82 136.44G 63.11M 2021 Transformer + CNN TransUnet 81.21 87.86 129.29G 93.23M 2022 Transformer + CNN UCTransNet 78.81 86.90 172.01G 66.24M 2023 Transformer + CNN CTO 82.49 89.17 123.35G 59.81M - Transformer + CNN BCS-SegNet 83.03 91.28 254.94G 25.17M

[0143] The data in Table 1 were obtained by training on the same dataset, using Adam as the learning rate optimizer, setting the momentum parameter to 0.9, using the cosine annealing learning rate decay strategy, with an initial learning rate of 1×10-4 and a minimum learning rate of 0.01 times the initial learning rate, for 800 epochs. mIoU is the mean Intersection over Union, representing the ratio of the intersection to the union of the ground truth and the predicted value. The Dice coefficient is the F1 coefficient, and FLOPS is the number of floating-point operations per second (Floating-point operations per second). As shown in Table 1, compared with other image segmentation methods, the BCS-SegNet of the present disclosure has higher mIoU and Dice. The mIoU is improved by 2.94% compared with the relatively good-performing Unet and by 0.54% compared with CTO. The Dice coefficient is improved by 4.09% compared with Unet and by 2.11% compared with CTO. The FLOPs are slightly larger than those of Unet and about twice that of CTO, but the number of parameters is much smaller than that of CTO and slightly larger than that of Unet. Therefore, considering the algorithm performance and segmentation effect comprehensively, the performance of BCS-SegNet is better than that of existing image segmentation algorithms.

[0144] Figure 10 A schematic diagram showing the segmentation result using the image semantic segmentation model introducing decoupled residual attention according to an embodiment of the present disclosure is schematically shown.

[0145] As Figure 10 shown, the result diagram includes six different endoscopic images, and the area ratio, position, and shape of the regions in each endoscopic image are different. Among them, the first two columns are the original endoscopic image and the label image respectively, and the images in the following columns are the segmentation results of semantic segmentation of the endoscopic image using the BCS-SegNet, Unet, DeepLabV3+, UCTransNet, TransUnet, Acial-DeepLab, and CTO algorithms of the present disclosure respectively.

[0146] According to embodiments of the present disclosure, due to the local perception and weight sharing of CNN, it is unable to effectively capture the global semantic information in images, resulting in a lack of segmentation robustness for large areas. Therefore, the segmentation effects of the medical image segmentation networks U-Net and DeepLabV3+ based on CNN are limited. In the first column, UNet and DeepLabV3+ are unable to segment the boundary in the lower right area, resulting in false detection problems. In contrast, both the networks based on Transformer and Transformer+CNN can segment the normal area in the lower right, and the segmentation effect of the BCS-SegNet of the present disclosure is closer to the labeled image. In the third column, the area includes more refined edge information. The boundary supervision decoding module in BCS-SegNet can effectively extract the lesion boundary, enabling BCS-SegNet to segment the fine lesion tips, with stronger detail detection and segmentation capabilities, and can also more completely segment all lesion areas. In the fourth column, the segmentation result of BCS-SegNet includes a slender area, which is not present in the segmentation results of other algorithms. The endoscopic image in the last column includes a large amount of boundary information. Compared with the segmentation results of other algorithms, BCS-SegNet can also effectively segment the boundary detail information. Therefore, compared with other algorithms, BCS-SegNet can more accurately identify and segment detail information.

[0147] Based on the above image semantic segmentation method introducing decoupled residual attention, the present disclosure also provides an image semantic segmentation device introducing decoupled residual attention. The following will be combined with Figure 11 to describe this device in detail.

[0148] Figure 11 Schematically shows a structural block diagram of an image semantic segmentation device introducing decoupled residual attention according to an embodiment of the present disclosure.

[0149] As Figure 11 shown, the image semantic segmentation device 1000 introducing decoupled residual attention of this embodiment includes a data input module 1110, a feature fusion module 1120, and a feature decoding module 1130.

[0150] The data input module 1110 is configured to input the image data of the endoscopic image into the encoder to obtain N encoded feature maps, where N is a positive integer, N≥3, and the i-th encoded feature map is obtained by inputting the (i - 1)-th encoded feature map into the encoder, and i is a positive integer, 1 < i ≤ N.

[0151] The feature fusion module 1120 is configured to respectively perform feature fusion on the first N - 2 encoded feature maps among the N encoded feature maps and the (N - 1)-th encoded feature map to obtain N - 2 fused feature maps.

[0152] The feature decoding module 1130 is configured to input N-2 fused feature maps, the (N-1)th encoded feature map, and the Nth encoded feature map into a decoder respectively to obtain a semantic segmentation image.

[0153] According to an embodiment of the present disclosure, the feature fusion module 1120 includes an encoded feature convolution unit, an encoded feature upsampling unit, an upsampled feature convolution unit, and a feature concatenation unit.

[0154] The encoded feature convolution unit is configured to perform a convolution operation on the jth encoded feature map to obtain the jth first dimension-reduced feature map, where j is a positive integer and j ≤ N-2.

[0155] The encoded feature upsampling unit is configured to perform upsampling on the (N-1)th encoded feature map by using a bilinear interpolation upsampling method with adjacent pixel weighting to obtain an upsampled feature map.

[0156] The upsampled feature convolution unit is configured to perform a convolution operation on the upsampled feature map to obtain a second dimension-reduced feature map.

[0157] The feature concatenation unit is configured to concatenate the jth first dimension-reduced feature map and the second dimension-reduced feature map in the channel dimension to obtain the jth fused feature map.

[0158] According to an embodiment of the present disclosure, the image semantic segmentation device 1000 introducing decoupled residual attention further includes a first decoding module and a second decoding module.

[0159] The first decoding module is configured to input the jth fused feature map into a decoder to obtain a first decoded feature corresponding to the jth encoded feature map.

[0160] The second decoding module is configured to input the Nth encoded feature map and the (N-1)th encoded feature map into a decoder to obtain a second decoded feature.

[0161] According to an embodiment of the present disclosure, the first decoding module includes an amplitude determination unit and a first decoding unit.

[0162] The amplitude determination unit is configured to obtain a gradient amplitude response according to a boundary supervision decoder.

[0163] The first decoding unit is configured to obtain a first decoded feature based on the gradient amplitude response and the jth fused feature map.

[0164] According to an embodiment of the present disclosure, the feature decoding module 1130 includes a feature concatenation unit.

[0165] The feature concatenation unit is configured to hierarchically concatenate the first decoded feature and the second decoded feature to obtain a semantic segmentation image.

[0166] According to an embodiment of the present disclosure, the image semantic segmentation device 1000 introducing decoupled residual attention further includes a self-attention feature determination module and a feature update module.

[0167] The self-attention feature determination module is configured to input the shallow encoded feature map into the decoupled residual self-attention module to obtain a self-attention feature map.

[0168] The feature update module is configured to update the shallow encoded feature map using the self-attention feature map.

[0169] According to an embodiment of the present disclosure, the self-attention feature determination module includes a feature projection unit, a position vector determination unit, and a self-attention feature determination unit.

[0170] The feature projection unit is configured to project the shallow encoded feature map through a projection mapping matrix to obtain a projection vector, where the projection vector includes a query vector, a key vector, and a value vector.

[0171] The position vector determination unit is configured to determine the position vectors of the projection vector, where the position vectors include a query position vector, a key position vector, and a value position vector.

[0172] The self-attention feature determination unit is configured to obtain a self-attention feature map based on the projection vector and the position vector.

[0173] According to an embodiment of the present disclosure, the self-attention feature determination unit includes a position determination subunit, a key determination subunit, a correlation determination subunit, a weight determination subunit, and a self-attention feature determination subunit.

[0174] The position determination subunit is configured to determine query correlation according to the query vector and the query position vector.

[0175] The key determination subunit is configured to determine key correlation according to the key vector and the key position vector.

[0176] The correlation determination subunit is configured to determine vector correlation according to the query vector and the key vector.

[0177] The weight determination subunit is configured to determine self-attention weights according to the query correlation, the key correlation, and the vector correlation.

[0178] The self-attention feature determination subunit is configured to determine a self-attention feature map according to the self-attention weights, the value vector, and the value position vector.

[0179] According to embodiments of the present disclosure, any multiple of the data input module 1110, the feature fusion module 1120, and the feature decoding module 1130 may be combined and implemented in one module, or any one of them may be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules may be combined with at least part of the functions of other modules and implemented in one module. According to embodiments of the present disclosure, at least one of the data input module 1110, the feature fusion module 1120, and the feature decoding module 1130 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on a substrate, a system in a package, an application specific integrated circuit (ASIC), or any other reasonable way of integrating or packaging circuits, etc., implemented by hardware or firmware, or implemented in any one of the three implementation manners of software, hardware, and firmware, or in an appropriate combination of any several of them. Alternatively, at least one of the data input module 1110, the feature fusion module 1120, and the feature decoding module 1130 may be at least partially implemented as a computer program module, and when the computer program module is run, corresponding functions may be executed.

[0180] Figure 12 A block diagram of an electronic device suitable for implementing an image semantic segmentation method introducing decoupled residual attention according to an embodiment of the present disclosure is schematically shown.

[0181] As Figure 12 shown, the electronic device 1200 according to an embodiment of the present disclosure includes a processor 1201, which may perform various appropriate actions and processes according to a program stored in a read only memory (ROM) 1202 or a program loaded from a storage section 1208 into a random access memory (RAM) 1203. The processor 1201 may include, for example, a general microprocessor (such as a CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (such as an application specific integrated circuit (ASIC)), etc. The processor 1201 may also include on board memory for caching purposes. The processor 1201 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.

[0182] In the RAM 1203, various programs and data required for the operation of the electronic device 1200 are stored. The processor 1201, the ROM 1202, and the RAM 1203 are connected to each other via a bus 1204. The processor 1201 performs various operations of the method flow according to the embodiments of the present disclosure by executing the programs in the ROM 1202 and / or the RAM 1203. It should be noted that the programs may also be stored in one or more memories other than the ROM 1202 and the RAM 1203. The processor 1201 may also perform various operations of the method flow according to the embodiments of the present disclosure by executing the programs stored in the one or more memories.

[0183] According to an embodiment of the present disclosure, the electronic device 1200 may further include an input / output (I / O) interface 1205, and the input / output (I / O) interface 1205 is also connected to the bus 1204. The electronic device 1200 may further include one or more of the following components connected to the input / output (I / O) interface 1205: an input portion 1206 including a keyboard, a mouse, etc.; an output portion 1207 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage portion 1208 including a hard disk, etc.; and a communication portion 1209 including a network interface card such as a LAN card, a modem, etc. The communication portion 1209 performs communication processing via a network such as the Internet. A drive 1210 is also connected to the input / output (I / O) interface 1205 as needed. A removable medium 1211, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1210 as needed so that a computer program read from it can be installed into the storage portion 1208 as needed.

[0184] The present disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or may exist separately without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the one or more programs are executed, the method according to the embodiments of the present disclosure is implemented.

[0185] According to an embodiment of the present disclosure, the computer-readable storage medium may be a non-volatile computer-readable storage medium, which may include, for example, but not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present disclosure, the computer-readable storage medium may include one or more memories other than the above-described ROM 1202 and / or RAM 1203 and / or ROM 1202 and RAM 1203.

[0186] An embodiment of the present disclosure further includes a computer program product, which includes a computer program that contains program code for executing the method shown in the flowchart. When the computer program product runs in a computer system, the program code is used to enable the computer system to implement the image semantic segmentation method introducing decoupled residual attention provided by the embodiment of the present disclosure.

[0187] When the computer program is executed by the processor 1201, it executes the above functions defined in the system / apparatus of the embodiment of the present disclosure. According to an embodiment of the present disclosure, the above-described systems, apparatuses, modules, units, etc. can be implemented by computer program modules.

[0188] In one embodiment, the computer program may rely on tangible storage media such as optical storage devices and magnetic storage devices. In another embodiment, the computer program may also be transmitted and distributed in the form of a signal on a network medium, and downloaded and installed through the communication part 1209, and / or installed from the removable medium 1211. The program code included in the computer program can be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0189] In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 1209, and / or installed from the removable medium 1211. When the computer program is executed by the processor 1201, it executes the above functions defined in the system of the embodiment of the present disclosure. According to an embodiment of the present disclosure, the above-described systems, devices, apparatuses, modules, units, etc. can be implemented by computer program modules.

[0190] In accordance with embodiments of the present disclosure, program code for executing the computer programs provided by the embodiments of the present disclosure can be written in any combination of one or more programming languages. Specifically, these computing programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. The programming languages include, but are not limited to, such as Java, C++, Python, the "C" language, or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., by connecting through the Internet using an Internet service provider).

[0191] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a portion of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and combinations of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0192] Those skilled in the art can understand that the features recited in the various embodiments and / or claims of the present disclosure can be combined or combined in various ways, even if such combinations or combinations are not explicitly recited in the present disclosure. In particular, without departing from the spirit and teachings of the present disclosure, the features recited in the various embodiments and / or claims of the present disclosure can be combined and combined in various ways. All such combinations and / or combinations fall within the scope of the present disclosure.

[0193] The embodiments of the present disclosure have been described above. However, these embodiments are merely for illustrative purposes and are not intended to limit the scope of the present disclosure. Although the embodiments have been described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. The scope of the present disclosure is defined by the appended claims and their equivalents. Without departing from the scope of the present disclosure, those skilled in the art can make various substitutions and modifications, and all such substitutions and modifications should fall within the scope of the present disclosure.

Claims

1. An image semantic segmentation method introducing decoupled residual attention, comprising: Inputting the image data of the endoscopic image into an encoder to obtain N encoded feature maps, where N is a positive integer, N≥3, and the i-th encoded feature map is obtained by inputting the (i - 1)-th encoded feature map into the encoder, and i is a positive integer, 1 < i ≤ N; Feature-fusing the first N - 2 encoded feature maps among the N encoded feature maps with the (N - 1)-th encoded feature map respectively to obtain N - 2 fused feature maps; Inputting the N - 2 fused feature maps, the (N - 1)-th encoded feature map and the N-th encoded feature map into a decoder respectively to obtain a semantic segmentation image.

2. The method according to claim 1, wherein, The step of feature-fusing the first N - 2 encoded feature maps among the N encoded feature maps with the (N - 1)-th encoded feature map respectively to obtain N - 2 fused feature maps includes: Performing a convolution operation on the j-th encoded feature map to obtain the j-th first dimension-reduced feature map, where j is a positive integer, j ≤ N - 2; Performing upsampling on the (N - 1)-th encoded feature map by using a bilinear interpolation upsampling method with adjacent pixel weighting to obtain an upsampled feature map; Performing a convolution operation on the upsampled feature map to obtain a second dimension-reduced feature map; Concatenating the j-th first dimension-reduced feature map and the second dimension-reduced feature map in the channel dimension to obtain the j-th fused feature map.

3. The method according to claim 2, further comprising: Inputting the j-th fused feature map into the decoder to obtain a first decoded feature corresponding to the j-th encoded feature map; Inputting the N-th encoded feature map and the (N - 1)-th encoded feature map into the decoder to obtain a second decoded feature.

4. The method according to claim 3, wherein The decoder includes a boundary supervision decoder; The step of inputting the j-th fused feature map into the decoder to obtain a first decoded feature corresponding to the j-th encoded feature map includes: Obtaining a gradient magnitude response according to the boundary supervision decoder; Based on the gradient magnitude response and the j-th fused feature map, obtaining the first decoded feature.

5. The method according to claim 3, wherein, The step of inputting the N - 2 fused feature maps, the (N - 1)-th encoded feature map and the N-th encoded feature map into the decoder respectively to obtain a semantic segmentation image includes: Hierarchically concatenating the first decoded feature and the second decoded feature to obtain the semantic segmentation image.

6. The method according to claim 1, further comprising: Inputting the shallow encoded feature map into a decoupled residual self-attention module to obtain a self-attention feature map; Updating the shallow encoded feature map by using the self-attention feature map.

7. The method according to claim 6, wherein, The N encoded feature maps include M shallow encoded feature maps, where M is a positive integer, M < N; The step of inputting the shallow encoded feature map into a decoupled residual self-attention module to obtain a self-attention feature map includes: Projecting the shallow encoded feature map through a projection mapping matrix to obtain a projection vector, where the projection vector includes a query vector, a key vector, and a value vector; Determine the position vectors of the projection vectors, where the position vectors include a query position vector, a keyword position vector, and a value position vector; Based on the projection vectors and the position vectors, obtain the self-attention feature map.

8. The method according to claim 7, wherein The obtaining the self-attention feature map based on the projection vectors and the position vectors includes: Determine the position correlation according to the query vector and the query position vector; Determine the keyword correlation according to the keyword vector and the keyword position vector; Determine the vector correlation according to the query vector and the keyword vector; Determine the self-attention weights according to the position correlation, the keyword correlation, and the vector correlation; Determine the self-attention feature map according to the self-attention weights, the value vector, and the value position vector.

9. An image semantic segmentation device introducing decoupled residual attention, comprising: A data input module, configured to input the image data of the endoscopic image into an encoder to obtain N encoded feature maps, where N is a positive integer, N≥3, and the i-th encoded feature map is obtained by inputting the (i - 1)-th encoded feature map into the encoder, and i is a positive integer, 1 < i ≤ N; A feature fusion module, configured to respectively perform feature fusion on the first N - 2 encoded feature maps among the N encoded feature maps and the (N - 1)-th encoded feature map to obtain N - 2 fused feature maps; A feature decoding module, configured to respectively input the N - 2 fused feature maps, the (N - 1)-th encoded feature map, and the N-th encoded feature map into a decoder to obtain a semantic segmentation image.

10. An electronic device, comprising: One or more processors; A memory, configured to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 8.