Skin wound image segmentation method, device, equipment and storage medium

CN113947574BActive Publication Date: 2025-05-23SUZHOU YUANHE CHUANGDA INTELLIGENT TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111171811.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-08
Publication Date
2025-05-23
Estimated Expiration
2041-10-08

AI Technical Summary

Technical Problem

In the prior art, only multi-scale feature information of the segmented image is obtained, resulting in poor results of the segmented image.

Method used

The skin wound image segmentation method based on the spatial attention mechanism and the channel attention mechanism is adopted to perform edge enhancement and spatial enhancement processing on the segmented image to be segmented, and the final segmented image is obtained through channel stitching.

Benefits of technology

By enhancing edge features and spatial relationship features, the quality and accuracy of segmented images are significantly improved and a better segmentation effect is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113947574B_ABST
    Figure CN113947574B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of image processing technology, and specifically to a method, device, equipment and storage medium for skin wound image segmentation, comprising the following steps: obtaining an image to be segmented containing a skin wound; performing edge enhancement processing on the image to be segmented based on a spatial attention mechanism and a channel attention mechanism to obtain a feature map after first edge feature enhancement, and performing spatial enhancement processing on the image to be segmented to obtain a feature map after first spatial feature enhancement; performing channel splicing on the feature map after the first edge feature enhancement and the feature map after the first spatial feature enhancement to obtain a segmented image. Since edge features contribute particularly important to segmentation, and the positions of objects in an image and the spatial relationships between objects are very important features in image segmentation, channel splicing using feature maps enhanced with edge features and feature maps enhanced with spatial features can make the final segmented image better.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular to a skin wound image segmentation method, device, equipment and storage medium. Background Art

[0002] Image segmentation refers to dividing an image into several non-overlapping regions based on features such as grayscale, color, texture, and shape, and making these features appear similar within the same region and significantly different between different regions.

[0003] Medical image segmentation methods are divided into traditional image segmentation methods and deep learning image segmentation methods. Traditional image segmentation methods include threshold-based segmentation, region-based segmentation, edge-based segmentation, and graph-based segmentation. With the rapid development of neural networks, deep learning image segmentation methods have been widely used in recent years, and have higher accuracy and generalization than traditional image segmentation methods.

[0004] After the introduction of the fully connected network, deep learning technology has been widely used in image segmentation. However, there is a problem with applying convolutional neural networks to image segmentation - the pooling layer loses the accuracy of position information while increasing the field of view. Later, with the introduction of network structures represented by U-Net, this problem was effectively solved, which in turn promoted the development of deep learning in image segmentation. U-Net network and its derivative networks are widely used in medical image segmentation due to their good feature extraction capabilities.

[0005] Skin accounts for about 16% of the body and is the largest organ in the human body. It contains a complex network of nerves and blood vessels. As the first line of defense of the human body, it covers the entire body surface and is easily damaged by physical, mechanical, biological and chemical factors. Skin injuries can be divided into acute wounds and chronic wounds according to the healing cycle of the wound. Acute wounds usually refer to traumas such as contusions and cuts, such as knife wounds, abrasions, gunshot wounds, chemical injuries, etc.; chronic wounds include lower limb venous ulcers, pressure ulcers, arterial ischemic ulcers and diabetic ulcers. Skin wound images have the following characteristics: a. The background interference in the image is serious; b. The color, shape, area and edge features of different wounds are quite different; c. The images are taken by real-life cameras and mobile phones, with different light, resolution, shooting angle and shooting quality. It can be seen that the segmentation of skin wound areas is extremely challenging.

[0006] The prior art discloses a method for automatically segmenting skin wounds using a lightweight network, which adds a spatial pyramid pooling module to the decoder to capture multi-scale feature information. Compared with traditional convolutional neural networks, this network replaces traditional convolutional layers with depthwise separable convolutional layers, reducing the computational cost by nearly k. 2times, where k is the convolution kernel size. Although this method has significantly improved computational efficiency, it only obtains multi-scale feature information of the segmented image, resulting in poor segmentation of the image. Summary of the invention

[0007] Therefore, the present invention aims to solve the technical problem that in the prior art, only the multi-scale feature information of the segmented image is obtained, resulting in poor segmented image effect, and thus provides a skin wound image segmentation method, comprising the following steps: obtaining an image to be segmented containing a skin wound;

[0008] Performing edge enhancement processing on the image to be segmented based on the spatial attention mechanism and the channel attention mechanism to obtain a feature map after the first edge feature is enhanced, and performing spatial enhancement processing on the image to be segmented to obtain a feature map after the first spatial feature is enhanced;

[0009] Channel stitching is performed on the feature map after the first edge feature enhancement and the feature map after the first spatial feature enhancement to obtain a first segmented image.

[0010] Preferably, the edge enhancement processing is performed on the image to be segmented based on the spatial attention mechanism and the channel attention mechanism to obtain a feature map after the first edge feature is enhanced, including:

[0011] Performing a first encoding and a first maximum pooling on the image to be segmented to obtain a first feature map, performing a second encoding and a second maximum pooling on the first feature map to obtain a second feature map, and performing a first edge enhancement process on the first feature map and the second feature map to obtain a first edge feature enhanced map;

[0012] Performing the m-1th edge feature enhancement process on the m+1th feature map to obtain the mth edge feature enhancement map; wherein the m+1th feature map is obtained by performing the m+1th encoding on the mth feature map obtained by the mth encoding and the mth maximum pooling, where m=2, 3, ... N;

[0013] When m=N, the obtained Nth edge feature enhancement image is the feature image after the first edge feature is enhanced.

[0014] Preferably, the edge enhancement processing is performed on the image to be segmented based on the spatial attention mechanism and the channel attention mechanism to obtain a feature map after the first edge feature is enhanced, including:

[0015] Performing a first edge enhancement process on the first feature map and the second feature map based on a spatial attention mechanism to obtain a first edge feature enhanced map;

[0016] Based on the spatial attention mechanism, a second edge enhancement process is performed on the first edge feature enhancement map and the third feature map to obtain a second edge feature enhancement map;

[0017] Based on the channel attention mechanism, the second edge feature enhancement map and the fourth feature map are subjected to a third edge enhancement process to obtain a third edge feature enhancement map;

[0018] Based on the channel attention mechanism, a fourth edge enhancement process is performed on the third edge feature enhancement map and the fifth feature map to obtain a feature map after the first edge feature is enhanced.

[0019] Preferably, performing an m-th edge enhancement process on the m-1th edge feature enhancement map and the m+1th feature map to obtain the mth edge feature enhancement map comprises:

[0020] Perform convolution and encoding processing on the m-1th edge feature enhancement map to obtain the m-1th edge feature encoding map, and perform convolution and up-sampling processing on the m+1th feature map to obtain the m+1th feature sampling map;

[0021] Perform channel splicing on the m-1th edge feature encoding map and the m+1th feature sampling map to obtain an mth edge feature splicing map; process the mth edge feature splicing map based on a spatial attention mechanism to obtain a spatial weight;

[0022] After the spatial weight passes through the activation function, it is multiplied with the m-1th edge feature encoding map to obtain the mth edge feature enhancement map.

[0023] Preferably, the processing of the mth edge feature splicing graph based on the spatial attention mechanism to obtain the spatial weight includes:

[0024] Performing average pooling on the mth edge feature splicing map to obtain a first pooling feature map, and performing maximum pooling on the mth edge feature splicing map to obtain a second pooling feature map; performing channel splicing on the first pooling feature map and the second pooling feature map to obtain a spliced ​​mth edge feature pooling map;

[0025] The mth edge feature pooling map is convolved to obtain a spatial weight.

[0026] Preferably, performing an m-th edge enhancement process on the m-1th edge feature enhancement map and the m+1th feature map to obtain the mth edge feature enhancement map comprises:

[0027] Perform convolution and encoding processing on the m-1th edge feature enhancement map to obtain the m-1th edge feature encoding map, and perform convolution and up-sampling processing on the m+1th feature map to obtain the m+1th feature sampling map;

[0028] Perform channel splicing on the m-1th edge feature encoding map and the m+1th feature sampling map to obtain an mth edge feature splicing map; process the mth edge feature splicing map based on a channel attention mechanism to obtain a channel weight;

[0029] After the channel weight passes through the activation function, it is multiplied with the m-1th edge feature encoding map to obtain the mth edge feature enhancement map.

[0030] Preferably, the processing of the mth edge feature splicing map based on the channel attention mechanism to obtain the channel weight includes: performing average pooling on the mth edge feature splicing map to obtain a third pooling feature map, and performing maximum pooling on the mth edge feature splicing map to obtain a fourth pooling feature map;

[0031] The third pooling feature map and the fourth pooling feature map are added together after passing through a multi-layer perceptron to obtain the channel weight.

[0032] Preferably, the image to be segmented is subjected to spatial enhancement processing to obtain a feature map after the first spatial feature is enhanced, including: when m=N, performing spatial enhancement processing on the m+1th feature map to obtain the m+1th spatial feature enhanced map; decoding the m+1th spatial feature enhanced map to obtain the feature map after the first spatial feature is enhanced.

[0033] Preferably, when m=N, performing spatial enhancement processing on the m+1th feature map to obtain the m+1th spatial feature enhancement map includes: performing convolution recombination on the m+1th feature map to obtain two feature matrices of different sizes;

[0034] Multiply the two feature matrices of different sizes, and reorganize them again to obtain the m+1th matrix feature map;

[0035] The m+1th matrix feature map and the m+1th feature map are channel-joined and linearly rectified to obtain the m+1th spatial feature enhancement map.

[0036] The present invention also provides an interactive skin wound image segmentation method, comprising the following steps:

[0037] The image to be segmented, the first Gaussian distance map and the second Gaussian distance map are combined to obtain a distance image to be segmented; wherein the first Gaussian distance map is a Gaussian distance map corresponding to the missed mark, and the second Gaussian distance map is a Gaussian distance map corresponding to the over-marked mark;

[0038] Performing edge enhancement processing on the distance image to be segmented based on a spatial attention mechanism and a channel attention mechanism to obtain a feature map after second edge feature enhancement;

[0039] The third Gaussian distance map is convolved to obtain a convolution segmentation map, the N+1th feature map and the convolution segmentation map are channel-joined to obtain a spliced ​​convolution map, multi-scale hole convolution and global pooling are respectively performed on the spliced ​​convolution map, and the convolution output is convolved to obtain a convolution pooling feature map; wherein the third Gaussian distance map is a Gaussian distance map corresponding to the center point of the first segmentation image, and the first segmentation image is obtained by using a skin wound image segmentation method;

[0040] Performing spatial enhancement processing on the convolutional pooling feature map to obtain a feature map after second spatial feature enhancement;

[0041] Channel stitching is performed on the feature map after the second edge feature is enhanced and the feature map after the second spatial feature is enhanced to obtain a second segmented image.

[0042] Preferably, the first Gaussian distance map is calculated by a first mathematical model, and the first mathematical model is:

[0043]

[0044]

[0045] Where D(x,y) is the shortest Euclidean distance from point (x,y) to all missed marked points;

[0046] And / or, the second Gaussian distance map is calculated by a second mathematical model, and the second mathematical model is:

[0047]

[0048]

[0049] Where D(m,n) is the shortest Euclidean distance from point (m,n) to all multi-point marker points;

[0050] The third Gaussian distance map is calculated by a third mathematical model, and the third mathematical model is:

[0051]

[0052]

[0053] Where D(p,q) is the shortest Euclidean distance from point (p,q) to the center point of the first segmented image.

[0054] The present invention also provides a skin wound image segmentation device, comprising:

[0055] An acquisition module, used for acquiring an image to be segmented including a skin wound surface;

[0056] A first processing module is used to perform edge enhancement processing on the image to be segmented based on a spatial attention mechanism and a channel attention mechanism to obtain a feature map after first edge feature enhancement, and to perform spatial enhancement processing on the image to be segmented to obtain a feature map after first spatial feature enhancement;

[0057] The first stitching module is used to perform channel stitching on the feature map after the first edge feature is enhanced and the feature map after the first spatial feature is enhanced to obtain a first segmented image.

[0058] The present invention also provides an interactive skin wound image segmentation device, comprising:

[0059] A second stitching module is used to stitch the image to be segmented, the first Gaussian distance map and the second Gaussian distance map to obtain a distance image to be segmented; wherein the first Gaussian distance map is a Gaussian distance map corresponding to the missed mark, and the second Gaussian distance map is a Gaussian distance map corresponding to the over-marked mark;

[0060] A second processing module is used to perform edge enhancement processing on the to-be-segmented distance image based on a spatial attention mechanism and a channel attention mechanism to obtain a feature map after second edge feature enhancement;

[0061] A convolution and pooling module is used to convolve the third Gaussian distance map to obtain a convolution segmentation map, channel-splicing the N+1th feature map and the convolution segmentation map to obtain a spliced ​​convolution map, and performing multi-scale hole convolution and global pooling on the spliced ​​convolution map, splicing the outputs and obtaining a convolution and pooling feature map after convolution; wherein the third Gaussian distance map is a Gaussian distance map corresponding to the center point of the first segmented image, and the first segmented image is obtained by using a skin wound image segmentation method; a third processing module is used to perform spatial enhancement processing on the convolution and pooling feature map to obtain a feature map after second spatial feature enhancement; a third splicing module is used to channel-splicing the feature map after the second edge feature enhancement and the feature map after the second spatial feature enhancement to obtain a second segmented image.

[0062] The present invention also provides a computer device, comprising: a memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the skin wound image segmentation method or the interactive skin wound image segmentation method by executing the computer instructions.

[0063] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute a skin wound image segmentation method or an interactive skin wound image segmentation method.

[0064] The technical solution of the present invention has the following advantages:

[0065] The skin wound image segmentation method provided by the present invention performs edge enhancement processing on the image to be segmented containing skin wounds based on the spatial attention mechanism and the channel attention mechanism, obtains the feature map after the first edge feature enhancement, performs spatial enhancement processing on the image to be segmented containing skin wounds, obtains the feature map after the first spatial feature enhancement, and performs channel splicing on the feature map after the first edge feature enhancement and the feature map after the first spatial feature enhancement to obtain the first segmentation image. Since the edge features are particularly important for segmentation, and the positions of the objects in the image and the spatial relationships between the objects are very important features in image segmentation, using the feature map enhanced by edge features and the feature map enhanced by spatial features for channel splicing can make the final formed segmentation image have a better effect. Brief Description of the Drawings

[0066] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the prior art. Obviously, the following-described drawings are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0067] Figure 1 It is a flowchart of the skin wound image segmentation method according to Embodiment 1 of the present invention;

[0068] Figure 2 It is a flowchart of calculating the feature map after the first edge feature enhancement in Embodiment 1 of the present invention;

[0069] Figure 3 For Figure 2 One of the flowcharts of step S202 in

[0070] Figure 4 For Figure 2 Another flowchart of step S202 in

[0071] Figure 5 It is a flowchart of calculating the feature map after the first spatial feature enhancement in Embodiment 1 of the present invention;

[0072] Figure 6 It is the network structure diagram of the skin wound image segmentation method according to Embodiment 1 of the present invention;

[0073] Figure 7 For Figure 6 The structural schematic diagram of the encoder in

[0074] Figure 8 For Figure 6 The structural schematic diagram of the decoder in

[0075] Fig. 9 for Figure 3 and Figure 4 Schematic diagram of the workflow;

[0076] Fig.10 for Figure 5 Schematic diagram of the workflow;

[0077] Fig.11 This is a flow chart of an interactive skin wound image segmentation method according to Embodiment 2 of the present invention;

[0078] Fig.12 This is a network structure diagram of the interactive skin wound image segmentation method according to Embodiment 2 of the present invention;

[0079] Fig.13 for Fig.12 Schematic diagram of the workflow of the spatial convolutional pooling pyramid module;

[0080] Fig.14 This is a structural block diagram of a skin wound image segmentation device according to Embodiment 3 of the present invention;

[0081] Fig.15 This is a structural block diagram of an interactive skin wound image segmentation device according to Embodiment 4 of the present invention;

[0082] Fig.16 This is a schematic structural diagram of a computer device according to Embodiment 5 of the present invention. DETAILED DESCRIPTION

[0083] The technical solution of the present invention will be described clearly and completely below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0084] In the description of the present invention, it should be noted that the terms "first", "second" and "third" are used for descriptive purposes only and cannot be understood as indicating or implying relative importance. In addition, the technical features involved in the different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0085] Skin is the first line of defense of the human body and is spread all over the human body. It is easy to be damaged by various accidents. Skin wound image segmentation is an important part of wound diagnosis and follow-up care. Before treating the patient, the surgeon can analyze the segmented image of the skin wound and provide quantitative parameters by measuring the wound area, which provides a basic guarantee for the subsequent treatment of the wound.

[0086] The prior art discloses a method for automatically segmenting skin wounds using a lightweight network. Although this method has a significant improvement in computational efficiency, it only obtains multi-scale feature information of the image to be segmented, resulting in poor results in the segmented image.

[0087] Example 1

[0088] This embodiment provides a skin wound image segmentation method. Figure 1 The flowchart is a flowchart illustrating how to perform edge enhancement and spatial enhancement processing on an image to be segmented containing a skin wound surface according to some embodiments of the present invention to finally obtain a segmented image. Although the process described below includes multiple operations that appear in a specific order, it should be clearly understood that these processes may also include more or fewer operations, and these operations may be performed sequentially or in parallel (for example, using a parallel processor or a multi-threaded environment).

[0089] This embodiment provides a skin wound image segmentation method for segmenting an image containing skin damage. The skin wound image segmentation method can be executed by a terminal device (hardware and software), such as a mobile phone, a computer, a cloud server, and a tablet computer. Figure 1 As shown, the skin wound image segmentation method comprises the following steps:

[0090] S101, obtaining an image to be segmented including a skin wound surface.

[0091] In the above implementation steps, the skin wound surface may be a skin wound caused by a knife wound, a gunshot wound, a chemical injury, a pressure ulcer, a diabetic ulcer, etc. The image to be segmented may be an image obtained by photographing the skin wound using an imaging device such as a camera, a mobile phone, and a video camera, or may be a frame image of a skin wound surface in a video.

[0092] S102. Perform edge enhancement processing on the image to be segmented based on the spatial attention mechanism and the channel attention mechanism to obtain a feature map after the first edge feature is enhanced, and perform spatial enhancement processing on the image to be segmented to obtain a feature map after the first spatial feature is enhanced.

[0093] In the above implementation steps, a pre-trained neural network model can be used for processing. Before performing image segmentation on the image to be segmented, the neural network model needs to be trained, and images containing skin wounds are collected and labeled to form a training set and a validation set. The training set is used to train the neural network model, and the validation set is used to validate the trained neural network model. The neural network with the accuracy that meets the requirements is selected for subsequent image segmentation. Figure 6-Figure 8As shown, the neural network model can be a fully convolutional network based on an encoder-decoder structure, wherein the encoder can be a ResNet34 structure, and the decoder and the encoder are connected by a jump connection to form a U-shaped neural network.

[0094] In the segmentation of skin wound images, the contribution of edge features to segmentation is particularly important and is an important basis for separating wound areas. Among them, edge features refer to the position, color, texture and other information of the wound edge in the image to be segmented. Based on the spatial attention mechanism and the channel attention mechanism, the edge enhancement processing is performed on the segmented image to obtain significant edge features and obtain the feature map after the first edge feature is enhanced.

[0095] Analysis of the skin wound image shows that the target wound area to be segmented is surrounded by the normal skin area. The location of the object in the image and the spatial relationship between the objects are very important features in image segmentation. The segmented image is spatially enhanced to allow the neural network model to learn and infer the global spatial relationship features of the target area, and obtain the feature map after the first spatial feature enhancement.

[0096] S103: performing channel stitching on the feature map after the first edge feature is enhanced and the feature map after the first spatial feature is enhanced to obtain a first segmented image.

[0097] In the above implementation steps, the first segmented image is obtained by channel splicing the feature map after the first edge feature is enhanced and the feature map after the first spatial feature is enhanced. Since the edge features and spatial relationship features in the feature map are enhanced, the first segmented image obtained by the final channel splicing can be better. In some embodiments, after the channel splicing, the first segmented image can be obtained by performing operations such as deconvolution and activation function (sigmoid).

[0098] In the above embodiment, the image to be segmented containing the skin wound surface is subjected to edge enhancement processing based on the spatial attention mechanism and the channel attention mechanism to obtain a feature map after the first edge feature is enhanced, the image to be segmented containing the skin wound surface is subjected to spatial enhancement processing to obtain a feature map after the first spatial feature is enhanced, and the feature map after the first edge feature enhancement and the feature map after the first spatial feature enhancement are channel-joined to obtain a first segmented image. Since the contribution of edge features to segmentation is particularly important, and the positions of objects in the image and the spatial relationships between objects are very important features in image segmentation, channel-joining using feature maps enhanced with edge features and feature maps enhanced with spatial features can make the final segmented image better.

[0099] In one or more embodiments, multiple encoders are connected in sequence in the neural network model, and an edge feature enhancement module can be added to the encoder part. The multiple edge feature enhancement modules are connected in sequence, and the output end of the encoder is connected to the input end of the edge feature enhancement module. The edge feature enhancement module is used to extract significant edge features from the feature map extracted by the encoder to obtain a feature map after the first edge feature is enhanced. Figure 2 As shown, the edge enhancement processing of the image to be segmented based on the spatial attention mechanism and the channel attention mechanism to obtain a feature map after the first edge feature is enhanced includes the following steps:

[0100] S201. Perform a first encoding and a first maximum pooling on the image to be segmented to obtain a first feature map, perform a second encoding and a second maximum pooling on the first feature map to obtain a second feature map, and perform a first edge enhancement process on the first feature map and the second feature map to obtain a first edge feature enhanced map.

[0101] In the above implementation steps, the neural network model includes a plurality of encoders connected in sequence, that is, the output end of the previous encoder is connected to the input end of the next encoder. The structure of the encoder can be as follows: Figure 7 As shown, the encoder 601 can be used to encode the segmented image, the first feature map, etc. After maximum pooling, encoding is performed to obtain the mth feature map. For example, the segmented image is subjected to the first maximum pooling and then the first encoding is performed to obtain the first feature map.

[0102] Among the plurality of edge feature enhancement modules connected in sequence, the input end of the first edge feature enhancement module is connected to the output ends of the two encoders. The first feature map and the second feature map are input into the first edge feature enhancement module, and the edge feature enhancement module obtains significant edge features from the first feature map and the second feature map to obtain a first edge feature enhancement map. In some embodiments, before the edge feature enhancement module processes the first feature map and the second feature map, the first feature map may be subjected to upsampling processing.

[0103] S202, performing an m-1th edge enhancement process on the m-1th edge feature enhancement map and the m+1th feature map to obtain an mth edge feature enhancement map; wherein the m+1th feature map is obtained by performing the m+1th encoding on the mth feature map obtained by the mth encoding and the mth maximum pooling, and m=2, 3, ...N.

[0104] In the above implementation steps, multiple edge feature enhancement modules are connected in sequence, that is, the output end of the previous edge feature enhancement module is connected to the input end of the next edge feature enhancement module. Except for the first edge feature enhancement module, the input end of the edge feature enhancement module is also connected to the output end of the corresponding encoder.

[0105] It should be noted that multiple encoders are connected in sequence in the neural network model, and the m+1th feature map is obtained by performing the m+1th encoding on the mth feature map obtained by the mth encoding and the mth maximum pooling, that is, the output of the previous encoder is the input of the next encoder. For example, the neural network model includes 5 encoders connected in sequence, the first encoder performs the first encoding on the image to be segmented after the maximum pooling to obtain the first feature map, the second encoder performs the second encoding on the first feature map after the maximum pooling to obtain the second feature map, the third encoder performs the third encoding on the second feature map after the maximum pooling to obtain the third feature map, the fourth encoder performs the fourth encoding on the third feature map after the maximum pooling to obtain the fourth feature map, and the fifth encoder encodes the fourth feature map after the maximum pooling to obtain the fifth feature map.

[0106] S203: When m=N, the obtained Nth edge feature enhancement image is the feature image after the first edge feature is enhanced.

[0107] In the above implementation steps, after the last edge feature enhancement module performs the Nth edge enhancement processing on the N-1th edge feature enhancement map and the N+1th feature map, the output Nth edge feature enhancement map is the feature map after the first edge feature enhancement.

[0108] For example, if Figure 6 As shown, four edge feature enhancement modules are connected in sequence, the image to be segmented 101 is first encoded to obtain a first feature map 102, and the first feature map 102 is second encoded to obtain a second feature map 103. The first feature map 102 and the second feature map 103 are input to the first edge feature enhancement module 201, and the first edge feature enhancement module 201 performs a first edge enhancement process on the first feature map 102 and the second feature map 103 to obtain a first edge feature enhancement map.

[0109] The second feature map 103 is encoded for the third time to obtain the third feature map 104, and the first edge feature enhancement map and the third feature map 104 are input into the second edge feature enhancement module 202. The second edge feature enhancement module 202 performs a second edge enhancement process on the first edge feature enhancement map and the third feature map 104 to obtain the second edge feature enhancement map.

[0110] The third feature map 104 is encoded for the fourth time to obtain the fourth feature map 105, and the second edge feature enhancement map and the fourth feature map 105 are input into the third edge feature enhancement module 203. The third edge feature enhancement module 203 performs a third edge enhancement process on the second edge feature enhancement map and the fourth feature map 105 to obtain the third edge feature enhancement map.

[0111] The fourth feature map 105 is encoded for the fifth time to obtain the fifth feature map, and the third edge feature enhancement map and the fifth feature map are input into the fourth edge feature enhancement module 204. The fourth edge feature enhancement module 204 performs the fourth edge enhancement processing on the third edge feature enhancement map and the fifth feature map to obtain the fourth edge feature enhancement map, which is the feature map after the first edge feature enhancement.

[0112] Among them, the first edge feature enhancement module 201 and the second edge feature enhancement module 202 can perform edge enhancement processing based on the spatial attention mechanism; the third edge feature enhancement module 203 and the fourth edge feature enhancement module 204 can perform edge enhancement processing based on the channel attention mechanism. The boundaries of low-level features are relatively obvious, and there is almost no semantic difference between different channels. The spatial attention mechanism can better focus on effective low-level features and obtain significant boundaries; while high-level features contain higher abstract semantics, the channel attention mechanism can give greater weight to the channels that play an important role in segmentation, thereby achieving better image segmentation effects.

[0113] It should be noted that those skilled in the art can reasonably select the number of edge feature enhancement modules according to actual conditions, for example, 5, 6, 7 or more. At the same time, those skilled in the art can also choose a spatial attention mechanism or a channel attention mechanism according to actual conditions. In some embodiments, the image to be segmented is also subjected to a convolution operation before the first encoding. For example, a convolution layer with a convolution kernel of 7×7 (Conv7×7) and a stride of 2 (stride2) is used to convolve the image to be segmented.

[0114] In one or more embodiments, Figure 3 As shown, the m-1th edge feature enhancement map and the m+1th feature map are subjected to the mth edge enhancement process to obtain the mth edge feature enhancement map, including the following steps:

[0115] S301, convolve and encode the m-1th edge feature enhancement map to obtain the m-1th edge feature encoding map, and convolve and upsample the m+1th feature map to obtain the m+1th feature sampling map.

[0116] In the above implementation steps, a convolution layer is used to convolve the m-1th edge feature enhancement map and the m+1th feature map. For example, a convolution layer with a convolution kernel of 1×1 (Conv 1×1) can be used for convolution. Figure 7 As shown, encoder 601 may be used to encode the m-1th edge feature enhancement map.

[0117] S302: Perform channel splicing on the m-1th edge feature coding map and the m+1th feature sampling map to obtain an mth edge feature splicing map.

[0118] In the above implementation steps, the m-1th edge feature encoding map and the m+1th feature sampling map are channel-joined to obtain the mth edge feature joint map.

[0119] S303: Process the mth edge feature splicing graph based on a spatial attention mechanism to obtain a spatial weight.

[0120] In the above implementation steps, the mth edge feature mosaic image is processed to obtain the spatial weight, including the following steps:

[0121] S3031, average pooling is performed on the mth edge feature splicing map to obtain a first pooling feature map, and maximum pooling is performed on the mth edge feature splicing map to obtain a second pooling feature map; S3032, channel splicing is performed on the first pooling feature map and the second pooling feature map to obtain a spliced ​​mth edge feature pooling map; S3033, convolution is performed on the mth edge feature pooling map to obtain a spatial weight. For example, a convolution layer with a convolution kernel of 7×7 (Conv7×7) can be used to convolve the mth edge feature pooling map to obtain a spatial weight.

[0122] S304: After the spatial weight passes through the activation function, it is multiplied with the m-1th edge feature encoding map to obtain the mth edge feature enhancement map.

[0123] In the above implementation steps, after the spatial weight passes through the activation (Sigmoid) function, it is multiplied with the m-1th edge feature encoding map to obtain the mth edge feature enhancement map. When the value of m is N, the mth edge feature enhancement map is the feature map after the first edge feature is enhanced. Among them, the multiplication with the m-1th edge feature encoding map is a pixel-level multiplication.

[0124] It should be noted that in the above steps S301 to S304, m is a natural number greater than or equal to 2, that is, in the edge feature enhancement modules connected in sequence, the input of the first edge feature enhancement module is different from the input of the subsequent edge feature enhancement module, and the input of the first edge feature enhancement module is the first feature map processed by the first encoder and the second feature map processed by the second encoder. The first edge feature enhancement module performs subsequent processing such as convolution and encoding on the first feature map, and performs subsequent processing such as convolution and upsampling on the second feature map.

[0125] like Figure 6 As shown, the m-1th edge feature enhancement map and the first feature map input to the edge feature enhancement module are called edge features, and the mth feature map input to the edge feature enhancement module is called segmentation feature. The size of the edge feature can be C e ×H×W, where H and W are the height and width of the image to be segmented, which can be 512×512; the number of channels Ce In the first edge feature enhancement module 201, the second edge feature enhancement module 202, the third edge feature enhancement module 203 and the fourth edge feature enhancement module 204, the size of the segmentation feature is C s ×H s ×W s , where the number of channels is C s In the first edge feature enhancement module 201, the second edge feature enhancement module 202, the third edge feature enhancement module 203 and the fourth edge feature enhancement module 204, the H s and W s The value of H is the same. s and W s In the first edge feature enhancement module 201, the second edge feature enhancement module 202, the third edge feature enhancement module 203 and the fourth edge feature enhancement module 204, the values ​​may be 128, 64, 32 and 16 respectively.

[0126] The edge features are convolved and encoded to obtain a size of C n ×H×W edge feature encoding map, where the number of channels is C n In the first edge feature enhancement module 201, the second edge feature enhancement module 202, the third edge feature enhancement module 203 and the fourth edge feature enhancement module 204, the number of pixels can be 64, 32, 16 and 8 respectively. The segmentation feature is convolved and up-sampled to obtain a feature sampling map with a size of 1×H×W. The channel splices the edge feature encoding map and the feature sampling map to obtain a size of (C n +1)×H×W edge feature splicing map.

[0127] For example, if Figure 6 and Fig. 9 As shown, the input of the second edge feature enhancement module 202 is the first edge feature enhancement map 106 and the third feature map 104, wherein the first edge feature enhancement map 106 is obtained by the first edge feature enhancement module 201 performing edge enhancement processing on the first feature map 102 and the second feature map 103. The first edge feature enhancement map 106 is convolved and encoded to obtain a first edge feature encoding map 1061, and the third feature map 104 is convolved and up-sampled to obtain a third feature sampling map 1041. The first edge feature encoding map 1061 and the third feature sampling map 1041 are channel-concatenated to obtain a second edge feature splicing map 107. The second edge feature splicing map 107 is processed using a spatial attention mechanism to obtain a spatial weight. The spatial weight is multiplied pixel-wise with the first edge feature encoding map 1061 after an activation function to obtain the first edge feature enhancement map.

[0128] Among them, the size of the second edge feature mosaic image 107 is (C n +1)×H×W, average pooling is performed on the second edge feature splicing map 107 to obtain the first pooling feature map, and maximum pooling is performed on the second edge feature splicing map 107 to obtain the second pooling feature map. The sizes of the first pooling feature map and the second pooling feature map are both 1×H×W. The first pooling feature map and the second pooling feature map are channel-concatenated to obtain a second edge feature pooling map with a size of 2×H×W. Finally, Conv7×7 can be used to convolve the second edge feature pooling map to obtain a spatial weight of 1×H×W.

[0129] In one or more embodiments, Figure 4 As shown, the m-1th edge feature enhancement map and the m+1th feature map are subjected to the mth edge enhancement process to obtain the mth edge feature enhancement map, including the following steps:

[0130] S401, convolve and encode the m-1th edge feature enhancement map to obtain the m-1th edge feature encoding map, and convolve and upsample the m+1th feature map to obtain the m+1th feature sampling map. For details, please refer to the relevant description of step S301, which will not be repeated here.

[0131] S402: Channel splicing is performed on the m-1th edge feature coding map and the m+1th feature sampling map to obtain the mth edge feature splicing map. For details, please refer to the relevant description of step S302, which will not be repeated here.

[0132] S403: Process the mth edge feature splicing graph based on a channel attention mechanism to obtain a channel weight.

[0133] In the above implementation steps, processing the mth edge feature splicing graph to obtain the channel weight may include the following steps:

[0134] S4031. Perform average pooling on the mth edge feature splicing map to obtain a third pooling feature map, and perform maximum pooling on the mth edge feature splicing map to obtain a fourth pooling feature map.

[0135] S4032: After passing through a multi-layer perceptron, the third pooling feature map and the fourth pooling feature map are added to obtain the channel weight, wherein the addition refers to pixel-level addition.

[0136] Among them, the multi-layer perceptron (MLP) can be composed of two convolutional layers with a convolution kernel of 1×1 and a ReLU activation function. Fig. 9 As shown, the size is (C n +1)×H×W of the mth edge feature concatenation image is average pooled to obtain a size of (C n+1)×1×1 third pooling feature map, after maximum pooling, the size is (C n +1)×1×1 fourth pooling feature map, the third pooling feature map and the fourth pooling feature map in the multi-layer perceptron, after convolution, the number of channels becomes 1 / 8 of the original, and the number of channels is restored to C after the activation function and convolution operation. n After passing through the multi-layer perceptron, the third pooling feature map and the fourth pooling feature map are added at the pixel level to obtain a size of C n ×1×1 feature weights. When two pooling operations are used simultaneously, the two aggregated channel features are located in the same semantic space. Therefore, a shared multi-layer perceptron is used for attention reasoning to save parameters and obtain the correlation between the two channels.

[0137] S404: After the channel weight passes through the activation function, it is multiplied with the m-1th edge feature encoding map to obtain the mth edge feature enhancement map.

[0138] In the above implementation steps, after the channel weight passes through the activation (Sigmoid) function, it is multiplied with the m-1th edge feature encoding map to obtain the mth edge feature enhancement map. When the value of m is N, the mth edge feature enhancement map is the feature map after the first edge feature is enhanced. Among them, the multiplication with the m-1th edge feature encoding map is a pixel-level multiplication.

[0139] For example, if Figure 6 and Fig. 9 As shown, the first edge feature encoding map 1061 and the third feature sampling map 1041 are channel-spliced ​​to obtain the second edge feature splicing map 107, and the second edge feature splicing map 107 is processed by the channel attention mechanism to obtain the channel weight, and the channel weight is multiplied with the first edge feature encoding map 1061 at the pixel level after passing through the activation function to obtain the second edge feature enhancement map.

[0140] In one or more embodiments, Figure 5 As shown, the image to be segmented is subjected to spatial enhancement processing to obtain a feature map after the first spatial feature is enhanced, comprising the following steps:

[0141] S501. When m=N, perform spatial enhancement processing on the m+1th feature map to obtain the m+1th spatial feature enhanced map.

[0142] In the above implementation steps, if Figure 6 As shown, the input end of the spatial relationship feature enhancement module 701 is connected to the output end of the last encoder, and the output end is connected to the input end of the first decoder. Performing spatial enhancement processing on the m+1th feature map may include the following steps:

[0143] S5011. Perform convolution reorganization on the m+1th feature map to obtain two feature matrices of different sizes; S5012. Multiply the two feature matrices of different sizes, and reorganize again to obtain the m+1th matrix feature map; S5013. Channel-join the m+1th matrix feature map and the m+1th feature map, and perform linear rectification to obtain the m+1th spatial feature enhancement map.

[0144] like Fig.10 As shown, the size of the m+1th feature map 301 is C r ×H r ×W r , for example, C r 512, H r and W r The m+1th feature map 301 is subjected to two Conv1×1 convolutions and reorganized to obtain H r W r ×C r and C r ×H r W r The feature matrix of size is multiplied pixel by pixel to obtain H r W r ×H r W r After the feature matrix is ​​obtained, the m+1th matrix feature map 302 is reorganized again, where the size of the m+1th matrix feature map 302 is H r W r ×H r ×W r The m+1th matrix feature map 302 and the m+1th feature map 301 are concatenated by the channel, and a ReLU function is used for linear rectification to eliminate negative spatial relationships, and finally the m+1th spatial feature enhancement map 303 is obtained. The size of the m+1th spatial feature enhancement map 303 can be (C r +H r W r )×H r ×W r .

[0145] S502: Decode the (m+1)th spatial feature enhancement map to obtain a feature map after the first spatial feature enhancement.

[0146] In the above implementation steps, after the spatial relationship feature enhancement module 701 outputs the m+1th spatial feature enhancement map, it is decoded to obtain the feature map after the first spatial feature enhancement. The spatial relationship feature enhancement module 701 combines the advantages of the non-local network and can capture the global spatial relationship features of the m+1th feature map, so that the image obtained by segmentation has a better effect.

[0147] For example, if Figure 6 and Figure 8 As shown, a decoder 602 may be used for decoding, and a plurality of decoders 602 are connected in sequence, that is, the output end of a previous decoder is connected to the input end of a subsequent decoder. Figure 6 The five decoders are connected in sequence, and the m+1th spatial feature enhancement map output by the spatial relationship feature enhancement module 701 is decoded to obtain the first decoding feature map 501, the first decoding feature map 501 is decoded to obtain the second decoding feature map 502, the second decoding feature map 502 is decoded to obtain the third decoding feature map 503, the third decoding feature map 503 is decoded to obtain the fourth decoding feature map 504, and the fourth decoding feature map 504 is decoded to obtain the fifth decoding feature map 505. At this time, the fifth decoding feature map 505 is the feature map after the first spatial feature is enhanced.

[0148] The last edge feature enhancement module outputs the feature map obtained after the first edge feature enhancement, and obtains the edge prediction result after Conv1×1 convolution. The edge prediction result is compared with the standard edge obtained by manual segmentation to calculate the edge loss.

[0149] Unbalanced data distribution is a major challenge in medical image segmentation. In order to optimize the neural network model in this embodiment, the problem of data imbalance can be effectively overcome. This embodiment uses the cross entropy loss (L bce ) and Dice loss (L dice ) is used as the segmentation loss function of the network, which is calculated by comparing the predicted image output by the network with the gold standard obtained by manual segmentation.

[0150] Since there is an imbalance between positive and negative samples in the edge result graph, this embodiment introduces the weighted cross entropy loss (L edge ) is used as the edge loss function, and the Canny operator is used to extract the gold standard of the edge, which is compared with the edge prediction map output by the network (i.e., the edge prediction result). The formula is as follows:

[0151] L total =L bce +L dice +L edge

[0152]

[0153]

[0154]

[0155] In the formula, t i represents the target pixel value in the gold standard, oi Represents the pixel value of the network's output prediction map, o i ′ represents the pixel value of the edge prediction map output by the network, n represents the number of all pixels in the input image, i represents the i-th pixel, and α represents the weighting coefficient. Among them, α can take the value of 0.25.

[0156] In the actual experiment, 552 skin wound images were used for the experiment, including burns, diabetic foot, venous or arterial ulcers, orthopedic wounds, pressure ulcers, and infected or necrotic toes. The images were randomly flipped left and right, flipped up and down, rotated from -30 degrees to 30 degrees, and added with additive Gaussian noise to amplify the images. First, the skin wound images were resampled to 512×512 pixels, and then the data set was randomly and evenly divided into five folds according to the wound type.

[0157] In order to objectively evaluate the performance of the skin wound image segmentation method of this embodiment, three evaluation indicators, namely, Dice coefficient, Jaccard index, and sensitivity (Sen), can be used. Assume that the wound surface pixels are positive samples, and the non-wound surface pixels are negative samples. The number of pixels predicted as positive in the positive samples is TP, the number of pixels predicted as negative in the positive samples is FN, the number of pixels predicted as positive in the negative samples is FP, and the number of pixels predicted as negative in the negative samples is TN. Then:

[0158]

[0159]

[0160]

[0161] In order to verify the effectiveness of the skin wound image segmentation method of this embodiment, the performance of this embodiment is compared with U-Net, CENet, and CPFNet. The results are shown in Table 1.

[0162] method Dice(%) Jaccard (%) Sen(%) Baseline 87.31±1.11 79.85±1.25 88.74±1.36 Baseline+EFA 88.86±1.21 82.00±1.47 90.43±0.87 Baseline+SFA 88.33±1.16 81.27±1.27 90.49±1.49 U-Net 75.58±1.73 65.47±1.94 81.33±2.45 CENet 88.07±0.72 80.99±0.84 89.57±1.30 CPFNet 87.77±1.00 80.56±1.19 89.76±1.65 This embodiment 89.16±0.79 82.16±1.06 90.76±0.84

[0163] Table 1

[0164] In order to verify the effectiveness of the skin wound image segmentation method of this embodiment, corresponding ablation experiments were carried out and compared with the baseline network (Baseline): A. The edge feature enhancement module was added to the baseline network, that is, "Baseline+EFA" in Table 1; B. The spatial relationship feature enhancement module was added to the baseline network, that is, "Baseline+SFA" in Table 1.

[0165] It can be seen from Table 1 that compared with the baseline network, the Baseline+EFA network and the Baseline+SFA network have improved Dice coefficient and Jaccard coefficient. Compared with the baseline network, the skin wound image segmentation method in this embodiment has an increase of 1.85% in Dice coefficient and 2.31% in Jaccard coefficient, and is significantly improved compared with the U-Net network, CENet network and CPFNet network.

[0166] Example 2

[0167] In practical applications, when the segmentation results of the neural network model are not satisfactory, it is often necessary to manually fine-tune the segmentation results, but relying on manual annotation to correct all segmentation errors is often time-consuming and laborious.

[0168] This embodiment provides an interactive skin wound image segmentation method. Fig.11 1 is a flowchart illustrating how to adjust a segmented image to obtain a more accurate segmented image according to some embodiments of the present invention. Although the processes described below include multiple operations that appear in a specific order, it should be clearly understood that these processes may also include more or fewer operations, and these operations may be performed sequentially or in parallel (e.g., using a parallel processor or a multi-threaded environment).

[0169] This embodiment provides an interactive skin wound image segmentation method, which is used to interactively adjust the segmentation results when segmentation errors occur, and optimize the results through a small amount of manual marking. Fig.11 As shown, the following steps are included:

[0170] S601 , assembling the image to be segmented, the first Gaussian distance map and the second Gaussian distance map to obtain a distance image to be segmented.

[0171] In the above implementation steps, the image to be segmented may be over-segmented and / or under-segmented when the skin wound image segmentation method of Example 1 is used to segment the image to be segmented. If the segmented image obtained has over-segmented or under-segmented images, it means that the segmented image is a defective segmented image. If the defective segmented image is analyzed, errors will occur, which will affect the diagnosis and care of the patient in the later stage.

[0172] When there are defects in the segmented image, the segmented image is manually analyzed and processed, and single-line marks are drawn on the missed areas and / or the multiple-divided areas. Among them, the first Gaussian distance map is the Gaussian distance map corresponding to the missed marks, and the second Gaussian distance map is the Gaussian distance map corresponding to the multiple-divided marks. For example, the image to be segmented, the first Gaussian distance map and the second Gaussian distance map are spliced ​​to obtain a distance image to be segmented with a size of 5×H×W.

[0173] S602: Perform edge enhancement processing on the distance image to be segmented based on the spatial attention mechanism and the channel attention mechanism to obtain a feature map after the second edge feature is enhanced.

[0174] In the above implementation steps, the image to be segmented, the first Gaussian distance map and the second Gaussian distance map are spliced ​​to obtain the distance image to be segmented. The distance image to be segmented can be input into a pre-trained neural network model, and the neural network model is used to segment the distance image to be segmented.

[0175] In order to obtain significant edge features in the distance image to be segmented, edge enhancement processing is performed on the distance image to be segmented to obtain a feature map after the second edge feature enhancement. For details, please refer to the relevant description of Example 1, which will not be repeated here.

[0176] S603, convolve the third Gaussian distance map to obtain a convolution segmentation map, channel-splice the N+1th feature map and the convolution segmentation map to obtain a spliced ​​convolution map, perform multi-scale hole convolution and global pooling on the spliced ​​convolution map, splice the outputs and obtain a convolution pooling feature map after convolution.

[0177] In the above implementation steps, if Fig.12 As shown, a spatial convolution pooling pyramid module 702 may be provided in the neural network model, the input end of the spatial convolution pooling pyramid module 702 is connected to the output end of the last encoder, and the output end is connected to the input end of the spatial relationship feature enhancement module 701.

[0178] The first segmented image is obtained by processing the skin wound image segmentation method of Example 1, and the third Gaussian distance map is the Gaussian distance map corresponding to the center point of the first segmented image. The third Gaussian distance map is convolved to obtain a convolution segmentation map. For example, the third Gaussian distance map can be input into 3 Conv3×3 convolution layers to output a convolution segmentation map of size 512×16×16. The N+1th feature map is the feature map output by the last encoder in the neural network model, where the value of N is greater than or equal to 2.

[0179] The N+1th feature map and the convolution segmentation map are concatenated to obtain a concatenated convolution map. Multi-scale dilated convolution and global pooling are performed on the concatenated convolution map. The feature maps output by the dilated convolution and global pooling are concatenated and passed through the convolution layer to obtain a convolution pooling feature map. It should be noted that the concatenated output refers to the feature map output by the concatenated dilated convolution and the feature map output by the global pooling.

[0180] For example, if Fig.13As shown in the figure, the channel splices the N+1th feature map and the convolution segmentation map to obtain a spliced ​​convolution map 7021, which is input into four dilated convolution layers with scales of 1, 6, 12, and 18 and a global pooling layer. Then the dilated convolution layer and the global pooling layer each output a feature map, and the five feature maps are spliced ​​to obtain a pooled convolution map 7022. The pooled convolution map 7022 is passed through an additional Conv1×1 convolution layer to obtain a convolutional pooled feature map 7023.

[0181] S604: Perform spatial enhancement processing on the convolutional pooling feature map to obtain a feature map after second spatial feature enhancement.

[0182] In the above implementation steps, if Fig.12 As shown, the output end of the spatial convolutional pooling pyramid module 702 is connected to the input end of the spatial relationship feature enhancement module 701, and the spatial convolutional pooling pyramid module 702 outputs the convolutional pooling feature map to the spatial relationship feature enhancement module 701. The spatial relationship feature enhancement module 701 performs spatial enhancement processing on the convolutional pooling feature map to obtain a feature map after the second spatial feature enhancement. For details, please refer to the relevant description of Example 1, which will not be repeated here.

[0183] S605: Perform channel stitching on the feature map after the second edge feature enhancement and the feature map after the second spatial feature enhancement to obtain a second segmented image. For details, please refer to the relevant description of Example 1, which will not be repeated here.

[0184] For example, if Fig.12 As shown, the first segmented image 403 has the defect of missing points, and a single line mark is drawn in the missed area and processed to obtain the first Gaussian distance map 402. The image to be segmented 101, the first Gaussian distance map 402 and the second Gaussian distance map 401 are combined to obtain the distance image to be segmented, wherein, since the segmented image does not have the defect of multiple points, the second Gaussian distance map 401 is a blank image.

[0185] The distance image to be segmented is input into a pre-trained neural network model, and the fifth feature map is output after encoding by multiple encoders, and the feature map after the second edge feature enhancement is obtained after edge enhancement processing by multiple edge feature enhancement modules.

[0186] The first segmented image 403 is passed through three Conv3×3 convolution layers to obtain a convolution segmentation map, and the channel concatenation convolution segmentation map and the fifth feature map are concatenated to obtain a concatenated convolution map. The concatenated convolution map is input into the spatial convolution pooling pyramid module 702, and multi-scale dilated convolution and global pooling are performed in the spatial convolution pooling pyramid module 702 to output a convolution pooling feature map.

[0187] The convolution pooling feature map is input to the spatial relationship feature enhancement module 701, which performs spatial enhancement processing on the convolution pooling feature map and obtains a second spatial feature enhanced feature map after decoding by the decoder. The second edge feature enhanced feature map and the second spatial feature enhanced feature map are channel-joined to obtain a second segmented image.

[0188] In one or more embodiments, the first Gaussian distance map is calculated by a first mathematical model, and the first mathematical model is:

[0189]

[0190]

[0191] Where D(x,y) is the shortest Euclidean distance from point (x,y) to all missed marking points, where the value of σ can be 10, 15 or 20. Point (x,y) is any point on the missed marking map.

[0192] The second Gaussian distance map is calculated by a second mathematical model, and the second mathematical model is:

[0193]

[0194]

[0195] Where D(m,n) is the shortest Euclidean distance from point (m,n) to all multi-point labeled points, and the value of σ can be 10, 15 or 20. Point (m,n) is any point on the multi-point labeled graph.

[0196] The third Gaussian distance map is calculated by a third mathematical model, and the third mathematical model is:

[0197]

[0198]

[0199] Where D(p,q) is the shortest Euclidean distance from point (p,q) to the center of the segmented image, and the value of σ can be 20, 25 or 30. Where point (p,q) is any point on the first segmented image.

[0200] Suppose the first segmented image is R(p,q), where the wound area is the set of all pixels with a value of 1 S = {(p,q)|R(p,q) = 1}, and K is the number of all pixels in S. Then the center point of the segmented image is:

[0201]

[0202]

[0203] When training and testing the neural network model, it is impossible to manually label each data set due to the large amount of data. Therefore, simulated manual interactive labeling can be used: compare the automatic segmentation results of the first step with the gold standard of manual segmentation to obtain over-scored and missed areas, extract the skeletons of these areas, and when the number of pixels contained in the skeleton is greater than 30, add these pixels to the simulated manual interactive labeling set.

[0204] In the actual experiment, 552 skin wound images were used for the experiment. To verify the effectiveness of the interactive skin wound image segmentation method of this embodiment, the interactive skin wound image segmentation method (Example 2) was compared with the skin wound image segmentation method (Example 1). In Table 2, FANet represents the skin wound image segmentation method, and IFANet represents the interactive skin wound image segmentation method.

[0205] method Dice(%) Jaccard (%) Sen(%) Baseline 87.31±1.11 79.85±1.25 88.74±1.36 FANet 89.16±0.79 82.16±1.06 90.76±0.84 IFANet 94.27±0.93 89.96±1.25 94.51±0.58

[0206] Table 2

[0207] It can be seen from Table 2 that, compared with the segmented image formed in Example 1, the segmented image formed in this embodiment has obvious improvements in Dice coefficient, Jaccard coefficient and Sen, and has a better segmentation effect, and can effectively remove over-divided areas and supplement under-divided areas.

[0208] Based on the baseline network ResNet34, the edge feature enhancement module is added to overcome the shortcomings of the current skin wound segmentation network, such as the difficulty in obtaining edge features and global context information, and the low utilization rate. The introduction of interactive methods has greatly improved the flexibility of the neural network, and the segmentation accuracy of the network has also been greatly improved through manual interaction.

[0209] Five-fold cross validation training is used. The specific method is as follows: when the first fold is used as a validation set, the other four folds of data are used as training sets to participate in the network training. The first fold of data is only used to evaluate the segmentation effect of the network model. When the second fold is used as a validation set, the other four folds are used as training sets to participate in the network training. The second fold of data is still used to evaluate the segmentation effect of the network model. The third, fourth, and fifth folds are similar to the above process. After five-fold training, five optimal models will be obtained. The corresponding data subsets (validation sets) are input into these trained network models to obtain the segmentation results.

[0210] Firstly, the neural network model used in the skin wound image segmentation method was trained, and the batch data size was set to 8, the learning rate was set to 0.01, and the number of iterations was set to 120 times; then based on the results of the neural network model, the neural network model used in the interactive skin wound image segmentation method was trained, and the batch data size was set to 4, the learning rate was set to 0.01, and the number of iterations was set to 200 times.

[0211] Example 3

[0212] This embodiment provides a skin wound image segmentation device for segmenting an image containing skin injuries, such as Fig.14 As shown, including:

[0213] The acquisition module 801 is used to acquire the image to be segmented including the skin wound surface. For details, please refer to the relevant description of step S101 in embodiment 1, which will not be repeated here.

[0214] The first processing module 802 is used to use a pre-trained neural network model to perform edge enhancement processing on the image to be segmented based on a spatial attention mechanism and a channel attention mechanism to obtain a feature map after the first edge feature is enhanced, and to perform spatial enhancement processing on the image to be segmented to obtain a feature map after the first spatial feature is enhanced. For details, please refer to the relevant description of step S102 in Example 1, which will not be repeated here.

[0215] The first splicing module 803 is used to perform channel splicing on the feature map after the first edge feature enhancement and the feature map after the first spatial feature enhancement to obtain a first segmented image. For details, please refer to the relevant description of step S103 in embodiment 1, which will not be repeated here.

[0216] In the above embodiment, the first processing module 802 performs edge enhancement processing on the image to be segmented containing the skin wound surface based on the spatial attention mechanism and the channel attention mechanism to obtain a feature map after the first edge feature is enhanced, and performs spatial enhancement processing on the image to be segmented containing the skin wound surface to obtain a feature map after the first spatial feature is enhanced, and the first splicing module 803 splices the feature map after the first edge feature is enhanced and the feature map after the first spatial feature is enhanced to obtain a first segmented image. Since the contribution of edge features to segmentation is particularly important, and the position of objects in the image and the spatial relationship between objects are very important features in image segmentation, channel splicing using feature maps enhanced by edge features and feature maps enhanced by spatial features can make the final segmented image better.

[0217] Example 4

[0218] The embodiment provides an interactive skin wound image segmentation device, which is used to interactively adjust the segmentation results when segmentation errors occur, and optimize the results through a small amount of manual marking. Fig.15 As shown, including:

[0219] The second stitching module 804 is used to stitch the image to be segmented, the first Gaussian distance map and the second Gaussian distance map to obtain the distance image to be segmented; wherein the first Gaussian distance map is a Gaussian distance map corresponding to the missed mark, and the second Gaussian distance map is a Gaussian distance map corresponding to the over-marked mark. For details, please refer to the relevant description of step S601 in embodiment 2, which will not be repeated here.

[0220] The second processing module 805 is used to perform edge enhancement processing on the distance image to be segmented based on the spatial attention mechanism and the channel attention mechanism to obtain a feature map after the second edge feature is enhanced. For details, please refer to the relevant description of step S602 in embodiment 2, which will not be repeated here.

[0221] The convolution pooling module 806 is used to convolve the third Gaussian distance map to obtain a convolution segmentation map, channel splice the N+1th feature map and the convolution segmentation map to obtain a spliced ​​convolution map, perform multi-scale hole convolution and global pooling on the spliced ​​convolution map, splice the output and obtain a convolution pooling feature map after convolution; wherein the third Gaussian distance map is the Gaussian distance map corresponding to the center point of the first segmentation image, and the first segmentation image is obtained by using the skin wound image segmentation method. For details, please refer to the relevant description of step S603 in Example 2, which will not be repeated here.

[0222] The third processing module 807 is used to perform spatial enhancement processing on the convolution pooling feature map to obtain a feature map after the second spatial feature enhancement. For details, please refer to the relevant description of step S604 in embodiment 2, which will not be repeated here.

[0223] The third splicing module 808 is used to perform channel splicing on the feature map after the second edge feature enhancement and the feature map after the second spatial feature enhancement to obtain a second segmented image. For details, please refer to the relevant description of step S605 in embodiment 2, which will not be repeated here.

[0224] Example 5

[0225] This embodiment provides a computer device, such as Fig.16 As shown, the computer device includes a processor 901 and a memory 902, wherein the processor 901 and the memory 902 can be connected via a bus or other means. Fig.16 The example of connecting through bus is taken in the following.

[0226] The processor 901 may be a central processing unit (CPU). The processor 901 may also be other general-purpose processors, digital signal processors (DSP), graphics processors (GPU), embedded neural network processors (NPU), or other dedicated deep learning coprocessors, application specific integrated circuits (ASIC), field programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, or a combination of the above chips.

[0227] The memory 902 is a non-transitory computer-readable storage medium that can be used to store non-transitory software programs, non-transitory computer executable programs and modules, such as the program instructions / modules corresponding to the skin wound image segmentation method or the interactive skin wound image segmentation method in the embodiment of the present invention (such as Fig.14 The processor 901 executes various functional applications and data processing of the processor by running the non-transitory software programs, instructions and modules stored in the memory 902, that is, the skin wound image segmentation method in embodiment 1 of the above method or the interactive skin wound image segmentation method in embodiment 2.

[0228] The memory 902 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required by at least one function; the data storage area may store data created by the processor 901, etc. In addition, the memory 902 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory 902 may optionally include a memory remotely arranged relative to the processor 901, and these remote memories may be connected to the processor 901 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0229] The one or more modules are stored in the memory 902, and when executed by the processor 901, the following is performed: Figure 1The skin wound image segmentation method in the illustrated embodiment or the interactive skin wound image segmentation method in embodiment 2.

[0230] In this embodiment, the memory 902 stores a program instruction or module of a skin wound image segmentation method. When the processor 901 executes the program instruction or module stored in the memory 902, the image to be segmented containing the skin wound is subjected to edge enhancement processing based on the spatial attention mechanism and the channel attention mechanism to obtain a feature map after the first edge feature is enhanced, the image to be segmented containing the skin wound is subjected to spatial enhancement processing to obtain a feature map after the first spatial feature is enhanced, and the feature map after the first edge feature enhancement and the feature map after the first spatial feature enhancement are channel-joined to obtain a first segmented image. Since the contribution of edge features to segmentation is particularly important, and the position of objects in the image and the spatial relationship between objects are very important features in image segmentation, channel-joining using feature maps enhanced by edge features and feature maps enhanced by spatial features can make the final segmented image better.

[0231] The embodiment of the present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions can execute the skin wound image segmentation method or the interactive skin wound image segmentation method in any of the above method embodiments. The storage medium can be a disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), a flash memory, a hard disk drive (HDD) or a solid-state drive (SSD), etc.; the storage medium can also include a combination of the above types of memories.

[0232] Obviously, the above embodiments are merely examples for the purpose of clear explanation, and are not intended to limit the implementation methods. For those skilled in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all the implementation methods here. The obvious changes or modifications derived therefrom are still within the scope of protection of the invention.

Claims

1. A skin wound image segmentation method, It is characterized in that The steps include: Acquire an image to be segmented that includes a skin wound surface; Based on the spatial attention mechanism and the channel attention mechanism, the image to be segmented is edge enhanced to obtain a feature map after the first edge feature is enhanced, and the image to be segmented is spatially enhanced to obtain a feature map after the first spatial feature is enhanced; wherein, the image to be segmented is edge enhanced based on the spatial attention mechanism and the channel attention mechanism to obtain a feature map after the first edge feature is enhanced, including: performing a first encoding and a first maximum pooling on the image to be segmented to obtain a first feature map, performing a second encoding and a second maximum pooling on the first feature map to obtain a second feature map, performing a first edge enhancement on the first feature map and the second feature map to obtain a first edge feature enhanced map; performing an m-1th edge enhancement on the m+1th feature map and the m+1th feature map to obtain an mth edge feature enhanced map; wherein the m+1th feature map is obtained by performing the m+1th encoding on the mth feature map obtained by the mth encoding and the mth maximum pooling, m=2,3,...N; when m=N, the obtained Nth edge feature enhanced map is the feature map after the first edge feature is enhanced; Channel stitching is performed on the feature map after the first edge feature enhancement and the feature map after the first spatial feature enhancement to obtain a first segmented image.

2. The skin wound image segmentation method according to claim 1, It is characterized in that The edge enhancement process is performed on the image to be segmented based on the spatial attention mechanism and the channel attention mechanism to obtain a feature map after the first edge feature is enhanced, including: Performing a first edge enhancement process on the first feature map and the second feature map based on a spatial attention mechanism to obtain a first edge feature enhanced map; Based on the spatial attention mechanism, a second edge enhancement process is performed on the first edge feature enhancement map and the third feature map to obtain a second edge feature enhancement map; Based on the channel attention mechanism, the second edge feature enhancement map and the fourth feature map are subjected to a third edge enhancement process to obtain a third edge feature enhancement map; Based on the channel attention mechanism, a fourth edge enhancement process is performed on the third edge feature enhancement map and the fifth feature map to obtain a feature map after the first edge feature is enhanced.

3. The skin wound image segmentation method according to claim 1 or 2, It is characterized in that The step of performing an m-th edge enhancement process on the m-1th edge feature enhancement map and the m+1th feature map to obtain the mth edge feature enhancement map comprises: Perform convolution and encoding processing on the m-1th edge feature enhancement map to obtain the m-1th edge feature encoding map, and perform convolution and up-sampling processing on the m+1th feature map to obtain the m+1th feature sampling map; Perform channel splicing on the m-1th edge feature encoding map and the m+1th feature sampling map to obtain an mth edge feature splicing map; Processing the mth edge feature splicing graph based on a spatial attention mechanism to obtain a spatial weight; After the spatial weight passes through the activation function, it is multiplied with the m-1th edge feature encoding map to obtain the mth edge feature enhancement map.

4. The skin wound image segmentation method according to claim 3, It is characterized in that The processing of the mth edge feature splicing graph based on the spatial attention mechanism to obtain the spatial weight includes: Performing average pooling on the mth edge feature splicing map to obtain a first pooling feature map, and performing maximum pooling on the mth edge feature splicing map to obtain a second pooling feature map; Perform channel splicing on the first pooling feature map and the second pooling feature map to obtain a spliced ​​mth edge feature pooling map; The mth edge feature pooling map is convolved to obtain a spatial weight.

5. The skin wound image segmentation method according to claim 1 or 2, It is characterized in that The step of performing an m-th edge enhancement process on the m-1th edge feature enhancement map and the m+1th feature map to obtain the mth edge feature enhancement map comprises: Perform convolution and encoding processing on the m-1th edge feature enhancement map to obtain the m-1th edge feature encoding map, and perform convolution and up-sampling processing on the m+1th feature map to obtain the m+1th feature sampling map; Perform channel splicing on the m-1th edge feature encoding map and the m+1th feature sampling map to obtain an mth edge feature splicing map; Processing the mth edge feature splicing graph based on a channel attention mechanism to obtain a channel weight; After the channel weight passes through the activation function, it is multiplied with the m-1th edge feature encoding map to obtain the mth edge feature enhancement map.

6. The skin wound image segmentation method according to claim 5, It is characterized in that The processing of the mth edge feature splicing graph based on the channel attention mechanism to obtain the channel weight includes: Performing average pooling on the mth edge feature splicing map to obtain a third pooling feature map, and performing maximum pooling on the mth edge feature splicing map to obtain a fourth pooling feature map; The third pooling feature map and the fourth pooling feature map are added together after passing through a multi-layer perceptron to obtain the channel weight.

7. The skin wound image segmentation method according to claim 1 or 2, It is characterized in that The image to be segmented is subjected to spatial enhancement processing to obtain a feature map after the first spatial feature is enhanced, including: When m=N, the m+1th feature map is spatially enhanced to obtain the m+1th spatial feature enhanced map; The m+1th spatial feature enhancement map is decoded to obtain a feature map after the first spatial feature enhancement.

8. The skin wound image segmentation method according to claim 7, It is characterized in that When m=N, performing spatial enhancement processing on the m+1th feature map to obtain the m+1th spatial feature enhancement map includes: Perform convolution on the m+1th feature map to obtain two feature matrices of different sizes; Multiply the two feature matrices of different sizes, and reorganize them again to obtain the m+1th matrix feature map; The m+1th matrix feature map and the m+1th feature map are channel-joined and linearly rectified to obtain the m+1th spatial feature enhancement map.

9. An interactive skin wound image segmentation method, It is characterized in that The steps include: The image to be segmented, the first Gaussian distance map and the second Gaussian distance map are combined to obtain a distance image to be segmented; wherein the first Gaussian distance map is a Gaussian distance map corresponding to the missed mark, and the second Gaussian distance map is a Gaussian distance map corresponding to the over-marked mark; Performing edge enhancement processing on the distance image to be segmented based on a spatial attention mechanism and a channel attention mechanism to obtain a feature map after second edge feature enhancement; Convolving the third Gaussian distance map to obtain a convolution segmentation map, channel-splicing the N+1th feature map and the convolution segmentation map to obtain a spliced ​​convolution map, performing multi-scale hole convolution and global pooling on the spliced ​​convolution map, splicing the outputs and convolving to obtain a convolution pooling feature map; wherein the third Gaussian distance map is a Gaussian distance map corresponding to the center point of the first segmented image, and the first segmented image is obtained by using the skin wound image segmentation method according to any one of claims 1 to 8; Performing spatial enhancement processing on the convolutional pooling feature map to obtain a feature map after second spatial feature enhancement; Channel stitching is performed on the feature map after the second edge feature is enhanced and the feature map after the second spatial feature is enhanced to obtain a second segmented image.

10. The interactive skin wound image segmentation method according to claim 9, It is characterized in that The first Gaussian distance map is calculated by a first mathematical model, and the first mathematical model is: In the formula, For point The shortest Euclidean distance to all missed markers; And / or, the second Gaussian distance map is calculated by a second mathematical model, and the second mathematical model is: In the formula, For point The shortest Euclidean distance to all multi-point markers; The third Gaussian distance map is calculated by a third mathematical model, and the third mathematical model is: In the formula, For point The shortest Euclidean distance to the center point of the first segmented image.

11. A skin wound image segmentation device, It is characterized in that include: An acquisition module, used for acquiring an image to be segmented including a skin wound surface; A first processing module is used to perform edge enhancement processing on the image to be segmented based on a spatial attention mechanism and a channel attention mechanism to obtain a feature map after the first edge feature is enhanced, and to perform spatial enhancement processing on the image to be segmented to obtain a feature map after the first spatial feature is enhanced; wherein the first processing module is specifically used to perform a first encoding and a first maximum pooling on the image to be segmented to obtain a first feature map, perform a second encoding and a second maximum pooling on the first feature map to obtain a second feature map, perform a first edge enhancement processing on the first feature map and the second feature map to obtain a first edge feature enhanced map; perform an m-1th edge enhancement processing on the m+1th feature map and the m+1th feature map to obtain an mth edge feature enhanced map; wherein the m+1th feature map is obtained by performing the m+1th encoding on the mth feature map obtained by the mth encoding and the mth maximum pooling, m=2,3,...N; when m=N, the obtained Nth edge feature enhanced map is the feature map after the first edge feature is enhanced; The first stitching module is used to perform channel stitching on the feature map after the first edge feature is enhanced and the feature map after the first spatial feature is enhanced to obtain a first segmented image.

12. An interactive skin wound image segmentation device, It is characterized in that include: A second stitching module is used to stitch the image to be segmented, the first Gaussian distance map and the second Gaussian distance map to obtain a distance image to be segmented; wherein the first Gaussian distance map is a Gaussian distance map corresponding to the missed mark, and the second Gaussian distance map is a Gaussian distance map corresponding to the over-marked mark; A second processing module is used to perform edge enhancement processing on the to-be-segmented distance image based on a spatial attention mechanism and a channel attention mechanism to obtain a feature map after second edge feature enhancement; A convolution pooling module, used to convolve the third Gaussian distance map to obtain a convolution segmentation map, channel splicing the N+1th feature map and the convolution segmentation map to obtain a spliced ​​convolution map, and performing multi-scale hole convolution and global pooling on the spliced ​​convolution map, splicing the output and convolving to obtain a convolution pooling feature map; wherein the third Gaussian distance map is a Gaussian distance map corresponding to the center point of the first segmented image, and the first segmented image is obtained by using the skin wound image segmentation method described in any one of claims 1 to 8; A third processing module is used to perform spatial enhancement processing on the convolution pooling feature map to obtain a feature map after second spatial feature enhancement; The third stitching module is used to perform channel stitching on the feature map after the second edge feature is enhanced and the feature map after the second spatial feature is enhanced to obtain a second segmented image.

13. A computer device, It is characterized in that include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the skin wound image segmentation method according to any one of claims 1 to 8 or the interactive skin wound image segmentation method according to any one of claims 9 to 10 by executing the computer instructions.

14. A computer-readable storage medium, It is characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the skin wound image segmentation method described in any one of claims 1-8 or the interactive skin wound image segmentation method described in any one of claims 9-10.

Citation Information

Patent Citations

  • Semantic image segmentation method and system based on edge enhancement

    CN111462126A