Lightweight segmentation method and device, and storage medium

By adopting a lightweight segmentation model in the medical image segmentation auxiliary system, using jump connections to transmit global features and advanced semantic features, the problem of high hardware requirements in existing systems is solved, and real-time segmentation in ordinary computer terminals is realized.

WO2025111770A1PCT designated stage expired Publication Date: 2025-06-05SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI +1

Patent Information

Application Number
PCT/CN2023/134539
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-28
Publication Date
2025-06-05

Smart Images

  • Figure CN2023134539_05062025_PF_FP_ABST
    Figure CN2023134539_05062025_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of picture segmentation. Disclosed are a lightweight segmentation method and device, and a storage medium. In the method, a brand-new lightweight segmentation model is created, a structure including an encoder module and a decoder module is used in the model, and a global feature and a high-level semantic feature are transmitted by means of skip connection, that is, a self-attention feature vector having the global feature and a deep convolution feature vector having the high-level semantic feature are only transmitted to the decoder module, and a process feature vector having a local feature and a low-level semantic feature is not transmitted to the decoder module, so that the number of channels of a shallow feature vector is reduced, thereby reducing the parameter quantity and the calculation quantity, and thus making the lightweight segmentation model in the present application perform picture segmentation on a common computer terminal to obtain a target surgical site image.
Need to check novelty before this filing date? Find Prior Art

Description

Lightweight segmentation method, device and storage medium Technical Field

[0001] The present application relates to the technical field of image segmentation, and in particular to a lightweight segmentation method, device, and storage medium. Background Art

[0002] Laparoscopic surgery is a widely used minimally invasive surgical procedure. During the procedure, the surgeon inserts a laparoscope equipped with a miniature camera into the abdominal cavity. Using digital camera technology, the images captured by the laparoscope are transmitted via optical fibers to a backend signal processing system and displayed in real time on a dedicated monitor. The surgeon then uses specialized laparoscopic instruments to perform the surgery, using the surgical field displayed on the monitor. Compared to traditional surgery, laparoscopic surgery is less invasive, causes less damage to surrounding tissues, and offers a faster recovery and less pain. Furthermore, most abdominal surgeries can be performed laparoscopically, making it widely popular with patients and garnering increasing attention. For computer-assisted laparoscopic surgery, using semantic segmentation techniques to accurately segment relevant organs and tissues is crucial. Semantic segmentation of medical images presents particular complexities and challenges. For example, labeled datasets are difficult to obtain, and optical imaging is often subject to irregularities such as occlusion, shadows, uneven illumination, and noise, making segmentation in medical images more difficult than in natural images.

[0003] Laparoscopic surgery assistance systems have very high requirements for real-time image segmentation. Although traditional deep learning models can ensure high segmentation accuracy, the amount of calculation and parameters required is huge. Only processors and data centers with high computing power can achieve real-time segmentation. In the actual situation where most hospitals lack powerful computing power, these models cannot meet the real-time requirements.

[0004] Summary of the Invention

[0005] The purpose of this application is to provide a lightweight segmentation method, device and storage medium.

[0006] In a first aspect, an embodiment of the present application provides a lightweight segmentation method, the method comprising:

[0007] Acquire real-time images;

[0008] Inputting the real-time image into a trained lightweight segmentation model to perform target segmentation to obtain a target surgical site image;

[0009] The lightweight segmentation model includes an encoder module and a decoder module. The step of inputting the real-time image into the trained lightweight segmentation model to perform target segmentation to obtain the target surgical site image further includes:

[0010] Inputting the real-time image into the encoder module for feature information extraction to obtain a first feature vector having low-level semantic features, a self-attention feature vector having global features, and at least one type of deep convolution feature vector having high-level semantic features;

[0011] The self-attention feature vector is output to the decoder module, and the deep convolution feature vector is output to the decoder module via a skip connection;

[0012] The decoder module segments the target surgical site image according to the plurality of first feature vectors having low-level semantic features, the deep convolution feature vector, and the self-attention feature vector.

[0013] Optionally, the encoder module includes a first convolutional layer, a first depthwise separable convolutional layer, a second depthwise separable convolutional layer, a first self-attention layer, and a second self-attention layer; the depthwise convolutional feature vector includes a first-class depthwise convolutional feature vector and a second-class depthwise convolutional feature vector;

[0014] The step of inputting the real-time image into the encoder module to extract feature information to obtain a first feature vector having low-level semantic features, a self-attention feature vector having global features, and at least one type of deep convolution feature vector having high-level semantic features further includes:

[0015] Inputting the real-time image into the first convolution layer to perform surface feature extraction to obtain a first convolution feature vector;

[0016] Inputting the first convolution feature vector into the first depthwise separable convolution layer for a first depthwise feature extraction to obtain the first type of depthwise convolution feature vector with high-level semantic features;

[0017] Inputting the first type of depth convolution feature vector into the second depth separable convolution layer for a second depth feature extraction to obtain the second type of depth convolution feature vector with high-level semantic features;

[0018] Inputting the second type of deep convolution feature vector into the first self-attention layer to perform a first self-attention calculation to obtain the first self-attention feature vector with global features;

[0019] The first self-attention feature vector is input into the second self-attention layer for performing a second self-attention calculation to obtain the self-attention feature vector with global features.

[0020] Optionally, the first convolutional layer includes a common convolutional layer, a first batch of processing layers, and a first activation function layer connected in sequence;

[0021] The common convolution layer is used to perform a convolution operation on the real-time image to obtain an initial feature vector;

[0022] The first batch of processing layers is used to normalize the initial feature vectors;

[0023] The first activation function layer is used to accelerate the normalization of the initial feature vector to obtain a normalized first convolution feature vector.

[0024] Optionally, the first depthwise separable convolutional layer and the second depthwise separable convolutional layer each separately include a first depthwise convolutional layer, a second batching layer, a maximum pooling layer, and a second activation function connected in sequence;

[0025] The first depth convolution layer is used to perform a depth convolution operation on the input convolution feature vector to obtain an initial depth convolution feature vector with high-level semantic features;

[0026] The second batch processing layer is used to normalize the initial depth convolution feature vector;

[0027] The maximum pooling layer is used to downsample the initial depth convolution feature vector to extract edge information;

[0028] The second activation function layer is used to accelerate the normalization of the initial depth convolution feature vector to obtain the first type of depth convolution feature vector with high-level semantic features or the second type of depth convolution feature vector with high-level semantic features after normalization.

[0029] Optionally, the first self-attention layer and the second self-attention layer both include a window-based multi-head attention module and a moving window-based multi-head self-attention module;

[0030] The window-based multi-head attention module is used to perform a single-window self-attention calculation on the input second-type deep convolution feature vector or the first self-attention feature vector to obtain an initial self-attention feature vector;

[0031] The moving window-based multi-head self-attention module is used to perform a moving window multi-head self-attention calculation on the input initial self-attention feature vector to obtain the first self-attention feature vector with global features or the second self-attention feature vector with global features.

[0032] Optionally, the decoder module includes a first depthwise separable convolutional layer, a first depthwise convolution-self-attention layer, a second depthwise convolution-self-attention layer, a second depthwise separable convolutional layer, and a convolutional output layer;

[0033] The decoder module further comprises the following steps:

[0034] The first depthwise separable convolutional layer is configured to preliminarily restore the spatial resolution of the second self-attention feature vector to obtain a third depthwise convolutional vector;

[0035] The first depthwise convolution-self-attention layer is jump-connected to the first self-attention layer and is used to perform secondary resolution recovery and feature fusion according to the third depthwise convolution vector and the first self-attention feature vector to obtain a fourth depthwise convolution vector;

[0036] The second depth convolution-self-attention layer is jump-connected to the second depth separable convolution layer and is used to perform resolution recovery and feature fusion based on the fourth depth convolution vector and the second type of depth convolution feature vector to obtain a fifth depth convolution vector;

[0037] The second depthwise separable convolution layer is jump-connected to the first depthwise separable convolution layer and is used to perform further resolution restoration and feature fusion according to the fifth depthwise convolution vector and the first type of depthwise convolution feature vector to obtain a sixth depthwise convolution vector;

[0038] The convolution output layer is used to obtain the target surgical site image after performing depth convolution on the sixth depth convolution vector.

[0039] Optionally, the first depth convolution-self-attention layer or the second depth convolution-self-attention layer includes a third self-attention layer, a second depth convolution layer, a third batching layer, a second resizing layer, and a third activation function layer:

[0040] A third self-attention layer is used to perform self-attention calculation on the input third depth convolution vector to obtain a third self-attention feature vector;

[0041] A second depthwise convolutional layer, configured to perform depthwise convolution on the third self-attention feature vector to obtain a third self-attention-depthwise convolutional feature vector;

[0042] The third batching layer is used to normalize the third self-attention-depth convolution feature vector;

[0043] The second resizing layer is configured to resize the third self-attention-depth convolution feature vector;

[0044] The third activation function layer is used to accelerate the normalization of the third self-attention-depth convolution feature vector to obtain the normalized third depth convolution feature vector with high-level semantic features or the fourth depth convolution feature vector with high-level semantic features.

[0045] Optionally, the lightweight segmentation model is trained by back propagation.

[0046] In a second aspect, an embodiment of the present application further provides a lightweight segmentation device, comprising:

[0047] A processor, a memory, a communication interface, and a communication bus, wherein the processor, the memory, and the communication interface communicate with each other via the communication bus;

[0048] The memory is used to store at least one executable instruction, and the executable instruction enables the processor to perform the operation of the lightweight segmentation method as described above.

[0049] In a third aspect, an embodiment of the present application further provides a storage medium storing at least one executable instruction. When the executable instruction is executed on a lightweight segmentation device / apparatus, the lightweight segmentation device / apparatus performs the operation of the lightweight segmentation method described above.

[0050] This solution creates a new lightweight segmentation model, adopts the structure of encoder module and decoder module in the model, and transmits global features and high-level semantic features through jump connections, that is, only the self-attention feature vector with global features and the deep convolution feature vector with high-level semantic features are transmitted, and the process feature vector with local features and low-level semantic features is not transmitted to the decoder module. Therefore, the number of channels of the shallow feature vector is reduced, thereby reducing the number of parameters and the amount of calculation, so that the lightweight segmentation model of this application can perform related operations on ordinary computer terminals, thereby solving the technical problem that the existing medical image segmentation auxiliary system has high hardware requirements and cannot achieve real-time segmentation on the terminal. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0052] FIG1 is a schematic diagram of a flow chart of a first embodiment of a lightweight segmentation method of the present application;

[0053] FIG2 is a flow chart of a second embodiment of the lightweight segmentation method of the present application;

[0054] FIG3 is a schematic diagram comparing surgical site images before and after segmentation using the lightweight segmentation method of the present application;

[0055] FIG3-a is a schematic diagram showing an image of a surgical site before segmentation using the lightweight segmentation method provided by the present invention;

[0056] FIG3-b is a schematic diagram showing an image of a surgical site after segmentation using the lightweight segmentation method provided by the present invention;

[0057] FIG4 is a flow chart of a third embodiment of the lightweight segmentation method of the present application;

[0058] FIG5 is a schematic diagram of the structure of the first self-attention layer or the second self-attention layer of the lightweight segmentation model in the lightweight segmentation method of the present application;

[0059] FIG6 is a schematic structural diagram of a lightweight segmentation model of the lightweight segmentation method of the present application;

[0060] FIG7 is a schematic diagram of the structure of the depth-separable convolution of the lightweight segmentation model in the lightweight segmentation method of the present application;

[0061] FIG8 is a schematic structural diagram of an embodiment of a lightweight segmentation device. DETAILED DESCRIPTION

[0062] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0063] This application provides a lightweight segmentation method that transmits global features and high-level semantic features through jump connections, thereby reducing the number of parameters and the amount of calculation, thereby solving the technical problem that the existing medical image segmentation auxiliary system has high hardware requirements and cannot achieve real-time segmentation on the terminal.

[0064] In one implementation scenario, as shown in FIG1 , this solution is implemented based on a camera component, a computer component, and a display component.

[0065] Among them, the camera component is used to obtain real-time images, the computer component is used to execute the lightweight segmentation method, and the display component is used to display the target surgical site image.

[0066] FIG1 shows a flowchart of a first embodiment of a lightweight segmentation method of the present invention, which is executed by a lightweight segmentation device. As shown in FIG1 , the method includes the following steps:

[0067] The method comprises:

[0068] Step S1, acquiring real-time images;

[0069] The real-time images can be acquired through cameras installed in various medical devices or other equipment.

[0070] Step S2: inputting the real-time image into a trained lightweight segmentation model to perform target segmentation to obtain an image of the target surgical site;

[0071] Since the real-time images obtained are in video or other formats, it is necessary to cut frames of the video or other files to obtain PNG images. In order to reduce the amount of calculation and parameters, the video or other formats of images are uniformly converted into PNG format, and a specification conversion function is set before or after the input of the lightweight segmentation model. For example, the input PNG formats of various specifications can be uniformly converted into 512*512, which facilitates the subsequent image segmentation of the lightweight segmentation model.

[0072] 2 , the lightweight segmentation model includes an encoder module and a decoder module. The step of inputting the real-time image into the trained lightweight segmentation model for target segmentation to obtain the target surgical site image further includes:

[0073] Step S21: inputting the real-time image into the encoder module to extract feature information to obtain a first feature vector having low-level semantic features, a self-attention feature vector having global features, and at least one type of deep convolution feature vector having high-level semantic features;

[0074] Among them, the feature information obtained by feature information extraction in this application includes local features, global features, low-level semantic features and high-level semantic features, and the above features have different meanings. Among them, low-level semantic features are surface parameters such as color, edge, brightness and texture, high-level semantic features are deep parameters such as the overall shape and category of the object, local features are small-scale parameters such as edge and texture, and global features are the overall color distribution, texture distribution, and relative positions and relationships between objects in the image. The first feature vector with low-level semantic features can be a feature vector obtained through ordinary convolution, deep convolution or self-attention calculation. The self-attention feature vector with global features is obtained through self-attention calculation, high-level semantic features are obtained through deep convolution, and low-level semantic features or local features are obtained through ordinary convolution. It should be noted that, for at least one category of deep convolution feature vectors with high-level semantic features mentioned in this application, the category of their high-level semantic features is different in the specific parameters obtained through deep convolution of different layers. The specific difference is that they are obtained from the input feature vectors of different layers. For the convenience of distinction, different categories are used to represent the deep convolution feature vectors of high-level semantic features obtained by different deep convolution layers. For example, the deep convolution feature vectors with high-level semantic features obtained in the Nth deep convolution layer are collectively referred to as the Nth category of deep convolution feature vectors, and the deep convolution feature vectors with high-level semantic features obtained in the N+1th deep convolution layer are collectively referred to as the N+1th category of deep convolution feature vectors, where N is a natural number greater than or equal to 0.

[0075] Step S22: the self-attention feature vector is output to the decoder module, and the deep convolution feature vector is output to the decoder module via a skip connection;

[0076] Among them, the jump connection at this time only transmits the self-attention feature vector with global features and the deep convolution feature vector with high-level semantic features, thereby reducing the number of channels of the shallow feature vector, and then the number of parameters and the amount of calculation, so that the lightweight segmentation model of this application can perform related operations on ordinary computer terminals, thereby solving the technical problem that the existing medical image segmentation auxiliary system has high hardware requirements and cannot achieve real-time segmentation on the terminal.

[0077] Step S23: The decoder module segments the target surgical site image according to the first feature vector with low-level semantic features, the deep convolution feature vector, and the self-attention feature vector.

[0078] Through the above scheme, a new lightweight segmentation model is created. The structure of the encoder module and the decoder module is adopted in the model. Through the combination of the first feature vector with low-level semantic features and at least one type of deep convolution feature vector with high-level semantic features, it can be ensured that the features acquired by the encoder module cover low-level semantic features, high-level semantic features and global features. The high-level semantic features are passed to the decoder module to assist in segmenting the target surgical site image to ensure that the segmentation of the target surgical site image will not fail due to series convergence or excessive convolution. The combination of ordinary convolution and depth convolution reduces the number of convolution layers and reduces the total computational complexity. Moreover, the present application transmits global features and high-level semantic features through jump connections, that is, only the self-attention feature vector with global features and the deep convolution feature vector with high-level semantic features are transmitted, and the process feature vector with local features and low-level semantic features is not transmitted to the decoder module. Therefore, the number of channels of the shallow feature vector is reduced, thereby reducing the number of parameters and the amount of calculation, so that the lightweight segmentation model of the present application can perform related operations on ordinary computer terminals, thereby solving the technical problem that the existing medical image segmentation auxiliary system has high hardware requirements and cannot realize real-time segmentation on the terminal. Referring to Figure 3, Figure 3-a shows a schematic diagram of the surgical site image before segmentation of the lightweight segmentation method provided by the present invention, and Figure 3-b shows a schematic diagram of the surgical site image after segmentation of the lightweight segmentation method provided by the present invention. Therefore, in the scheme of the application, the target organ can be completely segmented to assist doctors in completing related operations.

[0079] Furthermore, the transmission of global features and high-level semantic features through the above-mentioned jump connections can also overcome the problems of gradient disappearance and network degradation.

[0080] In one embodiment, the lightweight segmentation method, as shown in FIG6 , the encoder module includes a first convolutional layer 1, a first depth-separable convolutional layer 2, a second depth-separable convolutional layer 3, a first self-attention layer 4, and a second self-attention layer 5: the depth-convolutional feature vector includes a first-class depth-convolutional feature vector and a second-class depth-convolutional feature vector;

[0081] The step of inputting the real-time image into the encoder module for feature information extraction to obtain a first feature vector having low-level semantic features, a self-attention feature vector having global features, and at least one type of deep convolution feature vector having high-level semantic features, as shown in FIG4 , further includes:

[0082] Step S211: input the real-time image into the first convolution layer 1 to perform surface feature extraction to obtain a first convolution feature vector;

[0083] The first convolutional layer 1 extracts surface features to obtain a first convolutional feature vector. In this case, there are multiple first convolutional feature vectors, and the convolution performed is a normal convolution. The features extracted are local features and low-level semantic features. Low-level semantic features are surface parameters such as color, edges, brightness, and texture, while local features are parameters within a small range such as edges and texture.

[0084] Optionally, the first convolutional layer 1 includes a common convolutional layer, a first batch of processing layers and a first activation function layer connected in sequence;

[0085] The common convolution layer is used to perform a convolution operation on the real-time image to obtain an initial feature vector;

[0086] The convolution operation of a common convolutional layer can capture different local features of the input real-time image, such as edges and textures, and then output an initial feature vector with these features. Therefore, the initial feature vectors determined at this time can be multiple, determined by the number of convolution kernels in the current layer and the number of input feature vectors.

[0087] The first batch of processing layers is used to normalize the initial feature vectors;

[0088] Among them, the first batch of processing layers can be implemented using the Batch Normalization layer, which can normalize each mini-batch, which can not only accelerate the convergence of the model but also help alleviate the problems of gradient disappearance and gradient explosion.

[0089] The first activation function layer is used to accelerate the normalization of the initial feature vector to obtain a first convolution feature vector after normalization.

[0090] At this point, a nonlinear activation function is introduced to enhance the model's expressiveness. These layers allow for the extraction of more advanced and richer feature information. Optionally, the first activation function layer can be implemented using a ReLU function.

[0091] Step S212: Input the first convolution feature vector into the first depthwise separable convolution layer 2 to perform a first depth feature extraction to obtain a first type of depthwise convolution feature vector with high-level semantic features;

[0092] Optionally, the first depthwise separable convolutional layer 2 and the second depthwise separable convolutional layer 3 each separately include a first depthwise convolutional layer, a second batching layer, a maximum pooling layer, and a second activation function connected in sequence;

[0093] The first depth convolution layer is used to perform a depth convolution operation on the input convolution feature vector to obtain an initial depth convolution feature vector with high-level semantic features;

[0094] Among them, the operation performed by the first depthwise convolution layer is depthwise separable convolution. Like ordinary convolution, depthwise separable convolution has the ability to extract features. The depthwise convolution is responsible for capturing spatial features, so each time a depthwise separable convolution is performed, deeper, more advanced, and more abstract feature information can be extracted.

[0095] The second batch processing layer is used to normalize the initial depth convolution feature vector;

[0096] The second batching layer can be implemented using a batch normalization layer, which normalizes each mini-batch, accelerating model convergence and helping alleviate gradient vanishing and gradient exploding problems. The number of initial depthwise convolutional feature vectors can be multiple, determined by the number of convolution kernels in the current layer and the number of input feature vectors.

[0097] The maximum pooling layer is used to downsample the multiple initial depth convolution feature vectors to extract edge information;

[0098] The second activation function layer is used to accelerate the normalization of the initial deep convolution feature vector to obtain a normalized first-class deep convolution feature vector with high-level semantic features or a second-class deep convolution feature vector with high-level semantic features.

[0099] Among them, the second activation function layer can be implemented using the relu function.

[0100] Through the above process, deep convolution can replace traditional convolution. On the basis of obtaining more comprehensive high-level semantic features, the number of layers of ordinary convolution can be reduced, thereby improving the computing speed of the model.

[0101] Step S213: input the first type of deep convolutional feature vector into the second depthwise separable convolutional layer 3 to perform a second depth feature extraction to obtain a second type of deep convolutional feature vector with high-level semantic features;

[0102] Optionally, the second depthwise separable convolutional layer 3 includes a first depthwise convolutional layer, a second batching layer, a maximum pooling layer, and a second activation function connected in sequence;

[0103] The first depth convolution layer is used to perform a depth convolution operation on the input convolution feature vector to obtain an initial depth convolution feature vector with high-level semantic features;

[0104] Among them, the operation performed by the first depthwise convolution layer is depthwise separable convolution. Like ordinary convolution, depthwise separable convolution has the ability to extract features. The depthwise convolution is responsible for capturing spatial features, so each time a depthwise separable convolution is performed, deeper, more advanced, and more abstract feature information can be extracted.

[0105] The second batch processing layer is used to normalize the multiple initial depth convolution feature vectors;

[0106] Among them, the second batch processing layer can be implemented using the BN layer, which can normalize each mini-batch, which can not only accelerate the convergence of the model but also help alleviate the gradient disappearance and gradient explosion problems.

[0107] The maximum pooling layer is used to downsample the multiple initial depth convolution feature vectors to extract edge information;

[0108] The second activation function layer is used to accelerate the normalization of the initial deep convolution feature vector to obtain a normalized first-class deep convolution feature vector with high-level semantic features or a second-class deep convolution feature vector with high-level semantic features.

[0109] Among them, the second activation function layer can be implemented using the relu function.

[0110] Through the above process, deep convolution can replace traditional convolution. On the basis of obtaining more comprehensive high-level semantic features, the number of layers of ordinary convolution can be reduced, thereby improving the computing speed of the model.

[0111] Optionally, depthwise separable convolution includes channel-by-channel convolution and point-by-point convolution. As shown in Figure 7, depthwise separable convolution (DSC) consists of two parts: depthwise (DW) channel-by-channel convolution and pointwise (PW) point-by-point convolution. Channel-by-channel convolution is actually n*n convolution. In channel-by-channel convolution, a convolution kernel is responsible for only one channel, and a channel is convolved by only one convolution kernel. Because the number of output channels of the convolution operation is equal to the number of convolution kernels, a feature vector with n channels of 1 is obtained. The outputs of all convolution kernels are then concatenated to obtain an output feature vector with n channels, where n is a natural number greater than or equal to 1. Point-by-point convolution is actually 1×1 convolution, which is used to freely change the number of output feature vector channels. It also performs channel fusion on the feature vector obtained by the main channel convolution.

[0112] In the above implementation, depthwise convolution is responsible for capturing spatial features, while pointwise convolution is responsible for integrating information between channels. The number of convolution kernels in the pointwise convolution also changes the number of channels in the feature vector. Therefore, each depthwise separable convolution extracts deeper, more advanced, and more abstract feature information.

[0113] Depthwise convolution is responsible for capturing spatial features, while pointwise convolution is responsible for integrating information between channels. The number of convolution kernels in the pointwise convolution can also change the number of channels in the feature vector. Therefore, each depthwise separable convolution can extract deeper, more advanced, and more abstract feature information.

[0114] The computational complexity of MSA and SW-MSA can be roughly estimated as follows: Ω(MSA)=4hwC 2 +2(hw) 2 C Ω(W-MSA)=4hwC 2 +2M 2 HkDJ

[0115] Where h represents the height of the feature vector, w represents the width of the feature vector, C represents the number of channels of the feature vector, and M represents the size of each window.

[0116] In the above optional implementation scheme, the two depth-wise separable convolutions are both used to extract feature information, while reducing the number of parameters compared to ordinary convolution.

[0117] Step S214: input the second type of deep convolution feature vector into the first self-attention layer 4 to perform a first self-attention calculation to obtain a first self-attention feature vector with global features;

[0118] Optionally, the first self-attention layer 4 and the second self-attention layer 5 both include a window-based multi-head attention module and a moving window-based multi-head self-attention module;

[0119] The window-based multi-head attention module is used to perform single-window self-attention calculation on the input feature vector to obtain an initial self-attention feature vector;

[0120] The moving window-based multi-head self-attention module is used to perform a moving window multi-head self-attention calculation on the input initial self-attention feature vector to obtain the first self-attention feature vector with global features or the second self-attention feature vector with global features.

[0121] Step S215: Input the first self-attention feature vector into the second self-attention layer 5 for a second self-attention calculation to obtain the self-attention feature vector with global features.

[0122] Optionally, the first self-attention layer 4 and the second self-attention layer 5 both include a window-based multi-head attention module and a moving window-based multi-head self-attention module;

[0123] The window-based multi-head attention module is used to perform single-window self-attention calculation on the input feature vector to obtain an initial self-attention feature vector;

[0124] The moving window-based multi-head self-attention module is used to perform a moving window multi-head self-attention calculation on the input initial self-attention feature vector to obtain the first self-attention feature vector with global features or the second self-attention feature vector with global features.

[0125] In the above embodiment, the first self-attention layer 4 (Swin Transformer Blocks) or the second self-attention layer 5, as shown in Figure 5, is composed of two consecutive self-attention modules (Transformer). The first module uses a window-based multi-head attention module (Window-based Multihead Self-attention, W-MSA). Compared with the ordinary MSA module, W-MSA divides the feature vector into windows and performs self-attention calculations within the windows, which greatly reduces the amount of calculation. The second module uses a shifted window-based multihead self-attention module (SW-MSA). Information exchange is carried out between windows through a moving window mechanism, which solves the problem that information cannot be transferred between windows in the feature vector (dividing a feature vector into multiple channels (patches), and then transferring information between each patch and patch).

[0126] Optionally, the initial image is H*W, the first convolutional layer 1 performs a convolution operation to obtain a feature vector with parameters H / 2*W / 2*c1, the first depth-wise separable convolutional layer 2 performs a convolution operation to obtain a feature vector with parameters H / 4*W / 4*c2, the second depth-wise separable convolutional layer 3 performs a convolution operation to obtain a feature vector with parameters H / 8*W / 8*c3, the first self-attention layer 4 performs self-attention calculation to obtain a feature vector with parameters H / 16*W / 16*c4, and the second self-attention layer 5 performs self-attention calculation to obtain a feature vector with parameters H / 32*W / 32*c5, where H / 16, H / 8, H / 4, and H / 2 are the heights of the feature vectors, W / 16, W / 8, W / 4, and W / 2 are the widths of the feature vectors, and c4, c3, c2, and c1 are the number of image channels of the feature vectors.

[0127] In an optional embodiment, the decoder module includes a first depthwise separable convolutional layer 6, a first depthwise convolution-self-attention layer 7, a second depthwise convolution-self-attention layer 8, a second depthwise separable convolutional layer 9, and a convolutional output layer 10;

[0128] The decoder module segments the target surgical site image according to the first feature vector having low-level semantic features, the deep convolution feature vector, and the self-attention feature vector, as shown in FIG6 , further comprising:

[0129] The first depthwise separable convolutional layer 6 is configured to preliminarily restore the spatial resolution of the second self-attention feature vector to obtain a third depthwise convolutional vector;

[0130] The first depth convolution-self-attention layer 7 is jump-connected to the first self-attention layer and is used to perform secondary resolution recovery and feature fusion according to the third depth convolution vector and the global feature to obtain a fourth depth convolution vector;

[0131] The second depth convolution-self-attention layer 8 is jump-connected to the second depth separable convolution layer and is used to perform further resolution recovery and feature fusion according to the fourth depth convolution vector and the second high-level semantic feature to obtain a fifth depth convolution vector;

[0132] The second depthwise separable convolution layer 9 is jump-connected to the first depthwise separable convolution layer and is used to perform further resolution restoration and feature fusion according to the fifth depthwise convolution vector and the first high-level semantic feature to obtain a sixth depthwise convolution vector;

[0133] The convolution output layer 10 is used to obtain the target surgical site image after performing depthwise convolution on the sixth depthwise convolution vector.

[0134] Among them, the combined operation of the first deep convolution-self-attention layer 7 or the second deep convolution-self-attention layer 8 can gradually restore the spatial resolution of the image and extract relevant features. In the feature extraction process, deep convolution is used instead of traditional convolution to increase the speed, and Swin Transformer Blocks is introduced to extract global information and high-level semantic features, thereby ensuring the segmentation accuracy.

[0135] Optionally, the first depthwise separable convolution layer 6 performs a convolution operation to obtain a feature vector with parameters H / 16*W / 16*c4, the first depthwise convolution-self-attention layer 7 performs a convolution operation to obtain a feature vector with parameters H / 8*W / 8*c3, the second depthwise convolution-self-attention layer 8 performs a convolution operation to obtain a feature vector with parameters H / 4*W / 4*c2, the second depthwise separable convolution layer 9 performs a convolution operation to obtain a feature vector with parameters H / 2*W / 2*c1, and the convolution output layer 10 outputs a target surgical site image with parameters H*W. Here, H / 16, H / 8, H / 4, and H / 2 are the heights of the feature vectors, W / 16, W / 8, W / 4, and W / 2 are the widths of the feature vectors, and c4, c3, c2, and c1 are the number of image channels of the feature vectors.

[0136] It should be noted that in the present application, the feature vectors output by each layer are determined by the number of input feature vectors and the number and type of convolution kernels, and are not limited to one or more.

[0137] In an optional embodiment, the first depth convolution-self-attention layer 7 or the second depth convolution-self-attention layer 8 both include:

[0138] A third self-attention layer is used to perform self-attention calculation on the input third depth convolution vector to obtain a third self-attention feature vector;

[0139] Among them, the specific structure of the third self-attention layer is implemented with reference to the first self-attention layer 4 and the second self-attention layer 5, and the technical effects are the same, so they will not be repeated here.

[0140] A second depthwise convolutional layer, configured to perform a depthwise convolution on the third self-attention feature vector to obtain a third self-attention-depthwise convolution feature vector;

[0141] Among them, the specific structure of the second depth convolution layer is implemented with reference to the first depth convolution layer, and the technical effect is the same, so it will not be repeated here.

[0142] The third batching layer is used to normalize the third self-attention-depth convolution feature vector;

[0143] Among them, the specific structure of the third batch processing layer is implemented with reference to the first batch processing layer or the second batch processing layer, and the technical effects are the same, which will not be repeated here.

[0144] The third activation function layer is used to accelerate the normalization of the self-attention-third deep convolution feature vector to obtain the normalized third deep convolution feature vector with high-level semantic features or the fourth deep convolution feature vector with high-level semantic features.

[0145] Among them, the third self-attention layer uses the Swin Transformer Block to extract global features, and then uses the second deep convolution layer to extract local features, and fuses the two to achieve more accurate segmentation. At this time, deep convolution is used to extract the spatial features of the feature vector, and the third batching layer is used to normalize the feature vector, accelerate network training and enhance the generalization ability of the model. Upsampling is to reduce the spatial size of the feature vector, and the third activation function layer is to introduce nonlinearity to enhance the expression ability of the model. As a result, the input of the first deep convolution-self-attention layer 7 or the second deep convolution-self-attention layer 8 is the feature vector from the previous layer of the network, and the output of the first deep convolution-self-attention layer 7 or the second deep convolution-self-attention layer 8 is a feature vector larger than the input feature vector.

[0146] In this application, a feature vector generally refers to a feature map, which has attributes of position, size, and features.

[0147] The lightweight segmentation model is trained through back propagation.

[0148] At this time, model training updates the training parameters through the error between the output value and the true value. Specifically, backpropagation calculates the gradient of each layer parameter through reverse flow based on the loss function.

[0149] The above lightweight segmentation model significantly reduces the amount of floating-point operations and parameters while maintaining high segmentation accuracy, improves the inference speed, shortens the inference time, and is more in line with the needs of actual production implementation inference.

[0150] To further illustrate the beneficial effects of this solution, experiments were conducted on the CholecSeg8k dataset and compared with classic deep learning models and lightweight models. The results are shown in Table 1:

[0151] Table 1

[0152] In general, our model Dice is basically on par with other models and maintains a high level, but has a huge advantage in FLOPs and Params. Specifically, FLOPs is only 1 / 2000 of UNet.

[0153] The number of params is only 1 / 64 of that of UNet. Compared to TransUNet, our model uses only 1 / 1200 of the floating-point operations and 1 / 186 of the parameters. Compared to UNeXt, another lightweight segmentation network, our model surpasses them in all three parameters. Not only does it lead on Dice, but its number of parameters is only 1 / 3 and its floating-point operations are only 1 / 22 of UNeXt.

[0154] While the number of parameters and floating-point operations can be used to assess a model's computational complexity to a certain extent, inference time is a more accurate metric for truly reflecting a model's speed in real-world applications. Parameters and floating-point operations primarily measure the model's scale and computational requirements, but they do not directly account for factors such as the actual hardware speed, data transmission, and latency during computation. Therefore, to comprehensively evaluate model performance, inference time must also be considered as a metric. Calculating inference time is particularly important for lightweight models, as these models are often designed to run in environments with limited computing resources, such as mobile devices or embedded systems. To minimize the impact of latency caused by calling the timing function, we employ the following timing strategy: we read in an image, record the start time, then loop over the image 10,000 times, and record the end time at the end of the loop. This results in only two calls to the timing function for every 10,000 inferences, maximizing the time actually spent on inference. Since most hospitals' computing equipment does not have GPUs, we calculated the average inference time not only on server GPUs but also on CPUs to maximize the simulation of real hospital surgery scenarios. The average inference time for each model is shown in Table 2:

[0155] Table 2

[0156] From the above statistics, we can see that on the CPU, our inference time is only 47.49ms, which is 1 / 40 of Unet and 1 / 14 of DeepLabv3+. The inference speed is 3.5 times that of MobileNetV3 and 2.4 times that of UNext. This inference speed can meet the needs of real-time inference.

[0157] FIG8 is a schematic structural diagram of an embodiment of a lightweight segmentation device according to the present invention. The specific embodiment of the present invention does not limit the specific implementation of the lightweight segmentation device.

[0158] As shown in FIG8 , the lightweight segmentation device may include: a processor 402 , a communications interface 404 , a memory 406 , and a communication bus 408 .

[0159] Processor 402, communication interface 404, and memory 406 communicate with each other via communication bus 408. Communication interface 404 is used to communicate with other devices, such as clients or other server network elements. Processor 402 is used to execute program 410, which may specifically perform the steps described in the aforementioned embodiment of the lightweight segmentation method.

[0160] Specifically, the program 410 may include program code including computer-executable instructions.

[0161] Processor 402 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement embodiments of the present invention. The one or more processors included in the lightweight segmentation device may be processors of the same type, such as one or more CPUs, or may be processors of different types, such as one or more CPUs and one or more ASICs.

[0162] The memory 406 is used to store the program 410. The memory 406 may be a high-speed RAM memory, or may also include a non-volatile memory, such as at least one disk memory.

[0163] The program 410 may be specifically called by the processor 402 to enable the lightweight segmentation device to perform the operations of the above lightweight segmentation method.

[0164] It should be noted that, since the lightweight segmentation device of the present application can implement all embodiments of the lightweight segmentation method, the lightweight segmentation device of the present application has all the beneficial effects of the lightweight segmentation method, which will not be described in detail here.

[0165] An embodiment of the present invention provides a storage medium storing at least one executable instruction. When the executable instruction is executed on a lightweight segmentation device / apparatus, the lightweight segmentation device / apparatus executes the lightweight segmentation method in any of the above method embodiments.

[0166] It should be noted that, since the storage medium of the present application can implement all embodiments of the lightweight segmentation method, the storage medium of the present application has all the beneficial effects of the lightweight segmentation method, which will not be described in detail here.

[0167] The algorithms or displays provided herein are not inherently related to any particular computer, virtual system or other device. In addition, the embodiments of the present invention are not directed to any particular programming language.

[0168] In the description provided herein, numerous specific details are set forth. However, it is understood that embodiments of the present invention may be practiced without these specific details. Similarly, in order to streamline the present invention and aid in understanding one or more of the various inventive aspects, in the above description of exemplary embodiments of the present invention, various features of embodiments of the present invention are sometimes grouped together into a single embodiment, figure, or description thereof. The claims that follow the detailed description are hereby expressly incorporated into that detailed description, with each claim itself serving as a separate embodiment of the present invention.

[0169] Those skilled in the art will appreciate that the modules in the devices of the embodiments can be adaptively changed and installed in one or more devices different from the embodiments. The modules, units, or components in the embodiments can be combined into one module, unit, or component, and furthermore, they can be divided into multiple submodules, subunits, or subcomponents. Except that at least some of such features and / or processes or units are mutually exclusive.

[0170] It should be noted that the above embodiments illustrate rather than limit the invention, and that alternative embodiments may be devised by a person skilled in the art without departing from the scope of the appended claims. In the claims, any reference signs placed between brackets should not be construed as limiting the claims. The word "" does not exclude the presence of elements or steps not listed in the claims. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present invention may be implemented by means of hardware comprising several different elements and by means of appropriately programmed computers. In a unit claim enumerating several means, several of these means may be embodied by the same item of hardware. The use of the words first, second, and third etc. does not indicate any order. These words may be interpreted as names. The steps in the above embodiments should not be understood as limiting the order of execution unless otherwise specified.

Claims

1. A lightweight segmentation method, It is characterized in that The method comprises: Get real-time images; Inputting the real-time image into a trained lightweight segmentation model to perform target segmentation to obtain a target surgical site image; The lightweight segmentation model includes an encoder module and a decoder module. The step of inputting the real-time image into the trained lightweight segmentation model for target segmentation to obtain the target surgical site image further includes: Inputting the real-time image into the encoder module for feature information extraction to obtain a first feature vector having low-level semantic features, a self-attention feature vector having global features, and at least one type of deep convolution feature vector having high-level semantic features; The self-attention feature vector is output to the decoder module, and the deep convolution feature vector is output to the decoder module via a skip connection; The decoder module segments the target surgical site image according to the first feature vector having low-level semantic features, the deep convolution feature vector, and the self-attention feature vector.

2. The lightweight segmentation method according to claim 1, Features: The encoder module includes a first convolution layer, a first depth-separable convolution layer, a second depth-separable convolution layer, a first self-attention layer, and a second self-attention layer: the depth convolution feature vector includes a first type of depth convolution feature vector and a second type of depth convolution feature vector; The step of inputting the real-time image into the encoder module for feature information extraction to obtain a first feature vector having a low-level semantic feature, a self-attention feature vector having a global feature, and at least one type of deep convolution feature vector having a high-level semantic feature further comprises: Inputting the real-time image into the first convolution layer to extract surface features to obtain a first convolution feature vector; Inputting the first convolution feature vector into the first depth-separable convolution layer for a first depth feature extraction to obtain the first type of depth convolution feature vector with high-level semantic features; Inputting the first type of deep convolutional feature vector into the second deep separable convolutional layer for a second deep feature extraction to obtain the second type of deep convolutional feature vector with high-level semantic features; Inputting the second type of deep convolution feature vector into the first self-attention layer to perform a first self-attention calculation to obtain the first self-attention feature vector with global features; The first self-attention feature vector is input into the second self-attention layer for a second self-attention calculation to obtain the self-attention feature vector with global features.

3. The lightweight segmentation method according to claim 2, Features: The first convolutional layer includes a common convolutional layer, a first batch of processing layers and a first activation function layer connected in sequence; The common convolution layer is used to perform a convolution operation on the real-time image to obtain an initial feature vector; The first batch of processing layers is used to normalize the initial feature vector; The first activation function layer is used to accelerate the normalization of the initial feature vector to obtain the first convolution feature vector after normalization.

4. The lightweight segmentation method according to claim 2, Features: The first depthwise separable convolutional layer and the second depthwise separable convolutional layer each separately include a first depthwise convolutional layer, a second batching layer, a maximum pooling layer, and a second activation function connected in sequence; The first deep convolution layer is used to perform a deep convolution operation on the input convolution feature vector to obtain an initial deep convolution feature vector with high-level semantic features; The second batch processing layer is used to normalize the initial deep convolution feature vector; The maximum pooling layer is used to downsample the initial depth convolution feature vector to extract edge information; The second activation function layer is used to accelerate the normalization of the initial deep convolution feature vector to obtain the normalized The first type of deep convolution feature vector with high-level semantic features or the second type of deep convolution feature vector with high-level semantic features after normalization.

5. The lightweight segmentation method according to claim 2, Features: The first self-attention layer and the second self-attention layer both include a window-based multi-head attention module and a moving window-based multi-head self-attention module; The window-based multi-head attention module is used to perform a single-window self-attention calculation on the input second-type deep convolution feature vector or the first self-attention feature vector to obtain an initial self-attention feature vector; The moving window-based multi-head self-attention module is used to perform a moving window multi-head self-attention calculation on the input initial self-attention feature vector to obtain the first self-attention feature vector with global features or the second self-attention feature vector with global features.

6. The lightweight segmentation method according to claim 1, Features: The decoder module includes a first depthwise separable convolutional layer, a first depthwise convolution-self-attention layer, a second depthwise convolution-self-attention layer, a second depthwise separable convolutional layer, and a convolutional output layer; The decoder module segments the target surgical site image according to the first feature vector having low-level semantic features, the deep convolution feature vector, and the self-attention feature vector, further comprising: The first deep separable convolution layer is used to preliminarily restore the spatial resolution of the second self-attention feature vector to obtain a third deep convolution vector; The first deep convolution-self-attention layer is jump-connected to the first self-attention layer and is used to perform secondary resolution recovery and feature fusion according to the third deep convolution vector and the first self-attention feature vector to obtain a fourth deep convolution vector; The second deep convolution-self-attention layer is jump-connected to the second deep separable convolution layer, and is used to perform resolution recovery and feature fusion again according to the fourth deep convolution vector and the second type of deep convolution feature vector to obtain a fifth deep convolution vector; The second depth-separable convolution layer is jump-connected to the first depth-separable convolution layer, and is used to perform resolution recovery and feature fusion again according to the fifth depth-convolution vector and the first type of depth-convolution feature vector to obtain a sixth depth-convolution vector; The convolution output layer is used to obtain the target surgical site image after performing depth convolution on the sixth depth convolution vector.

7. The lightweight segmentation method according to claim 6, Features: The first deep convolution-self-attention layer or the second deep convolution-self-attention layer includes a third self-attention layer, a second deep convolution layer, a third batching layer, a second size adjustment layer and a third activation function layer: A third self-attention layer, used to perform self-attention calculation on the input third depth convolution vector to obtain a third self-attention feature vector; A second deep convolutional layer, used to perform deep convolution on the third self-attention feature vector to obtain a third self-attention-deep convolution feature vector; The third batching layer is used to normalize the third self-attention-depth convolution feature vector; The second size adjustment layer is used to adjust the size of the self-attention-third deep convolution feature vector; The third activation function layer is used to accelerate the normalization of the self-attention-third deep convolutional feature vector to obtain the normalized third deep convolutional feature vector with high-level semantic features or the fourth deep convolutional feature vector with high-level semantic features.

8. The lightweight segmentation method according to claim 1, Features: The lightweight segmentation model is trained by back propagation.

9. A lightweight segmentation device, It is characterized in that include: A processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other via the communication bus; The memory is used to store at least one executable instruction, and the executable instruction enables the processor to perform the operation of the lightweight segmentation method according to any one of claims 1-6.

10. A storage medium, It is characterized in that The storage medium stores at least one executable instruction. When the executable instruction is executed on the lightweight segmentation device / apparatus, the lightweight segmentation device / apparatus performs the operation of the lightweight segmentation method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Lightweight multi-scale feature fusion real-time image semantic segmentation method and system

    CN114445430A

  • Image segmentation method and device, storage medium and electronic device

    CN116402996A

  • Remote sensing image semantic segmentation improved algorithm based on context attention mechanism

    CN116912495A

  • Image semantic segmentation algorithm and system based on multi-channel deep weighted aggregation

    US20230316699A1

Cited By

  • Low-temperature boat hanging frame fault diagnosis method and system based on multi-mode dynamic fusion

    CN121117981A

  • Low-temperature davit frame fault diagnosis method and system based on multi-modal dynamic fusion

    CN121117981B