A light-weight catheter segmentation method based on temporal gating
Patent Information
- Application Number
- CN202610947596.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-29
- Publication Date
- 2026-10-09
- Estimated Expiration
- 2046-06-29
AI Technical Summary
传统方法依赖边缘检测算子和路径搜索算法,在复杂背景下精度不足
1、本发明基于LS卷积设计了全新的LS模块作为编码器核心单元,通过大核感知(LKP)捕获全局上下文并生成空间自适应动态权重,再利用该权重驱动小核聚合(SKA)对局部邻域进行分组动态卷积,在极致压缩模型参数量的同时,显著增强了对导丝细长、低对比度、易遮挡结构的特征捕捉能力。
Smart Images

Figure CN122473464B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer vision technology, and in particular to a lightweight guidewire segmentation method based on time-gated control. Background Technology
[0002] Percutaneous coronary intervention (PCI) is a core minimally invasive treatment for coronary artery stenosis / occlusion diseases such as coronary heart disease and acute myocardial infarction. The core of the procedure relies on the guidewire to precisely reach the lesion site and achieve vascular recanalization. To avoid damaging the vessel wall and causing serious complications such as perforation, the movement and positioning of the guidewire within the vessel must have extremely high precision. Therefore, precise guidewire segmentation under intraoperative X-ray sequences is a core prerequisite for successful PCI surgical navigation and robot-assisted operation. However, guidewire segmentation in clinical settings faces multiple severe challenges: the thin guidewire in X-ray images naturally suffers from severe class imbalance and low signal-to-noise ratio problems; the patient's heartbeat and respiration introduce a large number of motion artifacts; at the same time, the guidewire is easily obscured by anatomical structures such as blood vessels and ribs; similar structures can easily lead to misclassification of the guidewire; and in the case of dual guidewires in complex CTO lesions, there is also the segmentation problem caused by guidewire crossing and overlap.
[0003] Existing guidewire segmentation methods are mainly divided into traditional image processing methods and deep learning methods. Traditional methods rely on edge detection operators and path search algorithms, which lack accuracy in complex backgrounds. Among deep learning methods, models based on convolutional neural networks (CNNs), such as U-Net, have achieved some success, but are limited by local receptive fields and struggle to capture global dependencies. While Transformer-based models can effectively model the global context, their self-attention mechanism has enormous computational complexity and parameter count, making it difficult to meet the real-time requirements of clinical scenarios. Current lightweight networks, although reducing the number of parameters, often use single-frame images as input, completely ignoring the valuable temporal motion information between consecutive frames in X-ray fluoroscopy sequences, thus limiting further improvements in segmentation accuracy. Summary of the Invention
[0004] In view of the above, the main objective of this invention is to propose a time-gated lightweight guidewire segmentation method to solve the aforementioned technical problems.
[0005] This invention proposes a time-gated lightweight guidewire segmentation method, which includes the following steps: Step 1: Obtain the original X-ray guidewire image sequence and perform frame extraction on the original X-ray guidewire image sequence to obtain the input image sequence; further process the input image sequence to obtain the underlying spatial feature map; Step 2: Input the bottom-level spatial feature map into the encoder and perform temporal coding processing to obtain the coding spatial features of each level; recursively process the coding spatial features of each level to obtain the temporal enhancement features of each level. Step 3: Use the decoder to decode the temporal enhancement features of each level to obtain the decoded feature map at the original resolution; perform segmentation prediction on the decoded feature map at the original resolution to obtain the guide wire segmentation prediction map.
[0006] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. This invention designs a novel LS module as the core unit of the encoder based on LS convolution. It captures the global context and generates spatially adaptive dynamic weights through large kernel perception (LKP). Then, it uses these weights to drive small kernel aggregation (SKA) to perform grouped dynamic convolution on local neighborhoods. While maximally compressing the number of model parameters, it significantly enhances the feature capture ability of structures with thin guide wires, low contrast, and easy occlusion.
[0007] 2. This invention embeds a ConvGRU temporal encoder at the skip connections of each encoder level. Through update and reset gates, it performs gated fusion of the current frame's spatial features and historical hidden states. This preserves the pixel spatial structure while modeling the motion dependency of the guidewire between consecutive frames, thereby enhancing temporal continuity and segmentation stability in low-contrast, occluded, and overlapping scenarios. Furthermore, this temporal encoder employs a convolutional recursive structure, which has lower parameters and computational overhead compared to complex long-sequence modeling methods, making it suitable for real-time guidewire segmentation tasks. The interface is compatible with other mainstream temporal models, allowing for scheme adjustments without modifying the backbone network. This significantly improves adaptability to clinical scenarios with varying computing power and provides a general design approach for similar medical sequence segmentation tasks.
[0008] 3. This invention expands the input from a single-frame image to a short sequence of three consecutive frames, embeds a temporal gating module at the jump connection, and models the inter-frame continuity of the guidewire movement trajectory through a convolutional gated recurrent unit. Without compromising the lightweight characteristics of the backbone network, it fully exploits the amplifying effect of temporal dynamic information on guidewire segmentation, effectively improving the model's segmentation robustness in complex scenarios such as local guidewire occlusion and crossover, and providing reliable technical support for real-time guidewire positioning and navigation in interventional surgery. Attached Figure Description
[0009] Figure 1 This is a flowchart of the time-gated lightweight guidewire segmentation method proposed in this invention; Figure 2 This is a schematic diagram of the timing encoder structure for the lightweight guidewire segmentation method based on timing gating proposed in this invention; Figure 3This is a schematic diagram of the LS convolution module architecture of the time-gated lightweight guidewire segmentation method proposed in this invention. Detailed Implementation
[0010] Embodiments of the present invention are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.
[0011] These and other aspects of the embodiments of the present invention will become clear from the following description and accompanying drawings. In these descriptions and drawings, some specific embodiments of the present invention are specifically disclosed to illustrate some ways of implementing the principles of the embodiments of the present invention; however, it should be understood that the scope of the embodiments of the present invention is not limited thereto.
[0012] Please see Figure 1 This invention proposes a time-gated lightweight guidewire segmentation method, which includes the following steps: Step 1: Obtain the original X-ray guidewire image sequence and perform frame extraction on the original X-ray guidewire image sequence to obtain the input image sequence; further process the input image sequence to obtain the underlying spatial feature map; In step 1, the original X-ray guidewire image sequence is acquired, and frame extraction is performed on the original X-ray guidewire image sequence to obtain the input image sequence; the input image sequence is further processed to obtain the low-level spatial feature map, and the specific steps are as follows: Obtain a preset number of guidewire images from the original X-ray guidewire image sequence; arrange all guidewire images in chronological order to obtain the input image sequence; The batch dimension and frame number dimension of the input image sequence are merged to reshape it, resulting in a stacked feature map with merged dimensions. The stacked feature maps of the merged dimensions are processed by using a 3×3 convolution with a stride of 1 and padding of 1. Then, batch normalization and ReLU activation are performed sequentially to obtain the bottom-level spatial feature maps.
[0013] Furthermore, this step is the encoder stage, which consists of a stem module for preliminary feature extraction and five LS encoders arranged from top to bottom. At the same time, the stem module consists of a 3×3 convolution (stride of 1, padding of 1), batch normalization, and ReLU activation function, and the features output by the stem module will be used as the input of the first LS encoder.
[0014] Step 2: Input the bottom-level spatial feature map into the encoder and perform temporal coding processing to obtain the coding spatial features of each level; recursively process the coding spatial features of each level to obtain the temporal enhancement features of each level. Please see Figure 2 and Figure 3 In step 2, the bottom-level spatial feature map is input to the encoder and subjected to temporal coding to obtain the coding spatial features of each level; the coding spatial features of each level are recursively processed to obtain the temporal enhancement features of each level. The specific steps are as follows: Step 201: The bottom spatial feature map is sequentially processed by channel compression pointwise convolution, 7×7 depth separable convolution, channel restoration pointwise convolution, and output projection layer to obtain spatial adaptive dynamic weights. The underlying spatial feature map is subjected to grouped dynamic convolution processing to obtain channel aggregated feature representation; The aggregated feature representations of each channel and each spatial position are arranged according to the original space to obtain a complete aggregated feature map. Based on spatial adaptive dynamic weights, the complete aggregated feature map is batch normalized and then added element-wise with the underlying spatial feature map to obtain the aggregated output feature map. Channel attention is calculated on the aggregated output feature map to obtain intermediate features; The intermediate features are processed by residual connections using a feedforward network to obtain the final output feature map of the first level. The final output feature map of the first level is processed sequentially by convolution, max pooling, GELU activation function and batch normalization to obtain the coding space features of the first level. Step 202: Use the coding space features of the first level as the bottom space feature map, and repeat step 201 to obtain the coding space features of the second level. The second-level coding space features are used as the bottom-level spatial feature map. Step 201 is repeated to obtain the third-level coding space features. Using the third-level coding space features as the bottom-level spatial feature map, repeat step 201 to obtain the fourth-level coding space features. Using the fourth-level coding space features as the bottom-level spatial feature map, repeat step 201 to obtain the fifth-level coding space features. Step 203: The four-dimensional tensors of the coding space features of the first level, the second level, the third level, the fourth level, and the fifth level are reshaped into five-dimensional sequence forms, and then input into the corresponding level temporal encoders for recursive calculation to obtain the temporal enhancement features of the first level, the second level, the third level, the fourth level, and the fifth level.
[0015] Large-Small Convolution (LS) is the core component of the LS encoder block. Inspired by human visual mechanisms, it is a novel lightweight multi-scale convolutional structure for guidewire segmentation. Unlike standard convolution, LS convolution employs a two-stage design strategy of "first capturing the large, then focusing on the small." It first captures global contextual information through a large receptive field convolution, generating spatially adaptive dynamic weights. Then, the dynamic weights guide the small kernel convolution to focus on the fine local structure of the guidewire. With extremely low parameter count, it simultaneously achieves long-range dependent capture and high-frequency detail extraction (guidewire edges, textures, and curvature), which is highly compatible with the needs of guidewire segmentation tasks.
[0016] However, while standard large-kernel convolution can expand the receptive field, it leads to a sharp increase in the number of parameters. Dynamic convolution, on the other hand, can achieve feature extraction effects with a large receptive field under small kernel size by spatially adaptive weights. Therefore, this invention combines large-kernel depthwise separable convolution with grouped dynamic convolution to construct a two-level architecture of LS convolution. At the same time, it introduces an SE channel attention module to adaptively enhance key channel features, suppress redundant information, and further improve the ability to capture complex textures and fine structures of the guidewire.
[0017] The large kernel perception stage mainly generates context-dependent dynamic weights for each spatial location. Specifically, the low-level spatial feature map is sequentially processed by channel compression pointwise convolution, 7×7 depthwise separable convolution, channel restoration pointwise convolution, and output projection layer to obtain spatially adaptive dynamic weights. The corresponding relationship in this process is as follows: ; in, Represents the space-adaptive dynamic weights. This indicates the processing stage of the big-core sensing phase. This represents the bottom-level spatial feature map of the (m-1)th level; This represents the third 1×1 pointwise convolutional processing in the large kernel perception stage, which acts as a projection layer to achieve the process from... arrive Dimensional mapping; This indicates the second 1×1 pointwise convolution process in the large kernel perception stage, used to maintain the original number of channels unchanged; This indicates the first 1×1 pointwise convolution process in the big kernel perception stage, which is used to compress the number of channels to half of the original number of channels to reduce the computational load of subsequent operations. This indicates that the kernel size is Depth-separable convolution processing is used to capture a wide range of context at each spatial location, in this invention... ; The spatial size of the target dynamic convolution kernel is represented in this invention. ; This indicates the number of groups into which the channel is divided. This represents the channel dimension of the input feature map. Indicates batch, Represents the spatial height of the feature map. This represents the spatial width of the feature map.
[0018] In the small kernel aggregation stage, the global weight tensor generated by the large kernel perception is used to perform grouped dynamic convolution on the input features, achieving local fine feature enhancement of the guide wire under the guidance of the global context. The channel dimension of the input features is divided into G groups, each containing C / G channels. Channels within the same group share aggregation weights, further reducing the number of parameters and computational complexity. Specifically, in the process of performing grouped dynamic convolution on the low-level spatial feature map to obtain the channel aggregated feature representation, the following relationship exists: ; in, The channel aggregation feature representation of the i-th spatial location and the c-th channel. The size centered at the i-th spatial location on the bottom-level spatial feature map of any level is . ; This represents the i-th spatial location on the input feature map, i.e., the location of a pixel / feature point in the two-dimensional feature map; This represents the dynamic convolution kernel weights corresponding to the g-th channel at the i-th spatial location. This indicates a 3×3 convolution process; Based on spatial adaptive dynamic weights, the complete aggregated feature map is batch normalized and then added element-wise to the underlying spatial feature map to obtain the aggregated output feature map. The corresponding relationship in this process is as follows: ; in, This represents the aggregated output feature map obtained from the bottom-level spatial feature maps of level m-1. This represents the LS convolution operation. This indicates batch normalization processing. This indicates that the small core aggregation module is processing the data.
[0019] Channel attention is calculated on the aggregated output feature map to obtain intermediate features. The corresponding relationship in this process is as follows: ; in, This indicates that the intermediate features are obtained through LS convolution processing and SE channel attention operations in the m-th level LS encoder. This indicates channel attention calculation; A feedforward network is used to perform residual connection processing on the intermediate features to obtain the final output feature map of the first layer. The corresponding relationship in this process is as follows: ; in, This represents the final output feature map of the m-th level LS encoder. This indicates feedforward network processing. This means adding the residuals of the feedforward network output to the intermediate input features; The final output feature map of the first level is sequentially processed by convolution, max pooling (to adjust the number of channels and resolution to the input specifications required by the next LS encoder), GELU activation function, and batch normalization to obtain the coding space features of the first level. The corresponding relationship in this process is as follows: ; in, This represents the coding space features of the m-th level. This indicates the GELU activation function processing. This indicates a 3×3 standard convolution process.
[0020] This invention embeds plug-and-play time encoders in parallel across the five layers of the encoder. The time encoders adopt a universal and replaceable interface design, which can flexibly integrate various time modeling strategies such as ConvGRU, ConvLSTM, and TCN according to requirements. This allows for adaptation to different time information mining needs without changing the overall architecture of the backbone network. Please see Figure 2 As a preferred embodiment, the present invention uses a ConvGRU as a timing encoder, the structure of which is as follows: Figure 2 As shown; specifically, the four-dimensional tensors of the coding space features of the first, second, third, fourth, and fifth levels are reconstructed into a five-dimensional sequence form, and then input into the corresponding level's temporal encoder for recursive calculation to obtain the temporal enhancement features of the first, second, third, fourth, and fifth levels. The corresponding relationship in this process is as follows: ; in, This indicates the processing of the hyperbolic tangent activation function. This indicates channel splicing processing. Represents element-wise product. Indicates the time step index; This represents the update gate at time t, used to control the proportion of historical hidden states that are retained. Represents the coding space features at time t; It represents the hidden state at time t-1, storing the temporal information of the previous t-1 times; This represents the 2D convolution operation corresponding to the update gate; This represents the reset gate at time t, used to adjust the degree of influence of the historical hidden state on the candidate state; This indicates the 2D convolution operation corresponding to the reset gate. This represents the candidate state for model construction at time t. This represents the two-dimensional convolution operation corresponding to the candidate state. This represents the hidden state at time t.
[0021] Step 3: Use the decoder to decode the temporal enhancement features of each level to obtain the decoded feature map at the original resolution; perform segmentation prediction on the decoded feature map at the original resolution to obtain the guide wire segmentation prediction map. This step is the decoder stage. The decoder of this invention adopts a bottom-up four-layer structure, responsible for progressively restoring the spatiotemporal features compressed by the encoder to the original image resolution. Each decoder layer includes one upsampling operation and one skip connection with the corresponding layer's temporal encoder features. Specifically, the decoder decodes the temporal enhancement features of each layer to obtain a decoded feature map at the original resolution. Segmentation prediction is then performed on the decoded feature map at the original resolution to obtain a guidewire segmentation prediction map. The specific steps are as follows: The fifth-level temporal enhancement features and the fourth-level temporal enhancement features are concatenated along the channel dimension and then fused through grouped convolution to obtain the first fused feature. The first fused feature is subjected to two consecutive pointwise convolution processes to obtain the first layer of refined decoded features; After upsampling the first-layer refined decoding features, they are then concatenated and grouped convolutionally fused with the third-layer temporal enhancement features along the channel dimension to obtain the second fused feature. The second fused feature is subjected to two consecutive pointwise convolution processes to obtain the second layer of refined decoded features; After upsampling the refined decoding features of the second layer, they are then concatenated and grouped with the temporal enhancement features of the second layer along the channel dimension to obtain the third fused feature. The third fusion feature is subjected to two consecutive pointwise convolution processes to obtain the refined decoding feature of the third layer; After upsampling the refined decoding features of the third layer, they are then concatenated and grouped with the temporal enhancement features of the first layer along the channel dimension to obtain the fourth fused feature. The fourth fusion feature is subjected to two consecutive pointwise convolution processes to obtain the decoded feature map at the original resolution; A convolutional layer with a channel number mapped to 2 and a size of 1×1 is used to process the decoded feature map at the original resolution to obtain a binarized segmentation mask; the binarized segmentation mask is then used as the guide wire segmentation prediction map.
[0022] Furthermore, after upsampling the refined decoding features from the first layer, they are concatenated and grouped with the temporal enhancement features from the third layer along the channel dimension to obtain the second fused feature. The corresponding relationship in this process is as follows: ; in, Indicates fusion features, This indicates that grouped convolution operations are performed along the channel dimension on the concatenated features. This indicates that two features are concatenated along the channel dimension. This represents the temporal enhancement features of level 5-j. This represents the refined decoding features of the (j-1)th layer.
[0023] The second fused feature is subjected to two consecutive pointwise convolutional processes to obtain the refined decoded feature of the second layer. The corresponding relationship in this process is as follows: ; in, This represents intermediate features in the decoding process. This indicates a 1×1 pointwise convolution process. This indicates the decoding features.
[0024] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0025] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0026] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention. Therefore, the scope of protection of this patent should be determined by the appended claims.
Claims
1. A lightweight guidewire segmentation method based on time-gated control, characterized in that, The method includes the following steps: Step 1: Obtain the original X-ray guidewire image sequence and perform frame extraction on the original X-ray guidewire image sequence to obtain the input image sequence; further process the input image sequence to obtain the underlying spatial feature map; Step 2: Input the bottom-level spatial feature map into the encoder and perform temporal coding to obtain the coded spatial features of each level; The coding space features at each level are recursively processed to obtain the temporal enhancement features at each level. The specific steps are as follows: Step 201 involves sequentially performing channel compression pointwise convolution, 7×7 depthwise separable convolution, channel restoration pointwise convolution, and output projection layer processing on the bottom spatial feature map to obtain spatially adaptive dynamic weights. The corresponding relationship in this process is as follows: ; in, Represents the space-adaptive dynamic weights. This indicates the processing stage of the big-core sensing phase. This represents the bottom-level spatial feature map of the (m-1)th level. This indicates the third 1×1 pointwise convolutional processing in the large kernel sensing stage. This indicates the second 1×1 pointwise convolutional processing in the large kernel sensing stage. This indicates the first 1×1 pointwise convolutional processing in the large kernel sensing stage. This indicates that the kernel size is Depth-separable convolution processing; In the process of performing grouped dynamic convolution on the underlying spatial feature maps to obtain channel aggregated feature representations, the following relationship exists: ; in, The channel aggregation feature representation of the i-th spatial location and the c-th channel. The size centered at the i-th spatial location on the bottom-level spatial feature map of any level is . The neighborhood, Let represent the i-th spatial location on the bottom-level spatial feature map of any level. This represents the dynamic convolution kernel weights corresponding to the g-th channel at the i-th spatial location. This indicates a 3×3 convolution process; The underlying spatial feature map is subjected to grouped dynamic convolution processing to obtain channel aggregated feature representation; The aggregated feature representations of each channel and each spatial position are arranged according to the original space to obtain a complete aggregated feature map. Based on spatial adaptive dynamic weights, the complete aggregated feature map is batch normalized and then added element-wise with the underlying spatial feature map to obtain the aggregated output feature map. Channel attention is calculated on the aggregated output feature map to obtain intermediate features; The intermediate features are processed by residual connections using a feedforward network to obtain the final output feature map of the first level. The final output feature map of the first level is processed sequentially by convolution, max pooling, GELU activation function and batch normalization to obtain the coding space features of the first level. Step 202: Use the coding space features of the first level as the bottom space feature map, and repeat step 201 to obtain the coding space features of the second level. The second-level coding space features are used as the bottom-level spatial feature map. Step 201 is repeated to obtain the third-level coding space features. Using the third-level coding space features as the bottom-level spatial feature map, repeat step 201 to obtain the fourth-level coding space features. Using the fourth-level coding space features as the bottom-level spatial feature map, repeat step 201 to obtain the fifth-level coding space features. Step 203: The four-dimensional tensors of the coding space features of the first level, the second level, the third level, the fourth level, and the fifth level are reshaped into five-dimensional sequence forms, and then input into the corresponding level temporal encoder for recursive calculation to obtain the temporal enhancement features of the first level, the second level, the third level, the fourth level, and the fifth level. Step 3: Use the decoder to decode the temporal enhancement features of each level to obtain the decoded feature map at the original resolution; perform segmentation prediction on the decoded feature map at the original resolution to obtain the guide wire segmentation prediction map.
2. The lightweight guidewire segmentation method based on time-gated control according to claim 1, characterized in that, In step 1, the original X-ray guidewire image sequence is acquired, and frame extraction is performed on the original X-ray guidewire image sequence to obtain the input image sequence; the input image sequence is further processed to obtain the low-level spatial feature map, and the specific steps are as follows: Obtain a preset number of guidewire images from the original X-ray guidewire image sequence; arrange all guidewire images in chronological order to obtain the input image sequence; The batch dimension and frame number dimension of the input image sequence are merged to reshape it, resulting in a stacked feature map with merged dimensions. The stacked feature maps of the merged dimensions are processed by using a 3×3 convolution with a stride of 1 and padding of 1. Then, batch normalization and ReLU activation are performed sequentially to obtain the bottom-level spatial feature maps.
3. The lightweight guidewire segmentation method based on time-gated control according to claim 1, characterized in that, Based on spatial adaptive dynamic weights, the complete aggregated feature map is batch normalized and then added element-wise to the underlying spatial feature map to obtain the aggregated output feature map. The corresponding relationship in this process is as follows: ; in, This represents the aggregated output feature map obtained from the bottom-level spatial feature maps of level m-1. This represents the LS convolution operation. This indicates batch normalization processing. This indicates that the small core aggregation module is processing the data. Channel attention is calculated on the aggregated output feature map to obtain intermediate features. The corresponding relationship in this process is as follows: ; in, This indicates that the intermediate features are obtained through LS convolution processing and SE channel attention operations in the m-th level LS encoder. This indicates channel attention calculation.
4. The lightweight guidewire segmentation method based on time-gated control according to claim 3, characterized in that, A feedforward network is used to perform residual connection processing on the intermediate features to obtain the final output feature map of the first layer. The corresponding relationship in this process is as follows: ; in, This represents the final output feature map of the m-th level LS encoder. This indicates feedforward network processing. This means adding the residuals of the feedforward network output to the intermediate input features; The final output feature map of the first level is sequentially processed by convolution, max pooling, GELU activation function, and batch normalization to obtain the encoding space features of the first level. The corresponding relationship in this process is as follows: ; in, This represents the coding space features of the m-th level. This indicates the GELU activation function processing. This indicates a 3×3 standard convolution process.
5. The lightweight guidewire segmentation method based on time-gated control according to claim 4, characterized in that, The four-dimensional tensors of the coding space features at the first, second, third, fourth, and fifth levels are reconstructed into five-dimensional sequences, which are then input into the corresponding level's temporal encoders for recursive computation to obtain the temporal enhancement features at the first, second, third, fourth, and fifth levels. The corresponding relationships in this process are as follows: ; in, This indicates the processing of the hyperbolic tangent activation function. This indicates channel splicing processing. Represents element-wise product. Indicates the time step index. This represents the update gate at time t. This represents the coding space features at time t. This represents the hidden state at time t-1. This indicates the 2D convolution operation corresponding to the update gate. This represents the reset gate at time t. This indicates the 2D convolution operation corresponding to the reset gate. This represents the candidate state for model construction at time t. This represents the two-dimensional convolution operation corresponding to the candidate state. This represents the hidden state at time t.
6. The lightweight guidewire segmentation method based on time-gated control according to claim 5, characterized in that, In step 3, the temporal enhancement features of each level are decoded using a decoder to obtain a decoded feature map at the original resolution; the decoded feature map at the original resolution is then used for segmentation prediction to obtain a guidewire segmentation prediction map. The specific steps are as follows: The fifth-level temporal enhancement features and the fourth-level temporal enhancement features are concatenated along the channel dimension and then fused through grouped convolution to obtain the first fused feature. The first fused feature is subjected to two consecutive pointwise convolution processes to obtain the first layer of refined decoded features; After upsampling the first-layer refined decoding features, they are then concatenated and grouped convolutionally fused with the third-layer temporal enhancement features along the channel dimension to obtain the second fused feature. The second fused feature is subjected to two consecutive pointwise convolution processes to obtain the second layer of refined decoded features; After upsampling the refined decoding features of the second layer, they are then concatenated and grouped with the temporal enhancement features of the second layer along the channel dimension to obtain the third fused feature. The third fusion feature is subjected to two consecutive pointwise convolution processes to obtain the refined decoding feature of the third layer; After upsampling the refined decoding features of the third layer, they are then concatenated and grouped with the temporal enhancement features of the first layer along the channel dimension to obtain the fourth fused feature. The fourth fusion feature is subjected to two consecutive pointwise convolution processes to obtain the decoded feature map at the original resolution; A 1×1 convolutional layer that maps the number of channels to 2 is used to process the decoded feature map at the original resolution to obtain a binarized segmentation mask; the binarized segmentation mask is then used as the guide wire segmentation prediction map.
7. The lightweight guidewire segmentation method based on time-gated control according to claim 6, characterized in that, After upsampling the refined decoding features from the first layer, they are then concatenated and grouped with the temporal enhancement features from the third layer along the channel dimension to obtain the second fused feature. The corresponding relationship in this process is as follows: ; in, Indicates fusion features, This indicates that grouped convolution operations are performed along the channel dimension on the concatenated features. This means concatenating two features along the channel dimension. This represents the temporal enhancement features of level 5-j. This represents the refined decoding features of the (j-1)th layer.
8. The lightweight guidewire segmentation method based on time-gated control according to claim 7, characterized in that, The second fused feature is subjected to two consecutive pointwise convolutional processes to obtain the refined decoded feature of the second layer. The corresponding relationship in this process is as follows: ; in, This represents intermediate features in the decoding process. This indicates a 1×1 pointwise convolution process. This indicates the decoding characteristics output by the current decoding layer.
Citation Information
Patent Citations
Segmentation and view guidance in ultrasound imaging and associated devices, systems, and methods
CN113678167A
Video sequence guide wire segmentation method and device, electronic equipment and readable medium
CN114550033A