A fractal-driven laparoscopic Toldt's line localization method, device, and medium

By using a fractal-driven approach to perform feature modeling and multi-scale spatiotemporal feature fusion on laparoscopic video frames, the problem of difficulty in identifying the location of Toldt's lines inside the abdominal cavity in existing technologies has been solved, achieving high-precision identification in complex anatomical structures and low-contrast regions.

CN120635470BActive Publication Date: 2025-10-31THE SIXTH AFFILIATED HOSPITAL OF SUN YAT SEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511141992.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-15
Publication Date
2025-10-31
Estimated Expiration
2045-08-15

AI Technical Summary

Technical Problem

Existing deep learning video segmentation models trained on medical datasets are ineffective in identifying and segmenting Toldt's lines. They struggle to capture boundary cues and lack anatomical constraints, making it difficult to accurately identify the location of Toldt's lines inside the abdominal cavity.

Method used

A fractal-driven approach is adopted to model the features of laparoscopic video frames using a Toldt line segmentation model. The fractal dimension features are used to guide spatial attention to extract anatomical features. Multi-scale spatiotemporal feature fusion and aggregation are performed, and a multi-scale decoder is used to predict and identify the location of Toldt lines inside the abdominal cavity.

Benefits of technology

It improves the recognition accuracy in complex anatomical structures and low-contrast areas, and can more accurately track and identify the location of Toldt's lines, solving the problem of the difficulty in accurately identifying the location of Toldt's lines inside the abdominal cavity in existing technologies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120635470B_ABST
    Figure CN120635470B_ABST
Patent Text Reader

Abstract

This invention discloses a fractal-driven method, device, and medium for locating Toldt's line in laparoscopy. The method includes: acquiring several laparoscopic video frames; performing feature modeling on the several laparoscopic video frames according to a Toldt's line segmentation model to obtain multi-scale feature maps; performing multi-scale spatiotemporal feature fusion and aggregation processing on the multi-scale feature maps; inputting the processed results into a multi-scale decoder for prediction to obtain the location of Toldt's line inside the abdominal cavity. This invention proposes a fractal-driven method, device, and medium for locating Toldt's line in laparoscopy. By accurately capturing the anatomical features of Toldt's line through precise feature modeling, enhancing adaptability to the dynamic abdominal environment through multi-scale spatiotemporal feature fusion, and achieving high-precision prediction using a multi-scale decoder, the location of Toldt's line inside the abdominal cavity can be identified, thus solving the problem of accurately identifying the location of Toldt's line inside the abdominal cavity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing, and in particular to a fractal-driven method, apparatus, and medium for laparoscopic Toldt's line localization. Background Technology

[0002] A constant line, known as Toldt's line, runs along the lateral border of the ascending and descending colon where the visceral and parietal peritoneum meet. Laparoscopy provides medical staff with real-time visualization of the abdominal cavity's anatomical structures. Accurate identification of Toldt's line is crucial for surgical success in laparoscopic mesocolon resection surgery, but its variability in anatomical location and the complexity of surrounding tissues present challenges. Furthermore, with technological advancements, deep learning models can annotate regions of interest in images and videos, helping surgical staff quickly locate targets and reducing their workload. Current research focuses on efficient and high-precision detection and segmentation of intestinal polyps. Several high-quality image and video polyp segmentation datasets have been constructed, such as Kavsir-SEG and the CVC series. Based on these datasets, various polyp segmentation algorithms have been proposed, such as PraNet using reverse attention to enhance boundary cue capture, PNS-Net employing a regularized self-attention mechanism, and LDNet using a dynamic kernel generation and update mechanism. Additionally, Swin-UMamba utilizes a pre-trained architecture to improve segmentation accuracy. These algorithms have driven the development of intestinal polyp segmentation technology.

[0003] However, in existing technologies, deep learning video segmentation models trained on medical datasets perform poorly in identifying and segmenting Toldt's lines because their boundary features are blurred with the surrounding anatomical structures, increasing the segmentation difficulty. Models trained on polyp segmentation datasets struggle to capture Toldt's line boundary cues and lack anatomical constraints, resulting in a uniform strategy when dealing with different anatomical structures, making it unable to cope with the challenges of complex anatomical structures and low-contrast target areas. Therefore, existing technologies struggle to accurately identify the location of Toldt's lines within the abdominal cavity. Summary of the Invention

[0004] This invention provides a fractal-driven laparoscopic Toldt's line localization method, device, and medium to solve the problem of difficulty in accurately identifying the location of Toldt's line inside the abdominal cavity.

[0005] To achieve the above objectives, this application provides a fractal-driven laparoscopic Toldt's line localization method, comprising:

[0006] Acquire several laparoscopic video frames;

[0007] The Toldt line segmentation model is used to model the features of the several laparoscopic video frames to obtain multi-scale feature maps. The Toldt line segmentation model is established by quantifying the geometric characteristics of the Toldt line boundary through a preset algorithm and using fractal dimension features to guide spatial attention to extract anatomical features.

[0008] The multi-scale feature map is subjected to multi-scale spatiotemporal feature fusion and aggregation processing. The processing result is input into a multi-scale decoder for prediction to obtain the location of Toldt's line inside the abdominal cavity.

[0009] This invention first utilizes a Toldt's line segmentation model to perform feature modeling on laparoscopic video frames. For the model, the Toldt's line, as part of an anatomical structure, has unique geometric characteristics at its boundary. By precisely quantifying these geometric characteristics through a pre-defined algorithm, the model can learn the unique shape and orientation of the Toldt's line boundary. Guided by fractal dimension features, the model can focus on areas most relevant to the Toldt's line boundary, thus ignoring background noise and interference factors. This focusing ability helps improve the accuracy and stability of recognition. Spatiotemporal feature fusion and aggregation are performed on multi-scale feature maps. This process not only integrates spatial information at different scales but also captures dynamic changes in the temporal dimension, helping the model to more accurately track and identify the location of the Toldt's line when processing dynamic laparoscopic videos. Finally, the fused features are input into a multi-scale decoder for prediction. The multi-scale decoder can utilize rich feature information for fine segmentation prediction, thereby accurately outputting the location of the Toldt's line, helping the model maintain high accuracy when processing complex anatomical structures and low-contrast regions.

[0010] Compared to existing technologies, this invention captures the anatomical features of Toldt's lines through precise feature modeling, enhances adaptability to the dynamic abdominal environment by combining multi-scale spatiotemporal feature fusion, and then uses a multi-scale decoder to achieve high-precision prediction, thereby identifying the location of Toldt's lines inside the abdominal cavity. Therefore, it can solve the problem of accurately identifying the location of Toldt's lines inside the abdominal cavity.

[0011] As a preferred embodiment, the multi-scale feature map is subjected to multi-scale spatiotemporal feature fusion and aggregation processing. The processed result is then input into a multi-scale decoder for prediction to obtain the location of Toldt's line inside the abdominal cavity. Specifically:

[0012] Perform cross-temporal convolution on the multi-scale feature map to obtain a spatiotemporal fusion feature map;

[0013] The local fractal dimension of each pixel in the spatiotemporal fusion feature map is calculated using the fractal dimension estimation algorithm to obtain the fractal dimension feature map.

[0014] Multiple feature extractions are performed on the fractal dimension feature map to obtain a multi-scale enhanced feature map;

[0015] The multi-scale enhanced feature map is input into the multi-scale decoder for prediction to obtain the location of the Toldt line inside the abdominal cavity.

[0016] This preferred scheme performs cross-temporal convolution on multi-scale feature maps to obtain a spatiotemporal fusion feature map, which can capture the temporal dependencies between video frames and effectively fuse spatial and temporal features. The local fractal dimension of each pixel in the spatiotemporal fusion feature map is calculated using a fractal dimension estimation algorithm to obtain a fractal dimension feature map. Fractal dimension, as a geometric feature, can reflect the complexity and irregularity of anatomical structures. Multiple feature extraction is performed on the fractal dimension feature map to obtain a multi-scale enhanced feature map. This multi-path design can capture local detail features and global contextual information, enabling the construction of a more comprehensive and richer multi-scale enhanced feature map.

[0017] As a preferred approach, multiple feature extraction is performed on the fractal dimension feature map to obtain a multi-scale enhanced feature map, specifically:

[0018] The spatiotemporal fusion feature map and the fractal dimension feature map are input into the local texture-aware convolutional submodule of the Toldt line segmentation model for semantic extraction to obtain the first output feature map.

[0019] The first output feature map is input into the fractal-guided anatomy consistent attention submodule of the Toldt line segmentation model for filtering to obtain the multi-scale enhanced feature map.

[0020] In this preferred scheme, the local texture-aware convolutional submodule focuses on the local texture features of the anatomical structure, extracting high-level semantics through adaptive convolutional kernels to obtain the first output feature map. Meanwhile, the fractal-guided anatomical consistency attention submodule emphasizes the global consistency of the anatomical structure and the distinction between foreground and background, filtering features through an attention mechanism to obtain a multi-scale enhanced feature map. These two data processing methods capture information about the anatomical structure at different scales, thus enabling the construction of more comprehensive and richer multi-scale enhanced feature maps.

[0021] As a preferred embodiment, the spatiotemporal fusion feature map and the fractal dimension feature map are input into the local texture-aware convolutional submodule of the Toldt line segmentation model for semantic extraction to obtain the first output feature map, specifically:

[0022] The spatiotemporal fusion feature map and the fractal dimension feature map are input into the local texture-aware convolutional submodule of the Toldt line segmentation model;

[0023] The spatiotemporal fusion feature map and the fractal dimension feature map are concatenated along the channel dimension to generate a joint feature tensor;

[0024] High-level semantic extraction is performed on the joint feature tensor to obtain the expansion coefficient and offset direction;

[0025] The first output feature map of the local texture-aware convolutional submodule is calculated based on the dilation coefficient, the offset direction, and the spatiotemporal fusion feature map.

[0026] This preferred scheme generates a joint feature tensor by concatenating the spatiotemporal fusion feature map and the fractal dimension feature map along the channel dimension. The fractal dimension feature map provides detailed information about the texture of the anatomical structure, while the spatiotemporal fusion feature map contains high-level semantic information. This process achieves the fusion of multi-scale features. High-level semantic extraction is performed on the joint feature tensor to obtain the dilation coefficient and offset direction, realizing the transformation from low-level features to high-level semantics. The high-level semantic information helps the model understand the overall layout and interrelationships of the anatomical structure, providing strong support for subsequent segmentation tasks.

[0027] As a preferred embodiment, the first output feature map is input into the fractal-guided anatomy consistent attention submodule of the Toldt line segmentation model for filtering to obtain the multi-scale enhanced feature map, specifically:

[0028] The first output feature map is input into the fractal-guided anatomy consistent attention submodule of the Toldt line segmentation model;

[0029] The fractal dimension feature map is converted into a normalized distribution map, and the indexes of the target foreground region and the target background region are extracted from the distribution map to obtain an index set;

[0030] The first output feature map is flattened and a triplet embedding containing Query, Key and Value is generated by matrix projection.

[0031] The triplet embedding is filtered and attention features are extracted based on the index set to obtain the multi-scale enhanced feature map.

[0032] This preferred scheme converts the fractal dimension feature map into a normalized distribution map and extracts the indices of the target foreground region and the target background region from it to obtain an index set. This process can accurately distinguish between the foreground and background regions, providing accurate target regions for subsequent attention feature extraction.

[0033] As a preferred embodiment, the triplet embedding is subjected to data filtering and attention feature extraction based on the index set to obtain the multi-scale enhanced feature map, specifically as follows:

[0034] In the triplet embedding, the key-value pairs corresponding to the target foreground region and the target background region are filtered according to the index set to obtain an embedding set containing only the key and value;

[0035] Based on the Query and the embedding set, foreground attention results and background attention results are calculated. The foreground attention results and background attention results are then concatenated to obtain an attention feature map.

[0036] The number of channels in the attention feature map is adjusted to the number of channels in the spatiotemporal fusion feature map by convolutional compression. The channel-aligned attention feature map is then fused to the spatiotemporal fusion feature map through skip connections to obtain the multi-scale enhanced feature map.

[0037] This preferred approach filters key-value pairs corresponding to the target foreground and background regions using an index set. This demonstrates the need to maintain the global nature of the query. Therefore, by using the query and the filtered key-value pairs to calculate the foreground and background attention results, the model can focus more on key local anatomical structures under the guidance of global information. The number of channels in the attention feature map is adjusted by convolutional compression to align with the number of channels in the spatiotemporal fusion feature map. This adjustment method preserves the key information of the attention feature map while ensuring consistency in subsequent feature fusion.

[0038] As a preferred embodiment, the Toldt line segmentation model is established by quantifying the geometric properties of the Toldt line boundary using a preset algorithm and extracting anatomical features by using fractal dimension features to guide spatial attention. Specifically:

[0039] Acquire surgical video frames containing the Toldt's line region;

[0040] Mark the regions where Toldt's lines are located in the surgical video frames to obtain a laparoscopic Toldt's line segmentation dataset;

[0041] Based on the fractal dimension estimation algorithm and the preset module, the deep learning model is trained using the laparoscopic Toldt line segmentation dataset to obtain the Toldt line segmentation model; wherein, the fractal dimension estimation algorithm is used to calculate the fractal dimension, and the preset module is used to extract anatomical features by guiding spatial attention through fractal dimension features.

[0042] This preferred approach constructs a high-quality laparoscopic Toldt's line segmentation dataset by labeling the regions containing Toldt's lines in surgical video frames. This helps deep learning models learn accurate features of Toldt's lines, thereby improving recognition accuracy. Fractal dimension estimation algorithms typically have low computational complexity, and applying them to the training process of deep learning models does not significantly increase the computational burden. By guiding spatial attention through fractal dimension features, the model can more accurately focus on regions with significant fractal characteristics, thereby extracting more accurate and relevant anatomical features.

[0043] As a preferred approach, feature modeling is performed on the plurality of laparoscopic video frames based on the Toldt line segmentation model to obtain multi-scale feature maps, specifically:

[0044] The target continuous frames in the plurality of laparoscopic video frames are input into the Swin-Transformer backbone network encoder of the Toldt line segmentation model.

[0045] Based on the Swin-Transformer backbone network encoder, the local and global features of the target continuous frames are modeled through a sliding window attention mechanism to obtain the multi-scale feature map.

[0046] In this preferred embodiment, the sliding window attention mechanism can capture global information through the interaction between windows while processing local features. This enables the model to understand the overall structure while maintaining sensitivity to details, thereby accurately identifying key anatomical structures such as Toldt's lines in complex laparoscopic video frames.

[0047] This application also provides a fractal-driven laparoscopic Toldt's line positioning device, including a data module, a modeling module and a prediction module;

[0048] The data module is used to acquire several laparoscopic video frames.

[0049] The modeling module is used to perform feature modeling on the several laparoscopic video frames according to the Toldt line segmentation model to obtain multi-scale feature maps; wherein, the Toldt line segmentation model is established by quantifying the geometric characteristics of the Toldt line boundary through a preset algorithm and using fractal dimension features to guide spatial attention to extract anatomical features;

[0050] The prediction module is used to perform multi-scale spatiotemporal feature fusion and aggregation processing on the multi-scale feature map, and input the processing result into the multi-scale decoder for prediction to obtain the position of Toldt's line inside the abdominal cavity.

[0051] As a preferred embodiment, the prediction module includes a convolution unit, a computation unit, an extraction unit, and a prediction unit;

[0052] The convolutional unit is used to perform cross-temporal convolution on the multi-scale feature map to obtain a spatiotemporal fusion feature map.

[0053] The computing unit is used to calculate the local fractal dimension of each pixel position in the spatiotemporal fusion feature map according to the fractal dimension estimation algorithm, so as to obtain the fractal dimension feature map.

[0054] The extraction unit is used to perform multiple feature extraction on the fractal dimension feature map to obtain a multi-scale enhanced feature map.

[0055] The prediction unit is used to input the multi-scale enhanced feature map into the multi-scale decoder for prediction to obtain the position of the Toldt line inside the abdominal cavity.

[0056] As a preferred embodiment, the extraction unit includes an extraction subunit and a filtering subunit;

[0057] The extraction subunit is used to input the spatiotemporal fusion feature map and the fractal dimension feature map into the local texture-aware convolutional submodule of the Toldt line segmentation model for semantic extraction to obtain the first output feature map.

[0058] The filtering subunit is used to input the first output feature map into the fractal-guided anatomy consistent attention submodule of the Toldt line segmentation model for filtering to obtain the multi-scale enhanced feature map.

[0059] As a preferred embodiment, the extraction subunit specifically comprises:

[0060] The spatiotemporal fusion feature map and the fractal dimension feature map are input into the local texture-aware convolutional submodule of the Toldt line segmentation model;

[0061] The spatiotemporal fusion feature map and the fractal dimension feature map are concatenated along the channel dimension to generate a joint feature tensor;

[0062] High-level semantic extraction is performed on the joint feature tensor to obtain the expansion coefficient and offset direction;

[0063] The first output feature map of the local texture-aware convolutional submodule is calculated based on the dilation coefficient, the offset direction, and the spatiotemporal fusion feature map.

[0064] As a preferred embodiment, the filtering subunit specifically comprises:

[0065] The first output feature map is input into the fractal-guided anatomy consistent attention submodule of the Toldt line segmentation model;

[0066] The fractal dimension feature map is converted into a normalized distribution map, and the indexes of the target foreground region and the target background region are extracted from the distribution map to obtain an index set;

[0067] The first output feature map is flattened and a triplet embedding containing Query, Key and Value is generated by matrix projection.

[0068] The triplet embedding is filtered and attention features are extracted based on the index set to obtain the multi-scale enhanced feature map.

[0069] As a preferred embodiment, the triplet embedding is subjected to data filtering and attention feature extraction based on the index set to obtain the multi-scale enhanced feature map, specifically as follows:

[0070] In the triplet embedding, the key-value pairs corresponding to the target foreground region and the target background region are filtered according to the index set to obtain an embedding set containing only the key and value;

[0071] Based on the Query and the embedding set, foreground attention results and background attention results are calculated. The foreground attention results and background attention results are then concatenated to obtain an attention feature map.

[0072] The number of channels in the attention feature map is adjusted to the number of channels in the spatiotemporal fusion feature map by convolutional compression. The channel-aligned attention feature map is then fused to the spatiotemporal fusion feature map through skip connections to obtain the multi-scale enhanced feature map.

[0073] As a preferred embodiment, the modeling module includes a video unit, a labeling unit, and a training unit;

[0074] The video unit is used to acquire surgical video frames containing the Toldt line region;

[0075] The labeling unit is used to label the region where the Toldt line is located in the surgical video frame to obtain a laparoscopic Toldt line segmentation dataset;

[0076] The training unit is used to train the deep learning model on the laparoscopic Toldt line segmentation dataset according to the fractal dimension estimation algorithm and the preset module to obtain the Toldt line segmentation model; wherein, the fractal dimension estimation algorithm is used to calculate the fractal dimension, and the preset module is used to extract anatomical features by guiding spatial attention through fractal dimension features.

[0077] As a preferred embodiment, the modeling module includes an input unit and a construction unit;

[0078] The input unit is used to input the target continuous frames from the plurality of laparoscopic video frames into the Swin-Transformer backbone network encoder of the Toldt line segmentation model.

[0079] The construction unit is used to model the local and global features of the target continuous frames based on the Swin-Transformer backbone network encoder through a sliding window attention mechanism to obtain the multi-scale feature map.

[0080] This application also provides a storage medium storing a computer program, which is called and executed by a computer to implement the fractal-driven laparoscopic Toldt's line localization method as described above. Attached Figure Description

[0081] Figure 1 This is a schematic flowchart of a fractal-driven laparoscopic Toldt's line localization method provided in an embodiment of this application;

[0082] Figure 2 This is a schematic diagram illustrating the construction process of the Toldt's line segmentation model for laparoscopic video provided in this application embodiment;

[0083] Figure 3 This is a schematic diagram of the network architecture of FSA-Net provided in the embodiments of this application;

[0084] Figure 4 This is a flowchart of the fractal dimension estimation algorithm based on measure theory provided in the embodiments of this application;

[0085] Figure 5 This is a schematic diagram of the fractal guidance feature aggregation module provided in an embodiment of this application;

[0086] Figure 6 This is a schematic diagram of the FSA-Net workflow provided in the embodiments of this application;

[0087] Figure 7 This is a visual comparison of the Toldt line segmentation effects of different segmentation models provided in the embodiments of this application;

[0088] Figure 8 This is a comparison chart of the quantitative results of Toldt's line segmentation performance of different segmentation models provided in the embodiments of this application;

[0089] Figure 9 This is a schematic diagram of the structure of a fractal-driven Toldt's line positioning device for laparoscopy provided in an embodiment of this application. Detailed Implementation

[0090] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0091] In the description of this application, it should be understood that the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" and "second" may explicitly or implicitly include one or more of that feature. In the description of this application, unless otherwise stated, "several" means two or more.

[0092] The fractal-driven laparoscopic Toldt's line localization method provided in this application is mainly used to assist colorectal surgeons in accurately identifying the Toldt's line region inside the abdominal cavity during surgery.

[0093] Example 1:

[0094] Please see Figure 1 The embodiments of this application provide a fractal-driven laparoscopic Toldt's line localization method, including S1~S3, and the specific implementation steps are as follows:

[0095] S1. Acquire several laparoscopic video frames.

[0096] Step S1 in this embodiment of the application is specifically as follows:

[0097] Video frames are extracted from laparoscopic surgical video images at a rate of 2 frames per second. Then, the size of these frames is downsampled to 384×384 pixels. Next, data augmentation processing is performed on each frame, including scaling, rotation, flipping, and random noise addition, to finally generate a number of processed laparoscopic video frames.

[0098] S2. Based on the Toldt line segmentation model, feature modeling is performed on several laparoscopic video frames to obtain multi-scale feature maps. The Toldt line segmentation model is established by quantifying the geometric characteristics of the Toldt line boundary through a preset algorithm and using fractal dimension features to guide spatial attention to extract anatomical features.

[0099] Step S2 in this embodiment includes S2.1 to S2.2, where S2.1 is the process of establishing the Toldt line segmentation model, and S2.2 is the process of calculating multi-scale feature maps based on the Swin-Transformer backbone network encoder, specifically as follows:

[0100] S2.1 Acquire laparoscopic video data of mesocolon resection surgery and extract surgical video frames containing the Toldt's line region from the image data. In order to acquire laparoscopic image data of mesocolon resection surgery, data was first collected from the surgical procedures of 45 different patients. Each patient's surgery was recorded into several video segments with a duration of approximately 1 to 2 hours. Subsequently, short video segments with a duration of 15 to 30 seconds were selected from each long video segment by a preset model or by professional surgeons to ensure that each selected frame clearly showed the Toldt's line region. Finally, 150 high-quality laparoscopic surgical video segments were obtained. This data can lay a solid foundation for subsequent training of the Toldt's line segmentation model.

[0101] Frame-by-frame extraction was performed on all edited surgical video frames, exporting and saving each frame as a PNG image file. Subsequently, these extracted frame sequences were downsampled, reducing the original video frame rate from 30 FPS to 2 FPS (2 frames per second). After downsampling, a total of 5000 images were obtained, each clearly showing the Toldt's line region. Next, using a pre-defined model or manual annotation by experienced colorectal surgeons, the Toldt's line region was precisely defined by drawing polylines, resulting in a laparoscopic Toldt's line segmentation dataset. After annotation, the pre-defined model or colorectal surgery professors conducted quality assessments to ensure that all annotation results in the laparoscopic Toldt's line segmentation dataset met the pre-defined standards.

[0102] Based on the fractal dimension estimation algorithm and pre-defined modules, a deep learning model is trained using a laparoscopic Toldt's line segmentation dataset to obtain a Toldt's line segmentation model, which is a deep learning model based on a fractal-driven collaborative anatomical perception network (FSA-Net). The fractal dimension estimation algorithm is used to calculate the fractal dimension. The pre-defined modules include a Local Texture-Aware Convolution (LTC) sub-module and a Fractal-Guided Anatomical Consistent Attention (FAA) sub-module, which are used to extract anatomical features by guiding spatial attention through fractal dimension features, aiming to collaboratively extract internal and external anatomical features. Furthermore, the Toldt's line segmentation model includes a Swin-Transformer-based encoder (i.e., a Swin-Transformer backbone network encoder), a metric-theory-based fractal dimension estimation module, a fractal-guided feature aggregation module, and a Transformer-based multi-scale feature decoder.

[0103] Furthermore, performance tests of the Toldt line segmentation model show that it can process 30 to 40 frames per second, a processing speed that perfectly matches the 30 frames per second frame rate of laparoscopic cameras used in most current colorectal cancer surgeries, thus proving that the model has the ability to perform real-time segmentation of surgical video images.

[0104] For examples of this application, please refer to [link / reference]. Figure 2-3 , Figure 2 This is a schematic diagram illustrating the construction process of the Toldt's line segmentation model for laparoscopic video provided in this application, showing the construction process of the Toldt's line segmentation model;

[0105] Figure 3 This is a schematic diagram of the network architecture of FSA-Net provided in this application, which shows the network architecture of the Toldt line segmentation model.

[0106] In this embodiment S2.1, by labeling the regions where Toldt's lines are located in surgical video frames, a high-quality laparoscopic Toldt's line segmentation dataset can be constructed. This helps the deep learning model learn the accurate features of Toldt's lines, thereby improving recognition accuracy. Fractal dimension estimation algorithms generally have low computational complexity, and applying them to the training process of deep learning models will not significantly increase the computational burden. By guiding spatial attention through fractal dimension features, the model can more accurately focus on regions with significant fractal characteristics, thereby extracting more accurate and relevant anatomical features.

[0107] In summary, Toldt's lines segmentation in videos is extremely challenging due to significant inter-individual variations in their location and morphology, as well as low contrast with surrounding tissues and blurred boundaries. To address this, this embodiment innovatively proposes the FSA-Net framework, specifically designed for Toldt's white lines segmentation in surgical laparoscopic images. FSA-Net first introduces a novel fractal dimension estimation algorithm to ensure computational accuracy. Then, it integrates two sub-modules: Local Texture-Aware Convolution (LTC) and Fractal-Guided Anatomical Consistent Attention (FAA), to collaboratively capture anatomical features. Furthermore, this embodiment constructs the first laparoscopic Toldt's line segmentation dataset to support related research. Experimental validation on multiple datasets shows that FSA-Net significantly improves the segmentation accuracy of Toldt's lines, providing strong support for surgical procedures.

[0108] S2.2 Input the target continuous frames from several laparoscopic video frames into the Swin-Transformer backbone network encoder of the Toldt line segmentation model; wherein, the target continuous frames are a three-frame sequence consisting of key frames and their preceding and following adjacent frames;

[0109] Based on the Swin-Transformer backbone network encoder, a sliding window attention mechanism is used to model the local and global features of consecutive target frames, resulting in multi-scale feature maps. These multi-scale feature maps consist of four scales {F1, F2, F3, F4}, with resolutions decreasing sequentially to 1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the input size. Furthermore, the Swin-Transformer backbone network encoder combines sliding window attention and global attention, enabling effective simultaneous modeling of both global and local features of the image.

[0110] In this embodiment S2.2, the sliding window attention mechanism can capture global information through the interaction between windows while processing local features. This enables the model to understand the overall structure while maintaining sensitivity to details, thereby accurately identifying key anatomical structures such as Toldt's lines in complex laparoscopic video frames.

[0111] S3. Perform multi-scale spatiotemporal feature fusion and aggregation on the multi-scale feature map, and input the processed results into the multi-scale decoder for prediction to obtain the location of Toldt's line inside the abdominal cavity.

[0112] Step S3 in this embodiment includes S3.1 to S3.4, wherein S3.1 is the process of estimating the fractal dimension of multi-scale features to obtain a pixel-level fractal dimension feature map; S3.2 is the process of extracting the first output feature map based on the local texture-aware convolutional submodule; S3.3 is the process of extracting a multi-scale enhanced feature map based on the fractal-guided anatomical consistent attention submodule; and S3.4 is the process of predicting the location of Toldt's line inside the abdominal cavity based on the Transformer multi-scale feature decoder. Specifically:

[0113] S3.1 At each scale, a spatiotemporal fusion module based on 3D convolutional layers is used to convolve the feature maps of three consecutive frames in the multi-scale feature map across the time dimension to obtain the spatiotemporal fusion feature map F. i Among them, F i The subscript i represents different feature map scales; specifically, i = {1, 2, 3, 4}, corresponding to feature map scales of {1 / 4, 1 / 8, 1 / 16, 1 / 32} of the original input. Furthermore, the spatiotemporal fusion module based on 3D convolutional layers aims to effectively utilize temporal information between consecutive frames to improve the segmentation effect of Toldts lines. It is worth noting that 3D convolution, as a mature and widely used deep learning component, is often used to extract spatial features of 3D images and temporal information between consecutive video frames.

[0114] Based on the "fractal dimension estimation module based on measure theory" in the Toldt line segmentation model, the spatiotemporal fusion feature map F is calculated according to the fractal dimension estimation algorithm. i The local fractal dimension of each pixel location is used to obtain the pixel-level fractal dimension feature map D. i Among them, the fractal dimension characteristic diagram D i Let {D1,D2,D3,D4} be the size of the spatiotemporal fusion feature map F. i "Similarly; fractal dimension can effectively characterize the texture complexity of different anatomical structures; using fractal dimension features as constraints helps the model better perceive different anatomical structures."

[0115] For examples of this application, please refer to [link / reference]. Figure 4 , Figure 4 This is a flowchart of the fractal dimension estimation algorithm based on metric theory provided in this application. It shows the algorithm flow proposed in this embodiment for estimating the fractal dimension of an image at the pixel level, aiming to estimate the fractal dimension of a depth feature map more efficiently.

[0116] Wherein, Reflect Padding represents the reflection padding operation, used to add edge pixels to the feature map. B(Fi[H,W],ε) represents the ε×ε neighborhood centered at Fi[H,W], E[·] represents the mathematical expectation; q is the scaling exponent, which is set to 2 in this embodiment for ease of calculation during fractal dimension estimation; μ(x) represents the probability measure of pixel Fi, and Dis(·) represents the Euclidean distance between two pixels Fi; the specific calculation process is as follows:

[0117] ① Input feature map: Input a feature map Fi of size H×W.

[0118] ② Reflection fill: Perform reflection fill operation on feature map Fi to increase edge pixels.

[0119] ③ Define neighborhood: For each pixel in the feature map Fi, define a neighborhood B(Fi[H,W],ε) with the pixel as the center and a size of ε×ε.

[0120] ④ Calculate the probability measure: For each pixel, calculate its probability measure μ(x); and dμ(x) in the formula is used to represent the probability that each pixel in the neighborhood B(Fi[H,W],ε) is sampled, and its size is a fixed value.

[0121] ⑤ Calculate distance and expectation: Calculate the Euclidean distance Dis(·) between a pixel x in neighborhood B(Fi[H,W],ε) and any other pixel y in neighborhood B(Fi[H,W],ε). Furthermore, the scaling exponent q is not involved in the calculation of the mathematical expectation, but is used in (ε / E[·]). q-1 The calculation.

[0122] ⑥ Estimating fractal dimension: Using a measure-based fractal dimension estimation method, combined with the probability measure μ(x) and mathematical expectation E[·] calculated above, the fractal dimension of the feature map is estimated.

[0123] ⑦ Output: Output the estimated fractal dimension.

[0124] In this embodiment, S3.1, a cross-temporal convolution is performed on the multi-scale feature map to obtain a spatiotemporal fusion feature map, which can capture the temporal dependencies between video frames and effectively fuse spatial and temporal features. The local fractal dimension of each pixel in the spatiotemporal fusion feature map is calculated using a fractal dimension estimation algorithm to obtain a fractal dimension feature map; fractal dimension, as a geometric feature, can reflect the complexity and irregularity of anatomical structures.

[0125] While box-counting dimension is effective in estimating fractal dimensions for different structures, the feature maps extracted from the backbone of deep learning models often have high embedding dimensions, which significantly increases computational complexity. In contrast, the metric-based fractal dimension estimation method used in this embodiment, such as the Rényi generalized dimension, exhibits stronger robustness when dealing with such high-dimensional data and can significantly alleviate the curse of dimensionality caused by direct geometric overlay.

[0126] S3.2, merging the spatiotemporal feature map F i and fractal dimension characteristic map D i Simultaneously inputting a Local Texture-Aware Convolutional (LTC) submodule; wherein, the Local Texture-Aware Convolutional submodule belongs to the "Fractal-Guided Feature Aggregation Module" in the Toldt line segmentation model, and the "Fractal-Guided Feature Aggregation Module" also includes a Fractal-Guided Anatomical Consistent Attention (FAA) submodule, used for subsequent extraction of multi-scale enhanced feature maps; in addition, the "Fractal-Guided Feature Aggregation Module" combines the Local Texture-Aware Convolutional Module and the Fractal-Driven Anatomical Consistent Attention, which helps FSA-Net to collaboratively capture the contextual relationships within the anatomical structure and between different anatomical structures, thereby improving the model's ability to distinguish different anatomical structures;

[0127] Spatiotemporal fusion feature map F along the channel dimension i and fractal dimension characteristic map D i The components are concatenated to generate a joint feature tensor.

[0128] High-level semantic extraction is performed on the joint feature tensor to obtain the expansion coefficient E. i and offset direction A i ;

[0129] Based on the expansion coefficient E i Offset direction A i Spatiotemporal fusion feature map F i The first output feature map of the local texture-aware convolutional submodule is calculated.

[0130] For examples of this application, please refer to [link / reference]. Figure 5 , Figure 5 This is a schematic diagram of the fractal-guided feature aggregation module provided in this application, illustrating the use of the fractal-guided feature aggregation module to process the spatiotemporal fusion feature map F. i and fractal dimension characteristic map D i The process of data processing;

[0131] The local texture-aware convolutional submodule is designed to adaptively integrate features from different anatomical structures. Compared to standard deformable convolutional networks (DCNs), this embodiment can utilize anatomical knowledge embedded in fractal features to guide the offset of the convolutional kernel, thereby effectively avoiding the model learning unreasonable deformations.

[0132] Specific implementation as follows Figure 5 As shown in section “(b) Local Texture-Aware Convolution (LTC)”, the spatiotemporal fusion feature map F i and fractal dimension characteristic map D i As input, these are used together to predict the offset of the DCN kernel. The offset generator consists of two 3×3 depthwise separable convolutional blocks, which are responsible for predicting the dilation coefficient E. i and offset direction A i Then, using these predicted expansion coefficients and offset directions, the final offset O can be calculated. i After obtaining the offset, it is combined with the original spatiotemporal fusion feature map F. i and offset O i The output features (i.e., the first output feature map) are calculated through dilated and deformable convolution operations. In other words, the "dilated and deformable convolutional layer" used will accept the "spatiotemporal fusion feature map F". i and offset O i "As a parameter, the output feature map of the local texture-aware convolutional submodule is calculated.

[0133] In this embodiment, S3.2, a joint feature tensor is generated by concatenating the spatiotemporal fusion feature map and the fractal dimension feature map along the channel dimension. The fractal dimension feature map provides detailed information about the texture of the anatomical structure, while the spatiotemporal fusion feature map contains high-level semantic information. This process achieves the fusion of multi-scale features. High-level semantic extraction is performed on the joint feature tensor to obtain the dilation coefficient and offset direction, realizing the transformation from low-level features to high-level semantics. The high-level semantic information helps the model understand the overall layout and interrelationships of the anatomical structure, providing strong support for subsequent segmentation tasks.

[0134] S3.3 Input the first output feature map into the fractal-guided anatomy consistent attention submodule;

[0135] Using a linear transformation layer to transform the fractal dimension feature map D i Converted to normalized distribution map M i And from the distribution map M iThe indexes of the target foreground region and the target background region are extracted to obtain an index set. The target foreground region refers to the anatomical structure or pathological region that needs to be focused on, which is usually directly related to the diagnostic and treatment goals. For example, the Toldt's line-related region (as an anatomical boundary marker between the mesentery and retroperitoneal tissue in colon surgery, its accurate identification is crucial for surgical layer separation and vascular protection). The target background region refers to the interfering region that is not the core diagnostic and treatment goal, such as surgical instruments (which need to be masked in intraoperative images to avoid affecting the surgical field analysis) and distal tissues (such as adjacent organs or non-pathological tissues, which usually need to be excluded from the analysis).

[0136] The first output feature map is flattened and a triplet embedding containing Query, Key, and Value is generated by matrix projection.

[0137] In triple embedding, the key-value pairs corresponding to the target foreground region and the target background region are filtered according to the index set to obtain an embedding set containing only the key and value;

[0138] Based on the embedding set, according to the distribution graph M i Select the key-value pairs belonging to the foreground region, calculate the attention weights using global queries and foreground keys, and then aggregate them with foreground values ​​to generate the foreground attention result. Similarly, generate the background attention result. Concatenate the foreground attention result and the background attention result to obtain the attention feature map.

[0139] The number of channels in the attention feature map is adjusted to match the spatiotemporal fusion feature map F by compression using 1×1 convolutional layers. i The number of channels is determined, and the channel-aligned attention feature maps are fused to the spatiotemporal fusion feature map F via skip connections. i This updates the spatiotemporal fusion feature map F. i The multi-scale enhanced feature map output by the fractal-guided anatomy consistent attention submodule is obtained.

[0140] Specific implementation as follows Figure 5 As shown in the “(c) Fractal-guided Anatomical Consistent Attention (FAA)” section, the core of this module is a fractal-guided key-value pair filtering strategy designed to enhance the model’s ability to capture interanatomical information in a divide-and-conquer manner.

[0141] In this embodiment, S3.3, the fractal dimension feature map is converted into a normalized distribution map, and the indices of the target foreground region and the target background region are extracted from it to obtain an index set. This process can accurately distinguish between the foreground and background regions, providing accurate target regions for subsequent attention feature extraction.

[0142] Furthermore, filtering the key-value pairs corresponding to the target foreground and background regions using the index set demonstrates the need to maintain the globality of the query. Therefore, calculating the foreground and background attention results using the query and the filtered key-value pairs allows the model to focus more on key local anatomical structures under the guidance of global information. Adjusting the number of channels in the attention feature map through convolutional compression to align it with the number of channels in the spatiotemporal fusion feature map preserves the key information of the attention feature map while ensuring consistency in subsequent feature fusion.

[0143] Overall, in this embodiment, steps S3.2-S3.3 perform multiple feature extraction on the fractal dimension feature map to obtain a multi-scale enhanced feature map. This multi-path design can capture local detailed features and global contextual information, and can construct a more comprehensive and richer multi-scale enhanced feature map.

[0144] Furthermore, the local texture-aware convolutional submodule focuses on the local texture features of the anatomical structure, extracting high-level semantics through adaptive convolutional kernels to obtain the first output feature map. The fractal-guided anatomical consistency attention submodule, on the other hand, emphasizes the global consistency of the anatomical structure and the distinction between foreground and background, filtering features through an attention mechanism to obtain multi-scale enhanced feature maps. These two data processing methods capture information about the anatomical structure at different scales, thus enabling the construction of more comprehensive and richer multi-scale enhanced feature maps.

[0145] S3.4 Input the multi-scale enhanced feature map into the Transformer-based multi-scale feature decoder in the Toldt line segmentation model for prediction to obtain the location of the Toldt line inside the abdominal cavity;

[0146] The predicted Toldt lines are automatically plotted onto the original surgical video image using an algorithm to achieve a visualization effect of segmentation.

[0147] This visualization effect in embodiment S3.4 provides surgical medical staff with intuitive and clear reference information, which can help them make more accurate decisions during the operation, thereby enhancing the safety of the operation and improving the overall effect of the operation.

[0148] For examples of this application, please refer to [link / reference]. Figure 6-8 ;

[0149] Figure 6 This is a schematic diagram of the FSA-Net workflow provided in this application, illustrating the workflow of predicting the location of Toldt's lines inside the abdominal cavity based on the Toldt's line segmentation model (FSA-Net) in Embodiment 1.

[0150] Figure 7 This is a visualization comparison of the Toldt's line segmentation effects of different segmentation models provided in this application. It shows the Toldt's line segmentation of different segmentation models. The vertical image corresponding to "Ours" is the prediction result of the Toldt's line position inside the abdominal cavity in this embodiment 1.

[0151] Figure 8 This is a comparison chart of the quantification results of Toldt's line segmentation performance of different segmentation models provided in this application. It shows the quantification results of Toldt's line segmentation performance of different segmentation models. The horizontal data corresponding to "FSA-Net" represents the quantification results of Toldt's line segmentation performance within the abdominal cavity in Embodiment 1 of this application. Figure 7-8 As shown, the segmentation performance of FSA-Net in this embodiment is significantly better than that of existing endoscopic image segmentation models.

[0152] Overall, this embodiment has the following beneficial effects:

[0153] This application first utilizes a Toldt's line segmentation model to perform feature modeling on laparoscopic video frames. For the model, the Toldt's line, as part of an anatomical structure, has unique geometric characteristics at its boundary. By precisely quantifying these geometric characteristics through a pre-defined algorithm, the model can learn the unique shape and orientation of the Toldt's line boundary. Guided by fractal dimension features, the model can focus on areas most relevant to the Toldt's line boundary, thus ignoring background noise and interference factors. This focusing ability helps improve the accuracy and stability of recognition. Spatiotemporal feature fusion and aggregation are performed on multi-scale feature maps. This process not only integrates spatial information at different scales but also captures dynamic changes in the temporal dimension, helping the model to more accurately track and identify the location of the Toldt's line when processing dynamic laparoscopic videos. Finally, the fused features are input into a multi-scale decoder for prediction. The multi-scale decoder can utilize rich feature information for fine segmentation prediction, thereby accurately outputting the location of the Toldt's line, helping the model maintain high accuracy when processing complex anatomical structures and low-contrast regions.

[0154] Example 2:

[0155] Please see Figure 9 The embodiments of this application provide a fractal-driven laparoscopic Toldt's line positioning device, including a data module 10, a modeling module 20 and a prediction module 30;

[0156] Among them, data module 10 is used to acquire several laparoscopic video frames;

[0157] Modeling module 20 is used to perform feature modeling on several laparoscopic video frames based on the Toldt line segmentation model to obtain multi-scale feature maps. The Toldt line segmentation model is established by quantifying the geometric characteristics of the Toldt line boundary through a preset algorithm and using fractal dimension features to guide spatial attention to extract anatomical features.

[0158] The prediction module 30 is used to perform multi-scale spatiotemporal feature fusion and aggregation processing on the multi-scale feature map, and input the processing results into the multi-scale decoder for prediction to obtain the location of Toldt's line inside the abdominal cavity.

[0159] In one embodiment, data module 10 specifically comprises:

[0160] Video frames are extracted from laparoscopic surgical video images at a rate of 2 frames per second. Then, the size of these frames is downsampled to 384×384 pixels. Next, data augmentation processing is performed on each frame, including scaling, rotation, flipping, and random noise addition, to finally generate a number of processed laparoscopic video frames.

[0161] In one embodiment, the modeling module 20 includes a video unit, a labeling unit, a training unit, an input unit, and a construction unit. The video unit, labeling unit, and training unit constitute the process of establishing a Toldt line segmentation model, while the input unit and construction unit constitute the process of calculating multi-scale feature maps based on the Swin-Transformer backbone network encoder. Specifically:

[0162] The video unit is used to acquire image data of laparoscopic videos of mesocolon resection surgery and extract surgical video frames containing the Toldt's line region from the image data. In order to acquire laparoscopic image data of mesocolon resection surgery, data was first collected from the surgical procedures of 45 different patients. Each patient's surgery was recorded into several video segments with a duration of approximately 1 to 2 hours. Subsequently, short video segments with a duration of 15 to 30 seconds were selected from each long video segment by a preset model or by professional surgeons to ensure that each selected frame clearly showed the Toldt's line region. Finally, 150 high-quality laparoscopic surgical video segments were obtained. This data can lay a solid foundation for the subsequent training of the Toldt's line segmentation model.

[0163] The labeling unit is used to extract each frame of the edited surgical video, exporting and saving each frame as a PNG image file. Subsequently, these extracted frame sequences are downsampled, reducing the original video frame rate from 30 FPS to 2 FPS (2 frames per second). After downsampling, a total of 5000 images are obtained, each clearly showing the Toldt's line region. Next, using a pre-defined model or manual annotation by experienced colorectal surgeons, the Toldt's line region is precisely defined by drawing polylines, resulting in a laparoscopic Toldt's line segmentation dataset. After annotation, the pre-defined model or a colorectal surgery professor conducts a quality assessment to ensure that all annotation results in the laparoscopic Toldt's line segmentation dataset meet the pre-defined standards.

[0164] The training unit is used to train the deep learning model on the laparoscopic Toldt's line segmentation dataset according to the fractal dimension estimation algorithm and the preset modules, to obtain the Toldt's line segmentation model, which is a deep learning model based on the fractal-driven collaborative anatomical perception network (FSA-Net). The fractal dimension estimation algorithm is used to calculate the fractal dimension. The preset modules include a Local Texture-Aware Convolution (LTC) submodule and a Fractal-Guided Anatomical Consistent Attention (FAA) submodule, which are used to extract anatomical features by guiding spatial attention through fractal dimension features, aiming to collaboratively extract internal and external anatomical features. Furthermore, the Toldt's line segmentation model includes a Swin-Transformer-based encoder (i.e., a Swin-Transformer backbone network encoder), a metric theory-based fractal dimension estimation module, a fractal-guided feature aggregation module, and a Transformer-based multi-scale feature decoder.

[0165] Furthermore, performance tests of the Toldt line segmentation model show that it can process 30 to 40 frames per second, a processing speed that perfectly matches the 30 frames per second frame rate of laparoscopic cameras used in most current colorectal cancer surgeries, thus proving that the model has the ability to perform real-time segmentation of surgical video images.

[0166] For examples of this application, please refer to [link / reference]. Figure 2-3 , Figure 2 This is a schematic diagram illustrating the construction process of the Toldt's line segmentation model for laparoscopic video provided in this application, showing the construction process of the Toldt's line segmentation model;

[0167] Figure 3 This is a schematic diagram of the network architecture of FSA-Net provided in this application, which shows the network architecture of the Toldt line segmentation model.

[0168] In this embodiment, the video unit, labeling unit, and training unit construct a high-quality laparoscopic Toldt's line segmentation dataset by labeling the regions where Toldt's lines are located in surgical video frames. This helps the deep learning model learn the accurate features of Toldt's lines, thereby improving recognition accuracy. Fractal dimension estimation algorithms typically have low computational complexity, and applying them to the training process of deep learning models does not significantly increase the computational burden. By guiding spatial attention through fractal dimension features, the model can more accurately focus on regions with significant fractal characteristics, thereby extracting more accurate and relevant anatomical features.

[0169] In summary, Toldt's lines segmentation in videos is extremely challenging due to significant inter-individual variations in their location and morphology, as well as low contrast with surrounding tissues and blurred boundaries. To address this, this embodiment innovatively proposes the FSA-Net framework, specifically designed for Toldt's white lines segmentation in surgical laparoscopic images. FSA-Net first introduces a novel fractal dimension estimation algorithm to ensure computational accuracy. Then, it integrates two sub-modules: Local Texture-Aware Convolution (LTC) and Fractal-Guided Anatomical Consistent Attention (FAA), to collaboratively capture anatomical features. Furthermore, this embodiment constructs the first laparoscopic Toldt's line segmentation dataset to support related research. Experimental validation on multiple datasets shows that FSA-Net significantly improves the segmentation accuracy of Toldt's lines, providing strong support for surgical procedures.

[0170] The input unit is used to input the target continuous frames from several laparoscopic video frames into the Swin-Transformer backbone network encoder of the Toldt line segmentation model; wherein, the target continuous frames are a three-frame sequence consisting of keyframes and their preceding and following adjacent frames.

[0171] The building unit is used to model the local and global features of consecutive target frames based on the Swin-Transformer backbone network encoder through a sliding window attention mechanism, resulting in multi-scale feature maps. The multi-scale feature maps consist of four scales {F1, F2, F3, F4}, with the resolution decreasing sequentially to 1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the input size. Furthermore, the Swin-Transformer backbone network encoder combines sliding window attention and global attention, which can effectively model both global and local features of the image simultaneously.

[0172] In the input and construction units of this embodiment, the sliding window attention mechanism can capture global information through the interaction between windows while processing local features. This enables the model to understand the overall structure while maintaining sensitivity to details, thereby accurately identifying key anatomical structures such as Toldt's lines in complex laparoscopic video frames.

[0173] In one embodiment, the prediction module 30 includes a convolution unit, a computation unit, an extraction subunit, a filtering subunit, and a prediction unit.

[0174] The convolutional and computational units estimate the fractal dimension of multi-scale features to obtain pixel-level fractal dimension feature maps. The extraction sub-unit extracts the first output feature map based on the local texture-aware convolutional sub-module. The filtering sub-unit extracts multi-scale enhanced feature maps based on the fractal-guided anatomical consistent attention sub-module. The prediction unit predicts the location of Toldt's lines inside the abdominal cavity based on the Transformer multi-scale feature decoder.

[0175] The convolutional unit is used at each scale to perform cross-temporal convolution on the feature maps of three consecutive frames in the multi-scale feature map using a spatiotemporal fusion module based on 3D convolutional layers, resulting in a spatiotemporal fusion feature map F. i Among them, F i The subscript i represents different feature map scales; specifically, i = {1, 2, 3, 4}, corresponding to feature map scales of {1 / 4, 1 / 8, 1 / 16, 1 / 32} of the original input. Furthermore, the spatiotemporal fusion module based on 3D convolutional layers aims to effectively utilize temporal information between consecutive frames to improve the segmentation effect of Toldts lines. It is worth noting that 3D convolution, as a mature and widely used deep learning component, is often used to extract spatial features of 3D images and temporal information between consecutive video frames.

[0176] The computational unit is used to calculate the spatiotemporal fusion feature map F according to the fractal dimension estimation algorithm based on the "measure theory-based fractal dimension estimation module" in the Toldt line segmentation model. i The local fractal dimension of each pixel location is used to obtain the pixel-level fractal dimension feature map D. i Among them, the fractal dimension characteristic diagram D i Let {D1,D2,D3,D4} be the size of the spatiotemporal fusion feature map F. i "Similarly; fractal dimension can effectively characterize the texture complexity of different anatomical structures; using fractal dimension features as constraints helps the model better perceive different anatomical structures."

[0177] For examples of this application, please refer to [link / reference]. Figure 4 , Figure 4 This is a flowchart of the fractal dimension estimation algorithm based on metric theory provided in this application. It shows the algorithm flow proposed in this embodiment for estimating the fractal dimension of an image at the pixel level, aiming to estimate the fractal dimension of a depth feature map more efficiently.

[0178] Wherein, Reflect Padding represents the reflection padding operation, used to add edge pixels to the feature map. B(Fi[H,W],ε) represents the ε×ε neighborhood centered at Fi[H,W], E[·] represents the mathematical expectation; q is the scaling exponent, which is set to 2 in this embodiment for ease of calculation during fractal dimension estimation; μ(x) represents the probability measure of pixel Fi, and Dis(·) represents the Euclidean distance between two pixels Fi; the specific calculation process is as follows:

[0179] ① Input feature map: Input a feature map Fi of size H×W.

[0180] ② Reflection fill: Perform reflection fill operation on feature map Fi to increase edge pixels.

[0181] ③ Define neighborhood: For each pixel in the feature map Fi, define a neighborhood B(Fi[H,W],ε) with the pixel as the center and a size of ε×ε.

[0182] ④ Calculate the probability measure: For each pixel, calculate its probability measure μ(x); and dμ(x) in the formula is used to represent the probability that each pixel in the neighborhood B(Fi[H,W],ε) is sampled, and its size is a fixed value.

[0183] ⑤ Calculate distance and expectation: Calculate the Euclidean distance Dis(·) between a pixel x in neighborhood B(Fi[H,W],ε) and any other pixel y in neighborhood B(Fi[H,W],ε). Furthermore, the scaling exponent q is not involved in the calculation of the mathematical expectation, but is used in (ε / E[·]). q-1 The calculation.

[0184] ⑥ Estimating fractal dimension: Using a measure-based fractal dimension estimation method, combined with the probability measure μ(x) and mathematical expectation E[·] calculated above, the fractal dimension of the feature map is estimated.

[0185] ⑦ Output: Output the estimated fractal dimension.

[0186] In this embodiment, the convolutional and computational units perform cross-temporal convolutions on multi-scale feature maps to obtain spatiotemporal fusion feature maps. This captures the temporal dependencies between video frames and effectively fuses spatial and temporal features. The local fractal dimension of each pixel in the spatiotemporal fusion feature map is calculated using a fractal dimension estimation algorithm to obtain a fractal dimension feature map. Fractal dimension, as a geometric feature, can reflect the complexity and irregularity of anatomical structures.

[0187] While box-counting dimension is effective in estimating fractal dimensions for different structures, the feature maps extracted from the backbone of deep learning models often have high embedding dimensions, which significantly increases computational complexity. In contrast, the metric-based fractal dimension estimation method used in this embodiment, such as the Rényi generalized dimension, exhibits stronger robustness when dealing with such high-dimensional data and can significantly alleviate the curse of dimensionality caused by direct geometric overlay.

[0188] Extracting sub-units for spatiotemporal fusion feature maps F i and fractal dimension characteristic map D i Simultaneously inputting a Local Texture-Aware Convolutional (LTC) submodule; wherein, the Local Texture-Aware Convolutional submodule belongs to the "Fractal-Guided Feature Aggregation Module" in the Toldt line segmentation model, and the "Fractal-Guided Feature Aggregation Module" also includes a Fractal-Guided Anatomical Consistent Attention (FAA) submodule, used for subsequent extraction of multi-scale enhanced feature maps; in addition, the "Fractal-Guided Feature Aggregation Module" combines the Local Texture-Aware Convolutional Module and the Fractal-Driven Anatomical Consistent Attention, which helps FSA-Net to collaboratively capture the contextual relationships within the anatomical structure and between different anatomical structures, thereby improving the model's ability to distinguish different anatomical structures;

[0189] Extracting sub-units is also used to perform spatiotemporal fusion feature mapping F along the channel dimension. i and fractal dimension characteristic map D i The components are concatenated to generate a joint feature tensor.

[0190] The extracted sub-units are also used for high-level semantic extraction of the joint feature tensor to obtain the expansion coefficient E. i and offset direction A i ;

[0191] Extracting sub-units is also used based on the expansion coefficient E i Offset direction A i Spatiotemporal fusion feature map F i The first output feature map of the local texture-aware convolutional submodule is calculated.

[0192] For examples of this application, please refer to [link / reference]. Figure 5 , Figure 5 This is a schematic diagram of the fractal-guided feature aggregation module provided in this application, illustrating the use of the fractal-guided feature aggregation module to process the spatiotemporal fusion feature map F. i and fractal dimension characteristic map D i The process of data processing;

[0193] The local texture-aware convolutional submodule is designed to adaptively integrate features from different anatomical structures. Compared to standard deformable convolutional networks (DCNs), this embodiment can utilize anatomical knowledge embedded in fractal features to guide the offset of the convolutional kernel, thereby effectively avoiding the model learning unreasonable deformations.

[0194] Specific implementation as follows Figure 5 As shown in section “(b) Local Texture-Aware Convolution (LTC)”, the spatiotemporal fusion feature map F i and fractal dimension characteristic map D i As input, these are used together to predict the offset of the DCN kernel. The offset generator consists of two 3×3 depthwise separable convolutional blocks, which are responsible for predicting the dilation coefficient E. i and offset direction A i Then, using these predicted expansion coefficients and offset directions, the final offset O can be calculated. i After obtaining the offset, it is combined with the original spatiotemporal fusion feature map F. i and offset O i The output features (i.e., the first output feature map) are calculated through dilated and deformable convolution operations. In other words, the "dilated and deformable convolutional layer" used will accept the "spatiotemporal fusion feature map F". i and offset O i "As a parameter, the output feature map of the local texture-aware convolutional submodule is calculated.

[0195] In this embodiment, the extraction sub-unit generates a joint feature tensor by concatenating the spatiotemporal fusion feature map and the fractal dimension feature map along the channel dimension. The fractal dimension feature map provides detailed information about the texture of the anatomical structure, while the spatiotemporal fusion feature map contains high-level semantic information. This process achieves the fusion of multi-scale features. High-level semantic extraction is performed on the joint feature tensor to obtain the dilation coefficient and offset direction, realizing the transformation from low-level features to high-level semantics. The high-level semantic information helps the model understand the overall layout and interrelationships of the anatomical structure, providing strong support for subsequent segmentation tasks.

[0196] A filtering subunit is used to input the first output feature map into the fractal-guided dissection consistent attention submodule;

[0197] The filtering subunit is also used to apply a linear transformation layer to the fractal dimension feature map D. iConverted to normalized distribution map M i And from the distribution map M i The indexes of the target foreground region and the target background region are extracted to obtain an index set. The target foreground region refers to the anatomical structure or pathological region that needs to be focused on, which is usually directly related to the diagnostic and treatment goals. For example, the Toldt's line-related region (as an anatomical boundary marker between the mesentery and retroperitoneal tissue in colon surgery, its accurate identification is crucial for surgical layer separation and vascular protection). The target background region refers to the interfering region that is not the core diagnostic and treatment goal, such as surgical instruments (which need to be masked in intraoperative images to avoid affecting the surgical field analysis) and distal tissues (such as adjacent organs or non-pathological tissues, which usually need to be excluded from the analysis).

[0198] The filtering subunit is also used to flatten the first output feature map and generate a triplet embedding containing Query, Key and Value through matrix projection;

[0199] The filtering subunit is also used to filter the key-value pairs corresponding to the target foreground region and the target background region respectively in the triple embedding according to the index set, so as to obtain an embedding set containing only the key and value;

[0200] The filtering subunit is also used for filtering based on the embedding set, according to the distribution map M. i Select the key-value pairs belonging to the foreground region, calculate the attention weights using global queries and foreground keys, and then aggregate them with foreground values ​​to generate the foreground attention result. Similarly, generate the background attention result. Concatenate the foreground attention result and the background attention result to obtain the attention feature map.

[0201] The filtering subunit is also used to adjust the number of channels in the attention feature map to the spatiotemporal fusion feature map F through compression using a 1×1 convolutional layer. i The number of channels is determined, and the channel-aligned attention feature maps are fused to the spatiotemporal fusion feature map F via skip connections. i This updates the spatiotemporal fusion feature map F. i The multi-scale enhanced feature map output by the fractal-guided anatomy consistent attention submodule is obtained.

[0202] Specific implementation as follows Figure 5 As shown in the “(c) Fractal-guided Anatomical Consistent Attention (FAA)” section, the core of this module is a fractal-guided key-value pair filtering strategy designed to enhance the model’s ability to capture interanatomical information in a divide-and-conquer manner.

[0203] In this embodiment, the filtering subunit converts the fractal dimension feature map into a normalized distribution map and extracts the indices of the target foreground region and the target background region from it to obtain an index set. This process can accurately distinguish between the foreground and background regions, providing accurate target regions for subsequent attention feature extraction.

[0204] Furthermore, filtering the key-value pairs corresponding to the target foreground and background regions using the index set demonstrates the need to maintain the globality of the query. Therefore, calculating the foreground and background attention results using the query and the filtered key-value pairs allows the model to focus more on key local anatomical structures under the guidance of global information. Adjusting the number of channels in the attention feature map through convolutional compression to align it with the number of channels in the spatiotemporal fusion feature map preserves the key information of the attention feature map while ensuring consistency in subsequent feature fusion.

[0205] In general, this embodiment extracts multiple features from the fractal dimension feature map using extraction and filtering sub-units to obtain a multi-scale enhanced feature map. This multi-path design can capture local detailed features and global contextual information, and can construct a more comprehensive and richer multi-scale enhanced feature map.

[0206] Furthermore, the local texture-aware convolutional submodule focuses on the local texture features of the anatomical structure, extracting high-level semantics through adaptive convolutional kernels to obtain the first output feature map. The fractal-guided anatomical consistency attention submodule, on the other hand, emphasizes the global consistency of the anatomical structure and the distinction between foreground and background, filtering features through an attention mechanism to obtain multi-scale enhanced feature maps. These two data processing methods capture information about the anatomical structure at different scales, thus enabling the construction of more comprehensive and richer multi-scale enhanced feature maps.

[0207] The prediction unit is used to input the multi-scale enhanced feature map into the Transformer-based multi-scale feature decoder in the Toldt line segmentation model for prediction, so as to obtain the location of the Toldt line inside the abdominal cavity.

[0208] The prediction unit is also used to automatically plot the predicted Toldt lines onto the original surgical video image using an algorithm, so as to achieve a visualization effect of segmentation.

[0209] The visualization effect of the prediction unit in this embodiment provides surgical medical staff with intuitive and clear reference information, which can help them make more accurate decisions during surgery, thereby enhancing the safety of the surgery and improving the overall surgical outcome.

[0210] For examples of this application, please refer to [link / reference]. Figure 6-8 ;

[0211] Figure 6 This is a schematic diagram of the FSA-Net workflow provided in this application, illustrating the workflow of predicting the location of Toldt's lines inside the abdominal cavity based on the Toldt's line segmentation model (FSA-Net) in Embodiment 2.

[0212] Figure 7 This is a visualization comparison of the Toldt's line segmentation effects of different segmentation models provided in this application. It shows the Toldt's line segmentation situation of different segmentation models. Among them, the vertical image corresponding to "Ours" is the prediction result of the Toldt's line position inside the abdominal cavity in this embodiment 2.

[0213] Figure 8 This is a comparison chart of the quantification results of Toldt's line segmentation performance of different segmentation models provided in this application. It shows the quantification results of Toldt's line segmentation performance of different segmentation models. The horizontal data corresponding to "FSA-Net" represents the quantification results of Toldt's line segmentation performance within the abdominal cavity in Embodiment 2 of this application. Figure 7-8 As shown, the segmentation performance of FSA-Net in this embodiment is significantly better than that of existing endoscopic image segmentation models.

[0214] Overall, this embodiment has the following beneficial effects:

[0215] This application first utilizes a Toldt's line segmentation model to perform feature modeling on laparoscopic video frames. For the model, the Toldt's line, as part of an anatomical structure, has unique geometric characteristics at its boundary. By precisely quantifying these geometric characteristics through a pre-defined algorithm, the model can learn the unique shape and orientation of the Toldt's line boundary. Guided by fractal dimension features, the model can focus on areas most relevant to the Toldt's line boundary, thus ignoring background noise and interference factors. This focusing ability helps improve the accuracy and stability of recognition. Spatiotemporal feature fusion and aggregation are performed on multi-scale feature maps. This process not only integrates spatial information at different scales but also captures dynamic changes in the temporal dimension, helping the model to more accurately track and identify the location of the Toldt's line when processing dynamic laparoscopic videos. Finally, the fused features are input into a multi-scale decoder for prediction. The multi-scale decoder can utilize rich feature information for fine segmentation prediction, thereby accurately outputting the location of the Toldt's line, helping the model maintain high accuracy when processing complex anatomical structures and low-contrast regions.

[0216] Example 3:

[0217] This application provides a computer-readable storage medium including a stored computer program, wherein the computer program, when running, controls the device where the computer-readable storage medium is located to execute the fractal-driven laparoscopic Toldt's line localization method.

[0218] The fractal-driven laparoscopic Toldt's line localization method, when implemented as a software functional unit and used as an independent product, can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the above embodiments of the present invention can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.

[0219] The above are preferred embodiments of the present invention. It should be noted that, for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications are also considered to be within the scope of protection of the present invention.

Claims

1. A fractal-driven laparoscopic Toldt's line localization method, characterized in that, include: Acquire several laparoscopic video frames; The Toldt line segmentation model is used to model the features of the several laparoscopic video frames to obtain multi-scale feature maps. The Toldt line segmentation model is established by quantifying the geometric characteristics of the Toldt line boundary through a preset algorithm and using fractal dimension features to guide spatial attention to extract anatomical features. The multi-scale feature map is subjected to multi-scale spatiotemporal feature fusion and aggregation processing. The processed result is input into a multi-scale decoder for prediction to obtain the location of Toldt's line inside the abdominal cavity, specifically: Perform cross-temporal convolution on the multi-scale feature map to obtain a spatiotemporal fusion feature map; The local fractal dimension of each pixel in the spatiotemporal fusion feature map is calculated using the fractal dimension estimation algorithm to obtain the fractal dimension feature map. The spatiotemporal fusion feature map and the fractal dimension feature map are input into the local texture-aware convolutional submodule of the Toldt line segmentation model; the spatiotemporal fusion feature map and the fractal dimension feature map are concatenated along the channel dimension to generate a joint feature tensor; high-level semantic extraction is performed on the joint feature tensor to obtain the dilation coefficient and offset direction; the first output feature map of the local texture-aware convolutional submodule is calculated based on the dilation coefficient, the offset direction, and the spatiotemporal fusion feature map; the first output feature map is input into the fractal-guided anatomy consistent attention submodule of the Toldt line segmentation model for filtering to obtain a multi-scale enhanced feature map; The multi-scale enhanced feature map is input into the multi-scale decoder for prediction to obtain the location of the Toldt line inside the abdominal cavity.

2. The fractal-driven laparoscopic Toldt's line localization method as described in claim 1, characterized in that, The first output feature map is input into the fractal-guided anatomy consistent attention submodule of the Toldt line segmentation model for filtering to obtain the multi-scale enhanced feature map, specifically: The first output feature map is input into the fractal-guided anatomy consistent attention submodule of the Toldt line segmentation model; The fractal dimension feature map is converted into a normalized distribution map, and the indexes of the target foreground region and the target background region are extracted from the distribution map to obtain an index set; The first output feature map is flattened and a triplet embedding containing Query, Key and Value is generated by matrix projection. The triplet embedding is filtered and attention features are extracted based on the index set to obtain the multi-scale enhanced feature map.

3. The fractal-driven laparoscopic Toldt's line localization method as described in claim 2, characterized in that, Based on the index set, the triplet embeddings are subjected to data filtering and attention feature extraction to obtain the multi-scale enhanced feature map, specifically: In the triplet embedding, the key-value pairs corresponding to the target foreground region and the target background region are filtered according to the index set to obtain an embedding set containing only the key and value; Based on the Query and the embedding set, foreground attention results and background attention results are calculated. The foreground attention results and background attention results are then concatenated to obtain an attention feature map. The number of channels in the attention feature map is adjusted to the number of channels in the spatiotemporal fusion feature map by convolutional compression. The channel-aligned attention feature map is then fused to the spatiotemporal fusion feature map through skip connections to obtain the multi-scale enhanced feature map.

4. The fractal-driven laparoscopic Toldt's line localization method as described in claim 1, characterized in that, The Toldt line segmentation model is established by quantifying the geometric properties of the Toldt line boundary using a preset algorithm and extracting anatomical features by using fractal dimension features to guide spatial attention. Specifically: Acquire surgical video frames containing the Toldt's line region; Mark the regions where Toldt's lines are located in the surgical video frames to obtain a laparoscopic Toldt's line segmentation dataset; Based on the fractal dimension estimation algorithm and the preset module, the deep learning model is trained using the laparoscopic Toldt line segmentation dataset to obtain the Toldt line segmentation model; wherein, the fractal dimension estimation algorithm is used to calculate the fractal dimension, and the preset module is used to extract anatomical features by guiding spatial attention through fractal dimension features.

5. The fractal-driven laparoscopic Toldt's line localization method as described in claim 1, characterized in that, Based on the Toldt line segmentation model, feature modeling is performed on the several laparoscopic video frames to obtain multi-scale feature maps, specifically: The target continuous frames in the plurality of laparoscopic video frames are input into the Swin-Transformer backbone network encoder of the Toldt line segmentation model. Based on the Swin-Transformer backbone network encoder, the local and global features of the target continuous frames are modeled through a sliding window attention mechanism to obtain the multi-scale feature map.

6. A fractal-driven Toldt's line positioning device for laparoscopy, characterized in that, It includes a data module, a modeling module, and a prediction module; The data module is used to acquire several laparoscopic video frames. The modeling module is used to perform feature modeling on the several laparoscopic video frames according to the Toldt line segmentation model to obtain multi-scale feature maps; wherein, the Toldt line segmentation model is established by quantifying the geometric characteristics of the Toldt line boundary through a preset algorithm and using fractal dimension features to guide spatial attention to extract anatomical features; The prediction module is used to perform multi-scale spatiotemporal feature fusion and aggregation processing on the multi-scale feature map, and input the processing result into the multi-scale decoder for prediction to obtain the position of Toldt's line inside the abdominal cavity. The prediction module includes a convolution unit, a computation unit, an extraction unit, and a prediction unit. The convolutional unit is used to perform cross-temporal convolution on the multi-scale feature map to obtain a spatiotemporal fusion feature map. The computing unit is used to calculate the local fractal dimension of each pixel position in the spatiotemporal fusion feature map according to the fractal dimension estimation algorithm, so as to obtain the fractal dimension feature map. The extraction unit includes an extraction subunit and a filtering subunit. The extraction subunit is used to input the spatiotemporal fusion feature map and the fractal dimension feature map into the local texture-aware convolutional submodule of the Toldt line segmentation model; concatenate the spatiotemporal fusion feature map and the fractal dimension feature map along the channel dimension to generate a joint feature tensor; perform high-level semantic extraction on the joint feature tensor to obtain the dilation coefficient and offset direction; and calculate the first output feature map of the local texture-aware convolutional submodule based on the dilation coefficient, the offset direction, and the spatiotemporal fusion feature map. The filtering subunit is used to input the first output feature map into the fractal-guided anatomy consistent attention submodule of the Toldt line segmentation model for filtering to obtain a multi-scale enhanced feature map. The prediction unit is used to input the multi-scale enhanced feature map into the multi-scale decoder for prediction to obtain the position of the Toldt line inside the abdominal cavity.

7. The fractal-driven Toldt's line positioning device for laparoscopy as described in claim 6, characterized in that, The filtering subunit is specifically: The first output feature map is input into the fractal-guided anatomy consistent attention submodule of the Toldt line segmentation model; The fractal dimension feature map is converted into a normalized distribution map, and the indexes of the target foreground region and the target background region are extracted from the distribution map to obtain an index set; The first output feature map is flattened and a triplet embedding containing Query, Key and Value is generated by matrix projection. The triplet embedding is filtered and attention features are extracted based on the index set to obtain the multi-scale enhanced feature map.

8. The fractal-driven Toldt's line positioning device for laparoscopy as described in claim 7, characterized in that, Based on the index set, the triplet embeddings are subjected to data filtering and attention feature extraction to obtain the multi-scale enhanced feature map, specifically: In the triplet embedding, the key-value pairs corresponding to the target foreground region and the target background region are filtered according to the index set to obtain an embedding set containing only the key and value; Based on the Query and the embedding set, foreground attention results and background attention results are calculated. The foreground attention results and background attention results are then concatenated to obtain an attention feature map. The number of channels in the attention feature map is adjusted to the number of channels in the spatiotemporal fusion feature map by convolutional compression. The channel-aligned attention feature map is then fused to the spatiotemporal fusion feature map through skip connections to obtain the multi-scale enhanced feature map.

9. The fractal-driven Toldt's line positioning device for laparoscopy as described in claim 6, characterized in that, The modeling module includes a video unit, a labeling unit, and a training unit; The video unit is used to acquire surgical video frames containing the Toldt line region; The labeling unit is used to label the region where the Toldt line is located in the surgical video frame to obtain a laparoscopic Toldt line segmentation dataset; The training unit is used to train the deep learning model on the laparoscopic Toldt line segmentation dataset according to the fractal dimension estimation algorithm and the preset module to obtain the Toldt line segmentation model; wherein, the fractal dimension estimation algorithm is used to calculate the fractal dimension, and the preset module is used to extract anatomical features by guiding spatial attention through fractal dimension features.

10. The fractal-driven Toldt's line positioning device for laparoscopy as described in claim 6, characterized in that, The modeling module includes an input unit and a construction unit; The input unit is used to input the target continuous frames from the plurality of laparoscopic video frames into the Swin-Transformer backbone network encoder of the Toldt line segmentation model. The construction unit is used to model the local and global features of the target continuous frames based on the Swin-Transformer backbone network encoder through a sliding window attention mechanism to obtain the multi-scale feature map.

11. A storage medium, characterized in that, The storage medium stores a computer program, which is called and executed by a computer to implement a fractal-driven laparoscopic Toldt's line localization method as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Laparoscope image segmentation method and system based on Scaleform algorithm

    CN115311317A

  • Simulated tissue structures and methods

    US20160293055A1