Laparoscopic image liver landmark segmentation method based on SAM enhanced dual-decoder network
Through the SAM-enhanced dual decoder network, combined with internal and external liver consistency loss and deep supervision loss training, the accuracy and stability of liver landmark segmentation in laparoscopic environment is solved, and efficient and accurate liver and anatomical landmark segmentation is achieved to adapt to the application needs of different scenarios.
Patent Information
- Application Number
- CN202510541172.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-08-12
AI Technical Summary
The prior art is difficult to accurately segment anatomical landmarks on the liver surface in a laparoscopic environment, especially in complex situations such as light changes, instrument occlusion and tissue interference, and the calculation complexity is high, slow speed, insufficient generalization ability, and difficult to meet the needs of real-time application.
Using a dual decoder network based on SAM enhancement, combined with SAM encoder, liver decoder and landmark decoder, liver and landmark segmentation is optimized through internal and external liver consistency loss and deep supervision loss training, and a pre-trained SAM encoder is used to extract robust features, reduce inter-task interference, and design a comprehensive loss function to improve segmentation accuracy and stability.
It significantly improves the segmentation accuracy and robustness of liver and anatomical landmarks, can accurately identify landmarks in complex environments, improves the generalization ability and segmentation speed of the model, and adapts to the detection performance under different data sets and devices.
Smart Images

Figure CN120472156A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image segmentation in computer technology, and in particular to a laparoscopic image liver landmark segmentation method based on a SAM enhanced dual decoder network. Background Art
[0002] Accurate registration of 2D images with 3D models is a prerequisite for AR. This relies heavily on identifying anatomical landmarks on the liver surface in the image, such as the liver outline, falciform ligament, and hepatic ridge. However, in practice, the liver deforms due to breathing, instrument manipulation, and other factors, resulting in changes in the structure and position of these landmarks. Furthermore, accurate segmentation of these landmarks is challenging due to their small size, exposure to lighting variations, instrument occlusion, and tissue interference.
[0003] Existing methods for laparoscopic liver surface landmark segmentation typically convert the original coordinate-based landmark detection into a segmentation task, representing the liver surface landmarks as thick curves. However, since anatomical landmarks occupy a very small proportion of image pixels, this extremely unbalanced sample size makes reliable automatic segmentation difficult without the aid of liver structural information.
[0004] Current technologies use 2D image segmentation models such as ResUNet and SwinUNet for image segmentation. For example, the paper "Weighted res-unet for high-quality retina vessel segmentation" proposes an improved U-Net model to detect small blood vessels and segment the optic disc region in retinal vessel segmentation. However, this method directly uses laparoscopic images as input and is only used for the single task of segmenting / identifying liver surface landmarks. This method fails to consider the inherent and unique association between liver surface landmarks and the liver's appearance / shape, making it difficult to achieve stable and accurate landmark segmentation for identifying liver regions in images. Furthermore, D2GPLand is also a 2D image segmentation network, but its input requires a depth map estimated from a laparoscopic image. Stably estimating a depth map from a laparoscope is a difficult task, and inaccurate depth maps can severely impact the accuracy of landmark segmentation.
[0005] In summary, the current existing technology has the following shortcomings:
[0006] (1) Difficulty in coping with complex scenarios: Existing technologies perform poorly in complex laparoscopic environments and have difficulty dealing with illumination changes, instrument occlusion, and tissue interference. They are unable to effectively cope with changes in liver morphology and image quality, resulting in limited segmentation accuracy.
[0007] (2) High computational complexity and slow speed: Some existing methods use pre-trained complex encoders and additional input data (such as depth maps) to improve segmentation accuracy. Although this improves accuracy, it greatly increases computational complexity, resulting in slow processing speed, making it difficult to meet the needs of real-time applications and limiting its application.
[0008] (3) Limited generalization ability: Existing segmentation techniques are not adaptable enough to diverse scenarios such as different datasets, objects, and devices. When tested on laparoscopic liver images from different sources, the performance of many methods dropped significantly, indicating that they have difficulty coping with the differences brought about by different scenarios and cannot achieve stable segmentation results.
[0009] Therefore, there is currently a lack of a laparoscopic image segmentation method to solve or partially solve the above problems. Summary of the Invention
[0010] The purpose of the present invention is to overcome the defects of the above-mentioned prior art and provide a laparoscopic image liver landmark segmentation method based on a SAM enhanced dual decoder network to solve or partially solve the problem that the liver may be deformed due to factors such as breathing and operation, resulting in variable landmark structure and position, and the landmark size is small. In complex laparoscopic environments, such as lighting changes, instrument occlusion and tissue interference, the existing methods are difficult to accurately segment landmarks.
[0011] The purpose of the present invention can be achieved by the following technical solutions:
[0012] One aspect of the present invention provides a method for liver landmark segmentation in laparoscopic images based on a SAM enhanced dual decoder network, comprising the following steps:
[0013] Acquire the 2D laparoscopic key frame image to be segmented;
[0014] Extracting features using a SAM encoder based on the 2D laparoscopic key frame image;
[0015] Based on the features extracted by the SAM encoder, a liver decoder is used to extract liver contour information, obtain a liver mask, and implement liver segmentation;
[0016] Based on the features extracted by the SAM encoder, the landmark decoder is used to segment the falciform ligament, the liver ridge, and the liver contour to achieve segmentation of anatomical landmarks;
[0017] The SAM encoder, liver decoder and landmark decoder are pre-trained based on the liver internal and external consistency loss and the deep supervision loss.
[0018] As a preferred technical solution, the liver internal and external consistency loss is:
[0019]
[0020]
[0021] in, Loss of consistency between the liver and the outside, Dice in 、Dice out are the Dice losses in the liver inner and outer regions, respectively. liver is the predicted liver probability map output by the liver decoder, Represent the predicted probability maps of being inside the liver and outside the liver, respectively. Represent the true labels inside and outside the liver, respectively. land. is the predicted landmark probability map, represents the probability map of landmarks located inside the liver, represents the probability map of landmarks located outside the liver, Y liver is the true label in the sample, denote the true internal and external landmark annotations respectively, SPlit() denotes the split operation, and ∑ denotes the summation operation over all pixels.
[0022] As a preferred technical solution, the deep supervision loss is:
[0023]
[0024] in, is the depth supervision loss, N is the number of depth supervision layers, and denote the Dice loss and cross entropy loss of the i-th layer of the landmark decoder, respectively.
[0025] As a preferred technical solution, the training process also includes basic segmentation loss:
[0026]
[0027] in, is the basic segmentation loss, are Dice loss and cross entropy loss respectively, O is the probability map output by the liver decoder or landmark decoder, and Y is the corresponding true label.
[0028] As a preferred technical solution, the following steps are also included:
[0029] The SAM encoder, liver decoder, and landmark decoder were evaluated by calculating the Dice coefficient, Chamfer distance, and the average 2D chamfer distance of the test images, where the average 2D chamfer distance of the test images was calculated using the following formula:
[0030]
[0031] In the formula, CD (v,w) is the average 2D chamfer distance of all test images, CD i (v,w) is the average 2D chamfer distance of the test image i, N is the number of images in the test set, v and w are the pixel coordinate sets of the predicted segmented area and the true annotated area, respectively, |v| and |w| are the number of pixels in the predicted area and the true area, respectively.
[0032] As a preferred technical solution, any one of the liver decoder and the landmark decoder includes multiple layers of residual blocks and bilinear upsampling layers.
[0033] As a preferred technical solution, the SAM encoder is a pre-trained SAM-B encoder, and during the training process, the SAM-B encoder is end-to-end fine-tuned.
[0034] As a preferred technical solution, the training process includes the following steps:
[0035] Acquire an image dataset, normalize the images and resize them to a preset resolution, and augment the data by randomly inverting, rotating, and cropping them to increase the number of samples.
[0036] As a preferred technical solution, during the training process, the training is achieved by stochastic gradient descent with the goal of minimizing the loss function.
[0037] Another aspect of the present invention provides an electronic device comprising one or more processors, a memory, and one or more programs stored in the memory, wherein the one or more programs include instructions for executing the aforementioned laparoscopic image liver landmark segmentation method based on the SAM enhanced dual decoder network.
[0038] Compared with the prior art, the present invention has at least one of the following beneficial effects:
[0039] (1) Accurate segmentation and positioning: This application combines the SAM encoder + liver decoder / landmark decoder architecture, and a training scheme based on liver internal and external consistency loss and deep supervision loss, which significantly improves the segmentation accuracy of real-time 2D anatomical landmarks. In the test of the current L3D dataset, the overall Dice score reached 67.3%, the segmentation accuracy of the liver ridge, liver contour and falciform ligament was high, and the Chamfer distance was reduced, making the prediction results closer to manual annotation, and can accurately identify and locate landmarks, providing more reliable data for 3D-2D alignment and enhancing AR navigation effects.
[0040] (2) Strong generalization ability: This application shows good generalization ability in different datasets and complex scenarios. On the P2ILF dataset, despite many differences and challenges in the data, such as endoscope masking and large changes in liver surface features, the method still achieved good results. The overall Dice score was advantageous, the Chamfer distance increased the least, and it could maintain stable detection performance under different objects, devices and conditions.
[0041] (3) Adaptability to complex environments: It has strong adaptability in complex environments and can face problems such as lighting changes, instrument occlusion, and tissue interference. For example, in the key frames of the L3D dataset, when the lighting is poor, the liver is obscured by fat, or there are blood stains, it can still accurately identify landmarks, showing strong robustness and ensuring navigation accuracy.
[0042] (4) Optimizing model performance: Its key components effectively improve model performance. The dual decoder architecture reduces feature confusion between liver and landmark segmentation tasks, and collaborative optimization improves segmentation stability and accuracy. The consistency constraint inside and outside the liver strengthens the spatial alignment relationship between the liver and landmarks, making the segmentation results more consistent with the anatomical structure. The deep supervision mechanism enhances the model's multi-scale learning ability, improves training stability, and significantly improves the segmentation accuracy of specific challenging landmarks such as the falciform ligament. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 Flowchart of a method for liver landmark segmentation in laparoscopic images based on a SAM enhanced dual decoder network in an embodiment;
[0044] Figure 2 Schematic diagram of an electronic device in an embodiment. DETAILED DESCRIPTION
[0045] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0046] Example 1
[0047] In response to the problems existing in the aforementioned existing technologies, this embodiment provides a laparoscopic image liver landmark segmentation method based on a SAM enhanced dual decoder network. It uses a dual decoder architecture, one responsible for liver segmentation and the other focusing on landmark segmentation to reduce feature interference; with the help of pre-trained SAM encoder, robust features are extracted from the input key frames to enhance the model's adaptability to complex scenes; liver-guided consistency constraints are introduced, and a loss function is constructed by calculating the overlap of landmark segmentation in the areas inside and outside the liver to strengthen the spatial association between liver and landmark segmentation; a complete loss function consisting of basic segmentation loss, consistency loss and auxiliary deep supervision mechanism is designed to comprehensively optimize the model and ultimately achieve accurate and robust liver landmark segmentation.
[0048] See also Figure 1 , the method comprises the following steps:
[0049] Step S1: data preparation and processing.
[0050] Selecting an appropriate dataset and preprocessing it to prepare for subsequent model training is crucial for model performance. This example uses the L3D and P2ILF datasets, normalizing the images and adjusting their resolution to 1024×1024. Data augmentation strategies such as random flipping, rotation, and cropping are also employed during training.
[0051] Step S2: construct a model.
[0052] The SAM-D2Net model, built on the PyTorch 2.4.0 framework, consists of a pre-trained SAM encoder, a dual decoder structure, and liver internal and external consistency constraints. In the dual decoder architecture, the liver decoder extracts the overall liver contour and outputs a liver mask, while the landmark decoder focuses on segmenting the falciform ligament, liver ridge, and liver contour. Both decoders contain five layers of "Res-block layers" and bilinear upsampling. This structure optimizes the liver and landmark feature extraction processes, reduces inter-task interference, and utilizes the overall liver structure to optimize landmark segmentation.
[0053] By introducing a pre-trained SAM encoder, the model's generalization ability is enhanced, and high-level features are extracted from the input 2D laparoscopic key frames. Trained on a large-scale dataset, it can effectively cope with changes in liver morphology and image quality, providing strong support for subsequent segmentation.
[0054] To address the challenges of 2D anatomical landmark segmentation, a dual-decoder architecture is proposed. The liver decoder extracts the overall liver contour and outputs a liver mask, while the landmark decoder focuses on segmenting the falciform ligament, hepatic ridge, and liver contour. Both decoders utilize a five-layer "residual block" and a bilinear upsampling strategy, independently optimizing feature extraction to reduce cross-task interference while leveraging the overall liver structure to optimize landmark segmentation.
[0055] Step S3: training the model.
[0056] Training was performed on an NVIDIA RTX A6000 GPU using stochastic gradient descent (SGD) as the optimizer with an initial learning rate of 1e-2. 2 The Net uses the pre-trained SAM-B as the encoder and performs end-to-end fine-tuning during training. During training, the complete loss function consists of a basic segmentation loss (combining Dice loss and cross-entropy loss), a liver internal and external consistency loss, and a deep supervision loss. Model performance is improved by optimizing this loss function.
[0057] The complete loss function includes basic segmentation loss, liver internal and external consistency loss and deep supervision loss, namely
[0058]
[0059] The model is trained and optimized by minimizing this complete loss function.
[0060] (1) Basic segmentation loss.
[0061] Basic segmentation loss A weighted combination of Dice loss and cross entropy loss is used to optimize the liver segmentation and landmark segmentation results. The Dice loss calculation formula is:
[0062]
[0063] Among them, O is the probability map output by the network, Y is the corresponding true label, and the cross entropy loss is used to optimize the pixel-level classification accuracy. The two are added together to obtain the basic segmentation loss
[0064]
[0065] (2) Loss of consistency inside and outside the liver.
[0066] In order to strengthen the spatial alignment between the liver and the landmark, the consistency loss inside and outside the liver is defined
[0067] For input keyframe images The liver decoder and landmark decoder output liver probability map O respectivelyliver ∈[0,1] (H×W×2) and landmark probability map O land. ∈[0,1] (H×W×4) , whose corresponding ground truth labels are Y liver ∈{0,1} (H×W×2) and Y land. ∈{0,1} (H×W×4) .
[0068] First, the liver prediction results are divided into internal and external regions, and the calculation formula is:
[0069]
[0070] The Split() operation divides the channel, the first channel represents the probability of pixels inside the liver, and the second channel represents the probability of pixels outside the liver. represents the probability map of landmarks located inside the liver, and represents the probability map of landmarks located outside the liver. Similarly, and Representing true internal and external landmark annotations, respectively.
[0071] Based on this, the landmark segmentation probability maps located in the inner and outer areas of the liver are obtained respectively.
[0072]
[0073] in, represents the probability map of landmarks located inside the liver, and represents the probability map of landmarks located outside the liver. Similarly, and Representing true internal and external landmark annotations, respectively.
[0074] The Dice coefficient is used to evaluate the degree of overlap of the landmark segmentation results in the liver and outside areas. The formula is:
[0075]
[0076] Among them, ∑ represents the summation operation of all pixels. Finally, this study defines the consistency loss of the liver inside and outside as:
[0077]
[0078] By optimizing this loss, liver and landmark segmentation can be mutually promoted, thus improving segmentation accuracy.
[0079] (3) Deep Supervision Loss.
[0080] Due to the characteristics of anatomical landmarks in 2D images, in order to enhance the multi-scale learning ability and training stability of the model, an additional segmentation head is added after the output feature maps of the 2nd, 3rd, and 4th layers of the landmark decoder. The calculation formula is
[0081]
[0082] Among them, N is the number of layers of deep supervision (set to 3 in this study), and denote the Dice loss and cross entropy loss of the i-th layer respectively.
[0083] Basic segmentation loss Combining Dice loss and cross entropy loss ensures that both liver and landmark branches can be learned effectively.
[0084] Loss of liver-guided inside-out consistency That is, the above calculations are used to jointly optimize the liver and landmark segmentation results to ensure the logical consistency of the segmentation of the inner and outer regions of the liver.
[0085] Auxiliary deep supervision mechanism An auxiliary segmentation head is added to the middle layers of the landmark decoder (stages 2, 3, and 4). Auxiliary predictions are generated through a 1×1 convolutional layer and softmax activation. After resizing them to the same size as the labels, the loss is calculated using the same structure as the basic segmentation loss. The loss is averaged and incorporated into the overall loss, enhancing the stability of landmark segmentation in difficult cases. This complete loss function allows for comprehensive model optimization and improved segmentation performance.
[0086] Step S4: model evaluation.
[0087] The Dice coefficient and Chamfer distance are used as the main evaluation metrics. The Dice coefficient measures the overlap between the predicted results and the true annotations, which directly reflects the segmentation accuracy. The Chamfer distance measures the pixel-level spatial deviation between the predicted landmark segmentation area and the true annotation.
[0088] Where v and w represent the pixel coordinates of the predicted segmented region and the ground-truth labeled region, respectively, and |v| and |w| represent the number of pixels in the predicted region and the ground-truth region, respectively. In order to compare different methods, this study calculated the average 2D chamfer distance (CD) of all test images as:
[0089]
[0090] Where N is the number of images in the test set. In addition, the frames per second (FPS) are calculated to analyze the feasibility of the model application.
[0091] Step S5: Final prediction and output.
[0092] In practical applications, after obtaining the laparoscopic liver key frame image, it is input into the trained SAM-D 2 Net model. The model uses the SAM encoder to extract high-level features, uses a dual decoder to segment the liver and landmarks separately, and optimizes the segmentation results by combining the consistency constraints inside and outside the liver. This allows for the rapid and accurate detection of anatomical landmarks such as the falciform ligament, liver ridge, and liver surface contour.
[0093] In summary, this method has the following characteristics:
[0094] (1) Based on the dual decoder architecture of SAM, the pre-trained Segment Anything Model (SAM) is introduced as the encoder, and its powerful visual representation ability obtained in large-scale visual data training is used to provide robust deep features for liver and landmark segmentation. At the same time, the model adopts a dual decoder structure, and the liver decoder and landmark decoder are independently modeled to reduce interference between tasks. At the same time, the overall structural information of the liver is used to optimize landmark segmentation, thereby improving the adaptability and segmentation accuracy of the model in complex environments. SAM-enhanced dual decoder network (SAM-D 2 Net) architecture includes a unique dual decoder design and a combination with the SAM encoder, which not only absorbs the potential priors in the SAM encoder, but also the dual decoder strategy can alleviate the feature interference between different tasks.
[0095] (2) Liver internal and external consistency constraint: Considering that the distribution of anatomical landmarks is closely related to liver morphology, a liver internal and external consistency constraint is proposed. By dividing the liver prediction results into internal and external regions, the anatomical landmark segmentation results of these two regions are calculated separately, and the degree of overlap is evaluated using the Dice coefficient. The liver internal and external consistency loss is defined to strengthen the spatial alignment relationship between the liver and the landmarks, making the landmark segmentation results more consistent with the actual anatomical distribution and optimizing the overall segmentation effect.
[0096] (3) Complete loss function design: A complete loss function is designed that comprehensively considers the basic segmentation loss, the liver internal and external consistency loss, and the deep supervision mechanism. The basic segmentation loss is combined with the Dice loss and the cross-entropy loss to balance the class imbalance problem and optimize the liver and landmark segmentation; the liver internal and external consistency loss enhances spatial alignment; the deep supervision mechanism is introduced in multiple intermediate layers of the landmark decoder to enhance the model's multi-scale learning ability and improve training stability, thereby improving the accuracy, stability, and generalization ability of 2D anatomical landmark segmentation.
[0097] (4) Multi-dataset validation and generalization testing: Experimental evaluation was conducted using two publicly available datasets for laparoscopic liver resection 2D anatomical landmark detection, L3D and P2ILF. Training and testing on different datasets validated the model's performance and generalization capabilities, ensuring that the model maintains stable detection performance under different object and device conditions. Comparative experiments were also conducted on challenging video data to further verify the model's robustness in real-world complex scenarios.
[0098] Example 2
[0099] Based on Example 1, this embodiment provides an electronic device, including: one or more processors and a memory, wherein the memory stores one or more programs, and the one or more programs include instructions for executing the laparoscopic image liver landmark segmentation method based on the SAM enhanced dual decoder network as described in Example 1.
[0100] like Figure 2 As mentioned above, at the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory and a non-volatile memory, and may also include other hardware required for the business. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to achieve the above Figure 1 Of course, in addition to software implementation, the present invention does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc., that is, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.
[0101] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0102] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.
[0103] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and such modifications or substitutions are intended to be within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the scope of protection of the claims.
Claims
1. A method for liver landmark segmentation in laparoscopic images based on a SAM-enhanced dual-decoder network, characterized in that: The steps include: Acquire the 2D laparoscopic key frame image to be segmented; Extracting features using a SAM encoder based on the 2D laparoscopic key frame image; Based on the features extracted by the SAM encoder, a liver decoder is used to extract liver contour information, obtain a liver mask, and implement liver segmentation; Based on the features extracted by the SAM encoder, the landmark decoder is used to segment the falciform ligament, the liver ridge, and the liver contour to achieve segmentation of anatomical landmarks; The SAM encoder, liver decoder and landmark decoder are pre-trained based on the liver internal and external consistency loss and the deep supervision loss.
2. The method for liver landmark segmentation in laparoscopic images based on SAM enhanced dual decoder network according to claim 1, characterized in that: The loss of consistency between the liver inside and outside is: in, Loss of consistency between the liver and the outside, Dice in 、Dice out are the Dice losses of the liver inner and outer regions, respectively. liver is the predicted liver probability map output by the liver decoder, Represent the predicted probability maps of being inside the liver and outside the liver, respectively. Represent the true labels inside and outside the liver, respectively. land. is the predicted landmark probability map, represents the probability map of landmarks located inside the liver, represents the probability map of landmarks located outside the liver, Y liver is the true label in the sample, denote the true internal and external landmark annotations respectively, Split() denotes the division operation, and ∑ denotes the summation operation over all pixels.
3. The method for liver landmark segmentation in laparoscopic images based on SAM enhanced dual decoder network according to claim 1, characterized in that: The deep supervision loss is: in, is the depth supervision loss, N is the number of depth supervision layers, and denote the Dice loss and cross entropy loss of the i-th layer of the landmark decoder, respectively.
4. The method for liver landmark segmentation in laparoscopic images based on SAM enhanced dual decoder network according to claim 1, characterized in that: The training process also includes basic segmentation loss: in, is the basic segmentation loss, are Dice loss and cross entropy loss respectively, O is the probability map output by the liver decoder or landmark decoder, and Y is the corresponding true label.
5. The method for liver landmark segmentation in laparoscopic images based on SAM enhanced dual decoder network according to claim 1, characterized in that: The following steps are also included: The SAM encoder, liver decoder, and landmark decoder were evaluated by calculating the Dice coefficient, Chamfer distance, and the average 2D chamfer distance of the test images, where the average 2D chamfer distance of the test images was calculated using the following formula: In the formula, CD (v,w) is the average 2D chamfer distance of all test images, CD i (v,w) is the average 2D chamfer distance of the test image i, N is the number of images in the test set, v and w are the pixel coordinate sets of the predicted segmented area and the true annotated area, respectively, |v| and |w| are the number of pixels in the predicted area and the true area, respectively.
6. The method for liver landmark segmentation in laparoscopic images based on SAM enhanced dual decoder network according to claim 1, characterized in that: For any one of the liver decoder and the landmark decoder, multiple layers of residual blocks and bilinear upsampling layers are included.
7. The method for liver landmark segmentation in laparoscopic images based on SAM enhanced dual decoder network according to claim 1, characterized in that: The SAM encoder is a pre-trained SAM-B encoder. During the training process, the SAM-B encoder is end-to-end fine-tuned.
8. The method for liver landmark segmentation in laparoscopic images based on SAM enhanced dual decoder network according to claim 1, characterized in that: The training process includes the following steps: Acquire an image dataset, normalize the images and resize them to a preset resolution, and augment the data by randomly inverting, rotating, and cropping them to increase the number of samples.
9. The method for liver landmark segmentation in laparoscopic images based on SAM enhanced dual decoder network according to claim 1, characterized in that: During the training process, the training is achieved by stochastic gradient descent with the goal of minimizing the loss function.
10. An electronic device, characterized in that: The method comprises one or more processors, a memory and one or more programs stored in the memory, wherein the one or more programs include instructions for executing the laparoscopic image liver landmark segmentation method based on the SAM enhanced dual decoder network as described in any one of claims 1 to 9.