Surgical Tool Segmentation Using Synthetic Image Training

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for generating pixel-accurate segmentations of surgical tools in laparoscopic surgeries are costly and require extensive human annotation, while conventional models lack efficiency and accuracy.

Innovation Solution

A segmentation model is trained using a combination of real images annotated with bounding boxes and synthetic images generated from 3D models, leveraging machine learning models like DeepMAC and CycleGAN to enhance accuracy and reduce costs.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional machine learning models are used for surgical tool segmentation, then the system is simpler to implement, but the segmentation accuracy and pixel-level precision are insufficient

Engineering Contradiction:
Improvesegmentation accuracyVSAvoidmodel complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies segmentation by dividing the training process into two distinct stages: first training a bounding box detection model, then training a segmentation model that uses bounding box predictions as input. This multi-stage segmentation approach enables pixel-level precision while managing complexity through modular model design.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transitions from 2D bounding box annotations to pixel-level 2D segmentation masks, effectively adding a dimension of precision. By incorporating synthetic 3D-rendered images with ground truth segmentations, the system elevates the annotation quality from coarse bounding boxes to fine-grained pixel boundaries.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Measurement precision

If real surgical images with manual annotations are used for training, then the training data is more accurate, but the annotation cost and time consumption increase significantly

Engineering Contradiction:
Improveannotation accuracyVSAvoidannotation time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary action by pre-rendering synthetic surgical images with known ground truth segmentations before actual surgical procedures. This advance preparation creates a large corpus of training data with perfect annotations, eliminating the need for time-consuming manual annotation of real surgical images while maintaining high annotation accuracy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent creates synthetic copies of surgical scenes through 3D rendering, replicating realistic surgical environments, lighting conditions, and tool appearances without requiring actual surgical procedures. These synthetic copies serve as training data with automatically generated ground truth segmentations, replacing the need for manual annotation of real images.

Inventive Principle:
Principle #26Copying

3Manufacturing precision

If extensive manual annotation is performed to achieve pixel-accurate segmentations, then the segmentation quality improves, but the computational resources and costs increase

Engineering Contradiction:
Improvesegmentation qualityVSAvoidcomputational resources
Core Design Contradiction:
Manufacturing precisionVSQuantity of substance

Solution Approach 1:

The patent implements self-service by using synthetic 3D rendering to automatically generate ground truth segmentations without human intervention. The rendering engine inherently knows the precise boundaries of surgical tools in synthetic images, providing perfect annotations at minimal computational cost compared to manual pixel-level labeling of real images.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent changes the parameter of data source from real captured images to synthetic rendered images. This parameter change fundamentally alters the cost structure, as synthetic images provide unlimited training data with automatic ground truth segmentations, eliminating the resource-intensive manual annotation process while maintaining high segmentation quality.

Inventive Principle:
Principle #35Parameter changes

4Speed

If the model is trained to achieve real-time segmentation, then the processing speed increases, but the model complexity and training requirements increase

Engineering Contradiction:
Improveprocessing speedVSAvoidmodel complexity
Core Design Contradiction:
SpeedVSDevice complexity

Solution Approach 1:

The patent performs preliminary training using synthetic data that pre-teaches the model fundamental surgical tool appearances and segmentation patterns. This pre-training accelerates subsequent fine-tuning on real images and enables faster inference, as the model already possesses prior knowledge of surgical contexts before encountering actual surgical images.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent incorporates 3D-rendered synthetic images as an additional training dimension, teaching the model robust feature representations that generalize across varying lighting, angles, and surgical scenarios. This multi-dimensional training approach improves inference speed by reducing the model's need to process uncertain features during real-time operation.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS12541854B2Segmentation generation for image-related functions
Publication Date: 2026.02.03 VERILY LIFE SCIENCES LLC
  • US12541854B2 patent drawing
  • US12541854B2 patent drawing
  • US12541854B2 patent drawing

AI summary

A computer system may perform an image-related function using a segmentation (e.g., an image mask) that has been generated by a custom segmentation machine learning model. To begin, the system may receive image data corresponding to a surgical scene including a background that includes an anatomical feature and at least one surgical tool. The system may also generate segmentation data using the custom segmentation machine learning model based on inputting the first image data to the custom segmentation machine learning model. The system may also include generating a segmentation of the at least one surgical tool using the segmentation data. Once the segmentation has been generated, the system may perform an image-related function using the segmentation.