Endoscopic Treatment Tool Recognition With Synthetic Training Images
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The challenge in the medical field is the difficulty in collecting a large number of diverse image data for endoscopic image recognition due to ethical and practical constraints, particularly for recognizing various treatment tools used with endoscopes, which limits the accuracy of learning models.
Innovation Solution
An endoscopic image learning device and method that generates superimposed images by combining foreground images of treatment tools with background endoscopic images, using machine learning to create a learning model for accurate recognition, including data augmentation techniques like affine transformation and noise application.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a large number of diverse endoscopic images with treatment tools are collected for learning, then the accuracy of the learning model for recognizing treatment tools is improved, but it becomes difficult to collect such images due to medical practice constraints
Solution Approach 1:
The patent creates synthetic learning data by copying and combining foreground images of treatment tools with background endoscopic images. This allows generation of numerous diverse training images without actual medical practice, resolving the contradiction between needing large data quantities and the difficulty of collecting real clinical images.
Solution Approach 2:
The patent performs preliminary extraction of treatment tool images from source materials before combining them with background images. This preliminary preparation enables efficient generation of diverse training data without requiring actual endoscopic procedures, addressing both data quantity needs and collection constraints.
2Adaptability or versatility
If images of various treatment tools are collected for learning, then the learning model can recognize different treatment tools accurately, but the collection process requires intervention in medical practice which limits data availability
Solution Approach 1:
The patent segments treatment tool images from source materials as foreground elements, separating them from background endoscopic images. This segmentation enables independent collection and reuse of treatment tool images across multiple backgrounds, improving versatility without requiring repeated medical interventions.
Solution Approach 2:
The patent creates a universal learning dataset by combining extracted treatment tool foregrounds with various background endoscopic images. This universal approach enables the learning model to recognize multiple treatment tools across different contexts without requiring separate data collection for each tool or scenario.
3Reliability
If real endoscopic images with treatment tools are used for learning, then the learning model learns accurate real-world scenarios, but the amount of available learning data is limited due to collection constraints
Solution Approach 1:
The patent creates synthetic copies by combining authentic treatment tool foregrounds with real background endoscopic images. This copying approach maintains the authenticity of both foreground and background elements while generating unlimited combinations, resolving the contradiction between data authenticity and data volume.
Data Source
AI summary
An object is to provide an endoscopic image learning device, an endoscopic image learning method, an endoscopic image learning program, and an endoscopic image recognition device that appropriately learn a learning model for image recognition for recognizing an endoscopic image in which a treatment tool for an endoscope appears.The object is achieved by an endoscopic image learning device including an image generation unit and a machine learning unit. The image generation unit generates a superimposed image where a foreground image in which a treatment tool for an endoscope is extracted is superimposed on a background-endoscopic image serving as a background of the foreground image, and the machine learning unit performs the learning of a learning model for image recognition using the superimposed image.


