Medical Object Keypoint Localization With Synthetic 3D Training Images

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for training artificial intelligence to localize medical objects in images are tedious, time-consuming, and prone to errors, leading to limited generalization and performance issues due to manual data generation and labeling.

Innovation Solution

A method involving the use of synthetic training images generated from 3D model data, combined with a deep convolutional neural network, to accurately identify keypoints of medical objects, allowing for flexible and moving parts, and incorporating deformation parameters for precise positioning.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual placement and labeling of training images is used, then ground truth data can be obtained, but the process is time-consuming and error-prone

Engineering Contradiction:
Improveaccuracy of ground truth labelingVSAvoidtime for manual data generation and labeling
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent uses synthetic image generation to create copies of training data through computer-generated simulations rather than manual photography. The system renders virtual images of medical objects with automatically generated ground truth labels, eliminating the need for manual placement and labeling while maintaining high accuracy through precise virtual model positioning.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The synthetic data generation system automatically generates both the training images and their corresponding ground truth labels without human intervention. The rendering pipeline self-computes object positions, orientations, and keypoint locations from virtual models, making the entire data preparation process self-service and eliminating manual labor.

Inventive Principle:
Principle #25Self-service

2Adaptability or versatility

If limited manual training data is generated, then the studio approach can be completed, but the network cannot generalize to images out of training distribution

Engineering Contradiction:
Improvegeneralization capability of trained networkVSAvoidamount of training data
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The patent systematically varies parameters in the synthetic image generation process, including object positions, orientations, lighting conditions, camera angles, and background configurations. This creates diverse training images that cover a wide range of scenarios, enabling the network to generalize to unseen real-world images while maintaining a manageable data volume through efficient parameter sampling.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If manual ground truth labeling is performed, then reference values can be identified, but errors occur especially for visually hidden keypoints

Engineering Contradiction:
Improveaccuracy of keypoint identificationVSAvoiderror rate in labeling
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent creates virtual copies of medical objects in controlled digital environments where ground truth information is inherently known from the virtual models. This eliminates the ambiguity and errors associated with manually identifying hidden keypoints in real images, as the synthetic rendering process automatically tracks and records precise object states throughout the simulation.

Inventive Principle:
Principle #26Copying

Data Source

PatentEP4617921A1Method for evaluating images and method for training an artificial intelligence
Publication Date: 2025.09.17 KONINKLIJKE PHILIPS NV
  • EP4617921A1 patent drawingFigure 1
  • EP4617921A1 patent drawingFigure 2
  • EP4617921A1 patent drawingFigure 3

AI summary

The invention relates to a method (1) for evaluating images in a medical environment, wherein at least one image of a medical setting is acquired (S1), fed (S2) as input into a trained artificial intelligence and processed (S3). A plurality of keypoints (5) of medical objects (3) is received (S4) as an output from the trained artificial intelligence, and positioning information of the medical objects (3) in the medical setting is determined (S5) based on the received plurality of keypoints (5). The invention further relates to a method (2) for training an artificial intelligence, wherein 3D model data for a medical object (3) is acquired (S7), a plurality of keypoints (5) for the medical object (3) is determined (S8), a plurality of synthetic training images is generated (S9) based on the 3D model data, and the artificial intelligence is trained (S10) with the plurality of synthetic training images to recognize the plurality of keypoints (5).