Surgical Visualization Control Using Timed Voice and Eye Inputs

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Surgical visualization systems face inaccuracies due to noise in eye-tracking data and delays in speech recognition, leading to unreliable control processes.

Innovation Solution

A method that processes multi-modal user utterances, including linguistic commands and eye movements, to control surgical visualization systems, utilizing a combination of sensors to capture and process user inputs with reduced latency and improved accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If eye tracking and voice commands are used for control, then the surgical visualization system can be operated hands-free, but noise in eye-tracking data and delays in speech recognition lead to inaccurate control processes

Engineering Contradiction:
Improvehands-free operationVSAvoidcontrol accuracy
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent combines multiple control modalities (eye tracking, voice commands, and other input methods) into a unified control system. By merging these different input sources, the system can cross-validate signals and filter out noise from individual modalities, thereby maintaining hands-free operation while improving control accuracy and reliability.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system implements feedback mechanisms that continuously monitor the quality and reliability of control inputs from eye tracking and voice recognition. When noise or delays are detected, the system can adjust its response, request clarification, or switch to alternative control methods, ensuring reliable operation without compromising hands-free capability.

Inventive Principle:
Principle #23Feedback

2Adaptability or versatility

If speech recognition is used to process linguistic commands, then the system can interpret complex user intentions, but delays in speech recognition reduce real-time control capability

Engineering Contradiction:
Improvecommand interpretation capabilityVSAvoidcontrol delay
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The system performs preliminary processing of speech inputs by pre-processing acoustic signals and preparing recognition results before full command interpretation is required. This preliminary action reduces the effective delay by having parts of the speech recognition pipeline ready in advance, allowing faster response to time-critical control commands.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The speech recognition process is segmented into multiple parallel processing streams: a fast path for simple commands that provides quick response, and a slower path for complex commands that requires full interpretation. This segmentation allows the system to handle time-sensitive controls rapidly while still providing comprehensive command interpretation when needed.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20260000478A1Controlling surgical visualization systems using multi-modal user utterances
Publication Date: 2026.01.01 CARL ZEISS MEDITEC AG
  • US20260000478A1 patent drawing
  • US20260000478A1 patent drawing
  • US20260000478A1 patent drawing

AI summary

Provision is made for a computer-implemented method for controlling a surgical visualization system. A first user utterance of a first user utterance type is received, wherein this first user utterance extends over a time interval within a defined period of time. A multiplicity of second user utterances of at least one different second user utterance type are received, these being captured in a manner distributed within said period of time and varying within the period of time. The surgical visualization system is controlled using the first user utterance and at least one second user utterance that is prioritized based on temporal relationships of the second user utterances in relation to the time interval.