Endoscope Audio Recognition via Scene-Specific Dictionaries

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing endoscope systems face challenges in accurately recognizing audio inputs during medical examinations due to high error rates in word recognition across different scenes, which reduces operability.

Innovation Solution

An endoscope system equipped with an audio input device, image sensor, and processor that sets a tailored audio recognition dictionary based on the scene, allowing for improved audio recognition accuracy by only recognizing registered words and using image recognition to trigger and refine the dictionary settings.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If a general audio recognition dictionary is used for all scenes, then the system can recognize a wide range of words, but the recognition accuracy decreases due to mutual erroneous recognition between words

Engineering Contradiction:
Improveword recognition coverageVSAvoidaudio recognition accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent segments the audio recognition dictionary into multiple scene-specific dictionaries (e.g., examination room scene dictionary, treatment room scene dictionary). Each dictionary contains words and phrases relevant to its specific scene, thereby maintaining high recognition accuracy within each scene while collectively covering a wide range of medical scenarios.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system dynamically switches between different audio recognition dictionaries based on the detected scene. The processor automatically selects the appropriate dictionary corresponding to the current scene, making the recognition system adaptive and context-aware, thus resolving the contradiction between coverage and accuracy.

Inventive Principle:
Principle #15Dynamics

2Measurement precision

If a scene-specific audio recognition dictionary is used, then the audio recognition accuracy is improved, but the system complexity increases due to multiple dictionaries

Engineering Contradiction:
Improveaudio recognition accuracyVSAvoiddictionary management complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system employs image recognition technology to automatically identify the current scene and select the corresponding audio recognition dictionary without manual intervention. This self-service mechanism eliminates the need for complex manual dictionary management, thereby reducing system complexity while maintaining high recognition accuracy.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system uses feedback from image recognition results to automatically switch between different audio recognition dictionaries. The image recognition output serves as feedback that triggers the selection of the appropriate dictionary, creating a closed-loop system that simplifies dictionary management while ensuring accurate recognition.

Inventive Principle:
Principle #23Feedback

3Reliability

If audio recognition is performed continuously in all scenes, then no words are missed, but the operability is reduced due to increased error rates

Engineering Contradiction:
Improveword recognition completenessVSAvoidsystem operability
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The system dynamically adjusts the audio recognition process by switching between different scene-specific dictionaries based on the current scene. This dynamic adaptation ensures that the system maintains high recognition completeness while avoiding erroneous recognitions that would reduce operability, as each dictionary is optimized for its specific scene.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

Each audio recognition dictionary is tailored with local quality optimized for its specific scene, containing only the words and phrases relevant to that scene. This local optimization ensures high recognition completeness for scene-specific terms while minimizing erroneous recognitions, thereby maintaining excellent operability.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS20240188799A1Endoscope system, medical information processing apparatus, medical information processing method, medical information processing program, and recording medium
Publication Date: 2024.06.13 FUJIFILM CORP
  • US20240188799A1 patent drawing
  • US20240188799A1 patent drawing
  • US20240188799A1 patent drawing

AI summary

An embodiment according to the technique of the present disclosure is to provide an endoscope system, a medical information processing apparatus, a medical information processing method, a medical information processing program, and a recording medium capable of improving recognition accuracy of an audio input. The endoscope system according to an aspect of the present invention includes an audio input device; an image sensor that images a subject; and a processor, in which the processor acquires a plurality of medical images by causing the image sensor to image the subject in chronological order, accepts an input of an audio input trigger during capturing of the plurality of medical images, sets, in a case where the audio input trigger is input, an audio recognition dictionary according to the audio input trigger, and performs audio recognition on audio input to the audio input device after the setting, using the set audio recognition dictionary.