Endoscope Audio Recognition via Scene-Specific Dictionaries
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing endoscope systems face challenges in accurately recognizing audio inputs during medical examinations due to high error rates in word recognition across different scenes, which reduces operability.
Innovation Solution
An endoscope system equipped with an audio input device, image sensor, and processor that sets a tailored audio recognition dictionary based on the scene, allowing for improved audio recognition accuracy by only recognizing registered words and using image recognition to trigger and refine the dictionary settings.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If a general audio recognition dictionary is used for all scenes, then the system can recognize a wide range of words, but the recognition accuracy decreases due to mutual erroneous recognition between words
Solution Approach 1:
The patent segments the audio recognition dictionary into multiple scene-specific dictionaries (e.g., examination room scene dictionary, treatment room scene dictionary). Each dictionary contains words and phrases relevant to its specific scene, thereby maintaining high recognition accuracy within each scene while collectively covering a wide range of medical scenarios.
Solution Approach 2:
The system dynamically switches between different audio recognition dictionaries based on the detected scene. The processor automatically selects the appropriate dictionary corresponding to the current scene, making the recognition system adaptive and context-aware, thus resolving the contradiction between coverage and accuracy.
2Measurement precision
If a scene-specific audio recognition dictionary is used, then the audio recognition accuracy is improved, but the system complexity increases due to multiple dictionaries
Solution Approach 1:
The system employs image recognition technology to automatically identify the current scene and select the corresponding audio recognition dictionary without manual intervention. This self-service mechanism eliminates the need for complex manual dictionary management, thereby reducing system complexity while maintaining high recognition accuracy.
Solution Approach 2:
The system uses feedback from image recognition results to automatically switch between different audio recognition dictionaries. The image recognition output serves as feedback that triggers the selection of the appropriate dictionary, creating a closed-loop system that simplifies dictionary management while ensuring accurate recognition.
3Reliability
If audio recognition is performed continuously in all scenes, then no words are missed, but the operability is reduced due to increased error rates
Solution Approach 1:
The system dynamically adjusts the audio recognition process by switching between different scene-specific dictionaries based on the current scene. This dynamic adaptation ensures that the system maintains high recognition completeness while avoiding erroneous recognitions that would reduce operability, as each dictionary is optimized for its specific scene.
Solution Approach 2:
Each audio recognition dictionary is tailored with local quality optimized for its specific scene, containing only the words and phrases relevant to that scene. This local optimization ensures high recognition completeness for scene-specific terms while minimizing erroneous recognitions, thereby maintaining excellent operability.
Data Source
AI summary
An embodiment according to the technique of the present disclosure is to provide an endoscope system, a medical information processing apparatus, a medical information processing method, a medical information processing program, and a recording medium capable of improving recognition accuracy of an audio input. The endoscope system according to an aspect of the present invention includes an audio input device; an image sensor that images a subject; and a processor, in which the processor acquires a plurality of medical images by causing the image sensor to image the subject in chronological order, accepts an input of an audio input trigger during capturing of the plurality of medical images, sets, in a case where the audio input trigger is input, an audio recognition dictionary according to the audio input trigger, and performs audio recognition on audio input to the audio input device after the setting, using the set audio recognition dictionary.


