Voice-Triggered Keyframe Extraction in Radiographic Imaging
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In videofluoroscopic examinations of swallowing, observers face a heavy burden in searching for specific radiation images within a series of captured images, as the period of interest is often intertwined with other imaging periods, making it difficult to access desired images efficiently.
Innovation Solution
An information processing apparatus that utilizes voice recognition to generate image specification information based on specific voices or biological sounds during the imaging period, allowing for the association and easy retrieval of radiation images corresponding to these voices, thereby facilitating the extraction of keyframes or specific image frames within the series.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If a series of radiation images is captured continuously during the entire imaging period, then complete examination data is recorded, but it becomes difficult to access desired images efficiently and places a heavy burden on the observer
Solution Approach 1:
The patent extracts keyframe images from the continuous series of radiation images by detecting specific voices or biological sounds. Only images corresponding to these detected events are selected and stored as keyframes, separating the essential diagnostic information from the large volume of non-essential continuous imaging data.
Solution Approach 2:
The system performs preliminary detection of specific voices or biological sounds during the imaging process to identify which images should be retained. This preliminary action allows the system to pre-select keyframes before final storage, making subsequent access efficient without requiring manual search through all images.
2Reliability
If all radiation images are stored for later review, then no diagnostic information is lost, but storage requirements increase and retrieval time increases
Solution Approach 1:
The system extracts only the diagnostically relevant images by detecting specific voices or biological sounds that mark important events. This extraction process preserves all necessary diagnostic information while eliminating redundant images, thereby reducing storage requirements and retrieval time.
Solution Approach 2:
The patent introduces voice detection and biological sound detection as intermediary mechanisms between the continuous imaging process and the final image storage. These intermediaries act as filters that identify which images should be preserved, enabling efficient storage and retrieval without manual intervention.
3Measurement precision
If manual searching through continuous radiation images is performed to find the period of interest, then accurate localization is achieved, but observer burden and examination time increase
Solution Approach 1:
The system uses voice detection and biological sound detection as intermediary mechanisms to automatically identify and localize the period of interest. These intermediaries provide accurate temporal localization of diagnostic events without requiring manual searching, thereby maintaining measurement precision while significantly improving examination efficiency.
Solution Approach 2:
The system provides feedback by detecting specific voices or biological sounds and automatically marking the corresponding images as keyframes. This feedback mechanism enables the system to self-identify important diagnostic periods, eliminating the need for manual searching and improving productivity while maintaining accurate localization.
Data Source
AI summary
The information processing apparatus includes at least one processor. The processor acquires a series of radiation images captured by performing continuous irradiation with radiation, generates, in a case in which a voice uttered during an imaging period of the series of radiation images is a specific voice, image specification information for specifying a radiation image selected from among the series of radiation images according to a timing of occurrence of the voice, and associates the image specification information with the series of radiation images.


