Voice Transcription Search Accuracy via Temporal Filtering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional voice transcription techniques suffer from low accuracy due to all phrases from voice recognition processes being searched as targets, resulting in numerous candidate options.
Innovation Solution
An information processing apparatus that associates character strings with voice positional information, allowing for targeted retrieval of specific phrases by detecting played-back sections and filtering search targets based on temporal information, thereby improving search accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If all phrases from voice recognition process results are used as search targets, then comprehensive search coverage is achieved, but search accuracy deteriorates due to numerous candidate options
Solution Approach 1:
The patent segments the search process by dividing all phrases from voice recognition results into multiple groups based on their association with different played-back sections. Instead of treating all phrases as a single search pool, the system creates distinct groups where each group contains only phrases relevant to a specific time section, thereby reducing the number of candidates presented to the user and improving search accuracy.
Solution Approach 2:
The patent applies local quality by making different parts of the search results have different properties. Phrases are organized into groups with distinct characteristics based on their temporal association. Each group is tailored to a specific played-back section, providing locally optimized search results rather than a uniform list of all possible phrases, thus enhancing relevance and accuracy.
2Productivity
If comprehensive phrase retrieval is performed without temporal filtering, then all possible candidates are found, but user efficiency deteriorates due to excessive options
Solution Approach 1:
The patent extracts only the necessary subset of phrases from the complete voice recognition results by filtering out phrases that are not associated with the currently played-back section. This extraction process removes irrelevant candidates from the search results, presenting users with a focused list of only those phrases that are temporally relevant, thereby improving transcription efficiency without sacrificing comprehensive coverage of relevant options.
Solution Approach 2:
The patent performs preliminary filtering of phrases based on their association with played-back sections before presenting results to the user. By pre-organizing phrases into groups according to temporal relevance, the system eliminates the need for users to sift through irrelevant options, thus improving productivity by presenting only the most relevant candidates in advance.
Data Source
AI summary
According to an embodiment, an information processing apparatus includes a storage unit, a detector, an acquisition unit, and a search unit. The storage unit configured to store therein voice indices, each of which associates a character string included in voice text data obtained from a voice recognition process with voice positional information, the voice positional information indicating a temporal position in the voice data and corresponding to the character string. The acquisition unit acquires reading information being at least a part of a character string representing a reading of a phrase to be transcribed from the voice data played back. The search unit specifies, as search targets, character strings whose associated voice positional information is included in the played-back section information among the character strings included in the voice indices, and retrieves a character string including the reading represented by the reading information from among the specified character strings.


