Voice Transcription Search Accuracy via Temporal Filtering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional voice transcription techniques suffer from low accuracy due to all phrases from voice recognition processes being searched as targets, resulting in numerous candidate options.

Innovation Solution

An information processing apparatus that associates character strings with voice positional information, allowing for targeted retrieval of specific phrases by detecting played-back sections and filtering search targets based on temporal information, thereby improving search accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If all phrases from voice recognition process results are used as search targets, then comprehensive search coverage is achieved, but search accuracy deteriorates due to numerous candidate options

Engineering Contradiction:
Improvesearch accuracyVSAvoidnumber of search targets
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent segments the search process by dividing all phrases from voice recognition results into multiple groups based on their association with different played-back sections. Instead of treating all phrases as a single search pool, the system creates distinct groups where each group contains only phrases relevant to a specific time section, thereby reducing the number of candidates presented to the user and improving search accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies local quality by making different parts of the search results have different properties. Phrases are organized into groups with distinct characteristics based on their temporal association. Each group is tailored to a specific played-back section, providing locally optimized search results rather than a uniform list of all possible phrases, thus enhancing relevance and accuracy.

Inventive Principle:
Principle #3Local quality

2Productivity

If comprehensive phrase retrieval is performed without temporal filtering, then all possible candidates are found, but user efficiency deteriorates due to excessive options

Engineering Contradiction:
Improvetranscription efficiencyVSAvoidnumber of candidate phrases
Core Design Contradiction:
ProductivityVSQuantity of substance

Solution Approach 1:

The patent extracts only the necessary subset of phrases from the complete voice recognition results by filtering out phrases that are not associated with the currently played-back section. This extraction process removes irrelevant candidates from the search results, presenting users with a focused list of only those phrases that are temporally relevant, thereby improving transcription efficiency without sacrificing comprehensive coverage of relevant options.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent performs preliminary filtering of phrases based on their association with played-back sections before presenting results to the user. By pre-organizing phrases into groups according to temporal relevance, the system eliminates the need for users to sift through irrelevant options, thus improving productivity by presenting only the most relevant candidates in advance.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS9798804B2Information processing apparatus, information processing method and computer program product
Publication Date: 2017.10.24 TOSHIBA DIGITAL SOLUTIONS CORP
  • US9798804B2 patent drawing
  • US9798804B2 patent drawing
  • US9798804B2 patent drawing

AI summary

According to an embodiment, an information processing apparatus includes a storage unit, a detector, an acquisition unit, and a search unit. The storage unit configured to store therein voice indices, each of which associates a character string included in voice text data obtained from a voice recognition process with voice positional information, the voice positional information indicating a temporal position in the voice data and corresponding to the character string. The acquisition unit acquires reading information being at least a part of a character string representing a reading of a phrase to be transcribed from the voice data played back. The search unit specifies, as search targets, character strings whose associated voice positional information is included in the played-back section information among the character strings included in the voice indices, and retrieves a character string including the reading represented by the reading information from among the specified character strings.