Automatic Keyword Extraction for Voice Collation Verification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech recognition systems require manual preparation and registration of keywords for voice collation/verification, which can be labor-intensive and prone to errors, and lack robustness against malicious voice sharing or synthesis.
Innovation Solution
An information processing system that extracts keywords from conversation data and generates feature quantities related to voices, allowing for automatic keyword generation and association, reducing the need for manual preparation and enhancing security by dynamically generating keywords from conversation data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If manual preparation and registration of keywords is used for voice collation/verification, then system simplicity is maintained, but labor intensity increases and error-proneness worsens
Solution Approach 1:
The system automatically extracts keywords from conversation data and generates feature quantities without human intervention. The keyword extraction unit identifies important words from speech information, and the feature quantity extraction unit automatically creates voice特征 data, enabling the system to serve itself rather than requiring manual keyword preparation and registration
Solution Approach 2:
The system performs preliminary extraction of keywords and feature quantities from conversation data before verification is needed. By pre-processing the conversation data to identify keywords and generate their associated voice features, the system prepares verification data in advance, eliminating the need for manual keyword registration at the time of use
2Ease of operation
If pre-defined keywords are used for voice collation/verification, then verification process is simple, but security against malicious voice sharing or synthesis deteriorates
Solution Approach 1:
The system changes the verification parameter from static pre-defined keywords to dynamic keywords extracted from actual conversation data. By using keywords that are specifically identified from the target conversation and associating them with unique voice feature quantities, the system creates verification data that is specific to each conversation context, making it difficult to use with malicious voice sharing or synthesis while maintaining verification simplicity
3Device complexity
If manual keyword preparation is required, then system complexity is low, but productivity decreases due to labor-intensive processes
Solution Approach 1:
The system replaces the mechanical process of manual keyword preparation with automated information processing. The keyword extraction unit uses speech recognition and natural language processing to automatically identify keywords from conversation data, and the feature quantity extraction unit automatically generates voice特征 data, substituting human manual work with computational processes that significantly improve productivity while maintaining manageable system complexity
Data Source
AI summary
An information processing system includes: an acquisition unit that obtains conversation data including speech information on a plurality of people; a keyword extraction unit that extracts a keyword from the speech information; a feature quantity extraction unit that extracts a first feature quantity that is a feature quantity related to a voice when the keyword is said, from the speech information; and a generation unit that generates information for collation/verification, by associating the keyword with the first feature quantity. According to such an information processing system, it is possible to properly generate the information for collation/verification, from the conversation data.


