Key Phrase Identification Model for Audio Search
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current audio search systems cannot comprehensively and efficiently retrieve audio data based on user queries since they rely solely on title and text descriptions, failing to accurately identify key phrases within the audio content.
Innovation Solution
A method and apparatus for training a key phrase identification model using both natural language processing and manually labeled training samples to convert audio data into text, enabling the recognition of key phrases within audio data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If the system only uses title and text description for audio search, then the search system is simple to operate, but the search accuracy and comprehensiveness deteriorate
Solution Approach 1:
The system performs preliminary action by automatically generating text descriptions and extracting key phrases from audio content during the indexing phase, before the actual search occurs. This preprocessing enables the search system to query not only titles but also generated descriptions and key phrases, significantly improving search accuracy without complicating the user interface.
Solution Approach 2:
The patent introduces text descriptions and key phrases as intermediary elements between the audio content and the search query. These intermediaries bridge the gap between user search intent and audio content, allowing accurate search without requiring users to interact with complex audio analysis tools.
2Measurement precision
If the system performs comprehensive audio content analysis to improve search accuracy, then the search effectiveness improves, but the processing time and system complexity increase
Solution Approach 1:
The system performs comprehensive audio analysis in advance during the indexing phase, extracting text descriptions and key phrases before search operations. This preliminary processing shifts the computational burden to off-peak times, enabling fast search operations without real-time processing delays.
Solution Approach 2:
The patent segments the audio analysis process into distinct components: text description generation, key phrase extraction, and indexing. This segmentation allows each component to be optimized independently and processed efficiently, reducing overall processing time while maintaining comprehensive analysis.
3Measurement precision
If manual labeling is used for training data to improve model accuracy, then the key phrase identification precision improves, but the training time and labor cost increase
Solution Approach 1:
The training data is segmented into two distinct sources: manually labeled data for high-accuracy supervised learning, and automatically generated text descriptions for additional training samples. This segmentation allows the system to leverage the strengths of both approaches while minimizing their individual weaknesses.
Solution Approach 2:
The system makes manual labeling efforts serve multiple functions: they provide ground truth for supervised learning and simultaneously create templates for automatic key phrase extraction. This multi-functionality maximizes the value of manual labeling work, reducing the total time and labor required.
Data Source
AI summary
Embodiments of the present disclosure provide a method and an apparatus for identifying a key phrase in audio, a device and a computer readable storage medium. The method for identifying a key phrase in audio includes obtaining audio data to be identified. The method further includes identifying the key phrase in the audio data using a trained key phrase identification model. The key phrase identification model is trained based on first training data for identifying feature information of words in a first training text and second training data for identifying the key phrase in a second training text. In this way, embodiments of the present disclosure can accurately and efficiently identify key information in the audio data.


