Key Phrase Identification Model for Audio Search

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current audio search systems cannot comprehensively and efficiently retrieve audio data based on user queries since they rely solely on title and text descriptions, failing to accurately identify key phrases within the audio content.

Innovation Solution

A method and apparatus for training a key phrase identification model using both natural language processing and manually labeled training samples to convert audio data into text, enabling the recognition of key phrases within audio data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If the system only uses title and text description for audio search, then the search system is simple to operate, but the search accuracy and comprehensiveness deteriorate

Engineering Contradiction:
Improvesearch system simplicityVSAvoidsearch accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The system performs preliminary action by automatically generating text descriptions and extracting key phrases from audio content during the indexing phase, before the actual search occurs. This preprocessing enables the search system to query not only titles but also generated descriptions and key phrases, significantly improving search accuracy without complicating the user interface.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces text descriptions and key phrases as intermediary elements between the audio content and the search query. These intermediaries bridge the gap between user search intent and audio content, allowing accurate search without requiring users to interact with complex audio analysis tools.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If the system performs comprehensive audio content analysis to improve search accuracy, then the search effectiveness improves, but the processing time and system complexity increase

Engineering Contradiction:
Improvesearch effectivenessVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs comprehensive audio analysis in advance during the indexing phase, extracting text descriptions and key phrases before search operations. This preliminary processing shifts the computational burden to off-peak times, enabling fast search operations without real-time processing delays.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent segments the audio analysis process into distinct components: text description generation, key phrase extraction, and indexing. This segmentation allows each component to be optimized independently and processed efficiently, reducing overall processing time while maintaining comprehensive analysis.

Inventive Principle:
Principle #1Segmentation

3Measurement precision

If manual labeling is used for training data to improve model accuracy, then the key phrase identification precision improves, but the training time and labor cost increase

Engineering Contradiction:
Improvemodel accuracyVSAvoidtraining time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The training data is segmented into two distinct sources: manually labeled data for high-accuracy supervised learning, and automatically generated text descriptions for additional training samples. This segmentation allows the system to leverage the strengths of both approaches while minimizing their individual weaknesses.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system makes manual labeling efforts serve multiple functions: they provide ground truth for supervised learning and simultaneously create templates for automatic key phrase extraction. This multi-functionality maximizes the value of manual labeling work, reducing the total time and labor required.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11308937B2Method and apparatus for identifying key phrase in audio, device and medium
Publication Date: 2022.04.19 BEIJING BAIDU NETCOM SCI & TECH CO LTD
  • US11308937B2 patent drawing
  • US11308937B2 patent drawing
  • US11308937B2 patent drawing

AI summary

Embodiments of the present disclosure provide a method and an apparatus for identifying a key phrase in audio, a device and a computer readable storage medium. The method for identifying a key phrase in audio includes obtaining audio data to be identified. The method further includes identifying the key phrase in the audio data using a trained key phrase identification model. The key phrase identification model is trained based on first training data for identifying feature information of words in a first training text and second training data for identifying the key phrase in a second training text. In this way, embodiments of the present disclosure can accurately and efficiently identify key information in the audio data.