Voice Processing System Audio Type Classification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Voice enthusiasts find it difficult to quickly and efficiently find their preferred voices on the internet due to the low output efficiency of voice information transmission, and existing methods lack effective preprocessing for accurate voice matching.

Innovation Solution

A method and apparatus that classify user audio into audio type information using preprocessing techniques like noise reduction and blank removal, then determine matching audio type information based on user preferences and stored matching relationship information to adjust voice playback accordingly.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If voice information is transmitted without preprocessing, then the transmission process is simple, but the matching accuracy of voice types is low

Engineering Contradiction:
Improvematching accuracyVSAvoidpreprocessing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies preliminary action by performing preprocessing operations (noise reduction, blank removal) on voice information before matching. This advance preparation improves matching accuracy by ensuring the voice data is clean and properly formatted before comparison with stored voice types.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent segments the preprocessing task into distinct operations: noise reduction, blank removal, and voice type classification. Each operation handles a specific aspect of the data preparation, making the overall complex process more manageable and effective.

Inventive Principle:
Principle #1Segmentation

2Productivity

If voice matching is performed without classification, then the process is fast, but the efficiency of finding preferred voices is low

Engineering Contradiction:
Improvevoice matching efficiencyVSAvoidtime to find preferred voice
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent classifies voice information into different voice types as a preliminary action before matching. This classification creates organized categories that enable faster and more accurate matching of user preferences, improving overall productivity in finding preferred voices.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces voice type classification as an intermediary step between raw voice input and final matching. This intermediate categorization layer acts as a mediator that bridges the gap between unstructured voice data and structured matching requirements, enhancing efficiency.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Reliability

If preprocessing is applied to voice data, then the quality of voice information improves, but the processing time increases

Engineering Contradiction:
Improvevoice information qualityVSAvoidprocessing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent applies partial preprocessing by selectively performing noise reduction and blank removal operations only when necessary for improving voice quality. Not all preprocessing steps are applied to all voice data, balancing quality improvement with processing time considerations.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentEP3846164B1Method and apparatus for processing voice, electronic device, storage medium, and computer program product
Publication Date: 2023.01.04 BEIJING BAIDU NETCOM SCI & TECH CO LTD
  • EP3846164B1 patent drawingFigure 1
  • EP3846164B1 patent drawingFigure 2
  • EP3846164B1 patent drawingFigure 3

AI summary

A method and apparatus for processing a voice are provided. An implementation of the method may include: receiving a user audio sent by a user through a terminal; classifying the user audio, to obtain audio type information of the user audio; and determining, based on the audio type information and a preset matching relationship information, matching audio type information that matches the audio type information as target matching audio type information, the matching relationship information being used to represent a matching relationship between the audio type information and the matching audio type information.