AI Speech Up-Sampling Model Preserves Signal Fidelity

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing digital signal processing methods for improving speech sample frequency, such as interpolation and low pass filtering, result in insufficient generated speech information and loss of original speech quality, leading to poor user experience.

Innovation Solution

A method and device using artificial intelligence, specifically deep learning, to select and apply a pre-trained speech processing model for up-sampling digital speech signals, generating higher sample frequency signals without inserting fixed values or using low pass filters, thus preserving original speech information.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Speed

If digital signal processing is used to improve sample frequency by inserting fixed values and applying low pass filter, then the number of sampling points in unit time is improved, but the generated speech information is insufficient and original speech information is lost

Engineering Contradiction:
Improvesample frequencyVSAvoidspeech information
Core Design Contradiction:
SpeedVSLoss of information

Solution Approach 1:

The patent replaces traditional mechanical digital signal processing methods (fixed value insertion and low pass filtering) with an artificial intelligence-based neural network model. The neural network learns speech signal characteristics and generates up-sampled signals that preserve original information while adding realistic speech details, thereby substituting a more advanced computational approach for the conventional signal processing pipeline.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent changes the fundamental parameter of how up-sampling is achieved - instead of using fixed mathematical operations (interpolation formulas and filter coefficients), it employs a trained neural network model that adapts its parameters (weights and biases) based on learned speech patterns. This allows the system to maintain information fidelity while achieving the desired sample frequency increase.

Inventive Principle:
Principle #35Parameter changes

2Manufacturing precision

If a high sample frequency is used for speech signal, then the quality of speech is improved, but the requirements for processing speed and storage of digital signal processing system are higher

Engineering Contradiction:
Improvespeech qualityVSAvoidprocessing speed
Core Design Contradiction:
Manufacturing precisionVSProductivity

Solution Approach 1:

The patent applies preliminary action by pre-training the neural network model offline with large amounts of speech data. The model learns speech signal characteristics and up-sampling patterns in advance, so that during actual processing, the pre-trained model can quickly generate high-quality up-sampled signals without requiring real-time complex computations, thus resolving the conflict between quality and processing speed.

Inventive Principle:
Principle #10Preliminary action

3Quantity of substance

If fixed values are inserted between sampling points to improve sample frequency, then the number of sampling points is increased, but change in speech information is little and generated speech information is insufficient

Engineering Contradiction:
Improvenumber of sampling pointsVSAvoidgenerated speech information
Core Design Contradiction:
Quantity of substanceVSLoss of information

Solution Approach 1:

The patent replaces the mechanical insertion of fixed values (deterministic interpolation) with a neural network-based generation process. The neural network analyzes the input speech signal and generates new sampling points that contain realistic speech information based on learned patterns, rather than simply interpolating between existing points with fixed mathematical formulas.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent introduces the neural network model as an intermediary between the low-sample-frequency input signal and the high-sample-frequency output signal. This intermediary learns the complex mapping relationship and generates intermediate sampling points that contain meaningful speech information, bridging the gap between sparse input samples and dense output samples with realistic content.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS10360899B2Method and device for processing speech based on artificial intelligence
Publication Date: 2019.07.23 BAIDU ONLINE NETWORK TECH (BEIJIBG) CO LTD
  • US10360899B2 patent drawing
  • US10360899B2 patent drawing
  • US10360899B2 patent drawing

AI summary

The present disclosure provides a method and a device for processing a speech based on artificial intelligence. The method includes: receiving a speech processing request, in which the speech processing request includes a first digital speech signal and a first sample frequency corresponding to the first digital speech signal; selecting a target speech processing model from a pre-trained speech processing model base according to the first sample frequency; performing up-sampling processing on the first digital speech signal using the target speech processing model to generate a second digital speech signal having a second sample frequency, in which the second sample frequency is larger than the first sample frequency.