Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

6 results about "PESQ" patented technology

PESQ, Perceptual Evaluation of Speech Quality, is a family of standards comprising a test methodology for automated assessment of the speech quality as experienced by a user of a telephony system. It is standardized as ITU-T recommendation P.862 (02/01). Today, PESQ is a worldwide applied industry standard for objective voice quality testing used by phone manufacturers, network equipment vendors and telecom operators. Its usage requires a license.

A method for testing cloud game audio quality

This invention discloses a method for testing the audio quality of cloud games, comprising the following steps: An audio testing terminal acquires test audio played on a test client and a test server in real time, forming a first test audio and a second test audio, and acquires the corresponding source audio in real time; the audio testing terminal acquires the audio file size and track duration of the first test audio and the source audio within the same time period, respectively, and calculates the ratio of the PESQ value, similarity value, and audio file size of the first test audio to the source audio; the ratio of the audio file size, PESQ value, and similarity value of the first test audio to the source audio are compared with the ratio of the audio file size, PESQ value, and similarity value of the playback audio to the source audio when the source audio is played in high-quality, standard-quality, and smooth-quality playback software. This invention enables automatic testing of cloud game audio quality, with fast testing speed, effectively improving testing efficiency and reducing testing costs.
Owner:SHENZHEN RENDERBUS TECH

Lightweight lossless audio encoding and decoding method and system

The invention discloses a lightweight lossless audio encoding and decoding method and system, and belongs to the technical field of audio signal processing. The method comprises the following steps: converting an original audio waveform into a compact coding expression through an encoder formed by connecting a TConv unit, a convolution unit, a down-sampling unit and a Local-Transform unit in series; a single finite scalar quantizer is adopted to discretize the coded representation, and a mixed quantization strategy is introduced to generate a single and consistent discrete token stream; and finally, a high-fidelity audio waveform is reconstructed through a decoder formed by connecting a Local-Transform unit, a TEConv unit, an up-sampling unit and an audio recovery unit in series. Through unit structure optimization and joint loss function training, indexes such as SDR, PESQ, STOI, WER and the like under multiple different code rates are all superior to those of a baseline model, model parameters are only 10.31-11.29 M, collaboration of lightweight deployment and high-fidelity reconstruction is achieved, resource-constrained equipment is adapted to downstream scenes such as voice communication, ASR and the like, and good engineering practicability is achieved.
Owner:XI AN JIAOTONG UNIV

A speech enhancement method based on fusion network

ActiveCN116486826BData setNoise
In the low signal-to-noise ratio condition, aiming at the problems that the traditional neural network speech feature extraction is insufficient, and the speech enhancement effect needs to be improved, based on empirical mode decomposition (EMD), temporal convolution network (TCN) and gated convolution recurrent neural network (GCRN), and combining with feature fusion module (FFM), the application proposes a speech enhancement model of adaptive mean median empirical mode decomposition-multilayer gated feature fusion module convolutional recurrent neural network (ME-MGFCRN). The network model adopts the frequency learning strategy to learn the low frequency feature and the high frequency feature, that is, the TCN and the MGFCRN network are used to obtain the low frequency and the high frequency feature, and the two groups of features are processed through the FMM, so as to realize the speech enhancement in the feature mapping mode. The model proposed in the application carries out the ablation experiment and the comparison experiment on the data set, and uses the PESQ, fwSegSNR and STOI indexes to evaluate the speech enhancement effect. Research shows that under different noise environments and different signal-to-noise ratios, the model proposed in the application is improved compared with other baseline models, especially under the low signal-to-noise ratio condition of SNR of-5dB, the fwSegSNR and PESQ are improved by more than 0.86dB and 0.02 respectively compared with other baseline models.
Owner:HARBIN UNIV OF SCI & TECH

Bluetooth earphone based on real-time adaptive noise reduction algorithm

The invention provides a real-time self-adaptive noise reduction Bluetooth headset, which is characterized in that environment noise and ear canal residual noise are respectively acquired by a dual-microphone array at an included angle of 45 degrees, four scenes of traffic noise, human voice, wind noise and a quiet environment are identified in real time by an embedded lightweight convolutional neural network (parameter quantity is less than or equal to 5KB), and matched noise reduction parameters are dynamically loaded; an anti-phase sound wave is generated in combination with a self-adaptive feedback controller, system delay is compressed to 0.08 ms level by using an LMS filter with an adjustable step length mu, and a sound wave anti-phase physical limit is broken through; the noise reduction depth (static 40dB / walking 25dB) is dynamically adjusted according to the motion state through the six-axis sensor, and noise reduction is stopped within 0.5 second when the barometer detects that the earphone is taken off; the binaural exchanges noise data based on BLE 5.0, and executes a hybrid noise reduction strategy in an asymmetric scene. According to the scheme, the noise suppression depth is improved by 40% (the subway environment reaches 38dB), and the high-frequency phase deviation is lt; and the PESQ score of the binaural voice definition is 4.8 / 5.0, the power consumption of a motion scene is reduced by 35%, and the method is suitable for acoustic equipment such as TWS earphones and the like.
Owner:SHENZHEN PINSHENG IND CO LTD

A deep echo cancellation method based on cross-domain prior interaction gating and feature decoupling

PendingCN122314000ASaliency mapPESQ
This invention provides a deep echo cancellation method based on cross-domain prior interaction gating and feature decoupling, comprising: preprocessing the acquired near-end microphone signal and far-end reference signal to obtain a complex spectrum with uniform time-frequency resolution; constructing an echo cancellation model based on the complex spectrum using a cross-modal gating attention mechanism, a conditional Transformer, and a multi-branch decoder; and performing echo cancellation on the newly acquired mixed speech signal based on the echo cancellation model to obtain enhanced near-end speech. This method can still learn echo saliency maps through pseudo-labels / weak supervision in unlabeled scenarios, and inference only requires a single forward computation, thereby improving PESQ / STOI and ERLE and reducing speech distortion.
Owner:ANHUI UNIV