Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

13 results about "PESQ" patented technology

PESQ, Perceptual Evaluation of Speech Quality, is a family of standards comprising a test methodology for automated assessment of the speech quality as experienced by a user of a telephony system. It is standardized as ITU-T recommendation P.862 (02/01). Today, PESQ is a worldwide applied industry standard for objective voice quality testing used by phone manufacturers, network equipment vendors and telecom operators. Its usage requires a license.

Sound effect adjusting method for high-performance audio loudspeaker

The invention relates to the technical field of stereo devices, in particular to a sound effect adjusting method for a high-performance audio loudspeaker. The method comprises the following steps: firstly, inputting physical parameters of a loudspeaker unit, a cavity structure and a passive diaphragm, establishing a coupling acoustic model by adopting finite element simulation, and generating a theoretical frequency response curve; actual frequency response data of the sound box are acquired through analog-to-digital conversion and a sensor, and a core frequency band which has the greatest influence on sound quality is extracted by combining a Gaussian process regression method; extracting acoustic characteristic parameters based on a frequency response deviation function in the core frequency band, inputting the acoustic characteristic parameters to a PESQ model, and driving an LMS adaptive filtering algorithm to complete preliminary adjustment; and then carrying out secondary optimization on the full-band parameter space by adopting a particle swarm optimization algorithm, and finally obtaining an optimal sound effect adjustment result. The method has the advantages of being high in adjustment efficiency, high in parameter adaptive capacity, accurate in sound effect evaluation and the like, and is suitable for automatic sound effect optimization scenes of high-fidelity Bluetooth sound boxes, intelligent sound boxes and other products.
Owner:SHENZHEN ZHILIAN TECH CO LTD

A method for testing cloud game audio quality

This invention discloses a method for testing the audio quality of cloud games, comprising the following steps: An audio testing terminal acquires test audio played on a test client and a test server in real time, forming a first test audio and a second test audio, and acquires the corresponding source audio in real time; the audio testing terminal acquires the audio file size and track duration of the first test audio and the source audio within the same time period, respectively, and calculates the ratio of the PESQ value, similarity value, and audio file size of the first test audio to the source audio; the ratio of the audio file size, PESQ value, and similarity value of the first test audio to the source audio are compared with the ratio of the audio file size, PESQ value, and similarity value of the playback audio to the source audio when the source audio is played in high-quality, standard-quality, and smooth-quality playback software. This invention enables automatic testing of cloud game audio quality, with fast testing speed, effectively improving testing efficiency and reducing testing costs.
Owner:SHENZHEN RENDERBUS TECH

Lightweight lossless audio encoding and decoding method and system

The invention discloses a lightweight lossless audio encoding and decoding method and system, and belongs to the technical field of audio signal processing. The method comprises the following steps: converting an original audio waveform into a compact coding expression through an encoder formed by connecting a TConv unit, a convolution unit, a down-sampling unit and a Local-Transform unit in series; a single finite scalar quantizer is adopted to discretize the coded representation, and a mixed quantization strategy is introduced to generate a single and consistent discrete token stream; and finally, a high-fidelity audio waveform is reconstructed through a decoder formed by connecting a Local-Transform unit, a TEConv unit, an up-sampling unit and an audio recovery unit in series. Through unit structure optimization and joint loss function training, indexes such as SDR, PESQ, STOI, WER and the like under multiple different code rates are all superior to those of a baseline model, model parameters are only 10.31-11.29 M, collaboration of lightweight deployment and high-fidelity reconstruction is achieved, resource-constrained equipment is adapted to downstream scenes such as voice communication, ASR and the like, and good engineering practicability is achieved.
Owner:XI AN JIAOTONG UNIV

Method and system for training neural network

Disclosed herein is a method and system for training a neural network. According to one embodiment, the method includes receiving a noisy signal, generating a denoised output signal, determining a signal-to-distortion ratio (SDR) loss function based on the denoised output signal, determining a perceptual speech quality assessment (PESQ) loss function based on the denoised output signal, and optimizing a total loss function based on the perceptual speech quality assessment loss function and the signal-to-distortion ratio loss function.
Owner:SAMSUNG ELECTRONICS CO LTD

A speech enhancement method based on fusion network

ActiveCN116486826BData setNoise
In the low signal-to-noise ratio condition, aiming at the problems that the traditional neural network speech feature extraction is insufficient, and the speech enhancement effect needs to be improved, based on empirical mode decomposition (EMD), temporal convolution network (TCN) and gated convolution recurrent neural network (GCRN), and combining with feature fusion module (FFM), the application proposes a speech enhancement model of adaptive mean median empirical mode decomposition-multilayer gated feature fusion module convolutional recurrent neural network (ME-MGFCRN). The network model adopts the frequency learning strategy to learn the low frequency feature and the high frequency feature, that is, the TCN and the MGFCRN network are used to obtain the low frequency and the high frequency feature, and the two groups of features are processed through the FMM, so as to realize the speech enhancement in the feature mapping mode. The model proposed in the application carries out the ablation experiment and the comparison experiment on the data set, and uses the PESQ, fwSegSNR and STOI indexes to evaluate the speech enhancement effect. Research shows that under different noise environments and different signal-to-noise ratios, the model proposed in the application is improved compared with other baseline models, especially under the low signal-to-noise ratio condition of SNR of-5dB, the fwSegSNR and PESQ are improved by more than 0.86dB and 0.02 respectively compared with other baseline models.
Owner:HARBIN UNIV OF SCI & TECH

Voice quality objective evaluation method based on nonlinear regression improvement

The invention relates to a voice quality objective evaluation method based on nonlinear regression improvement, and the method comprises the steps: carrying out the preprocessing of an original voice signal and a distorted voice signal, and carrying out the time alignment; performing auditory conversion on the aligned signals to obtain auditory feature representation, and performing jitter processing; inputting the features subjected to jitter processing into a cognitive model to obtain a preliminary voice quality score; wherein the cognitive model is constructed based on an auditory psychology experiment result; and performing nonlinear regression correction on the initial voice quality score to obtain a final voice quality score. According to the method, a noise reconstruction and nonlinear regression module is added on the basis of the PESQ model, the predicted voice quality is remapped, experiments are performed on a plurality of data sets with or without noise, compared with an original PESQ model, the prediction score accuracy is improved by 82.5% at most, an experiment result proves that the model can effectively improve the prediction quality, and the prediction efficiency is improved. And the voice synthesis work can be more effectively assisted.
Owner:GUANGDONG UNIV OF PETROCHEMICAL TECH

A single-channel speech enhancement method based on multi-attention mechanism

The present invention relates to a single-channel speech enhancement method based on a multi-attention mechanism, and belongs to the technical field of audio signal processing. The present invention introduces a complex Conformer into a complex U-Net network to model the correlation between speech amplitude and phase, uses a three-dimensional attention mechanism to construct richer features to enhance the representation ability of the convolutional layer, and fuses speech detail features and deep features through a gated attention mechanism. The method can improve speech quality and intelligibility, and can be used for voice communication in noisy environments, command control, and the pre-processing part of speech-related tasks. Experimental results on public datasets show that the proposed method achieves evaluation results of 3.09, 4.28, 3.47, 3.72, and 95.07 on five objective evaluation indicators: PESQ, CSIG, CBAK, COVL, and STOI, respectively, and can effectively reduce noise and improve speech quality and intelligibility.
Owner:KUNMING UNIV OF SCI & TECH

Bluetooth earphone based on real-time adaptive noise reduction algorithm

The invention provides a real-time self-adaptive noise reduction Bluetooth headset, which is characterized in that environment noise and ear canal residual noise are respectively acquired by a dual-microphone array at an included angle of 45 degrees, four scenes of traffic noise, human voice, wind noise and a quiet environment are identified in real time by an embedded lightweight convolutional neural network (parameter quantity is less than or equal to 5KB), and matched noise reduction parameters are dynamically loaded; an anti-phase sound wave is generated in combination with a self-adaptive feedback controller, system delay is compressed to 0.08 ms level by using an LMS filter with an adjustable step length mu, and a sound wave anti-phase physical limit is broken through; the noise reduction depth (static 40dB / walking 25dB) is dynamically adjusted according to the motion state through the six-axis sensor, and noise reduction is stopped within 0.5 second when the barometer detects that the earphone is taken off; the binaural exchanges noise data based on BLE 5.0, and executes a hybrid noise reduction strategy in an asymmetric scene. According to the scheme, the noise suppression depth is improved by 40% (the subway environment reaches 38dB), and the high-frequency phase deviation is lt; and the PESQ score of the binaural voice definition is 4.8 / 5.0, the power consumption of a motion scene is reduced by 35%, and the method is suitable for acoustic equipment such as TWS earphones and the like.
Owner:SHENZHEN PINSHENG IND CO LTD

Communication voice quality evaluation method and device and electronic equipment

The invention discloses a communication voice quality evaluation method and device and electronic equipment. The method comprises the steps that a signal to be processed is acquired, the signal to be processed is divided into a target time frame with a preset length, and the preset length ranges from 50 milliseconds to 100 milliseconds; splitting the target time frame into a plurality of sub-blocks, and determining the Bark scale of each sub-block according to the scaling factor and the frequency of the target time frame; fusing the Bark scales of the plurality of sub-blocks to obtain a target Bark scale; and determining the target loudness of the target time frame according to the target Bark scale. According to the invention, the technical problem of inaccurate voice quality score caused by frame boundary mismatch and insufficient time domain masking effect when voice quality evaluation is carried out by using a traditional PESQ algorithm with short-time frame length due to relatively large signal propagation time delay in high-orbit satellite communication is solved.
Owner:CHINA TELECOM CORP LTD SATELLITE COMMUNICATIONS BRANCH

Speech enhancement method based on multi-modal fusion deep learning

The invention provides a speech enhancement model based on an encoder-decoder structure and fused with a generative adversarial network, and aims to solve the problem of speech quality improvement in a complex acoustic environment. The experiment is based on an AI-SHELL3 clean corpus data set, and after large noise and large reverberation interference are added, the model is adopted for processing. The result shows that the PESQ (objective evaluation index of speech quality) of the speech enhanced by the model is relatively improved by 1.26, which indicates that the designed model can effectively suppress strong noise interference and reduce reverberation influence, significantly improve the speech quality under severe acoustic conditions, and improve the speech quality. The effectiveness of the combination of the encoder-decoder and the generative adversarial network in the speech enhancement task is verified, and an efficient solution is provided for speech processing in a high-interference environment.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

A deep echo cancellation method based on cross-domain prior interaction gating and feature decoupling

PendingCN122314000ASaliency mapPESQ
This invention provides a deep echo cancellation method based on cross-domain prior interaction gating and feature decoupling, comprising: preprocessing the acquired near-end microphone signal and far-end reference signal to obtain a complex spectrum with uniform time-frequency resolution; constructing an echo cancellation model based on the complex spectrum using a cross-modal gating attention mechanism, a conditional Transformer, and a multi-branch decoder; and performing echo cancellation on the newly acquired mixed speech signal based on the echo cancellation model to obtain enhanced near-end speech. This method can still learn echo saliency maps through pseudo-labels / weak supervision in unlabeled scenarios, and inference only requires a single forward computation, thereby improving PESQ / STOI and ERLE and reducing speech distortion.
Owner:ANHUI UNIV

Acoustic quality evaluation device, acoustic quality evaluation method, and program

PCT designated stageWO2026009400A1Substation equipmentPublic address systemEnvironmental noise
The present invention achieves suitable acoustic quality evaluation of a loudspeaker communication system by objective evaluation without implementing a speaking test or a listening test, even in a call environment in which there is an ambient noise. An acoustic quality evaluation device according to the present disclosure comprises a first acoustic quality evaluation unit, an analysis unit, and a second acoustic quality evaluation unit. The first acoustic quality evaluation device presents an analytical reference signal, an analytical degradation signal, and an analytical ambient noise to an evaluator and acquires an analytical subjective evaluation value. An analysis device determines an analytical PESQ value from the analytical reference signal and the analytical degradation signal, and determines the relationship between the analytical subjective evaluation value and the analytical PESQ value. The second acoustic quality evaluation device determines a test PESQ value from a testing reference signal, a testing degradation signal, and a testing ambient noise, and in accordance with the relationship, calculates, from the test PESQ value, estimated subjective evaluation values of the testing degradation signal for the testing reference signal testing ambient noise obtained through comparison under the testing ambient noise.
Owner:NT T INC

An audiovisual speech noise reduction method based on a multi-modal gating enhancement model

The present invention discloses a method for audio-visual speech denoising based on a multimodal gated lifting model, comprising the following steps: separate storage of images and audio; preprocessing of audio and images; cropping of lip images and generation of speech spectrograms using a lip localization algorithm and a short-time Fourier transform; capturing and enhancing visual features and audio features using a hierarchical attention module and a dual-path spectrum enhancement module; gradually fusing visual features and audio features using a gated encoder; enhancing key audio-visual features using a time-frequency lifting module; estimating a pure speech spectrogram using a gated decoder; acquiring speech signals using an inverse short-time Fourier transform; and training or testing a network model. The present invention is highly robust and has a wide range of applications, and can achieve speech denoising in complex noisy environments. Compared with some mainstream denoising models, the present invention improves the SI‑SDR and PESQ evaluation indicators by approximately 15% and 19%, respectively.
Owner:XI AN JIAOTONG UNIV