Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

191 results about "Speech quality" patented technology

Measuring Speech Quality. In telecommunication, speech quality is an important contributing factor to the success of a product and to the success of the communication itself. High speech quality guarantees that the effort the users have to put forward in order to correctly perceive the communication is low.

Video translation method and system based on artificial intelligence

The invention discloses a video translation method and system based on artificial intelligence. The method relates to the technical field of video translation and comprises the following steps of original sound track extraction, target AI speaker adaptation, AI dubbing generation and mouth shape synchronization and video synthesis. According to the method, independent audio and video streams are obtained by adopting an audio and video separation technology, and multiple original sound tracks are extracted through a voice separation model; matching or generating an adaptive target AI speaker module in a preset tone library; converting the original language voice into a text, translating the text into a target language text, and synthesizing an AI dubbing audio track in combination with a target AI speaker module; and finally, the independent video stream and the multi-AI dubbing audio track are input into the mouth shape synchronization model to output a translated video, so that the timbre fitting degree, the voice quality and the voice consistency of the same speaker of AI dubbing are improved, and meanwhile, the resource utilization rate of video translation and the processing efficiency under batch tasks are improved. The problem that in the prior art, video translation is low in quality and efficiency is solved.
Owner:BEIJING DEEP LOGIC INTELLIGENT TECHNOLOGY CO LTD

Method and device for converting lip language into voice, computer storage medium and terminal

The invention discloses a lip language-to-voice conversion method and device, a computer storage medium and a terminal, and aims to solve the problems that a lip language-to-voice conversion technology cannot be deployed on terminal equipment and voice quality cannot meet application requirements. The method is combined with a neural vocoder which reduces calculation complexity and resource requirements, system configuration requirements are reduced while system parameters are reduced, a design basis is provided for deploying a lip language-to-speech method on terminal equipment, and context-related visual feature sequences and user audio embedding vectors are fused, so that the user audio-to-speech conversion efficiency is improved, and the user audio-to-speech conversion efficiency is improved. A Mel spectrum acoustic feature sequence used for being converted into a voice waveform is obtained, and a user audio vector fused with the Mel spectrum acoustic feature sequence is obtained, so that the output voice waveform is more consistent with the real voice of a user; according to the embodiment of the invention, technical support is provided for deploying and applying the lip language-to-speech method meeting the speech quality requirement on the terminal equipment.
Owner:BEIJING WATERTEK INFORMATION TECH

Far-field single-channel speech enhancement method

The invention relates to the technical field of speech enhancement, in particular to a far-field single-channel speech enhancement method based on an MFSE (Maximum Free Square Error) (Maximum Free Square Error) (Maximum Free Square Error) (Maximum Free Square Error) (Maximum Free Square Error) (Maximum Free Square Error)). The method comprises the following steps: step 1, processing a far-field voice signal to obtain a complex spectrogram of a noise voice signal; 2, inputting the compressed complex spectrogram into a feature encoder, processing the output of the complex spectrogram by N MamAttention blocks, and then sending the processed complex spectrogram into an amplitude mask decoder and a phase decoder to respectively predict a clean compressed amplitude mask and a phase spectrum; step 3, preheating and training the MamAttention model, and performing supervised confrontation training by taking the MamAttention model as a generator and the multi-resolution discriminator as a discriminator; and step 4, inputting test voice into the trained model to realize far-field single-channel voice enhancement. According to the method, the supervised adversarial training strategy and the MamAttention model are combined, so that the problems of signal attenuation, noise and reverberation interference in far-field voice are effectively solved, and the voice quality is remarkably improved in a scene that the distance of a loudspeaker exceeds 5 meters.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Voice quality optimization method and device, equipment, storage medium and product

The invention discloses a voice quality optimization method and device, equipment, a storage medium and a product, and relates to the technical field of communication, and the method comprises the steps: based on an RSRP sequence of each user equipment, recognizing each fast fading equipment with an RSRP fluctuation value exceeding a preset fluctuation threshold value from each user equipment through a preset recurrent neural network; determining a communication cell to which each fast fading device belongs according to the temporary identifier of each fast fading device, and determining the communication cell meeting a preset device scale condition and a fast fading device distribution condition as a fast fading cell; determining a voice problem cell from the fast fading cells according to the service quality identifier of each fast fading device in the fast fading cells; and voice quality optimization is carried out on voice problem equipment in the voice problem cell, and the voice problem equipment is fast fading equipment for executing voice services. Therefore, the influence of the fast fading phenomenon on the voice quality is reduced, and the voice quality is improved.
Owner:CHINA MOBILE GROUP JILIN BRANCH +1

Multi-speaker speech synthesis method and device, storage medium and computer equipment

The invention relates to the technical field of multi-speaker speech synthesis, and particularly provides a multi-speaker speech synthesis method and device, a storage medium and computer equipment. According to the method, customized structure expansion, engine transformation and optimization are carried out on the basis of an open source reasoning framework vLLM, so that the target vLLM framework obtained through transformation can face a speech synthesis scene, and efficient reasoning of a speech synthesis model is supported. Speech synthesis reasoning is realized by using the target vLLM framework and the LLM-based target speech synthesis model, so that speech synthesis can be accelerated under the condition of keeping speech quality.
Owner:SHANGHAI LINGGUANG ZHAXIAN TECHNOLOGY CO LTD

Zero sample speech synthesis method and device, computer equipment and storage medium

The invention relates to a zero-sample speech synthesis method and device, computer equipment, a storage medium and a program product, and the method comprises the steps: obtaining a target coding feature according to a reference speech and a target text; inputting the target coding features into a stream matching model to obtain a conditional velocity field and an unconditional velocity field; inputting the target coding feature into a prior model to obtain a prior speech feature; obtaining a prior generation flow field according to the prior voice features and the standard Gaussian noise; calculating a KL divergence value between the priori generated flow field and a preset real generated flow field, and taking a moment when the KL divergence value is smaller than or equal to a preset KL divergence threshold as an initial moment of the priori generated flow field; fusing the conditional velocity field and the unconditional velocity field to obtain a fused velocity field; and inputting the fusion velocity field from the initial moment to the target moment and the priori generated flow field into an ordinary differential equation solver to obtain the target speech features, thereby improving the speech quality of the synthesized speech.
Owner:CHINA SOUTHERN POWER GRID ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD

Voice wake-up interaction method and system based on microphone array

The invention discloses a voice wake-up interaction method and system based on a microphone array, and the method comprises the steps: carrying out VAD processing, so as to judge whether a target audio segment has voice or not; voice and noise source directions are obtained; selecting beam parameters of voice and noise in combination with a pre-designed fixed beam; whether GSC module processing is carried out or not is selected according to the difference between the beam parameters of the voice and the noise, so that an enhanced audio signal is obtained, a wake-up task is carried out, and a final voice wake-up result is obtained; and according to whether the wake-up is successful, determining whether to lock the voice beam direction in the current interaction stage for enhancement, thereby preventing interference of voice in other directions on subsequent interaction tasks. According to the voice enhancement mode based on the microphone array, sound source positioning can be realized, interference in a non-target direction can be suppressed, the voice quality in the target direction can be improved, the wake-up success rate can be effectively improved when the voice enhancement mode is applied to voice wake-up, and then the experience of back-end voice interaction is improved.
Owner:PANOVASIC TECHNOLOGY CO LTD

Frequency domain noise reduction and reduction method and system for synchronously collecting double groups of human voices

The invention relates to the technical field of voice communication, and discloses a frequency domain noise reduction and reduction method and system for double-group human voice synchronous acquisition, and the method comprises the steps: obtaining a synchronous signal through double groups of acquisition points, constructing a double-source frequency domain coordinate system, and mapping the coordinate system into an energy distribution point; calculating a spatial relationship to identify an emotional distortion state, and dividing four quadrants to generate a distortion type mark; dividing a frequency domain signal and executing targeted reconstruction and correction, and finally dividing a frequency band enhanced energy remarkably-reduced part; according to the method, the voice composite distortion under the extreme emotion can be accurately processed, and reliable voice quality guarantee is provided for accurate transmission of key information in an emergency communication scene.
Owner:GUANGZHOU CMX AUDIO CO LTD

NPU-based Chinese-English bilingual text-to-speech conversion method and system

The invention discloses a Chinese-English bilingual text-to-speech conversion method and system based on NPU, and belongs to the technical field of speech processing, and the method comprises the steps: carrying out the word segmentation processing and phoneme conversion of a Chinese-English mixed text based on the language type of each segment in the Chinese-English mixed text, and obtaining a text input sequence; performing vector combination on the phoneme ID sequence and the language ID sequence, and extracting text semantic features from a vector combination result to obtain text hidden variables; according to the text hidden variable, predicting the duration of each phoneme and a prior distribution parameter in an acoustic potential feature space; performing multi-level transformation on the phoneme alignment result and the prior distribution parameter to obtain a bilingual acoustic feature sequence; and converting the bilingual acoustic feature sequence into a corresponding voice waveform. Therefore, by implementing the method, the device and the system, the problems that the Chinese-English bilingual text occupies more resources in the voice conversion process and the output voice quality is lower in the prior art can be solved.
Owner:GUANGZHOU BAOLUN ELECTRONICS CO LTD

Speech synthesis style transfer method and device, computer device and storage medium

The present application relates to the field of speech synthesis, and particularly relates to a speech synthesis style migration method and device, computer equipment and a storage medium. The method comprises the following steps: inputting the text data to be converted into a speech synthesis encoder to obtain text encoding; obtaining multi-level speech style representation of the speech to be migrated; inputting the text data to be converted and the multi-level speech style representation into a style predictor to obtain a predicted speech style; performing multi-style layer regularization on the text encoding and the predicted speech style to obtain style-regularized text encoding; inputting the style-regularized text encoding into a pitch predictor to obtain rhythm change data; inputting the rhythm change data and the style-regularized text encoding into a speech synthesis decoder to obtain a mel-frequency spectrum diagram and generate target speech. The present application considers rhythm information and the style of words and phonemes, and the overall rhythm style of the speaker, thereby improving the speech quality of the target speech.
Owner:PING AN TECH (SHENZHEN) CO LTD

Audio processing method and device, server and storage medium

The invention provides an audio processing method and device, a server and a storage medium. The method comprises the following steps: receiving a to-be-adjusted original audio and a corresponding line; converting the lines into a corresponding text sequence; performing alignment operation on the original audio and the text sequence to obtain an audio-text corresponding relation; receiving a target text; determining a to-be-replaced first text from the text sequence according to the target text and frame information of the first text in the audio in the corresponding relation; replacing the first text with the target text by using the frame information to obtain an updated text sequence; and generating a new audio by using the updated text sequence. Therefore, the problems that the edited voice is stiff, the complex audio effect is poor, and the quality is possibly poor if the voice is modified too much can be solved, and the quality of the generated voice can be better controlled by controlling the durations of different granularities.
Owner:CHENGDU IQIYI INTELLIGENT INNOVATION TECH CO LTD

Voice processing method and device, equipment, storage medium and program product

The embodiment of the invention provides a voice processing method and device, equipment, a storage medium and a program product. The method comprises the following steps: acquiring a voice signal; performing signal processing on the voice signal to obtain a first time frequency signal; performing adaptive normalization processing on the first time-frequency signal to obtain a first feature map; performing two-dimensional modeling processing on the first feature map to obtain a second feature map; and performing synthesis processing on the second feature map to obtain an enhanced voice signal. The method can improve the voice quality.
Owner:RDA CHONGQING MICROELECTRONICS TECH CO LTD

System

An object of a system according to an embodiment is to analyze biological data of participants and automatically organize an effective breakout room.SOLUTION: A system includes a biological data collection section, an analysis section, a classification section, and an organization section. The biological data collection unit collects biological data such as facial expressions and voice quality from a camera or a microphone of a terminal of a participant. The analysis section analyzes the biological data collected by the biological data collection section. The classification component classifies the participants by character on the basis of the data analyzed by the analysis component. The organization unit automatically organizes a breakout room based on the characters classified by the classification unit.SELECTED DRAWING: Figure 1
Owner:SOFTBANK GROUP CORP

Ultra-low computing resource speech enhancement method based on Half-UNet architecture

The invention discloses an ultra-low computing resource speech enhancement method based on a Half-UNet framework. According to the method, a decoder of a UNet is simplified, a recurrent neural network module is arranged between an encoder and the decoder to construct a Half-UNet architecture, and in combination with feature fusion, adaptive frequency band division and power law compression penalty technologies, the calculation complexity and parameter quantity are greatly reduced while enhancing performance is ensured. The method comprises the following steps: carrying out frequency band combination on frequency spectrums of input noise voice by using a filter obtained by training of a self-adaptive frequency band division module, and reducing high-frequency characteristic redundancy; carrying out feature extraction by using Half-UNet, and reconstructing a frequency spectrum; a power law is used to compress penalty terms to enhance weak detail features, and the weak detail features are prevented from being submerged by strong noise features, so that the voice quality of the model in a low signal-to-noise ratio environment is improved. The method has the advantages that under the condition that only about 23.7 K parameters and 25.42 MMACs operand are needed, the voice enhancement effect equivalent to that of a large-scale deep model is achieved, and the method is particularly suitable for resource-limited equipment such as earphones and hearing aids.
Owner:EAST CHINA NORMAL UNIV +1

radio

PendingJP2026043102ATransmissionNoiseSpeech sound
To provide a radio capable of performing squelch control with little erroneous determination regardless of the voice quality of a speaker. [Solution] The radio 20 in this example comprises an SSB demodulation unit 21 that performs SSB demodulation processing on the received signal, a noise cancellation unit 22 that performs noise cancellation processing on the audio signal resulting from demodulation by the SSB demodulation unit 21, a pitch period calculation unit 23 that calculates the pitch period of the audio from the audio signal after noise cancellation processing, and a squelch control unit 24 that controls the opening and closing of the squelch of a squelch circuit 25 based on the pitch period calculated by the pitch period calculation unit 23.
Owner:KOKUSAI DENKI ELECTRIC INC

Voice quality evaluation method and device, equipment and storage medium

The invention discloses a voice quality evaluation method and device, equipment and a storage medium, and relates to the technical field of computers. The method comprises the steps of obtaining to-be-evaluated voice generated through text-to-voice conversion and style information for the to-be-evaluated voice, and extracting basic features of the to-be-evaluated voice; extracting text-to-speech features from the basic features; the text-to-speech features comprise any one or more of rhythm features, tone features and text matching features; calculating a style matching degree between an actual style embedding vector corresponding to the text-to-speech feature and a style embedding vector template corresponding to the style information; and predicting a speech degradation score based on the text-to-speech features, and determining speech quality based on the speech degradation score and the style matching degree. By extracting rhythm features, timbre features and text matching features, identifying specific quality defects in text-to-speech conversion; and in combination with the style matching degree, the problem of confusion of style differences and quality defects is solved, and the accuracy of voice quality scoring is improved.
Owner:MALANSHAN AUDIO & VIDEO LABORATORY

Voice data loss recovery processing method, device and equipment and storage medium

The application discloses a voice data loss recovery processing method and device, equipment and a storage medium, and relates to the technical field of voice transmission. The method comprises the following steps: performing frame processing on received voice message data to obtain frame voice data; when signal loss is detected, performing voice reconstruction on the frame voice data based on compressed sensing to obtain initial voice recovery data; and performing dynamic voice filtering on the initial voice recovery data through a preset filtering network to obtain target voice recovery data. The voice message data is first subjected to frame processing, the voice is reconstructed based on compressed sensing by using the sparsity of the voice signal when the signal is lost, and then dynamic filtering is performed through the preset filtering network, so that high-quality voice data is finally output. The method solves the problems of bandwidth waste caused by redundant transmission and poor effect of traditional interpolation filtering in the prior art, realizes the effects of saving bandwidth, efficiently recovering lost voice frames and improving voice quality, and meets the demand for high-quality voice communication.
Owner:SHENZHEN DINSTAR TECH

Multi-path driver passenger full duplex emergency call and stop announcement broadcast coexistence method

The invention discloses a multi-path driver passenger full duplex emergency call and stop announcement broadcast coexistence method, and relates to the technical field of audio processing. According to the scheme, station reporting broadcast and multi-path talkback signals are transmitted at the same time through a single MVB bus, coexistence of multi-path driver passenger emergency calls and station reporting broadcast and talkback of a driver with multiple passengers at the same time are achieved, hardware redundancy of a traditional discrete communication channel is omitted, the system complexity and cost are reduced, and the communication efficiency is improved. According to the weight, the multiple paths of talkback audio data are combined, the voice quality of the multiple paths of audio signals is improved, and signal conflicts are also avoided.
Owner:DALIAN HAITIAN IND TECH CO LTD

A speech enhancement method based on air-bone conduction dual-mode deep learning

This invention discloses a speech enhancement method based on air-bone dual-mode deep learning, comprising the following steps: Step 1, synchronously recording air-guided and bone-guided speech in a noise-free environment, adding environmental noise to the air-guided speech to construct a dataset, and dividing it into a training set, a test set, and a validation set; Step 2, segmenting the training set by cutting the speech data of the training set into multiple short speech segments of a fixed length; Step 3, constructing a neural network, including a high-order encoder, zero-order terms, high-order terms, and auxiliary post-processing filters; Step 4, model training; Step 5, model testing. The algorithm of this invention has low computational complexity, meets real-time requirements, has stable noise reduction performance, and provides high speech quality and intelligibility, making it applicable to portable devices.
Owner:杭州智元研究院有限公司

Voice quality detection method, computer equipment and computer storage medium

The embodiment of the invention discloses a voice quality detection method, computer equipment and a computer storage medium. Determining a first voice naturalness evaluation result of the target voice audio according to the acoustic attribute characteristics of the target voice audio, and determining a second voice naturalness evaluation result of the target voice audio according to the voice content structure characteristics of the target voice audio, and obtaining a voice naturalness evaluation result of the target voice audio based on the first voice naturalness evaluation result and / or the second voice naturalness evaluation result. Based on physical characteristics, content structures and other levels, the performance of the speech in the aspect of naturalness is objectively and accurately evaluated in a multi-dimensional manner. Quantitative analysis is directly carried out aiming at the physical characteristics and the content structure characteristics of the voice, and the association between the characteristics and the naturalness of the voice is visual and understandable, so that the method has better interpretability and debuggeability, and a clear guidance direction can be provided for algorithm optimization.
Owner:TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD

Two-stage lightweight echo cancellation method and system combining NKF and EMA-GT convolution

The invention discloses a two-stage lightweight echo cancellation method and system combining NKF and EMA-GT convolution. In the first stage of the method, the NKF is adopted to effectively remove linear echoes to obtain error signals after preliminary echo cancellation. In the second stage, grouping time convolution of EMA and a DPERNN network are combined to make full use of information in the first stage, and deeper feature extraction and voice signal reconstruction are carried out. In order to better reduce redundant information and improve the capability of a network to eliminate echoes, a TRA-based filter is designed, and the relation between a real part and an imaginary part in a complex spectrum is enhanced, so that the quality of reconstructed voice is improved.
Owner:GUANGZHOU MARITIME INST +2

Method and apparatus for performing speech enhancement, storage medium, device, and product

PendingUS20260004788A1Speech analysisNeural learning methodsPhonetic environmentNoise
A speech enhancement method, apparatus, and computer-readable storage medium for training neural networks to enhance speech quality. The method obtains a training set containing training samples, each comprising a sample reference speech, a sample comparison speech from the same sound-producing object, and a mixed speech combining interfering human voice, ambient noise, and the sample comparison speech. Sample voiceprint vectors are extracted from reference speech and sample audio features from mixed speech. A speech enhancement network processes these inputs to output predicted audio features, which are compared against comparison audio features to determine training loss values. The network's weight parameters are iteratively updated based on these loss values until training completion, enabling effective speech enhancement through voiceprint-guided processing.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Methods, devices, electronic equipment, and storage media for processing speech signals

This application proposes a method, apparatus, electronic device, and storage medium for processing speech signals. The method includes: acquiring first subframe speech data of the current environment and at least one historical subframe speech data preceding the first subframe speech data; predicting second subframe speech data following the first subframe speech data based on the first subframe speech data and / or at least one historical subframe speech data; concatenating the first subframe speech data and the second subframe speech data to obtain a first frame speech data; concatenating target historical subframe speech data from at least one historical subframe speech data with the first frame speech data to obtain a second frame speech data; and obtaining a target human voice signal from the first subframe speech data based on the first frame speech data and the second subframe speech data. This method eliminates the need for a set delay to wait for the acquisition of the second subframe speech data, avoids the time delay introduced by processing inter-frame overlapping data, reduces reverberation, and improves speech quality.
Owner:XIAOMI TECH (WUHAN) CO LTD +2

Voice quality testing including codec rate change detection

A disclosed method may include (i) detecting a codec rate change during a voice quality test of a testing UE that emulates a subscriber UE connecting to a base station in a mobile network such that a test segment during which the codec rate change occurred is identified and (ii) reducing, based on detecting that the codec rate change occurred during the test segment, a weight of the test segment within a voice quality assessment.
Owner:BOOST SUBSCRIBERCO LLC

Voice generation method and device based on preference alignment, equipment and medium

The invention relates to the technical field of speech synthesis, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a speech generation method and device based on preference alignment, equipment and a medium, and the method comprises the steps: obtaining a pre-training speech generation model and a preference training sample pair, constructing a preference alignment model and a non-preference alignment model, and taking the pre-training model as a reference; determining a loss function based on the preference samples and the non-preference samples, and updating model parameters to obtain a trained model; and receiving a target condition item and generating an agent prompt in a reasoning stage, fusing speed prediction results of the two types of models, and generating a target voice based on a fusion result in a stream matching process. According to the method, human preference signals are fused through a preference alignment mechanism, so that the naturalness and semantic consistency are considered in the speech generation process, and the speech quality and the personalized expression ability are improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Hearing aid control method and system for eliminating ear blockage effect

PendingCN121568021AHearing aids signal processingOcclusion effectSpeech sound
A hearing aid control method for eliminating an ear blockage effect comprises the following steps: acquiring an audio decibel recorded by a first microphone as a first decibel; the decibel interval where the first decibel falls is judged, the access mode of the microphones is selected based on the decibel interval, and the first access mode is that the first microphone and the second microphone are both electrically connected with the audio access interface, and the valve of the air hole of the hearing aid is in a closed state; the second access mode is that the first microphone is connected with the hearing aid control chip; and when the second access mode is executed, the valve opening degree of the hearing aid air hole is controlled based on the first decibel, and whether the second microphone is electrically connected with the audio access interface is controlled. According to the invention, the automatic opening and closing of the air hole are realized, while the ear plugging effect and the ear canal damp-heat problem are alleviated, the sound quality is ensured and the howling is effectively inhibited through an acoustic management mechanism of software and hardware collaboration, so that the hearing aid can give consideration to both comfort and definition in different use scenes.
Owner:FOSHAN VOHOM TECHNOLOGY CO LTD +1

Call voice real-time noise reduction method and system based on dynamic noise perception

The application discloses a call voice real-time noise reduction method and system based on dynamic noise perception, acquires a noisy original voice signal, carries out frame division and windowing pretreatment on the digital voice signal with noise, extracts the spectral characteristics of the noise, divides the noise into steady-state noise, non-steady-state noise and burst noise according to the spectral characteristics of the noise, constructs a lightweight neural network structure through model pruning and quantization, adjusts the parameters and strategies of the types of the steady-state noise, the non-steady-state noise and the burst noise to realize optimized reduction, carries out post-processing such as window removal and overlap addition on the voice signal after noise reduction, and outputs the voice signal after noise reduction; the application effectively removes the background noise in a complex noise environment through dynamic noise perception and an adaptive noise reduction strategy, significantly improves the voice quality, meets the real-time requirement, realizes real-time noise reduction processing under low delay conditions through lightweight model design, and meets the real-time requirement of voice communication and voice control.
Owner:JIANGXI RUI TECH CO LTD

Voice quality evaluation system and method

The invention relates to the technical field of voice signal processing, and provides a voice quality evaluation system and method. The system comprises an acoustic feature extraction module used for extracting acoustic information from an input voice signal and outputting a first time sequence feature; the semantic feature extraction module is used for extracting linguistic content and semantic information from the voice signal and outputting a second time sequence feature; the bidirectional cross-attention fusion module is used for carrying out interactive fusion on the first time sequence feature and the second time sequence feature to generate a cross-enhanced feature vector; and the score prediction module is used for predicting a voice quality score according to the cross enhancement feature vector. According to the method and the device, through the acoustic branch, the semantic branch and a bidirectional cross-attention fusion mechanism, deep interactive modeling of acoustic fidelity and semantic content is realized, systematic deviation in the prior art is effectively reduced, and the accuracy and generalization ability of a voice quality evaluation model in diversified and unknown scenes are remarkably improved.
Owner:INST OF ACOUSTICS CHINESE ACAD OF SCI

Dual-group human voice synchronous acquisition frequency domain noise reduction and restoration method and system

The application relates to the technical field of voice communication, and discloses a frequency domain noise reduction and restoration method and system for double-group human voice synchronous acquisition, wherein the method acquires a synchronous signal through double-group acquisition points, constructs a double-source frequency domain coordinate system and maps the double-source frequency domain coordinate system into energy distribution points; a spatial relationship is calculated to identify an emotional distortion state, four quadrants are divided to generate distortion type marks; the frequency domain signals are segmented according to the marks, targeted reconstruction and correction are performed, and finally, the energy of a frequency band is enhanced to significantly reduce a part; the application can accurately process voice composite distortion under extreme emotions, and provides reliable voice quality guarantee for accurate communication of key information in an emergency communication scene.
Owner:GUANGZHOU CMX AUDIO CO LTD

Robust speech enhancement method based on adaptive beam forming and sparse spectrum constraint

The invention discloses a robust speech enhancement method based on adaptive beam forming and sparse spectrum constraint. The robust speech enhancement method comprises the following steps: receiving a multichannel observation signal through a microphone array and constructing a signal model of a generalized sidelobe canceller structure; based on the signal model, a beam forming optimization model is constructed in combination with the least square criterion and the speech spectrum sparse constraint; performing iterative solution on the beam forming optimization model by adopting a Lagrange multiplier alternating direction method to obtain an adaptive filter weight vector in a generalized sidelobe canceller; and calculating and outputting an enhanced voice signal based on the weight vector. According to the method, the voice quality and intelligibility can be effectively improved in a strong interference and reverberation environment, interference and reverberation are remarkably suppressed, and the robustness of an algorithm is improved.
Owner:SOUTHWEAT UNIV OF SCI & TECH