Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

258 results about "Speech quality" patented technology

Measuring Speech Quality. In telecommunication, speech quality is an important contributing factor to the success of a product and to the success of the communication itself. High speech quality guarantees that the effort the users have to put forward in order to correctly perceive the communication is low.

Video translation method and system based on artificial intelligence

The invention discloses a video translation method and system based on artificial intelligence. The method relates to the technical field of video translation and comprises the following steps of original sound track extraction, target AI speaker adaptation, AI dubbing generation and mouth shape synchronization and video synthesis. According to the method, independent audio and video streams are obtained by adopting an audio and video separation technology, and multiple original sound tracks are extracted through a voice separation model; matching or generating an adaptive target AI speaker module in a preset tone library; converting the original language voice into a text, translating the text into a target language text, and synthesizing an AI dubbing audio track in combination with a target AI speaker module; and finally, the independent video stream and the multi-AI dubbing audio track are input into the mouth shape synchronization model to output a translated video, so that the timbre fitting degree, the voice quality and the voice consistency of the same speaker of AI dubbing are improved, and meanwhile, the resource utilization rate of video translation and the processing efficiency under batch tasks are improved. The problem that in the prior art, video translation is low in quality and efficiency is solved.
Owner:BEIJING DEEP LOGIC INTELLIGENT TECHNOLOGY CO LTD

Voice quality inspection method and device, computer equipment and storage medium

The invention discloses a voice quality inspection method and device, computer equipment and a storage medium, belongs to the technical field of artificial intelligence, and is applied to voice quality inspection scenes in the fields of finance, health medical care, old-age care and the like. The multi-modal feature fusion technology is introduced, the context semantic features of the text and the acoustic features of the voice are extracted, the emotion features are obtained by combining the pre-trained emotion recognition model, comprehensive understanding of the voice data from the three dimensions of semantics, acoustics and emotions is achieved, the deep fusion of the three feature vectors is carried out, and the voice recognition efficiency is improved. Compared with a traditional method which only depends on text or acoustic features, the method has the advantages that multi-modal features are realized by combining emotional features on the basis of the text or acoustic features, and information contained in voice content can be reflected more comprehensively and meticulously, so that the accuracy and practicability of voice quality inspection are improved, and the voice quality inspection efficiency is improved. And the requirements of application scenes such as intelligent customer service and voice auditing on high-quality automatic quality inspection are met.
Owner:PING AN TECH (BEIJING) CO LTD

Implementation method and device of multi-channel voiceprint recognition system

The invention relates to the technical field of voice recognition, in particular to an implementation method and device of a multi-channel voiceprint recognition system, and the implementation method comprises the steps of multi-channel data acquisition and synchronization, signal preprocessing and enhancement, feature extraction and fusion, model training, real-time deployment and adaptive optimization. Compared with the problems that a traditional multichannel voiceprint recognition system depends on a fixed beam forming algorithm and an independent clock synchronization module, the synchronization error is large, manual parameter adjustment is needed for noise suppression, and generalization is poor, hardware-level clock synchronization is achieved through a PTP protocol, and the accuracy of noise suppression is improved. The method combines an end-to-end neural network to automatically learn noise distribution and a sound source space position, dynamically generates a beam forming weight, can improve the voice quality in a complex noise scene without manual intervention, remarkably reduces the interference of a synchronization error on sound source positioning, and enables the precision and stability of far-field voice enhancement to reach a new level.
Owner:MINAMI ACOUSTICS LTD

Two-step mixed sound source separation and de-reverberation method

The invention relates to the technical field of voice signals, in particular to a two-step mixed sound source separation and de-reverberation method, which comprises the following steps of: firstly, carrying out separation network training by taking different types of signals as training targets in various separation networks for separating attention and the like; the invention provides an improved de-reverberation method based on a time convolution network-weight prediction error, multiple improvement strategies such as taking a scale invariance signal-to-noise interference ratio as a network loss function, adopting a transposition mechanism for an input signal and a mechanism for additionally adding a residual value to a network unit are used, and finally, the de-reverberation method based on the time convolution network-weight prediction error is obtained. And cascading the separation attention network with the best separation noise reduction effect with the time convolution-transpose-residual error-WPE network. According to the invention, the problem that the separation effect of the mixed audio signal with reverberation and noise signals is not good under the actual sound field condition is solved, compared with a single one-step or two-step existing separation, de-mixing and noise reduction network, the method is significantly improved, and the quality of the separated voice can be improved for later recognition and discrimination.
Owner:ZHONGBEI UNIV

Method and device for converting lip language into voice, computer storage medium and terminal

The invention discloses a lip language-to-voice conversion method and device, a computer storage medium and a terminal, and aims to solve the problems that a lip language-to-voice conversion technology cannot be deployed on terminal equipment and voice quality cannot meet application requirements. The method is combined with a neural vocoder which reduces calculation complexity and resource requirements, system configuration requirements are reduced while system parameters are reduced, a design basis is provided for deploying a lip language-to-speech method on terminal equipment, and context-related visual feature sequences and user audio embedding vectors are fused, so that the user audio-to-speech conversion efficiency is improved, and the user audio-to-speech conversion efficiency is improved. A Mel spectrum acoustic feature sequence used for being converted into a voice waveform is obtained, and a user audio vector fused with the Mel spectrum acoustic feature sequence is obtained, so that the output voice waveform is more consistent with the real voice of a user; according to the embodiment of the invention, technical support is provided for deploying and applying the lip language-to-speech method meeting the speech quality requirement on the terminal equipment.
Owner:BEIJING WATERTEK INFORMATION TECH

Far-field single-channel speech enhancement method

The invention relates to the technical field of speech enhancement, in particular to a far-field single-channel speech enhancement method based on an MFSE (Maximum Free Square Error) (Maximum Free Square Error) (Maximum Free Square Error) (Maximum Free Square Error) (Maximum Free Square Error) (Maximum Free Square Error)). The method comprises the following steps: step 1, processing a far-field voice signal to obtain a complex spectrogram of a noise voice signal; 2, inputting the compressed complex spectrogram into a feature encoder, processing the output of the complex spectrogram by N MamAttention blocks, and then sending the processed complex spectrogram into an amplitude mask decoder and a phase decoder to respectively predict a clean compressed amplitude mask and a phase spectrum; step 3, preheating and training the MamAttention model, and performing supervised confrontation training by taking the MamAttention model as a generator and the multi-resolution discriminator as a discriminator; and step 4, inputting test voice into the trained model to realize far-field single-channel voice enhancement. According to the method, the supervised adversarial training strategy and the MamAttention model are combined, so that the problems of signal attenuation, noise and reverberation interference in far-field voice are effectively solved, and the voice quality is remarkably improved in a scene that the distance of a loudspeaker exceeds 5 meters.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Speech recognition method and electronic equipment

The invention provides a voice recognition method and electronic equipment, and the method comprises the steps: the electronic equipment obtains source sound data in response to a wake-up operation of a user on a first application; determining a target voiceprint template based on the source sound data; the target voiceprint template is a voiceprint template with the highest matching degree with the source sound data in voiceprint templates of users corresponding to the electronic equipment; as the target voiceprint template is the voiceprint template which is most matched with the voiceprint of the user during the current speaking, the voice quality of the obtained target voice data can be improved by obtaining the target voice data according to the source voice data and the target voiceprint template, and then voice recognition is performed based on the target voice data to obtain the voice recognition result, so that the user experience is improved. The accuracy of a speech recognition result can be improved.
Owner:HONOR DEVICE CO LTD

Communication method and system for realizing private call

The invention provides a communication method for realizing a private call, which comprises the following steps that: a wireless silencer preprocesses acquired voice to improve the voice quality; establishing communication connection between the earphone and the wireless silencer; the earphone receives the voice data sent by the wireless silencer; the microphone of the earphone is turned off, and the earphone is switched to an audio output mode; the wireless silencer is worn at the mouth of a speaker so as to collect voice sent by the speaker, and first audio information obtained by converting the voice is sent to the earphone; the earphone transmits the first audio information to a call device; the earphone receives the second audio information from the call device, converts the second audio information into sound and transmits the sound to human ears; in the conversation process, the wireless silencer shields the voice sent by the speaker so as to prevent the voice from being spread to the surrounding environment, and private conversation is achieved.
Owner:ZHONGKE HUAYI (SHENZHEN) INTELLIGENT TECHNOLOGY CO LTD

Voice quality optimization method and device, equipment, storage medium and product

The invention discloses a voice quality optimization method and device, equipment, a storage medium and a product, and relates to the technical field of communication, and the method comprises the steps: based on an RSRP sequence of each user equipment, recognizing each fast fading equipment with an RSRP fluctuation value exceeding a preset fluctuation threshold value from each user equipment through a preset recurrent neural network; determining a communication cell to which each fast fading device belongs according to the temporary identifier of each fast fading device, and determining the communication cell meeting a preset device scale condition and a fast fading device distribution condition as a fast fading cell; determining a voice problem cell from the fast fading cells according to the service quality identifier of each fast fading device in the fast fading cells; and voice quality optimization is carried out on voice problem equipment in the voice problem cell, and the voice problem equipment is fast fading equipment for executing voice services. Therefore, the influence of the fast fading phenomenon on the voice quality is reduced, and the voice quality is improved.
Owner:CHINA MOBILE GROUP JILIN BRANCH +1

Driving safety analysis method based on voice emotion recognition and related device

The invention provides a driving safety analysis method based on voice emotion recognition and a related device, and the method comprises the steps: obtaining a driving voice data flow collected by a target vehicle in a driving environment, carrying out the emotion feature extraction of the driving voice data flow, and generating a multi-level emotion feature set; inputting the multi-level emotion feature set into a safety risk prediction model, generating a safety behavior scoring curve of the driver in a preset time period, and generating driving intervention prompt information corresponding to risk nodes according to the risk nodes exceeding a preset threshold in the safety behavior scoring curve, and the driving intervention prompt information is sent to the vehicle-mounted terminal of the target vehicle for real-time display. According to the invention, hardware deployment cost can be reduced, interference of voice quality fluctuation on an analysis result in a complex driving environment is ensured to be effectively suppressed, and effects of improving driving safety early warning accuracy and reducing traffic accident rate are realized.
Owner:GUIZHOU UNIVERSITY OF FINANCE AND ECONOMICS

Voice quality detection and evaluation method based on multi-modal fusion

The invention relates to the technical field of voice quality detection, in particular to a voice quality detection and evaluation method based on multi-modal fusion. The method comprises the following steps: carrying out short-time Fourier transform on a multi-modal fused noisy speech signal to obtain a plurality of noisy speech spectrums, and constructing a complex spectrum matrix of the noisy speech; calculating based on the complex spectrum matrix of the noisy voice to obtain an actual value voice feature matrix corresponding to the multi-modal fused noisy voice signal, inputting the actual value voice feature matrix into the multi-modal fused voice signal reconstruction analysis model, and outputting an optimal actual value voice feature; training a deep network by taking the optimal real-value speech feature as a target to realize speech enhancement; the prior signal-to-noise ratio fusing the specific person information is calculated based on the enhanced voice signal, and quality detection and evaluation are performed on the voice signal based on the prior signal-to-noise ratio, so that the reliability and the accuracy of performing quality detection and evaluation on the multi-modal fused voice signal can be improved.
Owner:JIANGSU BAIYING INFORMATION TECH CO LTD

Unsupervised learning voice quality evaluation method and device, equipment and medium

The invention relates to the technical field of intelligent voice, can be applied to the fields of finance and medical treatment, and discloses an unsupervised learning voice quality evaluation method, device, equipment and medium, and the method comprises the steps: extracting multi-dimensional voice features in a voice signal; constructing a voice feature model based on an unsupervised learning auto-encoder structure, learning quality representation of voice features, comprehensively evaluating the quality of the multi-dimensional voice features according to a learning result, and generating a quality score of the voice features; comparing the clear voice data, the distorted voice data and the unlabeled voice data, and optimizing the voice feature model according to a comparison result; and training the optimized voice feature model, obtaining the difference between the unlabeled voice data and the voice features in the optimized voice feature model, carrying out fine tuning on the trained voice feature model by taking the difference as a fine tuning demand, and outputting a new quality score of the voice features.
Owner:PING AN TECH (SHENZHEN) CO LTD

Multi-speaker speech synthesis method and device, storage medium and computer equipment

The invention relates to the technical field of multi-speaker speech synthesis, and particularly provides a multi-speaker speech synthesis method and device, a storage medium and computer equipment. According to the method, customized structure expansion, engine transformation and optimization are carried out on the basis of an open source reasoning framework vLLM, so that the target vLLM framework obtained through transformation can face a speech synthesis scene, and efficient reasoning of a speech synthesis model is supported. Speech synthesis reasoning is realized by using the target vLLM framework and the LLM-based target speech synthesis model, so that speech synthesis can be accelerated under the condition of keeping speech quality.
Owner:SHANGHAI LINGGUANG ZHAXIAN TECHNOLOGY CO LTD

Voice quality inspection method and device, equipment, storage medium and program product

The embodiment of the invention discloses a voice quality inspection method and device, equipment, a storage medium and a program product. The method comprises the following steps: acquiring voice data; performing emotion recognition on the voice data based on the acoustic features of the voice data to obtain an emotion recognition result; performing voice recognition on the voice data to obtain voice text data; and constructing a first prompt word according to the emotion recognition result and the voice text data, and processing the first prompt word by using a pre-trained text quality inspection model to obtain a voice quality inspection result. According to the embodiment of the invention, the quality inspection accuracy of the voice data can be improved.
Owner:CHINA MOBILE SHANGHAI ICT CO LTD +2

Zero sample speech synthesis method and device, computer equipment and storage medium

The invention relates to a zero-sample speech synthesis method and device, computer equipment, a storage medium and a program product, and the method comprises the steps: obtaining a target coding feature according to a reference speech and a target text; inputting the target coding features into a stream matching model to obtain a conditional velocity field and an unconditional velocity field; inputting the target coding feature into a prior model to obtain a prior speech feature; obtaining a prior generation flow field according to the prior voice features and the standard Gaussian noise; calculating a KL divergence value between the priori generated flow field and a preset real generated flow field, and taking a moment when the KL divergence value is smaller than or equal to a preset KL divergence threshold as an initial moment of the priori generated flow field; fusing the conditional velocity field and the unconditional velocity field to obtain a fused velocity field; and inputting the fusion velocity field from the initial moment to the target moment and the priori generated flow field into an ordinary differential equation solver to obtain the target speech features, thereby improving the speech quality of the synthesized speech.
Owner:CHINA SOUTHERN POWER GRID ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD

Voice wake-up interaction method and system based on microphone array

The invention discloses a voice wake-up interaction method and system based on a microphone array, and the method comprises the steps: carrying out VAD processing, so as to judge whether a target audio segment has voice or not; voice and noise source directions are obtained; selecting beam parameters of voice and noise in combination with a pre-designed fixed beam; whether GSC module processing is carried out or not is selected according to the difference between the beam parameters of the voice and the noise, so that an enhanced audio signal is obtained, a wake-up task is carried out, and a final voice wake-up result is obtained; and according to whether the wake-up is successful, determining whether to lock the voice beam direction in the current interaction stage for enhancement, thereby preventing interference of voice in other directions on subsequent interaction tasks. According to the voice enhancement mode based on the microphone array, sound source positioning can be realized, interference in a non-target direction can be suppressed, the voice quality in the target direction can be improved, the wake-up success rate can be effectively improved when the voice enhancement mode is applied to voice wake-up, and then the experience of back-end voice interaction is improved.
Owner:PANOVASIC TECHNOLOGY CO LTD

Frequency domain noise reduction and reduction method and system for synchronously collecting double groups of human voices

The invention relates to the technical field of voice communication, and discloses a frequency domain noise reduction and reduction method and system for double-group human voice synchronous acquisition, and the method comprises the steps: obtaining a synchronous signal through double groups of acquisition points, constructing a double-source frequency domain coordinate system, and mapping the coordinate system into an energy distribution point; calculating a spatial relationship to identify an emotional distortion state, and dividing four quadrants to generate a distortion type mark; dividing a frequency domain signal and executing targeted reconstruction and correction, and finally dividing a frequency band enhanced energy remarkably-reduced part; according to the method, the voice composite distortion under the extreme emotion can be accurately processed, and reliable voice quality guarantee is provided for accurate transmission of key information in an emergency communication scene.
Owner:GUANGZHOU CMX AUDIO CO LTD

Inference acceleration method for stream matching TTS model, electronic equipment and storage medium

The embodiment of the invention discloses a reasoning acceleration method for a stream matching TTS model, electronic equipment and a storage medium, and the method comprises the steps: carrying out the reasoning of input data through a TTS model, and recording a sampling track; performing qualitative analysis on the recorded sampling track to form a sampling strategy after preliminary pruning; reasoning the input data again by the TTS model by using the sampling strategy after the preliminary pruning; evaluating the quality of the generated voice in a mode of combining objective indexes and subjective evaluation, and judging whether the quality of the generated voice meets a preset standard or not; when the sampling step number reaches the standard, judging whether the sampling step number can be further reduced on the current basis or not; if the sampling step number can be further reduced, the previous reasoning and analysis processes are repeated after corresponding steps are reduced; when the generation quality does not reach the standard, returning part of the steps which are pruned before, and adjusting a sampling strategy; and when the voice quality reaches the standard and the step number cannot be further reduced, storing a sampling strategy for the TTS model.
Owner:SHANGHAI JIAOTONG UNIV

NPU-based Chinese-English bilingual text-to-speech conversion method and system

The invention discloses a Chinese-English bilingual text-to-speech conversion method and system based on NPU, and belongs to the technical field of speech processing, and the method comprises the steps: carrying out the word segmentation processing and phoneme conversion of a Chinese-English mixed text based on the language type of each segment in the Chinese-English mixed text, and obtaining a text input sequence; performing vector combination on the phoneme ID sequence and the language ID sequence, and extracting text semantic features from a vector combination result to obtain text hidden variables; according to the text hidden variable, predicting the duration of each phoneme and a prior distribution parameter in an acoustic potential feature space; performing multi-level transformation on the phoneme alignment result and the prior distribution parameter to obtain a bilingual acoustic feature sequence; and converting the bilingual acoustic feature sequence into a corresponding voice waveform. Therefore, by implementing the method, the device and the system, the problems that the Chinese-English bilingual text occupies more resources in the voice conversion process and the output voice quality is lower in the prior art can be solved.
Owner:GUANGZHOU BAOLUN ELECTRONICS CO LTD

Intelligent speech synthesis system based on electronic artificial throat audio input signal

PendingCN120412534ASpeech recognitionSpeech synthesisLarynxElectrolarynx
The invention discloses an intelligent speech synthesis system based on an electronic artificial throat audio input signal. The intelligent speech synthesis system is characterized by comprising a signal processing module, a speech recognition module, a speech synthesis module and a user interface module. Compared with the prior art, the method has the advantages that STFT is used for converting the electronic throat signals into frequency domain features, meanwhile, a Conformer model is adopted for semantic extraction, natural voice is generated in combination with FastSpeech2 and Hi Fi-GAN, and emotion expression and voice quality are improved.
Owner:GUANGXI UNIV FOR NATITIES +1

Speech synthesis style transfer method and device, computer device and storage medium

The present application relates to the field of speech synthesis, and particularly relates to a speech synthesis style migration method and device, computer equipment and a storage medium. The method comprises the following steps: inputting the text data to be converted into a speech synthesis encoder to obtain text encoding; obtaining multi-level speech style representation of the speech to be migrated; inputting the text data to be converted and the multi-level speech style representation into a style predictor to obtain a predicted speech style; performing multi-style layer regularization on the text encoding and the predicted speech style to obtain style-regularized text encoding; inputting the style-regularized text encoding into a pitch predictor to obtain rhythm change data; inputting the rhythm change data and the style-regularized text encoding into a speech synthesis decoder to obtain a mel-frequency spectrum diagram and generate target speech. The present application considers rhythm information and the style of words and phonemes, and the overall rhythm style of the speaker, thereby improving the speech quality of the target speech.
Owner:PING AN TECH (SHENZHEN) CO LTD

An Automatic Mongolian Speech Quality Assessment Method Based on Hierarchical Transfer Learning

The present invention discloses a Mongolian automatic speech quality assessment method based on hierarchical transfer learning, which includes the following steps: pre-training an English speech self-supervised model and an English speech quality assessment model to obtain a trained English speech self-supervised model and a trained English speech quality assessment model; performing transfer learning on the trained English speech self-supervised model and the trained English speech quality assessment model to obtain a trained self-supervised model and a trained speech quality assessment model; using the trained speech self-supervised model and the BERT model to extract the feature vectors of Mongolian speech and the text features in the corresponding text; fusing the speech features and the text features into sentence-level semantic features f, and sending f into the trained speech quality assessment model to obtain the MOS score z corresponding to the speech signal, completing the automatic assessment of Mongolian speech quality, pioneering an automatic assessment method for Mongolian speech quality and filling the gap in this field.
Owner:INNER MONGOLIA UNIVERSITY

Audio signal processing method, device, equipment and storage medium

The present disclosure relates to an audio signal processing method, apparatus, device and storage medium. The present disclosure obtains a first feature map by encoding a first complex spectrum map corresponding to the original audio signal, and obtains a second feature map by processing the time series and frequency series corresponding to the first feature map respectively. Based on the second feature map, complex ratio masking and complex spectrum mapping are simultaneously learned. Thus, the amplitude spectrum and phase of the original audio signal are enhanced simultaneously by combining masking prediction and spectrum prediction, thereby improving the speech enhancement effect. Speech enhancement can greatly improve speech quality, solve noise interference problems, and improve speech recognition effects.
Owner:ALIBABA DAMO (HANGZHOU) TECH CO LTD

Audio processing method and device, server and storage medium

The invention provides an audio processing method and device, a server and a storage medium. The method comprises the following steps: receiving a to-be-adjusted original audio and a corresponding line; converting the lines into a corresponding text sequence; performing alignment operation on the original audio and the text sequence to obtain an audio-text corresponding relation; receiving a target text; determining a to-be-replaced first text from the text sequence according to the target text and frame information of the first text in the audio in the corresponding relation; replacing the first text with the target text by using the frame information to obtain an updated text sequence; and generating a new audio by using the updated text sequence. Therefore, the problems that the edited voice is stiff, the complex audio effect is poor, and the quality is possibly poor if the voice is modified too much can be solved, and the quality of the generated voice can be better controlled by controlling the durations of different granularities.
Owner:CHENGDU IQIYI INTELLIGENT INNOVATION TECH CO LTD

Call voice real-time noise reduction method and system based on dynamic noise perception

The invention discloses a call voice real-time noise reduction method and system based on dynamic noise perception. The method comprises the following steps: collecting an original voice signal with noise; performing framing and windowing preprocessing on the digital voice signal with noise; extracting spectrum features of the noise, and dividing the noise into steady-state noise, unsteady-state noise and burst noise according to the spectrum features of the noise; quantitatively constructing a lightweight neural network structure through model pruning; if yes, performing parameter and strategy longitude adjustment on the types of the steady-state noise, the unsteady-state noise and the burst noise to realize optimization reduction; performing post-processing of window removal and overlapping addition on the voice signal after noise reduction, and outputting the voice signal after noise reduction; according to the method, background noise in a complex noise environment is effectively removed through dynamic noise perception and a self-adaptive noise reduction strategy, and the voice quality is remarkably improved; real-time noise reduction processing is achieved under the low-delay condition through lightweight model design, and the real-time requirements of voice communication and voice control are met.
Owner:JIANGXI RUI TECH CO LTD

Voice processing method and device, equipment, storage medium and program product

The embodiment of the invention provides a voice processing method and device, equipment, a storage medium and a program product. The method comprises the following steps: acquiring a voice signal; performing signal processing on the voice signal to obtain a first time frequency signal; performing adaptive normalization processing on the first time-frequency signal to obtain a first feature map; performing two-dimensional modeling processing on the first feature map to obtain a second feature map; and performing synthesis processing on the second feature map to obtain an enhanced voice signal. The method can improve the voice quality.
Owner:RDA CHONGQING MICROELECTRONICS TECH CO LTD

System

An object of a system according to an embodiment is to analyze biological data of participants and automatically organize an effective breakout room.SOLUTION: A system includes a biological data collection section, an analysis section, a classification section, and an organization section. The biological data collection unit collects biological data such as facial expressions and voice quality from a camera or a microphone of a terminal of a participant. The analysis section analyzes the biological data collected by the biological data collection section. The classification component classifies the participants by character on the basis of the data analyzed by the analysis component. The organization unit automatically organizes a breakout room based on the characters classified by the classification unit.SELECTED DRAWING: Figure 1
Owner:SOFTBANK GROUP CORP

Ultra-low computing resource speech enhancement method based on Half-UNet architecture

The invention discloses an ultra-low computing resource speech enhancement method based on a Half-UNet framework. According to the method, a decoder of a UNet is simplified, a recurrent neural network module is arranged between an encoder and the decoder to construct a Half-UNet architecture, and in combination with feature fusion, adaptive frequency band division and power law compression penalty technologies, the calculation complexity and parameter quantity are greatly reduced while enhancing performance is ensured. The method comprises the following steps: carrying out frequency band combination on frequency spectrums of input noise voice by using a filter obtained by training of a self-adaptive frequency band division module, and reducing high-frequency characteristic redundancy; carrying out feature extraction by using Half-UNet, and reconstructing a frequency spectrum; a power law is used to compress penalty terms to enhance weak detail features, and the weak detail features are prevented from being submerged by strong noise features, so that the voice quality of the model in a low signal-to-noise ratio environment is improved. The method has the advantages that under the condition that only about 23.7 K parameters and 25.42 MMACs operand are needed, the voice enhancement effect equivalent to that of a large-scale deep model is achieved, and the method is particularly suitable for resource-limited equipment such as earphones and hearing aids.
Owner:EAST CHINA NORMAL UNIV +1

radio

PendingJP2026043102ATransmissionNoiseSpeech sound
To provide a radio capable of performing squelch control with little erroneous determination regardless of the voice quality of a speaker. [Solution] The radio 20 in this example comprises an SSB demodulation unit 21 that performs SSB demodulation processing on the received signal, a noise cancellation unit 22 that performs noise cancellation processing on the audio signal resulting from demodulation by the SSB demodulation unit 21, a pitch period calculation unit 23 that calculates the pitch period of the audio from the audio signal after noise cancellation processing, and a squelch control unit 24 that controls the opening and closing of the squelch of a squelch circuit 25 based on the pitch period calculated by the pitch period calculation unit 23.
Owner:KOKUSAI DENKI ELECTRIC INC

Voice quality evaluation method and device, equipment and storage medium

The invention discloses a voice quality evaluation method and device, equipment and a storage medium, and relates to the technical field of computers. The method comprises the steps of obtaining to-be-evaluated voice generated through text-to-voice conversion and style information for the to-be-evaluated voice, and extracting basic features of the to-be-evaluated voice; extracting text-to-speech features from the basic features; the text-to-speech features comprise any one or more of rhythm features, tone features and text matching features; calculating a style matching degree between an actual style embedding vector corresponding to the text-to-speech feature and a style embedding vector template corresponding to the style information; and predicting a speech degradation score based on the text-to-speech features, and determining speech quality based on the speech degradation score and the style matching degree. By extracting rhythm features, timbre features and text matching features, identifying specific quality defects in text-to-speech conversion; and in combination with the style matching degree, the problem of confusion of style differences and quality defects is solved, and the accuracy of voice quality scoring is improved.
Owner:MALANSHAN AUDIO & VIDEO LABORATORY