Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

55 results about "Speech spectrum" patented technology

Speech spectrum - the average sound spectrum for the human voice. acoustic spectrum, sound spectrum - the distribution of energy as a function of frequency for a particular sound source.

Information processing method and system for voice conversion

The invention relates to the technical field of voice signal processing and synthesis, and particularly discloses an information processing method and system for voice conversion, and the method comprises the steps: extracting a language feature vector of each word from an input text; calculating posterior probabilities of the words in different languages by combining a Bayesian reasoning mechanism, and generating language attribution confidence coefficient characteristic values; a Monte Carlo sampling method is adopted to carry out multiple times of context sensitive simulation, and pronunciation path selection probability distribution characteristic values are generated; further fusing the feature values into a multi-language pronunciation decision vector, inputting the multi-language pronunciation decision vector into a multi-language end-to-end speech synthesis model, calling a phoneme mapping rule of a corresponding language and an acoustic parameter prediction module, and generating a high-quality target speech spectrogram; and finally, dynamically adjusting Bayesian prior distribution and a language recognition threshold according to the output speech spectrogram.
Owner:SHANDONG POLYTECHNIC COLLEGE

Electric power emerging business risk keyword identification and early warning method, system and device based on NLP technology and medium

The invention discloses an electric power emerging business risk keyword identification and early warning method, system and device based on an NLP technology and a medium, and belongs to the technical field of electric power risk, and the method comprises the steps: obtaining multi-source customer service text data, and extracting risk keywords and corresponding voice spectrum data from the multi-source customer service text data; inputting the risk keyword into a power service knowledge graph, and generating a first risk research and judgment result through knowledge reasoning, the first risk research and judgment result comprising a risk traceability path; performing acoustic feature analysis on the speech spectrum data, extracting acoustic feature parameters, calculating an emotional excitation index, and generating a second risk identification result; performing weighted fusion on the first risk research and judgment result and the second risk identification result to generate a comprehensive risk score; and when the comprehensive risk score reaches a preset threshold value, generating a structured early warning work order, and carrying out data binding on the structured early warning work order, the original customer service work order and the call record. According to the invention, the efficiency of risk management and control is improved.
Owner:GUIZHOU POWER GRID CO LTD

Voice coding and decoding method, device, equipment and medium

ActiveCN121054006ASpeech analysisProduct quantizationDecoding methods
The invention relates to the technical field of artificial intelligence, can be applied to the fields of financial science and technology and medical science and technology, and discloses a voice coding and decoding method, device, equipment and medium. A continuous vector is obtained by encoding the data through an encoder trained by a first-stage mirror image architecture; segmenting a continuous vector into sub-vectors, matching the sub-vectors with corresponding sub-coding dictionaries through a quantizer to obtain sub-indexes, and combining the sub-indexes to generate an overall index; analyzing the overall index to obtain sub-indexes, calling sub-discrete vectors and splicing the sub-discrete vectors into a discrete vector; and reconstructing the discrete vector through a decoder trained by a second-stage non-mirror-image architecture to obtain a target speech spectrum feature, and converting the target speech spectrum feature into a target speech signal. According to the method, double-stage training is adopted, the first-stage mirror image architecture guarantees the coding stability, the second-stage non-mirror image architecture improves the decoding flexibility, and the product quantization technology is combined, so that the calculation efficiency and the storage overhead are balanced, and meanwhile, the voice reconstruction quality is improved.
Owner:平安科技(上海)有限公司

Speech spectrum reconstruction method and system combining time domain half-wave rectification and weighted Gaussian mixture model decoder, terminal and medium

The invention discloses a speech spectrum reconstruction method and system combining time-domain half-wave rectification and a weighted Gaussian mixture model decoder, a terminal and a medium, and relates to the technical field of speech signal processing, the method comprises the following steps: acquiring a low-resolution audio, executing time-domain half-wave rectification, and executing weighted Gaussian mixture model decoding; obtaining a mixed amplitude spectrum based on the amplitude spectrum of the low-resolution audio and the amplitude spectrum of the rectification signal; outputting GRU features based on an encoder, inputting the GRU features into a weighted Gaussian mixture model decoder, and generating a plurality of groups of Gaussian component parameters through three parallel linear layers and applying three constraint designs; and calculating a frame-level frequency distribution weight, combining the frame-level frequency distribution weight with the mixed amplitude spectrum to obtain a model output amplitude spectrum, inputting the model output amplitude spectrum and the amplitude spectrum of the low-resolution audio into a frequency band guide masking module to obtain a final amplitude spectrum, and reconstructing an audio time domain waveform based on the final amplitude spectrum and the phase of the low-resolution audio. According to the method, high-fidelity and high-frequency reconstruction can be realized while the complexity of the model is greatly reduced.
Owner:ELEVOC TECH CO LTD

Bluetooth earphone call intelligent noise reduction method based on cloud collaboration

The invention discloses a Bluetooth earphone call intelligent noise reduction method based on cloud cooperation, and the method comprises the following steps: S1, obtaining a call voice signal and an environment state parameter, carrying out the preprocessing, and constructing a voice spectrum sequence and an environment feature sequence; s2, performing noise estimation and suppression processing on the speech spectrum sequence by using a DCCRN model to obtain a spectrum feature sequence; s3, constructing a cooperative processing sequence based on the environment feature sequence and the spectrum feature sequence; s4, using a Conv-TasNet model to execute speech enhancement on the co-processing sequence, and generating a target spectrum sequence; s5, performing phase reconstruction and inverse frequency spectrum transformation based on the target frequency spectrum sequence to generate a target voice signal; and S6, constructing a loss function according to the target voice signal, and updating the Conv-TasNet model. According to the method, the DCCRN model and the like are fused, and the method has the advantages of high noise reduction precision, high environment adaptability and sustainable optimization.
Owner:深圳市美迪声科技有限公司

Radio analog telephone communication encryption method, device, equipment and medium

The invention discloses a radio analog telephone communication encryption method, device and equipment and a medium, and the method comprises the steps: generating a frequency spectrum segment number and a pseudo-random sorting sequence through a preset encryption algorithm; and carrying out controllable segmentation on the total spectrum of the analog voice according to a continuous sequence from low frequency to high frequency, carrying out sequence recombination on each sub-spectrum according to a pseudo-random sequence, finally combining and outputting encrypted voice, and recovering the original voice by utilizing the same pseudo-random sequence reverse recombination at a receiving end. Therefore, the time-frequency continuity and statistical characteristics of the voice spectrum in the traditional radio analog communication are fundamentally changed. According to the method, frequency spectrums are recombined as a whole according to a pseudo-random rule, so that an eavesdropper can only obtain disordered voice signals after the frequency spectrums are rearranged even if the eavesdropper intercepts the signals, and the difficulty of passive cracking means such as frequency spectrum analysis, filtering reduction and voice recognition is remarkably increased; the problem that in the prior art, radio analog telephone communication is insufficient in encryption and easy to crack is solved.
Owner:SHENZHEN TIANHAI COMM CO LTD

Low-power communication method and apparatus for wireless audio device, soc chip, and storage medium

PCT designated stageWO2026175268A1Source encodingData stream
A low-power communication method and apparatus for a wireless audio device, an SoC chip, and a storage medium. The low-power communication method for a wireless audio device comprises: sampling acoustic features of heterogeneous audio data from multiple target audio devices, calculating speech spectra corresponding to device characteristics, extracting multi-device acoustic features, and performing multi-frequency sub-band speech enhancement, heterogeneous segmentation and grouping of the multiple target audio devices, check value calculation, and differentiated encoding on each extracted acoustic feature sequence to obtain a multi-source encoded speech segment; dividing the multi-source encoded speech segment into transmission units, and performing cooperative signal modulation, cooperative resource allocation, and data stream transmission on each transmission unit to obtain a transmitted speech data stream; and acquiring the transmitted speech data stream transmitted with low power, and performing audio recovery on the transmitted speech data stream to obtain a target audio data stream transmitted by the multiple target audio devices, thereby reducing the communication power consumption of the multiple target audio devices and ensuring the audio transmission quality.
Owner:SHENZHEN HESHENGCHENG TECHNOLOGY CO LTD

Speech synthesis model training method, speech synthesis method, device, and product

The application relates to a speech synthesis model training method, a speech synthesis method, equipment and products. The speech synthesis model training method comprises the following steps: obtaining training speech spectrum information corresponding to a training speech sample of a training speech object; inputting the training speech spectrum information into a first coding module in a speech synthesis model to be trained, obtaining a first coding vector corresponding to the training speech spectrum information through the first coding module, and determining training vector distribution parameters corresponding to the first coding vector; obtaining a training object timbre feature corresponding to the training speech object based on the training vector distribution parameters, obtaining first synthesized speech according to the training object timbre feature and training text information corresponding to the training speech sample; obtaining a first model loss value according to the difference between the first synthesized speech and the training speech sample; and adjusting model parameters according to the first model loss value to obtain a pre-trained speech synthesis model, which can effectively improve the similarity between the timbre of the synthesized speech and the timbre of the speaker.
Owner:TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD

Multi-point recording interference method and system

The invention relates to the technical field of information security, and discloses a multi-point recording interference method and system, and the method comprises the steps: obtaining environment information, and determining a basic interference volume according to the environment information; acquiring a sound source-node distance vector and sound field uniformity, and converting the sound source-node distance vector and the sound field uniformity into a spatial distribution coefficient; acquiring a voice initial instantaneous slope and an interference-voice cross-correlation function peak value, and converting the slope and the peak value into a dynamic response coefficient; receiving a target voice, obtaining the spectral characteristics and the voice spectrum centroid of the target voice, and converting the spectral characteristics and the voice spectrum centroid into a spectral efficiency coefficient; reading the basic interference volume, the spatial distribution coefficient, the dynamic response coefficient and the spectral efficiency coefficient, and generating an adjustment instruction; according to the invention, through collaborative analysis of the environment reference coefficient, the spatial distribution coefficient, the dynamic response coefficient and the spectrum efficiency coefficient, multi-dimensional adaptive matching of the interference volume and the environment state, the spatial layout, the voice dynamic characteristic and the spectrum characteristic is realized, and the flexibility is extremely high.
Owner:BEIJING HUAZHONG CHUANGSHI TECH DEV CO LTD

Speech enhancement method based on dynamic convolution and narrowband conformer

The present application relates to the technical field of speech processing, and particularly relates to a speech enhancement method based on dynamic convolution and narrowband Conformer, which comprises a training stage and a test stage and can realize high-quality speech enhancement.The speech enhancement model proposed in the present application is composed of a generator and a discriminator, the narrowband Conformer network is used in the generator to improve the extraction ability of the model to speech spectrum information, and the dynamic convolution is further used to replace the traditional convolution, so that the parameter quantity and the calculation quantity of the model are greatly reduced, the noise reduction effect is improved, and the operation efficiency of the algorithm and the stability and reliability of the model are effectively improved.
Owner:NANJING UNIV OF POSTS & TELECOMM

Method and electronic device for reducing echo residue

The present application discloses a method and an electronic device for reducing echo residue. The method for reducing echo residue can be applied to the electronic device, and comprises: performing echo cancellation on a voice input signal according to an echo reference signal to obtain an echo cancellation signal; converting the echo reference signal into a reference spectrum signal of each frame by fast Fourier transform; converting the echo cancellation signal into a voice spectrum signal of each frame by fast Fourier transform; obtaining an a priori signal-to-noise ratio of a current frame by using the reference spectrum signal of the current frame and the voice spectrum signal of the current frame according to an additive noise principle; filtering the voice spectrum signal of the current frame by a Wiener filter coefficient of the current frame determined by the a priori signal-to-noise ratio of the current frame to obtain a target spectrum signal of each frame; and converting the target spectrum signal of each frame by inverse fast Fourier transform to obtain a target voice signal. Therefore, the method for reducing echo residue can accurately filter out residual echo.
Owner:ALI CORP

Speech signal processing method and device, electronic equipment and storage medium

The present disclosure relates to a speech signal processing method and device, electronic equipment and storage medium, and relates to the technical field of audio. The method comprises: acquiring a multi-channel far-field speech signal collected for a target sound source; inputting the multi-channel far-field speech signal into a preset multi-channel filtering model to perform speech enhancement processing to obtain target speech spectrum information; the preset multi-channel filtering model is a filtering model obtained by adaptively updating an original multi-channel filtering model according to the multi-channel far-field speech signal, or is a filtering model obtained by adaptively constructing according to a prediction result of a first neural network model for the multi-channel far-field speech signal; and inputting the target speech spectrum information into a preset near-field speech generation model to perform speech reconstruction processing to obtain a near-field speech signal corresponding to the multi-channel far-field speech signal. The technical scheme provided by the embodiments of the present disclosure can realize the transformation of speech hearing sensation and improve the quality and intelligibility of speech.
Owner:BEIJING DAJIA INTERNET INFORMATION TECH CO LTD

Artificial intelligence-based audio generation method, apparatus, device, and storage medium

The application relates to the field of artificial intelligence and discloses an audio generation method, device and equipment based on artificial intelligence and a storage medium, the method comprising the following steps: obtaining to-be-converted audio data, and obtaining a target field identifier corresponding to expected sound quality and / or expected emotion; performing cepstrum feature conversion based on a speech spectrum on the to-be-converted audio data to obtain cepstrum feature data of the to-be-converted audio data; inputting the cepstrum feature data and the target field identifier into a speech conversion model to encode the cepstrum feature data at an encoding layer, randomly sample the encoded vector at a resampling layer, and decode and reconstruct the vector sampled at a decoding layer based on the target field identifier to obtain target audio data; and on the basis of effectively removing sound rhythm information of a sounder, the application effectively converts the audio data into audio data with expected rhythm style information, thereby improving an audio conversion effect.
Owner:PING AN TECH (SHENZHEN) CO LTD

Multimodal-based speech synthesis method and device, equipment and storage medium

The application relates to the technical field of artificial intelligence, and discloses a speech synthesis method and device based on multiple modes, equipment and a storage medium, which comprises the following steps: preprocessing a text to be synthesized to obtain character sequence information, character-level graph sequence information and word-level graph sequence information as input sequence information; encoding the character sequence information to obtain a time domain coding vector; encoding the character-level graph sequence information and the word-level graph sequence information to obtain a first space domain coding vector and a second space domain coding vector; performing first cross-modal attention calculation on the time domain coding vector and the first space domain coding vector to obtain a first decoding vector; performing second cross-modal attention calculation on the first decoding vector and the second space domain coding vector to obtain a second decoding vector; and obtaining a speech spectrum graph according to the second decoding vector to generate synthesized speech. The application guarantees the prosody and accuracy of the synthesized speech and effectively improves the level of financial services.
Owner:PING AN TECH (SHENZHEN) CO LTD

Frequency spectrum shaping parameter derivation method in LC3 coding framework based on neural network

The invention relates to the technical field of voice signal processing and audio coding, and discloses a frequency spectrum shaping parameter derivation method in an LC3 coding framework based on a neural network. The method comprises the following steps: performing frequency band energy analysis on an input audio frame, calculating energy of each frequency band according to a preset frequency band division rule, and obtaining multi-frequency band energy information representing spectrum distribution characteristics of the current frame; based on the multi-band energy information, combining a frame-level coding parameter and a transient detection mark, constructing an input feature vector for neural network reasoning, and predicting a spectrum shaping parameter consistent with an LC3 standard by a neural network; and the prediction parameters are further quantized and coded according to the LC3 standard and are used for subsequent spectrum shaping processing, so that on the premise of not changing an LC3 standard bit stream structure and decoder end behaviors, the improvement of a spectrum shaping parameter derivation mode is realized, and the stability and adaptability of coded speech spectrum features are improved.
Owner:GUANGDONG UNIV OF TECH

Speech-based linear prediction codec post-processing method, device, equipment and medium

The invention discloses a speech-based linear prediction codec post-processing method and device, equipment and a medium, and relates to the technical field of computers, and the method comprises the steps: obtaining original decoding parameters of all speech frames obtained by processing an original speech signal through a standard linear prediction codec; performing parameter enhancement on the original decoding parameters by using a preset parameter enhancement model corresponding to each original decoding parameter to obtain enhanced parameters; obtaining an original voice spectrum based on the original decoding parameter, obtaining an enhanced voice spectrum based on the enhanced parameter, and splicing the original voice spectrum and the enhanced voice spectrum to obtain a spliced voice spectrum; processing the spliced speech spectrum by using a preset spectrum enhancement network to obtain a depth filter coefficient; and performing depth filtering on the enhanced speech spectrum based on the depth filter coefficient to obtain a target speech spectrum, and obtaining a target speech signal based on the target speech spectrum. The voice quality can be improved.
Owner:MALANSHAN AUDIO & VIDEO LABORATORY

Voice state intelligent classification method based on voice spectrum characteristics and reinforcement learning optimization mechanism

The invention belongs to the technical field of artificial intelligence and medical speech analysis, and discloses a voice state intelligent classification method based on speech spectrum features and a reinforcement learning optimization mechanism. The method comprises the following steps: carrying out data preprocessing, acoustic feature extraction and feature optimization on a tested voice sample, and mapping the preprocessed tested voice sample into a two-dimensional Mel spectrogram; realizing voice state classification by using a multi-scale feature extraction structure comprising a first convolution branch, a second convolution branch and a third convolution branch and a bidirectional time sequence feature learning structure; meanwhile, a reinforcement learning optimization mechanism is introduced, feature selection, model structure configuration, training hyper-parameters and an updating strategy are subjected to self-adaptive optimization, confidence coefficient calibration, uncertainty judgment and interpretability result generation are combined, and a final voice state judgment result, calibrated confidence coefficient and an acoustic attention area are output. The voice state classification accuracy, stability and interpretability can be improved.
Owner:DALIAN UNIV OF TECH

AI digital human intelligent creation management method and system

PendingCN122289480AFeature extractionAnimation
This application discloses an AI digital human intelligent creation management method and system, relating to the field of artificial intelligence technology. The method includes: determining the digital human image and speech content based on user-input digital human creation instructions; extracting features from the speech content to obtain speech spectrum features and emotional feature parameters; generating an initial lip-sync sequence based on the speech spectrum features; correcting the initial lip-sync sequence based on the emotional feature parameters to obtain a target lip-sync sequence; driving the digital human image based on the target lip-sync sequence and emotional feature parameters to generate a digital human animation; and creating an AI digital human based on the digital human animation and speech content. This application, by jointly driving the digital human image with the corrected target lip-sync sequence and emotional feature parameters, enables the generated AI digital human to adaptively follow the emotional fluctuations of the speech in both lip-sync dynamics and overall performance, enhancing the realism and vividness of the digital human video.
Owner:WUHAN BAOJI NEW MEDIA TECHNOLOGY CO LTD

Robot emotion accompanying system based on deep learning

The invention relates to the technical field of robot emotion accompanying, and discloses a robot emotion accompanying system based on deep learning. The system comprises an emotional state recognition module, an emotional demand analysis module and an emotional interaction strategy generation module. The emotional state recognition module extracts facial micro-expression features, voice spectrum features and body movement track features based on the user interaction data flow and performs multi-modal feature fusion to generate a user emotional state vector; the emotion demand analysis module retrieves an emotion demand knowledge graph based on the vector, matches an emotion demand tag, calculates an emotion demand compactness index, and generates a demand analysis result; and the emotion interaction strategy generation module calls the basic strategy template, adjusts the strategy parameter weight according to the demand closeness index, and generates a personalized emotion interaction strategy instruction set. The system can comprehensively and accurately identify the user emotion, accurately analyze the emotion demand, generate a personalized interaction strategy, and improve the emotion accompanying quality and the user experience.
Owner:HANGZHOU HAISANG HEALTH TECHNOLOGY DEVELOPMENT CO LTD

Sound source positioning and voice wake-up method and device

The invention provides a sound source positioning method and device and a voice wake-up method and device. The sound source positioning method comprises the following steps: acquiring a sound signal acquired by a microphone array; based on a cabin opening and closing state of the vehicle, extracting a speech spectrum feature or a spatial feature of the sound signal as a positioning feature of the sound signal; and carrying out sound source positioning on the sound signals based on the positioning features. According to the method and the device provided by the invention, directional selection is carried out on the positioning characteristics for sound source positioning based on the opening and closing state of the cabin of the vehicle, so that the positioning characteristics obtained based on directional selection can reflect the discrimination between the sound in the vehicle and the sound outside the vehicle in the opening and closing state of the cabin when being applied to sound source positioning of sound signals; therefore, the sound source positioning method can stably output an accurate and reliable positioning result in any cabin opening and closing state, so that the sound inside the vehicle and the sound outside the vehicle are effectively distinguished, and a technical support is provided for defending malicious wake-up outside the vehicle.
Owner:IFLYTEK CO LTD

A multi-modal interaction data processing method and system for children's picture book reading

The application provides a kind of multi-modal interaction data processing method and system of children's picture book reading, it is related to data processing technical field, the method includes: to multi-modal original interaction data set is processed in parallel, extracts and outputs visual semantic feature, interactive speech semantic feature and interactive behavior semantic feature;In visual semantic feature and interactive speech semantic feature, dynamically determine at least three key semantic anchor points, i.e., visual key semantic anchor point corresponds to the core role space coordinate of current picture book page and key text area center point, interactive speech key semantic anchor point corresponds to the core emotion frame timestamp in speech spectrum and question word position;Based on the key semantic anchor point, construct dynamic semantic geometric relationship graph in cross-modal joint feature space.The application can realize the accurate adaptive fusion of cross-modal feature, effectively improve the accuracy of children's picture book reading interaction data processing and the adaptability of interactive response.
Owner:XIAMEN SANDU EDUCATION TECH CO LTD

A method, apparatus, device and medium for training a speech generation model

The application belongs to the field of artificial intelligence, and relates to a training method of a speech generation model, comprising the following steps: obtaining reference timbre spectrum, phoneme information and speech spectrum of a target object; training a preset initial speech generation model based on the reference timbre spectrum, the phoneme information and the speech spectrum to obtain model parameters; and adjusting parameters of a multi-timbre feature extraction network, a phoneme feature extraction network, a prosody feature discretization network, a time sequence alignment module, an attention fusion module and a speech reconstruction decoding network of the initial speech generation model based on the model parameters to construct the speech generation model. The application also provides an apparatus, a device and a medium. In addition, the application also relates to blockchain technology, and speech training data and model parameters can be stored in a blockchain. The application can realize decoupling of timbre and prosody information, and flexibly adjust the timbre and prosody information to generate synthesized speech with diversity and flexibility.
Owner:PING AN TECH (SHENZHEN) CO LTD

Multi-modal emotion recognition method for driver

The invention relates to the technical field of intelligent driving, in particular to a multi-modal emotion recognition method for a driver, and the method comprises the steps: synchronously collecting the driving data of the driver, and carrying out the preprocessing, so as to obtain a standard voice signal, a standard text, and standard video data; extracting speech spectrum features by using a mixed MFCC method; a FastText method is adopted to extract semantic features of the standard text; processing the standard video data frame by frame to extract HOG features of the face of the driver; based on the voice spectrum features, the semantic features of the standard text and the HOG features of the driver face, three independent sub-networks are used for preliminary emotion inference, the driver emotion is classified through a decision-making layer fusion strategy, and finally the driver emotion state is output. According to the invention, the technical problems of inaccurate identification result and poor reliability caused by the fact that a data source is one-sided and is easily interfered by a driving environment when driver emotion identification is carried out by depending on single modal data in the prior art are solved.
Owner:CHINA FAW CO LTD +1

Voice modification

ActiveUS12670918B2TimbreSpeech sound
A computing system that receives an audio waveform representing speech from an individual and produces as output a modified version of the audio waveform that maintains the speaker's speech characteristics as well as prosody for specific utterances (e.g., voice timbre, intonation, timing, intensity). The system uses a bottleneck-based autoencoder with speech spectrograms as input and output. To produce the output audio waveform, the system includes a reconstruction error-based loss function with two additional loss functions. The second loss function is speaker “real vs fake” discriminator that penalizes for the output not sounding like the speaker. The third loss function is a speech intelligibility scorer that penalizes the output for speech that is difficult for the target population to understand. The produced modified audio waveform is an enhanced speech output that delivers speech m a target accent without sacrificing the personality of the speaker.
Owner:SRI INTERNATIONAL

Noise suppression for speech enhancement

A noise suppression method includes transforming a time-domain input signal into an input spectrum that is the spectrum of the input signal, the input signal comprising speech components and noise components, and the input spectrum comprising a speech spectrum that is the spectrum of the speech components and a noise spectrum that is the spectrum of the noise components, smoothing magnitudes of the input spectrum to provide a smoothed-magnitude input spectrum, and estimating basic suppression filter coefficients from the input spectrum and the smoothed input spectrum. The method further includes determining noise suppression filter coefficients from the estimated basic suppression filter coefficients and a spectral correlation factor, the spectral correlation factor indicating whether speech is present in the input signal or not, filtering the input spectrum based on the noise suppression filter coefficients to generate an output spectrum; and transforming the output spectrum into a time-domain output signal.
Owner:HARMAN BECKER AUTOMOTIVE SYST GMBH

Ambient noise compensation in teleconferencing

A method of compensating for ambient noise during a teleconference may involve estimating, by a control system, a current speech spectrum corresponding to a speech of a remote teleconference participant; estimating, by the control system, a current noise spectrum corresponding to ambient noise in a local environment in which a local teleconference participant is located; calculating, by the control system, a current speech intelligibility index (SII) based at least in part on the current speech spectrum and the current noise spectrum; determining, by the control system and based at least in part on the current SII, whether to adjust a local audio system used by the local teleconference participant, where the determining involves evaluating the current SII according to one or more target SII parameters; and updating at least one of the one or more target SII parameters in response to a user input corresponding to the playback volume change.
Owner:DOLBY LABORATORIES LICENSING CORP

Multi-speaker chinese speech synthesis method based on dense connection delay neural network

The application discloses a multi-speaker Chinese speech synthesis method based on a dense connection time delay neural network, wherein a speaker encoder module in a multi-speaker Chinese speech synthesis network based on the dense connection time delay neural network extracts speaker embedding from a reference speech spectrum; the speaker encoder module is simple in structure and small in parameter quantity; the extracted speaker embedding fuses multi-level information; therefore, the speaker embedding can be optimized together with other modules in the multi-speaker Chinese speech synthesis network, the training process is simplified, and the speaker embedding more suitable for a speech synthesis task can be extracted; secondly, the output of a text encoder module of the multi-speaker Chinese speech synthesis network is taken as a key and a value, the output of the speaker encoder module is taken as a query, and the key, the value and the query are input into an encoder's scaling dot product attention mechanism to generate a conditional text representation as an input of a decoder, so that the speaker embedding can effectively control the style in the synthesized speech and improve the naturalness and similarity of the synthesized speech.
Owner:NANJING UNIV

System and method for authenticating users in a computing system

In response to receiving a voice call from a user, a new voice spectrogram is generated based on the voice of the calling user. A plurality of phonetic indicators are extracted from the new voice spectrogram and compared to phonetic indicators of a plurality of historic voice spectrograms associated with respective users. When a historic voice spectrogram includes one or more of the phonetic indicators extracted from the new voice spectrogram, it is determined that the identity of the calling user is authenticated. On the other hand, when none of the historic voice spectrograms include the one or more of the phonetic indicators extracted from the new voice spectrogram, it is determined that the identity of the calling user is not authenticated.
Owner:BANK OF AMERICA CORP

Cross-lingual corpus synthesis method, speech synthesis model training method, and related devices

The application provides a cross-language corpus synthesis method, a speech synthesis model training method and related equipment. The cross-language corpus synthesis method comprises the following steps: obtaining a cross-language text, a target speaker embedding vector, and a language embedding vector corresponding to each language included in the cross-language text; determining a character embedding vector corresponding to each character included in the cross-language text; determining a speech spectrum corresponding to the cross-language text according to the language embedding vector corresponding to each language included in the cross-language text, the character embedding vector corresponding to each character included in the cross-language text, and the target speaker embedding vector; and composing a cross-language corpus from the speech spectrum and the cross-language text. The application can synthesize speech spectrums of the same speaker switching between languages, thereby obtaining a cross-language corpus composed of the speech spectrums and the cross-language text, so that a speech synthesis model with higher naturalness of synthesized speech can be constructed based on the obtained cross-language corpus subsequently.
Owner:UNIV OF SCI & TECH OF CHINA

Speech synthesis processing method, apparatus, and related device

The application belongs to the technical field of financial technology, and provides a speech synthesis processing method and device and related equipment. In order to solve the problem of low similarity between synthesized speech and real human speech in the prior art, the text to be converted into speech is obtained, the prompt text is obtained, and the discrete semantic token is determined. Then, based on the preset text-to-speech large language model, the alignment relationship between the text and the discrete semantic token is established according to the prompt text, the target discrete semantic token corresponding to the text is obtained, and based on the preset condition flow matching model, the speech spectrum feature corresponding to the target discrete semantic token is determined. Then, based on the preset speech synthesis decoder, the speech spectrum feature is converted into a speech signal to obtain the speech. The similarity between the synthesized speech and the real human speech can be improved. For example, for self-service voice services of insurance or banking businesses, the synthesized speech obtained by using the above method has high similarity with the real human speech.
Owner:PING AN TECH (SHENZHEN) CO LTD