Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

29 results about "Phonetic representation" patented technology

Phonetic representation, or more commonly phonetic transcription is the representation of speech sounds using symbols in phonetic alphabet such as IPA, X-SAMPA, Kirshenbaum for linguistic studies and for learning the pronunciation of languages. Among these systems, the International Phonetic Alphabet has been the most widely used so far, whose symbols are printed in most dictionaries and books on linguistics.

Training method and system for self-supervised training predictor for speech separation

ActiveCN115762557Bimprove performanceSelf-supervised training feature improvementSpeech recognitionSound sourcesVoice source
The embodiment of the present application provides a training method and system of a self-supervised training predictor for speech separation. The method comprises: extracting self-supervised training features of each single person voice source speech by using a pre-training model; extracting shallow features for speech representation and deep features for context information in the self-supervised training features, and determining the shallow features and the deep features of each single person voice source speech as training labels of the self-supervised training predictor; inputting training mixed speech generated by each single person voice source speech into the self-supervised training predictor to obtain estimated features of each single person voice source speech; and training the self-supervised training predictor based on a loss function determined based on the estimated features and the training labels corresponding to each single person voice source speech. The self-supervised training predictor is trained and applied in a speech separation model in the embodiment of the present application, so that the accuracy of the self-supervised training features is improved, the performance of the speech separation system is improved, and the model parameters and the calculation complexity are reduced.
Owner:AISPEECH CO LTD

Emotive text-to-speech with auto detection of emotions

A method of providing emotive text-to-speech includes obtaining input text characterizing a natural language response generated by an assistant LLM to a query input by a user during a conversation between the user and the assistant LLM, and processing, using the assistant LLM, the input text conditioned on an emotion detection task prompt to predict, as output from the assistant LLM, an emotional state of the natural language response. The method also includes determining, based on the emotional state of the natural language response predicted as output from the assistant LLM, an emotional embedding for the input text and instructing a TTS model to process the input text and the emotional embedding to generate a synthesized speech representation of the natural language response conveying the emotional state of the natural language response as specified by the emotional embedding.
Owner:GOOGLE LLC

Method, apparatus, device and storage medium for text-to-speech conversion

According to embodiments of this disclosure, a method, apparatus, device, and storage medium for text-to-speech conversion are provided. The method includes generating a predicted speech representation of the target text read by a first speaker, based on a target text to be converted and a first timbre of a first speaker. The predicted speech representation indicates speech features that vary over time. The method further includes generating a predicted time-frequency representation of the target text read by a second speaker, based on the predicted speech representation and a second timbre of a second speaker. The predicted time-frequency representation indicates the speech signal strength that varies over time at different frequencies. The method further includes converting the predicted time-frequency representation into audio of the target text read by the second speaker. This reduces the difficulty of prediction and improves the sound quality of the generated audio.
Owner:BEIJING YOUZHUJU NETWORK TECH CO LTD

System and method for secure processing of speech signals using pseudo-speech representations

A method, computer program product, and computing system for processing a speech signal. A sensitive portion of the speech signal is identified. A pseudo-speech representation of the sensitive portion is generated using a voice converter system. Speech processing is performed on the speech signal and the pseudo-speech representation of the sensitive portion using a speech processing system.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Scalable model specialization framework for speech model personalization

A method for speech conversion includes obtaining a speech conversion model configured to convert input utterances of human speech directly into corresponding output utterances of synthesized speech. The method further includes receiving a speech conversion request including input audio data corresponding to an utterance spoken by a target speaker associated with atypical speech and a speaker identifier uniquely identifying the target speaker. The method includes activating, using the speaker identifier, a particular sub-model for biasing the speech conversion model to recognize a type of the atypical speech associated with the target speaker identified by the speaker identifier. The method includes converting, using the speech conversion model biased by the activated particular sub-model, the input audio data corresponding to the utterance spoken by the target speaker associated with atypical speech into output audio data corresponding to a synthesized canonical fluent speech representation of the utterance spoken by the target speaker.
Owner:GOOGLE LLC

Emotion regulation instruction recommendation method, device, equipment and medium

The application relates to the fields of artificial intelligence and medical health technology, and discloses a mood regulation instruction recommendation method, device, equipment and medium, wherein the method comprises the following steps: adopting a pre-trained dialogue semantic representation model to generate a dialogue semantic representation vector of psychological counseling data; adopting a pre-trained facial expression representation model to generate a micro-expression representation vector of the psychological counseling data; adopting a pre-trained voice representation model to generate a voice representation vector of the psychological counseling data; adopting a pre-trained abnormal mood comprehensive recognition model to generate a comprehensive representation vector of the dialogue semantic representation vector, the micro-expression representation vector and the voice representation vector; and adopting a pre-trained mood regulation instruction recommendation model to recommend a mood regulation instruction for the comprehensive representation vector and a personal information feature vector corresponding to a counseling object, so as to obtain a recommendation result. Therefore, the accuracy of the recommendation result is improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Digital human mouth shape prediction system and method based on variational auto-encoder in smart agriculture

ActiveCN121236246AClimate change adaptationSemantic analysisPattern recognitionSoft tissue deformation
The invention provides a digital human mouth shape prediction system and method based on a variational auto-encoder in smart agriculture, and relates to the technical field of artificial intelligence, and the method comprises the steps: receiving a voice signal of a smart agriculture scene, extracting multi-dimensional voice features, and fusing the multi-dimensional voice features to generate voice representation; inputting the voice representation into a model, and decoding the voice representation into a mouth shape key point sequence; performing semantic enhancement on the mouth shape key point sequence by combining agricultural proper nouns in the input voice and a preset mouth shape template, and simultaneously inhibiting mouth shape jitter and abnormal frames; and mapping the enhanced mouth shape key point sequence to a three-dimensional digital human face model, converting the enhanced mouth shape key point sequence into face soft tissue deformation data in combination with a face parameter of a target user, calculating a skeleton node position and a rotation parameter according to a preset corresponding relationship between the mouth shape and the skeleton, and generating a digital human mouth shape animation synchronized with the input voice through combined rendering. The stable and accurate voice synchronous digital human mouth shape can be generated in an intelligent agricultural scene.
Owner:XIAODUO INTELLIGENT TECH (BEIJING) CO LTD

Training method of universal voice representation model and related device

The invention discloses a universal voice representation model training method and a related device, and relates to the technical field of artificial intelligence, and the method comprises the steps: obtaining a voice sample and a mask voice sample corresponding to the voice sample, inputting the voice sample into a teacher model, obtaining a teacher voice representation, inputting the mask voice sample into a student model, and obtaining a student voice representation; obtaining a first student voice representation, generating global knowledge distillation loss and local comparison loss according to the teacher voice representation and the first student voice representation, and updating target parameters of the student model by using the global knowledge distillation loss and the local comparison loss; and the trained student model is used as a general voice representation model. According to the method, knowledge distillation and comparative learning are combined, so that the student model can extract multi-aspect voice information, and therefore, the method can adapt to various voice tasks, and the universality is higher. The characteristics of the mask area can be accurately reconstructed through double knowledge migration, so that the accuracy of voice representation extracted by the student model is higher.
Owner:ANHUI LINGDONG GENERAL ROBOT TECHNOLOGY CO LTD

Method, device and storage medium for speech processing

According to an embodiment of the disclosure, a method, apparatus, device and computer-readable storage medium for speech processing are provided. The method includes: acquiring a speech feature sequence corresponding to a speech sample, the speech feature in the speech feature sequence corresponding to a speech frame in the speech sample. For the target speech feature in the speech feature sequence, one or more subsequent speech tokens respectively corresponding to the one or more subsequent speech features are generated based on the target speech feature and the one or more preceding speech features and according to the speech encoding model. A speech encoding model is trained based on the one or more subsequent speech tokens. This training approach enables the speech encoding model to learn high-quality speech representations.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD +1

Unsupervised data selection via discrete speech representation for automatic speech recognition

A method includes obtaining a corpus of unlabeled training data including a plurality of spoken utterances, each corresponding spoken utterance of the plurality of spoken utterances includes audio data characterizing the corresponding spoken utterance. The method also includes receiving a target domain. The method also includes selecting, using a contrastive data selection model, a subset of the utterances from the corpus of unlabeled training data that correspond to the target domain. The method includes training an automatic speech recognition (ASR) model on the subset of utterances.
Owner:GOOGLE LLC

Scalable model specialization framework for speech model personalization

A method for speech conversion includes obtaining a speech conversion model configured to convert input utterances of human speech directly into corresponding output utterances of synthesized speech. The method further includes receiving a speech conversion request including input audio data corresponding to an utterance spoken by a target speaker associated with atypical speech and a speaker identifier uniquely identifying the target speaker. The method includes activating, using the speaker identifier, a particular sub-model for biasing the speech conversion model to recognize a type of the atypical speech associated with the target speaker identified by the speaker identifier. The method includes converting, using the speech conversion model biased by the activated particular sub-model, the input audio data corresponding to the utterance spoken by the target speaker associated with atypical speech into output audio data corresponding to a synthesized canonical fluent speech representation of the utterance spoken by the target speaker.
Owner:GOOGLE LLC

A digital human mouth shape prediction system and method based on a variational autoencoder in smart agriculture

ActiveCN121236246BClimate change adaptationSemantic analysisPattern recognitionSoft tissue deformation
The application provides a digital human mouth shape prediction system and method based on a variational autoencoder in smart agriculture, and relates to the technical field of artificial intelligence. The application receives a voice signal of a smart agriculture scene, extracts multi-dimensional voice features and fuses to generate a voice representation. The voice representation is input into a model and decoded into a mouth shape key point sequence. The agricultural specific terms in the input voice are combined with a preset mouth shape template to perform semantic enhancement on the mouth shape key point sequence, while suppressing mouth shape shaking and abnormal frames. The enhanced mouth shape key point sequence is mapped to a three-dimensional digital human face model, combined with face parameters of a target user to convert into face soft tissue deformation data, and then the preset corresponding relationship between the mouth shape and the skeleton is used to calculate the position and rotation parameters of the skeleton nodes. After combination and rendering, a digital human mouth shape animation synchronized with the input voice is generated, which can generate stable and accurate voice synchronized digital human mouth shapes in a smart agriculture scene.
Owner:XIAODUO INTELLIGENT TECH (BEIJING) CO LTD

End-to-end speaker recognition method based on multi-scale SincNet and CGAN

The application belongs to the technical field of speaker recognition, and specifically discloses an end-to-end speaker recognition method based on multi-scale SincNet and CGAN. The multi-scale SincNet is introduced, so that important information is avoided to be lost in manual feature conversion, the multi-scale SincNet captures low-level speech representation of three channels in a waveform according to three customized filter banks, the SincGAN model better captures important narrowband speaker features, the end-to-end recognition is realized based on the improved conditional generative adversarial network, the speaker is recognized by using a small amount of training sentences, and the loss function of the application comprises an adversarial loss of a classic GAN and a classification cross-entropy loss of a classification task. Experimental results show that on TIMIT and LIBRISPEECH corpus, the model of the application shows better performance, and when lacking training data, the model of the application shows stronger robustness than a baseline method.
Owner:SHANDONG UNIV OF SCI & TECH

Method and electronic device for domain adaptation with self-supervised speech representation in speech enhancement

A method for generating a customized speech enhancement model includes: obtaining noisy-clean speech data from a source domain; obtaining noisy voice data from the target domain; acquiring original voice data; training, using the noisy-clean speech data, the noisy speech data, and the original speech data, a customized SE model based on at least one of: self-supervised representation-based adaptation (SSRA), set mapping, or self-supervised adaptation loss; generating a customized SE model by denoising the noisy speech data using a trained customized SE model; and providing the customized SE model to a user device to use the de-noised noisy speech data.
Owner:SAMSUNG ELECTRONICS CO LTD

Zero-Shot Cross-Lingual Voice Transfer for Text-To-Speech

A method for performing zero-shot voice transfer using text-to-speech (TTS) includes receiving an input text sequence characterizing an utterance and receiving a reference speech representation characterizing a reference utterance spoken by a target speaker. The method also includes generating an encoded textual representation for the input text sequence, processing, using a speaker encoder, the reference speech representation to generate a speaker representation characterizing voice characteristics of the target speaker and learning fine-grained embedding vectors based on the speaker representation to obtain a final embedding vector. The method also includes predicting a duration and upsampling the encoded textual representation into an upsampled output. The method also includes generating a synthesized speech representation based on the upsampled output and the final embedding vector and generating a time-domain audio waveform of the input text sequence that clones a voice of the target speaker.
Owner:GOOGLE LLC

Controllable zero-sample speech conversion method, device, equipment and medium

ActiveCN121034280BPersonalizationEngineering
The application relates to the technical field of voice semantics, can be applied to a business system platform such as financial technology and medical health, and discloses a controllable zero-sample voice conversion method, device, equipment and medium, the method comprises the following steps: performing self-supervised voice learning on unlabeled voice data to obtain a self-supervised voice representation, extracting a content feature vector and a prosody style vector of the self-supervised voice representation, and converting the content feature vector and the prosody style vector into discrete content tokens and discrete prosody tokens; generating a mask for the discrete prosody tokens to obtain target prosody tokens; obtaining reference voice of a target user; extracting a user style embedding in the reference voice; performing flow matching on the discrete content tokens, the target prosody tokens and the user style embedding; generating a target mel-frequency spectrum graph; and performing voice waveform reconstruction and optimization on the target mel-frequency spectrum graph to obtain a zero-sample voice conversion result. The application realizes the zero-sample voice conversion problem of personalization, high fidelity and consistent style under the condition of unlabeled voice data.
Owner:PING AN TECH (SHENZHEN) CO LTD

Voice keyword recognition system based on comparative learning pre-training

A voice keyword recognition system based on comparative learning pre-training comprises the following steps that firstly, a large amount of paired voice-text data easy to obtain is sent into a designed pre-training framework based on ternary comparative learning for pre-training; the objective of the invention is to train and obtain a voice encoder capable of encoding an original voice signal into a universal voice representation and a query encoder capable of encoding an identification keyword into a universal pronunciation representation by optimizing three types of losses including voice element comparison loss, voice-phoneme element comparison loss and phoneme element comparison loss. And then, respectively obtaining universal voice representation and pronunciation representation from a small amount of keyword data on the labeled target data set through a pre-trained voice encoder and a pre-trained query encoder, and sending the two representations into a classifier network for final decision making. The method solves the problem that voice data of a large number of specific keywords is difficult to collect when keywords need to be customized in voice keyword recognition.
Owner:EAST CHINA UNIV OF SCI & TECH

Training a speech recognition model based on cross-lingual speech synthesis

A method (800) for training a speech recognition model (200) includes obtaining a multilingual text-to-speech (TTS) model (310). The method also includes generating a native-synthetic speech representation (306) for an input text sequence (302) in a first language, the native-synthetic speech representation conditioned on speaker characteristics (304) of a native speaker of the first language. The method also includes generating a cross-lingual-synthetic speech representation for the input text sequence in the first language, the cross-lingual-synthetic speech representation conditioned on speaker characteristics of a native speaker of a different, second language. The method also includes generating first and second speech recognition results (312) for the native-synthetic speech representation and the cross-lingual-synthetic speech representation. The method also includes determining a consistency loss term (352) based on the first and second speech recognition results, and updating parameters of the speech recognition model based on the consistency loss term.
Owner:GOOGLE LLC

Method and apparatus for processing input utterances by a speech recognition system

A method and apparatus process input utterances by a speech recognition system. The method and apparatus are implemented by a computer of a speech recognition system to process an utterance that is received as an input. The method includes processing the utterance by a rule-based natural language-understanding engine. The method further includes, when the rule-based natural language-understanding engine fails to process the utterance, converting a representation of the utterance and allowing a machine learning-based natural language-understanding engine to process the utterance by using a large language model (LLM) agent. The method further includes processing the utterance with a converted representation by the machine learning-based natural language-understanding engine.
Owner:HYUNDAI MOTOR CO LTD +1

Language conversion for conversational ai systems and applications

In various examples, first textual data may be applied to a first MLM to generate an intermediate speech representation (e.g., a frequency-domain representation), the intermediate audio representation and a second MLM may be used to generate output data indicating second textual data, and parameters of the second MLM may be updated using the output data and ground truth data associated with the first textual data. The first MLM may include a trained Text-To-Speech (TTS) model and the second MLM may include an Automatic Speech Recognition (ASR) model. A generator from a generative adversarial networks may be used to enhance an initial intermediate audio representation generated using the first MLM and the enhanced intermediate audio representation may be provided to the second MLM. The generator may include generator blocks that receive the initial intermediate audio representation to sequentially generate the enhanced intermediate audio representation.
Owner:NVIDIA CORP

A psychoacoustic speech masking method for active defense against voice cloning

PendingCN122337245AMasking thresholdTime frequency decomposition
This application discloses a psychoacoustic speech masking method for proactive defense against speech cloning, relating to the field of speech information security. The method includes: performing time-frequency decomposition on the original reference speech using time-frequency analysis to obtain speech information in the speech representation space; adding a look-ahead perturbation to the speech information to obtain look-ahead speech; determining a speaker confusion loss using a speaker coding network based on the original reference speech and the look-ahead speech; determining a total masking threshold using a psychoacoustic masking model based on the speech information in the speech representation space; determining a speech quality constraint loss based on the perturbation power spectrum according to the total masking threshold; determining a total loss based on the speaker confusion loss and the speech quality constraint loss; generating a final perturbation using an iterative optimization algorithm based on the total loss; and generating a final protected reference speech based on the final perturbation and the original reference speech. This application weakens the ability to extract the identity of the real speaker, thereby reducing the risk of speech cloning and voiceprint impersonation.
Owner:QINGHAI UNIV FOR NATITIES

System and method of reinforcing general purpose natural language models with acquired subject matter

In a contact center apparatus, a method for generating recognized text based upon verbal data, comprising: receiving verbal data from a user; applying a natural language understanding engine the verbal data to generate phonetic representation text, the phonetic representation text configured as a phonetic representation of the verbal data; applying a language reinforcement engine to the phonetic representation text to generate recognized text, the recognized text identifying a word associated with the verbal data; and directing the user to a contact center resource based upon the recognized text.
Owner:SWAMPFOX TECH INC

Accent conversion method and device based on discrete speech representation, equipment and medium

The invention relates to the technical field of voice processing, and particularly discloses an accent conversion method and device based on discrete voice representation, equipment and a medium. According to the method, the to-be-converted voice is discretized to obtain the voice unit which only retains phoneme information, and then the voice unit is mapped into the target accent by using the accent conversion model, so that redundant features are removed, the model is enabled to concentrate on accent conversion, and the accuracy of accent conversion is improved; and secondly, the target voice is synthesized in combination with the tone characteristics of the speaker, so that the tone of the speaker is reserved, and the naturalness of the target voice after accent conversion is improved. The method is applied to financial services such as telemarketing, telephone banking and telephone return visit and medical systems such as remote inquiry, patient follow-up visit and medical education, the voice accent of a speaker in communication participants can be accurately converted into a target accent familiar with a listener, the tone of the speaker is kept, and the communication efficiency is effectively improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

METHOD AND DEVICE FOR PROCESSING INPUT OUTPUTS BY A SPEECH RECOGNITION SYSTEM

A method and a device process input utterances through a speech recognition system. The method and the device are implemented by a computer of a speech recognition system to process an utterance received as input. The method includes processing the utterance by a rule-based natural language understanding engine. The method further includes, if the rule-based natural language understanding engine cannot process the utterance, converting a representation of the utterance and enabling a machine learning-based natural language understanding engine to process the utterance using a large language model agent (LLM agent). The method further includes processing the utterance with a converted representation by the machine learning-based natural language understanding engine.
Owner:HYUNDAI MOTOR CO LTD +1

Speech recognition method and apparatus, and electronic device

A speech recognition method and apparatus, and an electronic device. The speech recognition method includes: acquiring (S302) speech representation vectors and speaker representation vectors corresponding to to-be-recognized speech frame data; performing (S304) parallel speech frame text activation value prediction on the speech representation vectors, and when prediction results indicate that speech frame text activation values reach a firing threshold, generating a firing vector according to speech representation vectors within a range of the firing threshold; determining (S306) a corresponding text for the firing vector and a speaker corresponding to the text according to the firing vector and the speaker representation vectors. The accuracy of speech recognition and speaker marking is improved.
Owner:ALIBABA INNOVATION PRIVATE LIMITED

Generative data for conversational ai systems and applications

In various examples, first textual data may be applied to a first MLM to generate an intermediate speech representation (e.g., a frequency-domain representation), the intermediate audio representation and a second MLM may be used to generate output data indicating second textual data, and parameters of the second MLM may be updated using the output data and ground truth data associated with the first textual data. The first MLM may include a trained Text-To-Speech (TTS) model and the second MLM may include an Automatic Speech Recognition (ASR) model. A generator from a generative adversarial networks may be used to enhance an initial intermediate audio representation generated using the first MLM and the enhanced intermediate audio representation may be provided to the second MLM. The generator may include generator blocks that receive the initial intermediate audio representation to sequentially generate the enhanced intermediate audio representation.
Owner:NVIDIA CORP

Voice processing method and device, equipment and storage medium

The embodiment of the invention provides a voice processing method and device, equipment and a storage medium. The method comprises the following steps: acquiring a voice feature sequence corresponding to a voice sample, wherein voice features in the voice feature sequence correspond to voice frames in the voice sample; and for a target voice feature in the voice feature sequence, generating one or more post-order voice marks corresponding to the one or more post-order voice features based on the target voice feature and the one or more pre-order voice features by using a voice coding model. A speech encoding model is trained based on the one or more post-order speech markers. The training mode can enable the speech coding model to learn high-quality speech representation.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD +1

Extending multilingual speech synthesis with zero supervision of discovered data

PendingCN121753095ABiological modelsSpeech recognitionSpeech codeAcoustics
A method (600) includes receiving training data (301) including a plurality of sets of training utterances (310) each associated with a respective language. Each training utterance includes a corresponding reference speech representation (304) paired with a corresponding input text sequence (302). For each training utterance, the method includes generating a corresponding encoded text representation for a corresponding input text sequence (312, 313), generating a corresponding speech code for a corresponding reference speech representation (314), generating a shared encoder output (332, 334), and outputting the shared encoder output (332, 334). And determine a text-to-speech (TTS) loss based on the corresponding encoded text representation, the corresponding speech coding, and the shared encoder output (305). The method further comprises training the TTS model (501) based on the TTS loss determined for the training utterances in each set of training utterances to teach the TTS model how to learn to synthesize speech in each of the respective languages.
Owner:GOOGLE LLC