Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

51 results about "Phonetic representation" patented technology

Phonetic representation, or more commonly phonetic transcription is the representation of speech sounds using symbols in phonetic alphabet such as IPA, X-SAMPA, Kirshenbaum for linguistic studies and for learning the pronunciation of languages. Among these systems, the International Phonetic Alphabet has been the most widely used so far, whose symbols are printed in most dictionaries and books on linguistics.

Audio and video depth forgery detection method based on quality perception and multi-scale alignment

The invention discloses an audio and video depth forgery detection method based on quality perception and multi-scale alignment, and the method comprises the following steps: coding a synchronous audio and video sequence, and obtaining a frame-level visual feature, a facial action unit and a phoneme-level voice representation; a visual quality evaluation module is introduced to generate a spatial reliability mask, and quality weighting is carried out on the visual features; designing a global-local multi-scale cross-modal alignment mechanism, performing bidirectional cross-attention modeling on voice and face dynamic synchronization globally, and performing physiological coupling alignment on phonemes and face action units locally; and an uncertainty perception reasoning and calibration scheme is provided, adaptive temperature scaling is carried out according to quality and consistency, and uncertainty calibration is carried out by self-supervision loss. According to the method, the problems of insufficient robustness and excessive self-confidence misjudgment of an existing method in a low-quality video and high-synchronization counterfeit scene are solved, and the cross-dataset generalization capability and the actual deployment reliability are remarkably improved.
Owner:NANJING UNIV OF SCI & TECH

Realtime AI sign language recognition with avatar

Disclosed herein are method and system aspects for translating between a sign language and a target language and presenting such translations. For example, a method receives input language data and translates the input language data into sign language grammar. The method retrieves phonetic representations that correspond to the sign language grammar from a sign language database and generates coordinates from the phonetic representations using a generative network. The phonetic representations are digital representations of individual signs created through manual input of body configuration information corresponding to the individual signs. Further, the method renders an avatar that moves between the coordinates. In another example, a bidirectional communication system allows for realtime communication between a signing entity and a non-signing entity.
Owner:SIGN-SPEAK INC

Children voice expression error recognition and correction method based on comparative learning

PendingCN121011207ASpeech analysisSpeech developmentFalse recognition
The invention discloses a children voice expression error recognition and correction method based on comparative learning, and the method comprises the steps: carrying out the preprocessing of an inputted children voice signal, obtaining a logarithmic Mel spectrum feature sequence, and converting the logarithmic Mel spectrum feature sequence into voice semantic coding features through an improved Transform encoder; on the basis of an age-adaptive positive and negative sample selection mechanism, voice features are optimized by using a comparative learning method, and an enhanced voice representation vector is obtained; constructing a multi-modal fusion network, combining an enhanced voice vector and BERT language model features, realizing adaptive fusion through multi-head cross attention and a gating mechanism, adopting a bidirectional LSTM to design an error positioning module to recognize the position and the type of a pronunciation error, and using a joint loss function to execute end-to-end training; standard correction audio is generated according to the error type, and correction guidance is provided for children. According to the invention, high-precision recognition, positioning and correction of children's speech expression errors are realized, and technical support is provided for children's language development.
Owner:HECHEN ZONGHENG INFORMATION TECH CO LTD

Low-latency conversational large language models

A method includes receiving a transcription of an utterance, processing, using a first model, the transcription to generate a first text segment that represents an initial portion of a response to the utterance, processing, using a TTS system, the first text segment to generate a first synthesized speech representation, and providing, for audible output, the first synthesized speech representation. The method also includes providing, to a second model different from the first model, the transcription and the first text segment, the second model comprising an LLM configured to process the transcription and the first text segment to generate a second text segment that represents a remaining portion of the response to the utterance. The method further includes obtaining a second synthesized speech representation generated from the second text segment, and providing, for audible output by the user device, the second synthesized speech representation.
Owner:GOOGLE LLC

Control device, handling system, determining device, control method and computer program

UndeterminedDE102026106796A1Object basedEngineering
According to one embodiment, a control device (3; 3-2) recognizes an object area from image information that corresponds to an object as a target in automated handling operations. The control device generates a handling strategy for the object and outputs a speech representation of the handling strategy. The control device constructs the object area based on a speech representation of an object shape encompassed in the handling strategy. The control device evaluates the handleability of the object based on geometric information about the object area. The control device evaluates the adaptability of a handling tool (14) to the object based on the speech representation of the object shape.Based on the manageability and adaptability of the handling tool (14), the control device determines whether handling of the object is possible and / or generates a handling position and a posture of the handling tool (14).
Owner:TOSHIBA KANAGAWA

Controllable zero sample voice conversion method, device, equipment and medium

The invention relates to the technical field of voice semantics, can be applied to business system platforms of financial science and technology, medical treatment and health and the like, and discloses a controllable zero sample voice conversion method, device, equipment and medium, the method comprises the following steps: carrying out self-supervised voice learning on unlabeled voice data to obtain self-supervised voice representation; the method comprises the following steps of: extracting a content feature vector and a rhythm style vector represented by self-supervised speech, converting the content feature vector and the rhythm style vector into a discrete content token and a discrete rhythm token, performing mask generation on the discrete rhythm token to obtain a target rhythm token, obtaining reference speech of a target user, extracting user style embedding in the reference speech, and obtaining a target user. And performing stream matching on the discrete content token, the target rhythm token and the user style embedding to generate a target Mel spectrogram, and performing voice waveform reconstruction and optimization on the target Mel spectrogram to obtain a zero sample voice conversion result. According to the invention, under the condition of no annotated voice data, personalized, high-fidelity and style-consistent zero-sample voice conversion is realized.
Owner:PING AN TECH (SHENZHEN) CO LTD

Training method and system for self-supervised training predictor for speech separation

ActiveCN115762557Bimprove performanceSelf-supervised training feature improvementSpeech recognitionSound sourcesVoice source
The embodiment of the present application provides a training method and system of a self-supervised training predictor for speech separation. The method comprises: extracting self-supervised training features of each single person voice source speech by using a pre-training model; extracting shallow features for speech representation and deep features for context information in the self-supervised training features, and determining the shallow features and the deep features of each single person voice source speech as training labels of the self-supervised training predictor; inputting training mixed speech generated by each single person voice source speech into the self-supervised training predictor to obtain estimated features of each single person voice source speech; and training the self-supervised training predictor based on a loss function determined based on the estimated features and the training labels corresponding to each single person voice source speech. The self-supervised training predictor is trained and applied in a speech separation model in the embodiment of the present application, so that the accuracy of the self-supervised training features is improved, the performance of the speech separation system is improved, and the model parameters and the calculation complexity are reduced.
Owner:AISPEECH CO LTD

Emotive text-to-speech with auto detection of emotions

A method of providing emotive text-to-speech includes obtaining input text characterizing a natural language response generated by an assistant LLM to a query input by a user during a conversation between the user and the assistant LLM, and processing, using the assistant LLM, the input text conditioned on an emotion detection task prompt to predict, as output from the assistant LLM, an emotional state of the natural language response. The method also includes determining, based on the emotional state of the natural language response predicted as output from the assistant LLM, an emotional embedding for the input text and instructing a TTS model to process the input text and the emotional embedding to generate a synthesized speech representation of the natural language response conveying the emotional state of the natural language response as specified by the emotional embedding.
Owner:GOOGLE LLC

Method, apparatus, device and storage medium for text-to-speech conversion

According to embodiments of this disclosure, a method, apparatus, device, and storage medium for text-to-speech conversion are provided. The method includes generating a predicted speech representation of the target text read by a first speaker, based on a target text to be converted and a first timbre of a first speaker. The predicted speech representation indicates speech features that vary over time. The method further includes generating a predicted time-frequency representation of the target text read by a second speaker, based on the predicted speech representation and a second timbre of a second speaker. The predicted time-frequency representation indicates the speech signal strength that varies over time at different frequencies. The method further includes converting the predicted time-frequency representation into audio of the target text read by the second speaker. This reduces the difficulty of prediction and improves the sound quality of the generated audio.
Owner:BEIJING YOUZHUJU NETWORK TECH CO LTD

Hybrid language models for conversational AI systems and applications

In various examples, first textual data may be applied to a first MLM to generate an intermediate speech representation (e.g., a frequency-domain representation), the intermediate audio representation and a second MLM may be used to generate output data indicating second textual data, and parameters of the second MLM may be updated using the output data and ground truth data associated with the first textual data. The first MLM may include a trained Text-To-Speech (TTS) model and the second MLM may include an Automatic Speech Recognition (ASR) model. A generator from a generative adversarial networks may be used to enhance an initial intermediate audio representation generated using the first MLM and the enhanced intermediate audio representation may be provided to the second MLM. The generator may include generator blocks that receive the initial intermediate audio representation to sequentially generate the enhanced intermediate audio representation.
Owner:NVIDIA CORP

Voice authentic identification method and device, computer equipment and storage medium

The invention relates to the technical field of artificial intelligence, finance and medical health, and discloses a voice authentication method and device, computer equipment and a storage medium. The method comprises the following steps: acquiring to-be-identified voice information; inputting the voice information to be subjected to authentic identification into the self-supervised large model for feature extraction to obtain an extraction result; inputting the potential speech representation features into a classification model for deep coding to obtain a coding result, converting the coding result into low-dimensional representation, and calculating the probability distribution of speech authenticity to obtain a classification result; and outputting a classification result. By implementing the method provided by the invention, the capability of identifying the forged voice in high-imitation audio and noise environments can be remarkably improved, and reliable guarantee is provided for voice safety in key fields such as finance, medical health and the like.
Owner:PING AN TECH (SHENZHEN) CO LTD

A method and system for audio semantic retrieval of law enforcement recorders based on retrieval enhancement

The present invention relates to the technical fields of speech recognition and natural language processing, and specifically discloses a method and system for semantic retrieval of law enforcement recorder audio based on retrieval enhancement. The system comprises: a data acquisition module for acquiring audio data and text queries; a speech adapter module for projecting the audio data into a text embedding space to obtain a speech representation; a cross-modal retriever for performing cross-modal retrieval of the speech representation and text query to obtain speech tokens; a speech language model for obtaining text hypotheses; a query generation module for extracting query fragments that may contain entity names; an entity retrieval module for searching an entity database based on the query fragments to obtain relevant entity names; a context construction module for constructing contextual information; and a large language model for obtaining semantic retrieval results. The present invention improves the recognition and retrieval accuracy of entity names and key information in law enforcement recorder audio.
Owner:先进计算与关键软件(信创)海河实验室 +1

System and method for secure processing of speech signals using pseudo-speech representations

A method, computer program product, and computing system for processing a speech signal. A sensitive portion of the speech signal is identified. A pseudo-speech representation of the sensitive portion is generated using a voice converter system. Speech processing is performed on the speech signal and the pseudo-speech representation of the sensitive portion using a speech processing system.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Scalable model specialization framework for speech model personalization

A method for speech conversion includes obtaining a speech conversion model configured to convert input utterances of human speech directly into corresponding output utterances of synthesized speech. The method further includes receiving a speech conversion request including input audio data corresponding to an utterance spoken by a target speaker associated with atypical speech and a speaker identifier uniquely identifying the target speaker. The method includes activating, using the speaker identifier, a particular sub-model for biasing the speech conversion model to recognize a type of the atypical speech associated with the target speaker identified by the speaker identifier. The method includes converting, using the speech conversion model biased by the activated particular sub-model, the input audio data corresponding to the utterance spoken by the target speaker associated with atypical speech into output audio data corresponding to a synthesized canonical fluent speech representation of the utterance spoken by the target speaker.
Owner:GOOGLE LLC

Emotion regulation instruction recommendation method, device, equipment and medium

The application relates to the fields of artificial intelligence and medical health technology, and discloses a mood regulation instruction recommendation method, device, equipment and medium, wherein the method comprises the following steps: adopting a pre-trained dialogue semantic representation model to generate a dialogue semantic representation vector of psychological counseling data; adopting a pre-trained facial expression representation model to generate a micro-expression representation vector of the psychological counseling data; adopting a pre-trained voice representation model to generate a voice representation vector of the psychological counseling data; adopting a pre-trained abnormal mood comprehensive recognition model to generate a comprehensive representation vector of the dialogue semantic representation vector, the micro-expression representation vector and the voice representation vector; and adopting a pre-trained mood regulation instruction recommendation model to recommend a mood regulation instruction for the comprehensive representation vector and a personal information feature vector corresponding to a counseling object, so as to obtain a recommendation result. Therefore, the accuracy of the recommendation result is improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Digital human mouth shape prediction system and method based on variational auto-encoder in smart agriculture

ActiveCN121236246AClimate change adaptationSemantic analysisPattern recognitionSoft tissue deformation
The invention provides a digital human mouth shape prediction system and method based on a variational auto-encoder in smart agriculture, and relates to the technical field of artificial intelligence, and the method comprises the steps: receiving a voice signal of a smart agriculture scene, extracting multi-dimensional voice features, and fusing the multi-dimensional voice features to generate voice representation; inputting the voice representation into a model, and decoding the voice representation into a mouth shape key point sequence; performing semantic enhancement on the mouth shape key point sequence by combining agricultural proper nouns in the input voice and a preset mouth shape template, and simultaneously inhibiting mouth shape jitter and abnormal frames; and mapping the enhanced mouth shape key point sequence to a three-dimensional digital human face model, converting the enhanced mouth shape key point sequence into face soft tissue deformation data in combination with a face parameter of a target user, calculating a skeleton node position and a rotation parameter according to a preset corresponding relationship between the mouth shape and the skeleton, and generating a digital human mouth shape animation synchronized with the input voice through combined rendering. The stable and accurate voice synchronous digital human mouth shape can be generated in an intelligent agricultural scene.
Owner:XIAODUO INTELLIGENT TECH (BEIJING) CO LTD

Training method of universal voice representation model and related device

The invention discloses a universal voice representation model training method and a related device, and relates to the technical field of artificial intelligence, and the method comprises the steps: obtaining a voice sample and a mask voice sample corresponding to the voice sample, inputting the voice sample into a teacher model, obtaining a teacher voice representation, inputting the mask voice sample into a student model, and obtaining a student voice representation; obtaining a first student voice representation, generating global knowledge distillation loss and local comparison loss according to the teacher voice representation and the first student voice representation, and updating target parameters of the student model by using the global knowledge distillation loss and the local comparison loss; and the trained student model is used as a general voice representation model. According to the method, knowledge distillation and comparative learning are combined, so that the student model can extract multi-aspect voice information, and therefore, the method can adapt to various voice tasks, and the universality is higher. The characteristics of the mask area can be accurately reconstructed through double knowledge migration, so that the accuracy of voice representation extracted by the student model is higher.
Owner:ANHUI LINGDONG GENERAL ROBOT TECHNOLOGY CO LTD

Method, device and storage medium for speech processing

According to an embodiment of the disclosure, a method, apparatus, device and computer-readable storage medium for speech processing are provided. The method includes: acquiring a speech feature sequence corresponding to a speech sample, the speech feature in the speech feature sequence corresponding to a speech frame in the speech sample. For the target speech feature in the speech feature sequence, one or more subsequent speech tokens respectively corresponding to the one or more subsequent speech features are generated based on the target speech feature and the one or more preceding speech features and according to the speech encoding model. A speech encoding model is trained based on the one or more subsequent speech tokens. This training approach enables the speech encoding model to learn high-quality speech representations.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD +1

Unsupervised data selection via discrete speech representation for automatic speech recognition

A method includes obtaining a corpus of unlabeled training data including a plurality of spoken utterances, each corresponding spoken utterance of the plurality of spoken utterances includes audio data characterizing the corresponding spoken utterance. The method also includes receiving a target domain. The method also includes selecting, using a contrastive data selection model, a subset of the utterances from the corpus of unlabeled training data that correspond to the target domain. The method includes training an automatic speech recognition (ASR) model on the subset of utterances.
Owner:GOOGLE LLC

Scalable model specialization framework for speech model personalization

A method for speech conversion includes obtaining a speech conversion model configured to convert input utterances of human speech directly into corresponding output utterances of synthesized speech. The method further includes receiving a speech conversion request including input audio data corresponding to an utterance spoken by a target speaker associated with atypical speech and a speaker identifier uniquely identifying the target speaker. The method includes activating, using the speaker identifier, a particular sub-model for biasing the speech conversion model to recognize a type of the atypical speech associated with the target speaker identified by the speaker identifier. The method includes converting, using the speech conversion model biased by the activated particular sub-model, the input audio data corresponding to the utterance spoken by the target speaker associated with atypical speech into output audio data corresponding to a synthesized canonical fluent speech representation of the utterance spoken by the target speaker.
Owner:GOOGLE LLC

Emotive Text-To-Speech with Auto Detection of Emotions

A method of providing emotive text-to-speech includes obtaining input text characterizing a natural language response generated by an assistant LLM to a query input by a user during a conversation between the user and the assistant LLM, and processing, using the assistant LLM, the input text conditioned on an emotion detection task prompt to predict, as output from the assistant LLM, an emotional state of the natural language response. The method also includes determining, based on the emotional state of the natural language response predicted as output from the assistant LLM, an emotional embedding for the input text and instructing a TTS model to process the input text and the emotional embedding to generate a synthesized speech representation of the natural language response conveying the emotional state of the natural language response as specified by the emotional embedding.
Owner:GOOGLE LLC

A digital human mouth shape prediction system and method based on a variational autoencoder in smart agriculture

ActiveCN121236246BClimate change adaptationSemantic analysisPattern recognitionSoft tissue deformation
The application provides a digital human mouth shape prediction system and method based on a variational autoencoder in smart agriculture, and relates to the technical field of artificial intelligence. The application receives a voice signal of a smart agriculture scene, extracts multi-dimensional voice features and fuses to generate a voice representation. The voice representation is input into a model and decoded into a mouth shape key point sequence. The agricultural specific terms in the input voice are combined with a preset mouth shape template to perform semantic enhancement on the mouth shape key point sequence, while suppressing mouth shape shaking and abnormal frames. The enhanced mouth shape key point sequence is mapped to a three-dimensional digital human face model, combined with face parameters of a target user to convert into face soft tissue deformation data, and then the preset corresponding relationship between the mouth shape and the skeleton is used to calculate the position and rotation parameters of the skeleton nodes. After combination and rendering, a digital human mouth shape animation synchronized with the input voice is generated, which can generate stable and accurate voice synchronized digital human mouth shapes in a smart agriculture scene.
Owner:XIAODUO INTELLIGENT TECH (BEIJING) CO LTD

End-to-end speaker recognition method based on multi-scale SincNet and CGAN

The application belongs to the technical field of speaker recognition, and specifically discloses an end-to-end speaker recognition method based on multi-scale SincNet and CGAN. The multi-scale SincNet is introduced, so that important information is avoided to be lost in manual feature conversion, the multi-scale SincNet captures low-level speech representation of three channels in a waveform according to three customized filter banks, the SincGAN model better captures important narrowband speaker features, the end-to-end recognition is realized based on the improved conditional generative adversarial network, the speaker is recognized by using a small amount of training sentences, and the loss function of the application comprises an adversarial loss of a classic GAN and a classification cross-entropy loss of a classification task. Experimental results show that on TIMIT and LIBRISPEECH corpus, the model of the application shows better performance, and when lacking training data, the model of the application shows stronger robustness than a baseline method.
Owner:SHANDONG UNIV OF SCI & TECH

Method and electronic device for domain adaptation with self-supervised speech representation in speech enhancement

A method for generating a customized speech enhancement model includes: obtaining noisy-clean speech data from a source domain; obtaining noisy voice data from the target domain; acquiring original voice data; training, using the noisy-clean speech data, the noisy speech data, and the original speech data, a customized SE model based on at least one of: self-supervised representation-based adaptation (SSRA), set mapping, or self-supervised adaptation loss; generating a customized SE model by denoising the noisy speech data using a trained customized SE model; and providing the customized SE model to a user device to use the de-noised noisy speech data.
Owner:SAMSUNG ELECTRONICS CO LTD

Zero-Shot Cross-Lingual Voice Transfer for Text-To-Speech

A method for performing zero-shot voice transfer using text-to-speech (TTS) includes receiving an input text sequence characterizing an utterance and receiving a reference speech representation characterizing a reference utterance spoken by a target speaker. The method also includes generating an encoded textual representation for the input text sequence, processing, using a speaker encoder, the reference speech representation to generate a speaker representation characterizing voice characteristics of the target speaker and learning fine-grained embedding vectors based on the speaker representation to obtain a final embedding vector. The method also includes predicting a duration and upsampling the encoded textual representation into an upsampled output. The method also includes generating a synthesized speech representation based on the upsampled output and the final embedding vector and generating a time-domain audio waveform of the input text sequence that clones a voice of the target speaker.
Owner:GOOGLE LLC

A voice cloning method, device, equipment and storage medium thereof

The embodiment of the application belongs to the field of research and development design and audio processing technology, is applied to different pronunciation object voice conversion processing scenes, and relates to a voice cloning method, device and equipment and a storage medium thereof, voice coding sequences are respectively obtained by inputting source voice and reference voice into a pre-trained self-supervised voice representation model; the voice coding sequence of the source voice is discretized; the discretized voice coding is sequentially used as a query coding, and a target number of most similar codes are screened out from the voice coding sequence of the reference voice; according to the most similar codes and the discretization sequence of all query codes, voice coding sequence cloning is performed to obtain cloned voice coding sequences; the cloned voice coding sequences are input into a preset vocoder to output audio waveforms. The application is applied to voice intelligent customer service and virtual voice generation scenes, and personalized voice selection services are more simply and low-costly provided for customer service or customers.
Owner:PING AN TECH (SHENZHEN) CO LTD

A voice conversion method, device, equipment and computer readable storage medium

The application discloses a speech conversion method and device, equipment and a computer readable storage medium, based on a conditional VAE framework, using an encoder to extract the timbre information and content information of the speech respectively, and then using a generator to synthesize the speech, and in order to decouple the speech into content information and timbre information, a pre-trained self-supervised speech representation learning model is introduced to improve the content decoupling capability, and a feature vector x containing source content information C output by the self-supervised speech representation learning model is input into a flow model to obtain a feature vector x ssl As a conditional constraint content information C decoupling, using the forward and reverse characteristics of the flow model, using two KL divergence criteria to guide the model parameter connection, ensuring the lossless of the content information, and modeling the timbre information as a standard normal distribution to reduce the timbre leakage problem, realizing the improvement of the decoupling capability of the content information, and solving the technical problem of timbre leakage.
Owner:PING AN TECH (SHENZHEN) CO LTD

Controllable zero-sample speech conversion method, device, equipment and medium

ActiveCN121034280BPersonalizationEngineering
The application relates to the technical field of voice semantics, can be applied to a business system platform such as financial technology and medical health, and discloses a controllable zero-sample voice conversion method, device, equipment and medium, the method comprises the following steps: performing self-supervised voice learning on unlabeled voice data to obtain a self-supervised voice representation, extracting a content feature vector and a prosody style vector of the self-supervised voice representation, and converting the content feature vector and the prosody style vector into discrete content tokens and discrete prosody tokens; generating a mask for the discrete prosody tokens to obtain target prosody tokens; obtaining reference voice of a target user; extracting a user style embedding in the reference voice; performing flow matching on the discrete content tokens, the target prosody tokens and the user style embedding; generating a target mel-frequency spectrum graph; and performing voice waveform reconstruction and optimization on the target mel-frequency spectrum graph to obtain a zero-sample voice conversion result. The application realizes the zero-sample voice conversion problem of personalization, high fidelity and consistent style under the condition of unlabeled voice data.
Owner:PING AN TECH (SHENZHEN) CO LTD

Using aligned text and speech representations to train automatic speech recognition models without transcribed speech data

A method includes receiving training data that includes unspoken textual utterances in a target language. Each unspoken textual utterance not paired with any corresponding spoken utterance of non-synthetic speech. The method also includes generating a corresponding alignment output for each unspoken textual utterance using an alignment model trained on transcribed speech utterance in one or more training languages each different than the target language. The method also includes generating a corresponding encoded textual representation for each alignment output using a text encoder and training a speech recognition model on the encoded textual representations generated for the alignment outputs. Training the speech recognition model teaches the speech recognition model to learn how to recognize speech in the target language.
Owner:GOOGLE LLC

Voice keyword recognition system based on comparative learning pre-training

A voice keyword recognition system based on comparative learning pre-training comprises the following steps that firstly, a large amount of paired voice-text data easy to obtain is sent into a designed pre-training framework based on ternary comparative learning for pre-training; the objective of the invention is to train and obtain a voice encoder capable of encoding an original voice signal into a universal voice representation and a query encoder capable of encoding an identification keyword into a universal pronunciation representation by optimizing three types of losses including voice element comparison loss, voice-phoneme element comparison loss and phoneme element comparison loss. And then, respectively obtaining universal voice representation and pronunciation representation from a small amount of keyword data on the labeled target data set through a pre-trained voice encoder and a pre-trained query encoder, and sending the two representations into a classifier network for final decision making. The method solves the problem that voice data of a large number of specific keywords is difficult to collect when keywords need to be customized in voice keyword recognition.
Owner:EAST CHINA UNIV OF SCI & TECH