Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

169 results about "Prosody" patented technology

In linguistics, prosody is concerned with those elements of speech that are not individual phonetic segments (vowels and consonants) but are properties of syllables and larger units of speech, including linguistic functions such as intonation, tone, stress, and rhythm. Such elements are known as suprasegmentals.

Face-translator: end-to-end system for speech-translated lip-synchronized and voice preserving video generation

A neural end-to-end system is provided for the face and voice preserving translation of videos. The system is a pipeline of multiple models that produces a video of the original speaker speaking in the target language with modified lip movement to match the target speech, while preserving emphases and prosody of the original speech, and voice characteristics of the original speaker. The pipeline starts with automatic speech recognition including emphasis detection, followed by the translation model. The translated text is then synthesized by a Text-to-Speech model that recreates the original emphases in the target sentence. The resulting synthetic speech is then converted back to the original speakers' voice using a voice conversion model. Finally, to synchronize the lips of the speaker with the translated audio, a generative model generates frames of adapted lip movements which are combined with the audio to produce the final output. The disclosure further describes several use-cases and configurations that apply these techniques to video conferencing, dubbing, low-bandwidth transmission, speech enhancement and assistive technology for the hearing impaired.
Owner:WAIBEL ALEXANDER

Rhythm migration method and device, electronic equipment and storage medium

The invention relates to the technical field of voice processing, and provides a rhythm migration method and device, electronic equipment and a storage medium, and the method comprises the steps: obtaining a decoupled rhythm feature based on a source rhythm voice, and a decoupled tone feature based on the voice of a target speaker, the decoupled rhythm feature represents the rhythm of the source rhythm voice, and the decoupled tone feature represents the tone of the target speaker; the decoupled timbre features represent the timbre of the voice of the target speaker; generating a target voice vector sequence based on the text features of the target text, the voice features of the voice of the target speaker, the decoupled rhythm features and the decoupled timbre features; and synthesizing a target audio based on the target voice vector sequence. According to the method and the device, the decoupled rhythm features and the decoupled timbre features are acquired, and the target voice is generated based on the features, so that the problem of feature mixing is effectively relieved, the timbre purity of the target speaker in cross-person rhythm migration is ensured, the expressive force of rhythm migration is improved, and the synthesized audio is more natural and vivid.
Owner:IFLYTEK CO LTD

Text-to-speech processing

A speech-processing system may be configured to generate expressive synthesized speech. The system may include a prosody prediction model that generates a combination of durations and acoustic representations that may be based on the content of the text as well as additional context information. The model may be trained to predict a joint probability between linguistic representations (e.g., derived from text) and combined duration / acoustic representations (e.g., derived from audio). At inference, the model can process linguistic representations derived from text to predict combined duration / acoustic representations. In some implementations, the model may process additional information; for example, semantic embeddings output by a language model based on the text. In another example, the model may receive a speaker embedding representing voice characteristics of a particular speaker. A decoder may process the durations and acoustic representations output by the model to generate audio data representing the synthesized speech and representing expressive prosodic variation.
Owner:AMAZON TECH INC

Systems, devices, and methods for dynamic synchronization of a prerecorded vocal backing track to a live vocal performance

Disclosed are systems, methods, and devices, that overcome timing and self-expression limitations experienced by vocalists when using prerecorded vocal backing tracks to enhance live performances. The disclosed system, devices, and methods, dynamically synchronizes prerecorded vocal backing tracks with a live vocal stream by extracting vocal elements, such as phonemes, vector embeddings, or vocal audio spectra, from the live vocal performance in real-time. These extracted vocal elements are matched against corresponding timestamped vocal elements previously derived from the prerecorded vocal backing track, enabling precise real-time adjustment and alignment of the backing track timing to the live performance. Additionally, the system enhances expressive performance by identifying prosody factors, such as pitch, vibrato, accent, stress, dynamics, and level, in the live vocal performance, and dynamically adjusting corresponding prerecorded prosody factors within predefined ranges. This maintains naturalness and spontaneity in the vocalist's live performance, overcoming traditional limitations associated with prerecorded vocal backing tracks.
Owner:EIDOL CORP

Intelligent communication content adaptation system and method based on user intention recognition

The invention discloses an intelligent communication content adaptation system and method based on user intention recognition. The system integrates and processes texts, voices, emoticons and unstructured behavior data through a multi-modal input analysis module; the deep learning intention recognition engine adopts a triple attention mechanism, current input and weighted historical interaction features are fused, and multi-level intention classification including basic operation, semantic targets and emotion driving is output; the real-time emotion state analysis module fuses acoustics, semantics and physiological indexes to generate a dynamic emotion matrix; the content adaptation decision engine is based on the intention confidence and the emotional state; and the multi-channel output optimization module performs cooperative adjustment on the speech synthesis rhythm, the text abstract and the visual interface according to the speech synthesis rhythm and the text abstract. According to the method, the perception precision of the deep intention and emotion of the user in a complex scene is remarkably improved, the individuation and multi-channel adaptive optimization of the response content are realized, and the communication efficiency and the user experience are effectively improved.
Owner:CHINA INFOMRAITON CONSULTING & DESIGNING INST CO LTD

Cross-language voice interaction method and system based on multi-modal semantic understanding, and storage medium

The invention discloses a cross-language voice interaction method and system based on multi-modal semantic understanding and a storage medium, and relates to the technical field of artificial intelligence and natural language processing. The method comprises the following steps: acquiring a source language speech stream, performing parallel double-flow feature extraction, extracting text semantic features through a mixed language code recognition model based on a unified phoneme space, and extracting acoustic features containing rhythm information at the same time; performing alignment fusion on the text and the acoustic features by using a cross-modal attention mechanism to generate multi-modal semantic representation; analyzing the explicit intention and the implicit emotion based on the representation, and generating a reply strategy and an emotion control label of the target language; and finally, synthesizing a target voice with a corresponding emotion color. According to the invention, the bottleneck that the traditional cascade architecture loses side language information is broken through, the precise understanding and strategic feedback of complex contexts such as Chinese-English mixed language codes (Code-mixing), anti-quiescence, hesitation and the like are realized, and the method is particularly suitable for transnational business negotiation and international customer service scenes.
Owner:杭州智慧沟通智能科技有限公司

Speech emotion recognition method and system based on multi-modal feature fusion

The invention discloses a speech emotion recognition method and system based on multi-modal feature fusion, and relates to the technical field of speech emotion recognition. The speech emotion recognition method and system based on multi-modal feature fusion comprises the following steps: S1, collecting a speech emotion data set, and carrying out label unified coding and normalization processing; s2, frame-level acoustics and construction of frequency spectrum, rhythm and sound quality emotion features are carried out; s3, a speech emotion representation generation method fusing multi-sub-mode depth coding and a gating cooperative attention mechanism; s4, performing random forest weight initialization and two-order variation grey wolf mapping evaluation; and S5, voice emotion recognition and operation feedback adaptive updating are carried out. According to the method, the feature selection efficiency and the emotion classification accuracy in voice emotion recognition are effectively improved, and the problems that existing voice emotion feature selection is single in stage and single in index, emotion retention and real-time performance are difficult to consider while dimension reduction is performed, and the overall performance is limited are solved.
Owner:HUNAN XIAOYU ZHIHE TECHNOLOGY CO LTD

Parallel tacotron non-autoregressive and controllable TTS

A method for training a non-autoregressive TTS model includes receiving training data that includes a reference audio signal and a corresponding input text sequence. The method also includes encoding the reference audio signal into a variational embedding that disentangles the style / prosody information from the reference audio signal and encoding the input text sequence into an encoded text sequence. The method also includes predicting a phoneme duration for each phoneme in the input text sequence and determining a phoneme duration loss based on the predicted phoneme durations and a reference phoneme duration. The method also includes generating one or more predicted mel-frequency spectrogram sequences for the input text sequence and determining a final spectrogram loss based on the predicted mel-frequency spectrogram sequences and a reference mel-frequency spectrogram sequence. The method also includes training the TTS model based on the final spectrogram loss and the corresponding phoneme duration loss.
Owner:GOOGLE LLC

Synthetic speech processing related to prosody prediction

A speech-processing system receives input data representing text. A prosody prediction component processes the input data to determine prosody embedding data corresponding to prosody of the text. A decoder processes the prosody embedding data and phoneme encoded data derived from the input data to determine audio output data corresponding to the text and the prosody.
Owner:AMAZON TECH INC

Response system and method for user voice interruption in voice interaction process

A response system for user voice interruption in a voice interaction process comprises an information importance mapping (ICM) generation module, an interruption response configuration file (BRP) generation module, an acoustic feature extraction module, a response control module, a semantic understanding module, a rhythm generation module, a transient response trigger and a response manager, according to the method, interruption intentions can be distinguished, the context importance can be evaluated, and differentiated responses with low delay, certainty and controllable sound characteristics can be realized.
Owner:SHANGHAI JIAOTONG UNIV

Multi-user emotion recognition and digital human feedback method based on behavior data

The invention discloses a multi-user emotion recognition and digital human feedback method based on behavior data, and belongs to the field of artificial intelligence emotion calculation. Sensing user approaching, identifying and confirming the identity of a registered user, distributing an identifier, distributing an identifier for a visitor, and establishing an independent session channel in a multi-user scene; capturing time sequence dialogue behavior data containing voice rhythm features, dialogue interaction modes and linguistic features in real time; inputting the data into a sequence information processing model based on a state space theory and provided with a data dependence selection mechanism and an emotion analysis model spliced by static context features, and outputting multi-dimensional emotion and cognitive state tags; generating an emotion support strategy through a decision-making model in combination with the personal file of the user; and generating a multi-mode control instruction, and driving the digital human to synchronously present special effects such as expressions and the like. According to the method, privacy concerns are eliminated through non-intrusive collection, the method is adaptive to multi-user scenes, deep cognitive states can be recognized, and interaction effectiveness is improved for a long time.
Owner:SICHUAN UNIV JINCHENG INST

Spoken English pronunciation quality evaluation method based on multi-mode speech feature analysis

ActiveCN121528247ASpeech analysisFeature extractionModal voice
The invention belongs to the technical field of speech analysis, and discloses a spoken English pronunciation quality evaluation method based on multi-modal speech feature analysis, which comprises the following steps: acquiring a spoken English speech signal of a target user, and performing multi-domain decomposition on the speech signal to obtain multi-modal speech features; performing time-frequency domain corresponding relation analysis and feature extraction on the multi-mode speech features to obtain a pronunciation detail feature sequence; performing multi-scale matching on the pronunciation detail feature sequence and a preset standard pronunciation template, constructing a multi-dimensional representation model based on a multi-scale matching result, and calculating a fine-grained quality score of a phoneme unit in each multi-dimensional representation model in combination with rhythm and rhythm parameters in the multi-modal speech features, the rhythm coherence score and the overall fluency score are fused to generate a comprehensive pronunciation quality evaluation result and a visual diagnosis report of pronunciation deviation; the oral English pronunciation evaluation method realizes comprehensive and refined evaluation of oral English pronunciation, and provides a scientific guidance basis for personalized language learning.
Owner:ZHANG ZHOU HALTH VOCATIONAL COLLEGE

A spoken dialogue data processing method based on an agent

The application discloses a kind of oral conversation data processing methods based on agent, belong to artificial intelligence big data processing technical field, including receiving the voice input file sent by user end, voice signal is converted into text data, while extracting emotional characteristics, generate emotional label and voice prosody characteristics, text and emotional information are combined to form dialogue anchor point information, construct path graph, calculate edge weight;According to the similarity of current dialogue anchor point and historical path node, construct propagation matrix, carry out intelligent path reasoning, calculate path score, ensure that emotional fluctuation influences task path selection, agent selects optimal path according to path score to carry out subsequent task processing, and real-time correction is carried out after user feedback;Through emotional analysis technology, real-time adjust the anchor point information of current conversation, and adjust the reasoning process according to feedback, form closed-loop adaptive adjustment mechanism, to optimize the response quality of system, enhance the sense of trust of user to system.
Owner:SHANDONG LINGCHAO SOFTWARE TECH CO LTD

Multi-language intelligent dubbing generation system based on sound cloning and emotion migration

The invention discloses a multi-language intelligent dubbing generation system based on sound cloning and emotion migration. The multi-language intelligent dubbing generation system comprises a sound cloning module, a cross-language synthesis module, an emotion migration module and a lip shape synchronization module. By constructing a few-sample speaker encoder, a cross-language rhythm migration module, a fine-grained emotion control module and a video lip shape synchronization module, end-to-end automatic generation from original dubbing audio to multi-language target dubbing is realized, and tone consistency, emotion authenticity and picture synchronism are kept.
Owner:JIANGSU HOPERUN SOFTWARE CO LTD

Children language narrative ability evaluation method and tool based on multi-modal analysis

The invention discloses a children's language narrative ability assessment method and tool based on multi-modal analysis, and relates to the technical field of natural language processing, and the technical scheme is characterized in that audio data and video data of a testee during a narrative process are collected, and voice rhythm features, text language features and visual behavior features are extracted for fusion; the method comprises the steps of inputting a multi-modal scoring model to generate multi-dimensional quantitative scores of a macrostructure, a microstructure, a language organization and a pragmatic function of the language narrative ability of a testee child, generating a structured evaluation report for the testee child according to the scores, and giving personalized intervention suggestions. According to the method, multi-modal data such as voice, language texts and visual behaviors are collected, so that the narrative ability of the children can be comprehensively analyzed from multiple dimensions such as macrostructures, microstructures, language organizations and pragmatic functions, and the method is closer to the real natural expression state of the children; therefore, the comprehensive language ability and communication performance of children can be reflected more comprehensively and truly.
Owner:TONGJI HOSPITAL ATTACHED TO TONGJI MEDICAL COLLEGE HUAZHONG SCI TECH

ASV system risk assessment method and system based on multi-dimensional pronunciation characterization decoupling and fusion

PendingCN121354597ASpeech synthesisFeature extractionSpeaking style
The invention provides an ASV system risk assessment method and system based on multi-dimensional pronunciation characterization decoupling and fusion, and the method comprises the steps: obtaining a reference audio of a target user, and carrying out the decoupling extraction of a volume feature vector, a pitch feature vector, and a speaking style feature vector from the reference audio through a multi-dimensional pronunciation feature extraction network; a text to be verified is converted into a phoneme sequence, the phoneme sequence is input into a pronunciation feature prediction network for predicting a frame-level dynamic pronunciation feature sequence based on the phoneme sequence, and three pronunciation feature vectors are injected into the network through an adaptive instance normalization mechanism to dynamically modulate a prediction process. Generating a frame-level dynamic pronunciation feature sequence containing time sequence rhythm change; inputting a VAE-GAN synthesis trunk, re-injecting the three pronunciation feature vectors through an adaptive instance normalization mechanism, and generating a test voice sample consistent with the voiceprint characteristics of the target user; and initiating an identity verification query for the ASV system, and calculating safety indexes of different user groups according to a verification result so as to evaluate the risk of the ASV system.
Owner:FUJIAN NORMAL UNIV

TTS audio super-division method and device based on text semantics, equipment and medium

The invention discloses a TTS audio super-division method and device based on text semantics, equipment and a medium, and relates to the technical field of audio super-division, and the method comprises the steps: obtaining an initial TTS audio and initial text data, and carrying out the preprocessing of the initial TTS audio and initial text data, so as to obtain a preprocessed audio and a preprocessed text; performing feature extraction on the preprocessed text to obtain semantic features and rhythm features, and performing feature extraction on the preprocessed audio to obtain audio features; and fusing the semantic feature, the rhythm feature and the audio feature to obtain a fused feature, and generating a target TTS audio corresponding to the initial TTS audio by using the fused feature to realize audio super-division of the initial TTS audio. By introducing text features and rhythm features into audio super-division, the problem of semantic expression distortion caused by a pure audio driving method is solved.
Owner:MALANSHAN AUDIO & VIDEO LABORATORY

Anonymization privacy protection method and system for voice information retention

The embodiment of the invention provides an anonymization privacy protection method and system for voice information retention. The method comprises the following steps: extracting speaker embedding of an original audio, eliminating the tone of the speaker, and keeping semantic and rhythm speaker irrelevant features; embedding and inputting a speaker into a speaker anonymous module matched with a three-stage stream based on a U-Net architecture to obtain anonymous embedding; and combining irrelevant features of the speaker with anonymous embedding by using a pre-trained voice reconstruction model to generate anonymized voice with tone privacy. According to the embodiment of the invention, voice anonymization facing content privacy and tone privacy reserved by voice information is realized, the effectiveness of the anonymized voice generated by using the method in a downstream task is superior to that of a baseline model, meanwhile, the privacy of a speaker is also guaranteed, and safe use of data is realized.
Owner:SHANGHAI JIAOTONG UNIV

English pronunciation error correction training method based on speech recognition

The invention relates to the technical field of speech recognition and processing, in particular to an English pronunciation error correction training method based on speech recognition, and the method comprises the following steps: S1, collecting a speech signal generated by a learner in a pronunciation training process, digitalizing the speech signal, associating the digitalized speech signal with a target standard text, and generating an original audio data record with a timestamp; according to the invention, phoneme-level decoding is carried out on the voice signal by using the recurrent neural network acoustic model, and accurate alignment of the pronunciation of the learner and the standard phoneme sequence is realized in combination with the dynamic time warping algorithm, so that pronunciation errors such as misreading, missed reading and increased reading can be accurately identified; meanwhile, acoustic features such as Mel frequency cepstrum coefficient, pitch and fundamental frequency are extracted to be quantitatively compared with a standard native language pronunciation database, multi-dimensional evaluation covering accuracy, integrity, fluency and rhythm is generated, and the accuracy and systematicness of oral English pronunciation error correction are remarkably improved.
Owner:吕丽沙

Voice data generation method based on large model and method for training large model

The invention provides a voice data generation method based on a large model and a method for training the large model, and relates to the technical field of artificial intelligence, in particular to the technical fields of voice generation, intelligent customer service, video production and the like. The large model-based voice data generation method comprises the following steps of: receiving a rhythm description text and a voice text, wherein the rhythm description text describes a pronunciation rhythm intention of a plurality of text words in the voice text; semantic fusion is conducted on the rhythm description text and the voice text through a large model, rhythm fusion features are obtained, sub-features in the rhythm fusion features represent the pronunciation rhythm of the voice segments for the text words, and the to-be-generated target voice data comprise the voice segments; and according to the rhythm fusion feature and a specified pronunciation attribute related to the specified object, generating target voice data which represents that the specified object pronunciates according to the pronunciation rhythm intention and corresponds to the voice text.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Electronic apparatus, terminal apparatus and controlling method thereof

An electronic apparatus, a terminal apparatus, and a controlling method thereof. The electronic apparatus includes an input interface; and a processor including a prosody module configured to extract an acoustic feature and a vocoder module configured to generate a speech waveform, wherein the processor is configured to: receive a text input using the input interface; identify a first acoustic feature from the text input using the prosody module, wherein the first acoustic feature corresponds to a first sampling rate; generate a modified acoustic feature corresponding to a modified sampling rate different from the first sampling rate, based on the identified first acoustic feature; and generate a plurality of vocoder learning models by training the vocoder module based on the first acoustic feature and the modified acoustic feature.
Owner:SAMSUNG ELECTRONICS CO LTD

Voice synthesis model training method, voice synthesis method, and related device

The application provides a speech synthesis model training method, a speech synthesis method and related equipment. The method comprises the following steps: obtaining training data, the training data comprising target speech and a phoneme sequence corresponding to the target speech; preprocessing the target speech to determine a target mel-frequency spectrum; inputting the phoneme sequence into a speech synthesis model for synthesis processing to obtain a predicted mel-frequency spectrum; dividing and pairing the target mel-frequency spectrum and the predicted mel-frequency spectrum according to a sound rule of the target speech to obtain N spectrum segment pairs; using N discriminators in an adversarial discriminant model to perform adversarial generation training on the speech synthesis model based on the N spectrum segment pairs; and using the trained speech synthesis model to synthesize text into synthesized speech. The technical problems of ambiguous pronunciation and excessive smoothness are solved. The synthesized speech is clear, the pronunciation is more natural, the rhythm and prosody are better, and the sound is closer to that of a real person.
Owner:MASHANG CONSUMER FINANCE CO LTD

Artificial intelligence-based speech synthesis method and device, computer equipment and medium

The application is suitable for the technical field of speech synthesis, and particularly relates to a speech synthesis method and device based on artificial intelligence, computer equipment and a medium. The application extracts a text feature vector of a target text through a feature extraction model, predicts the text feature vector through a stress predictor, outputs a stress prediction vector, adds the stress prediction vector to the text feature vector to obtain a text stress vector, predicts the text feature vector through a pause predictor, outputs a pause prediction vector, adds the pause prediction vector to the text feature vector to obtain a text pause vector, predicts the text stress vector and the text pause vector through a prosody predictor, outputs a text prosody vector, matches the text prosody vector with a phoneme sequence of the target text, obtains a phoneme sequence with prosody labels, performs speech conversion on the phoneme sequence with prosody labels, obtains synthesized speech, and through the prediction of stress, pause and prosody, the expressiveness, naturalness and accuracy of the synthesized speech are improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Voice dictation with audio large language model

A method comprising receiving audio data 102, generating a transcription 151 comprising a sequence of terms 152, such as “Buy some tomatoes and bananas. Change tomatoes to potatoes”, parallel processing the audio data and the transcription using a multimodal large language model (LLM) 150 to identify one or more revision terms 152R, for example “Change”, specifying a revision action to perform on at least one other term in the sequence, in this instance “tomatoes”, and modifying the transcription 151M accordingly – “Buy some potatoes and bananas”. Identifying the revision term(s) may be based on a corresponding user intent 154 determined for each respective term in the sequence, for example, the user 10 does not intend the final transcription to include “tomatoes”. For each term in the sequence, parallel processing may comprise correlating its speech characteristics 156 such as pitch, tone or prosody information determined from the audio data with its corresponding linguistic context 158 determined from the transcription. Transcription correction may be based on a revision token inserted into the sequence, the token indicating an N number of terms for replacement and their corresponding replacement terms. User context data 104 may be obtained to tailor the LLM to a particular user. [Figure 1A]
Owner:GOOGLE LLC

Intelligent control method and system of massage equipment based on rhythm information

The invention discloses an intelligent control method and system for massage equipment based on rhythm information, and relates to the technical field of voice recognition processing. The massage equipment is intelligently controlled through rhythm information, synchronous matching of music rhythm, user respiration and voice signals is achieved, massage actions are more natural and harmonious, and comfort and immersion can be remarkably improved; meanwhile, space positioning and pressure response adjustment are combined, the accuracy and individuation of massage parts and force are guaranteed, and the equipment can dynamically adjust a control strategy according to the state of a user.
Owner:LIANYUNGANG FIRST PEOPLES HOSPITAL +1

Voice quality evaluation method and device, equipment and storage medium

The invention discloses a voice quality evaluation method and device, equipment and a storage medium, and relates to the technical field of computers. The method comprises the steps of obtaining to-be-evaluated voice generated through text-to-voice conversion and style information for the to-be-evaluated voice, and extracting basic features of the to-be-evaluated voice; extracting text-to-speech features from the basic features; the text-to-speech features comprise any one or more of rhythm features, tone features and text matching features; calculating a style matching degree between an actual style embedding vector corresponding to the text-to-speech feature and a style embedding vector template corresponding to the style information; and predicting a speech degradation score based on the text-to-speech features, and determining speech quality based on the speech degradation score and the style matching degree. By extracting rhythm features, timbre features and text matching features, identifying specific quality defects in text-to-speech conversion; and in combination with the style matching degree, the problem of confusion of style differences and quality defects is solved, and the accuracy of voice quality scoring is improved.
Owner:MALANSHAN AUDIO & VIDEO LABORATORY

Voice synthesis defect correction method and device for vq tts model and storage medium

The application relates to the technical field of voice synthesis, in particular to a voice synthesis defect correction method and device for a VQTTS model and a storage medium, which comprises the following steps: acquiring synthesized text, positioning a text with a synthesis defect position as a defect phrase T; using a large language model to generate M texts W containing the defect phrase T; using VQTTS to perform voice synthesis on the text W; if it is judged that the synthesized voice does not have defects, intercepting a segment and adding the segment to a set K; adding a (T, K) pair to a system data set; acquiring an input text W to be synthesized, generating M Oracle vectors with a length of K; if the defect phrase T is a substring of the input text W, updating the Oracle vector; using VQTTS and a Beam Search algorithm to generate a corrected prosody label sequence; selecting an optimal discrete prosody label sequence and generating voice. The application can correct synthesis defects without updating the model, and solves the technical problem of voice synthesis defect repair.
Owner:AISPEECH CO LTD

Controllable zero sample voice conversion method, device, equipment and medium

The invention relates to the technical field of voice semantics, can be applied to business system platforms of financial science and technology, medical treatment and health and the like, and discloses a controllable zero sample voice conversion method, device, equipment and medium, the method comprises the following steps: carrying out self-supervised voice learning on unlabeled voice data to obtain self-supervised voice representation; the method comprises the following steps of: extracting a content feature vector and a rhythm style vector represented by self-supervised speech, converting the content feature vector and the rhythm style vector into a discrete content token and a discrete rhythm token, performing mask generation on the discrete rhythm token to obtain a target rhythm token, obtaining reference speech of a target user, extracting user style embedding in the reference speech, and obtaining a target user. And performing stream matching on the discrete content token, the target rhythm token and the user style embedding to generate a target Mel spectrogram, and performing voice waveform reconstruction and optimization on the target Mel spectrogram to obtain a zero sample voice conversion result. According to the invention, under the condition of no annotated voice data, personalized, high-fidelity and style-consistent zero-sample voice conversion is realized.
Owner:PING AN TECH (SHENZHEN) CO LTD

Voice conversion method and apparatus

The application discloses a voice conversion method and device. The method comprises the following steps: aligning a frame-level acoustic feature sequence with a phoneme-level text feature sequence to generate a phoneme-level acoustic feature sequence of attention content information; generating a phoneme-level hidden variable sequence carrying content and acoustic information simultaneously based on the phoneme-level text feature sequence, the phoneme-level acoustic feature sequence and a target object identifier; inputting the phoneme-level hidden variable sequence into a trained duration prediction network to obtain a predicted duration sequence corresponding to the phoneme-level hidden variable sequence; expanding the duration of each phoneme-level hidden variable in the phoneme-level hidden variable sequence based on the predicted duration sequence to obtain a frame-level hidden variable sequence; and generating target audio corresponding to the target object identifier based on the frame-level hidden variable sequence. The application can not only retain the emotion of the source audio and not reveal the timbre, but also generate audio closer to the prosody and timbre of the target speaker.
Owner:SHANGHAI HODE INFORMATION TECH CO LTD

A speech prosody recognition method, system, device and storage medium

This invention discloses a speech prosody recognition method, system, device, and storage medium. First, the voice from a customer service call is captured to obtain a dialogue speech signal. Then, the dialogue speech signal is preprocessed to obtain a preprocessed dialogue speech file. Next, the preprocessed dialogue speech file is vectorized to obtain a corresponding feature matrix. The feature matrix is ​​input into a trained prosody model to obtain the model calculation result. Then, based on Mandarin and dialect templates, corresponding template thresholds are obtained. Using the model calculation result and template thresholds, the prosody recognition result of the feature matrix is ​​obtained. Finally, based on the prosody recognition result, the feature matrix is ​​processed by text mapping to obtain the dialogue word order text. This invention effectively improves the recognition accuracy and efficiency for speech containing dialects.
Owner:BEIJING GARUI INTELLIGENT TECH GRP CO LTD