Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

217 results about "Prosody" patented technology

In linguistics, prosody is concerned with those elements of speech that are not individual phonetic segments (vowels and consonants) but are properties of syllables and larger units of speech, including linguistic functions such as intonation, tone, stress, and rhythm. Such elements are known as suprasegmentals.

Face-translator: end-to-end system for speech-translated lip-synchronized and voice preserving video generation

A neural end-to-end system is provided for the face and voice preserving translation of videos. The system is a pipeline of multiple models that produces a video of the original speaker speaking in the target language with modified lip movement to match the target speech, while preserving emphases and prosody of the original speech, and voice characteristics of the original speaker. The pipeline starts with automatic speech recognition including emphasis detection, followed by the translation model. The translated text is then synthesized by a Text-to-Speech model that recreates the original emphases in the target sentence. The resulting synthetic speech is then converted back to the original speakers' voice using a voice conversion model. Finally, to synchronize the lips of the speaker with the translated audio, a generative model generates frames of adapted lip movements which are combined with the audio to produce the final output. The disclosure further describes several use-cases and configurations that apply these techniques to video conferencing, dubbing, low-bandwidth transmission, speech enhancement and assistive technology for the hearing impaired.
Owner:WAIBEL ALEXANDER

Tone conversion method and device based on cultural semantics, equipment and medium

The invention relates to the technical field of artificial intelligence, can be applied to the field of medical health, and discloses a timbre conversion method, device, equipment and medium based on cultural semanteme, the method comprises the following steps: constructing a cultural semantic timbre library comprising a semantic label and timbre characteristic parameter mapping relationship, the semantic label characterizing emotional semanteme of a target timbre, and the timbre characteristic parameter mapping relationship between the semantic label and the timbre characteristic parameter; the timbre characteristic parameters comprise a pitch range, rhythm rhythm and a harmonic structure; performing feature extraction based on text, image and audio multi-mode information to obtain semantic keywords, visual emotion features and audio acoustic features; performing attention weight fusion on the features through a multi-modal fusion deep learning model, and dynamically adjusting model parameters in combination with a semantic timbre library to generate a target timbre; and finally, intelligent conversion from the multi-mode information to the adaptive tone is realized. Through semantic-driven multi-modal feature collaborative optimization, the defect that timbre conversion machinery is stiff and lacks emotional expression is overcome, and the integrating degree of timbre expression and semantic scenes is improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Speech cloning system and method fusing rhythm characteristics

The invention discloses a voice cloning system and method fusing rhythm characteristics, belongs to the technical field of voice synthesis and natural language, and is applied to the aspect of fine-grained rhythm control in zero-sample voice synthesis. The implementation method comprises the following steps of: 1, extracting rhythm features and audio features of an audio file, and further respectively acquiring pause features, speed features and tone features in the rhythm features by sequentially adopting transcriptional text inverse coding, syllable-level speed registration quantization and pitch sequence feature splicing modes; 2, fusing the features of the audio files in a manner of discarding feature screening without guidance of a classifier; 3, generating a target Mel spectrogram based on conditional flow matching; generating a target audio file from a to-be-cloned audio file through the trained voice cloning model controlled by the fusion rhythm; compared with the prior art, fine-grained rhythm control of tone and rhythm feature decoupling is realized in zero-sample speech synthesis, so that intonation accuracy based on a context scene is improved.
Owner:BEIJING INST OF TECH

Education large model evaluation method oriented to spoken language practice scene

The invention discloses a spoken language practice scene-oriented education large model evaluation method, which comprises the following steps of: obtaining voice data and a corresponding standard text, interacting with a test model through the voice data, and obtaining a recognition text, audio data and an answer text provided by the test model; the standard text and the recognition text are compared and evaluated, and the voice recognition accuracy is obtained; carrying out pronunciation accuracy, fluency and rhythm evaluation on the audio data based on multi-modal feature fusion to obtain an audio score; performing grammar accuracy evaluation on the answer text based on a dynamic punishment mechanism and the constructed prompt large model, and performing question degree evaluation on the answer text based on the constructed prompt large model to obtain an answer text score; and comprehensively summarizing the voice recognition accuracy, the audio score and the answer text score to obtain a final evaluation result.
Owner:BEIJING NORMAL UNIVERSITY

Rhythm migration method and device, electronic equipment and storage medium

The invention relates to the technical field of voice processing, and provides a rhythm migration method and device, electronic equipment and a storage medium, and the method comprises the steps: obtaining a decoupled rhythm feature based on a source rhythm voice, and a decoupled tone feature based on the voice of a target speaker, the decoupled rhythm feature represents the rhythm of the source rhythm voice, and the decoupled tone feature represents the tone of the target speaker; the decoupled timbre features represent the timbre of the voice of the target speaker; generating a target voice vector sequence based on the text features of the target text, the voice features of the voice of the target speaker, the decoupled rhythm features and the decoupled timbre features; and synthesizing a target audio based on the target voice vector sequence. According to the method and the device, the decoupled rhythm features and the decoupled timbre features are acquired, and the target voice is generated based on the features, so that the problem of feature mixing is effectively relieved, the timbre purity of the target speaker in cross-person rhythm migration is ensured, the expressive force of rhythm migration is improved, and the synthesized audio is more natural and vivid.
Owner:IFLYTEK CO LTD

Speech synthesis model training method, speech synthesis method, electronic device, and storage medium

The present disclosure provides a speech synthesis model training method. The method can comprise: acquiring initial training data; selecting a plurality of first text segments, second text segments, and third text segments from coherent text, and acquiring first audio segments, second audio segments, and third audio segments from coherent audio; detecting whether the plurality of audio segments and text segments meet a stitching condition; on the basis of the text sequence, stitching, by alternatively arranging text and audio, the plurality of text segments and audio segments meeting the stitching condition to generate combined training data, so as to generate a combined training data set; on the basis of the combined training data set, training the initial speech generation model to obtain a trained speech synthesis model, wherein during the training, prosodic, tonal, and / or emotional features from the second text segments and / or the third text segments are extracted. Further disclosed is a speech synthesis method. The present disclosure allows for use of the context of the text and audio to realize text comprehension to capture prosody, tonality, and emotion in the context.
Owner:SHANGHAI XIYU JIZHI TECH CO LTD

Enhanced wireless communication handover management system

To effectively resolve the issue of inappropriate handovers during communication sessions, methods, systems, and machine-readable mediums which utilize speech features captured by microphones in the original wireless peripheral device and / or the wireless peripheral device to which the communication session is to be handed over to determine if a handover should proceed or should be reversed. Speech features refer to the composite attributes of spoken language that encompass both acoustic and linguistic features. Acoustic features characterize the sound properties of speech and include, but are not limited to, timbre, pitch, intonation, speaking rate, articulation, prosody, melody, spectral features, formant frequencies, and the like. Linguistic features pertain to the actual content conveyed, comprising words, phrases, syntax, and semantics. This includes the analysis of words, phrases, syntax, and semantics to understand the context and continuity of the conversation.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Text-to-speech processing

A speech-processing system may be configured to generate expressive synthesized speech. The system may include a prosody prediction model that generates a combination of durations and acoustic representations that may be based on the content of the text as well as additional context information. The model may be trained to predict a joint probability between linguistic representations (e.g., derived from text) and combined duration / acoustic representations (e.g., derived from audio). At inference, the model can process linguistic representations derived from text to predict combined duration / acoustic representations. In some implementations, the model may process additional information; for example, semantic embeddings output by a language model based on the text. In another example, the model may receive a speaker embedding representing voice characteristics of a particular speaker. A decoder may process the durations and acoustic representations output by the model to generate audio data representing the synthesized speech and representing expressive prosodic variation.
Owner:AMAZON TECH INC

Intelligent reminding system and method for medicine use in elderly department based on voice interaction

The invention provides a voice interaction-based intelligent medicine use reminding system and method for the elderly department. The voice interaction-based intelligent medicine use reminding method comprises the following steps of: extracting phoneme characteristics of a target elderly in each historical dialogue round from each interaction voice; determining the pronunciation loss of the target old person in each historical dialogue round according to the phoneme similarity between the interactive voices in each historical dialogue round and the rhythm feature of each interactive voice, and further determining the semantic constraint condition of the target old person in each historical dialogue round; performing multi-intention recognition on the interaction voice of the target old person in the current dialogue round to obtain different semantic intentions of the target old person in the current dialogue round; further determining the semantic integrity of the target old person in the current dialogue round; and correcting the voice content of the target old person in the current dialogue round through the semantic integrity and the intention confidence of each semantic intention. By adopting the scheme of the invention, the fuzzy and distorted voice content of the elderly can be corrected in the voice interaction process.
Owner:CHONGQING NO 3 PEOPLES HOSPITAL

Systems, devices, and methods for dynamic synchronization of a prerecorded vocal backing track to a live vocal performance

Disclosed are systems, methods, and devices, that overcome timing and self-expression limitations experienced by vocalists when using prerecorded vocal backing tracks to enhance live performances. The disclosed system, devices, and methods, dynamically synchronizes prerecorded vocal backing tracks with a live vocal stream by extracting vocal elements, such as phonemes, vector embeddings, or vocal audio spectra, from the live vocal performance in real-time. These extracted vocal elements are matched against corresponding timestamped vocal elements previously derived from the prerecorded vocal backing track, enabling precise real-time adjustment and alignment of the backing track timing to the live performance. Additionally, the system enhances expressive performance by identifying prosody factors, such as pitch, vibrato, accent, stress, dynamics, and level, in the live vocal performance, and dynamically adjusting corresponding prerecorded prosody factors within predefined ranges. This maintains naturalness and spontaneity in the vocalist's live performance, overcoming traditional limitations associated with prerecorded vocal backing tracks.
Owner:EIDOL CORP

Intelligent communication content adaptation system and method based on user intention recognition

The invention discloses an intelligent communication content adaptation system and method based on user intention recognition. The system integrates and processes texts, voices, emoticons and unstructured behavior data through a multi-modal input analysis module; the deep learning intention recognition engine adopts a triple attention mechanism, current input and weighted historical interaction features are fused, and multi-level intention classification including basic operation, semantic targets and emotion driving is output; the real-time emotion state analysis module fuses acoustics, semantics and physiological indexes to generate a dynamic emotion matrix; the content adaptation decision engine is based on the intention confidence and the emotional state; and the multi-channel output optimization module performs cooperative adjustment on the speech synthesis rhythm, the text abstract and the visual interface according to the speech synthesis rhythm and the text abstract. According to the method, the perception precision of the deep intention and emotion of the user in a complex scene is remarkably improved, the individuation and multi-channel adaptive optimization of the response content are realized, and the communication efficiency and the user experience are effectively improved.
Owner:CHINA INFOMRAITON CONSULTING & DESIGNING INST CO LTD

Cross-language voice interaction method and system based on multi-modal semantic understanding, and storage medium

The invention discloses a cross-language voice interaction method and system based on multi-modal semantic understanding and a storage medium, and relates to the technical field of artificial intelligence and natural language processing. The method comprises the following steps: acquiring a source language speech stream, performing parallel double-flow feature extraction, extracting text semantic features through a mixed language code recognition model based on a unified phoneme space, and extracting acoustic features containing rhythm information at the same time; performing alignment fusion on the text and the acoustic features by using a cross-modal attention mechanism to generate multi-modal semantic representation; analyzing the explicit intention and the implicit emotion based on the representation, and generating a reply strategy and an emotion control label of the target language; and finally, synthesizing a target voice with a corresponding emotion color. According to the invention, the bottleneck that the traditional cascade architecture loses side language information is broken through, the precise understanding and strategic feedback of complex contexts such as Chinese-English mixed language codes (Code-mixing), anti-quiescence, hesitation and the like are realized, and the method is particularly suitable for transnational business negotiation and international customer service scenes.
Owner:杭州智慧沟通智能科技有限公司

Speech emotion recognition method and system based on multi-modal feature fusion

The invention discloses a speech emotion recognition method and system based on multi-modal feature fusion, and relates to the technical field of speech emotion recognition. The speech emotion recognition method and system based on multi-modal feature fusion comprises the following steps: S1, collecting a speech emotion data set, and carrying out label unified coding and normalization processing; s2, frame-level acoustics and construction of frequency spectrum, rhythm and sound quality emotion features are carried out; s3, a speech emotion representation generation method fusing multi-sub-mode depth coding and a gating cooperative attention mechanism; s4, performing random forest weight initialization and two-order variation grey wolf mapping evaluation; and S5, voice emotion recognition and operation feedback adaptive updating are carried out. According to the method, the feature selection efficiency and the emotion classification accuracy in voice emotion recognition are effectively improved, and the problems that existing voice emotion feature selection is single in stage and single in index, emotion retention and real-time performance are difficult to consider while dimension reduction is performed, and the overall performance is limited are solved.
Owner:HUNAN XIAOYU ZHIHE TECHNOLOGY CO LTD

Audio speech signal analysis for fraud detection

A device, system and method for analyzing audio speech signals to detect fraudulent calls to a contact center comprising splitting an audio recording of a call in real-time into a foreground speech signal attributed to a main speaker and a background audio signal, extracting audio features from the foreground speech signal and background audio signal, inputting the extracted audio features into an ensemble model comprising multiple different machine learning models co-trained to cumulatively detect fraud, wherein the multiple different machine learning models include: a speaker audio model to detect audio speech anomalies, a speaker intent model to classify intent of the main speaker, a synthetic voice detection model to identify a non-human entity, and a prosody model to detect voice intonation of the main speaker. A prediction may be output, by the ensemble model, indicating whether the call is fraudulent.
Owner:MORGAN STANLEY SERVICES GROUP INC

Parallel tacotron non-autoregressive and controllable TTS

A method for training a non-autoregressive TTS model includes receiving training data that includes a reference audio signal and a corresponding input text sequence. The method also includes encoding the reference audio signal into a variational embedding that disentangles the style / prosody information from the reference audio signal and encoding the input text sequence into an encoded text sequence. The method also includes predicting a phoneme duration for each phoneme in the input text sequence and determining a phoneme duration loss based on the predicted phoneme durations and a reference phoneme duration. The method also includes generating one or more predicted mel-frequency spectrogram sequences for the input text sequence and determining a final spectrogram loss based on the predicted mel-frequency spectrogram sequences and a reference mel-frequency spectrogram sequence. The method also includes training the TTS model based on the final spectrogram loss and the corresponding phoneme duration loss.
Owner:GOOGLE LLC

Synthetic speech processing related to prosody prediction

A speech-processing system receives input data representing text. A prosody prediction component processes the input data to determine prosody embedding data corresponding to prosody of the text. A decoder processes the prosody embedding data and phoneme encoded data derived from the input data to determine audio output data corresponding to the text and the prosody.
Owner:AMAZON TECH INC

Response system and method for user voice interruption in voice interaction process

A response system for user voice interruption in a voice interaction process comprises an information importance mapping (ICM) generation module, an interruption response configuration file (BRP) generation module, an acoustic feature extraction module, a response control module, a semantic understanding module, a rhythm generation module, a transient response trigger and a response manager, according to the method, interruption intentions can be distinguished, the context importance can be evaluated, and differentiated responses with low delay, certainty and controllable sound characteristics can be realized.
Owner:SHANGHAI JIAOTONG UNIV

Multi-user emotion recognition and digital human feedback method based on behavior data

The invention discloses a multi-user emotion recognition and digital human feedback method based on behavior data, and belongs to the field of artificial intelligence emotion calculation. Sensing user approaching, identifying and confirming the identity of a registered user, distributing an identifier, distributing an identifier for a visitor, and establishing an independent session channel in a multi-user scene; capturing time sequence dialogue behavior data containing voice rhythm features, dialogue interaction modes and linguistic features in real time; inputting the data into a sequence information processing model based on a state space theory and provided with a data dependence selection mechanism and an emotion analysis model spliced by static context features, and outputting multi-dimensional emotion and cognitive state tags; generating an emotion support strategy through a decision-making model in combination with the personal file of the user; and generating a multi-mode control instruction, and driving the digital human to synchronously present special effects such as expressions and the like. According to the method, privacy concerns are eliminated through non-intrusive collection, the method is adaptive to multi-user scenes, deep cognitive states can be recognized, and interaction effectiveness is improved for a long time.
Owner:SICHUAN UNIV JINCHENG INST

Spoken English pronunciation quality evaluation method based on multi-mode speech feature analysis

The invention belongs to the technical field of speech analysis, and discloses a spoken English pronunciation quality evaluation method based on multi-modal speech feature analysis, which comprises the following steps: acquiring a spoken English speech signal of a target user, and performing multi-domain decomposition on the speech signal to obtain multi-modal speech features; performing time-frequency domain corresponding relation analysis and feature extraction on the multi-mode speech features to obtain a pronunciation detail feature sequence; performing multi-scale matching on the pronunciation detail feature sequence and a preset standard pronunciation template, constructing a multi-dimensional representation model based on a multi-scale matching result, and calculating a fine-grained quality score of a phoneme unit in each multi-dimensional representation model in combination with rhythm and rhythm parameters in the multi-modal speech features, the rhythm coherence score and the overall fluency score are fused to generate a comprehensive pronunciation quality evaluation result and a visual diagnosis report of pronunciation deviation; the oral English pronunciation evaluation method realizes comprehensive and refined evaluation of oral English pronunciation, and provides a scientific guidance basis for personalized language learning.
Owner:ZHANG ZHOU HALTH VOCATIONAL COLLEGE

Voice synthesis method, device and equipment based on potential rhythm of diffusion and medium

The invention relates to the technical field of speech synthesis, financial science and technology and medical health, and discloses a speech synthesis method, device and equipment of potential rhythm based on diffusion and a medium, and the method comprises the steps: generating a reference rhythm vector corresponding to each frequency band of an initial audio according to a Mel spectrogram and a preset real phoneme duration; according to the text hidden representation and the speaker hidden representation, constructing an analysis phoneme duration corresponding to the reference rhythm vector; constructing an analysis rhythm vector by using the text hidden representation, the speaker hidden representation and the analysis phoneme duration; generating a potential rhythm vector according to the reference rhythm vector and the analysis rhythm vector; and performing speech synthesis by using the text hidden representation, the speaker hidden representation and the potential rhythm vector to obtain a synthesized speech. The time step number is remarkably reduced through model analysis, the generation speed is increased, and meanwhile, the smoothness of the generated voice is remarkably improved through rhythm vector quantization.
Owner:PING AN TECH (SHENZHEN) CO LTD

A spoken dialogue data processing method based on an agent

The application discloses a kind of oral conversation data processing methods based on agent, belong to artificial intelligence big data processing technical field, including receiving the voice input file sent by user end, voice signal is converted into text data, while extracting emotional characteristics, generate emotional label and voice prosody characteristics, text and emotional information are combined to form dialogue anchor point information, construct path graph, calculate edge weight;According to the similarity of current dialogue anchor point and historical path node, construct propagation matrix, carry out intelligent path reasoning, calculate path score, ensure that emotional fluctuation influences task path selection, agent selects optimal path according to path score to carry out subsequent task processing, and real-time correction is carried out after user feedback;Through emotional analysis technology, real-time adjust the anchor point information of current conversation, and adjust the reasoning process according to feedback, form closed-loop adaptive adjustment mechanism, to optimize the response quality of system, enhance the sense of trust of user to system.
Owner:SHANDONG LINGCHAO SOFTWARE TECH CO LTD

Multi-language intelligent dubbing generation system based on sound cloning and emotion migration

The invention discloses a multi-language intelligent dubbing generation system based on sound cloning and emotion migration. The multi-language intelligent dubbing generation system comprises a sound cloning module, a cross-language synthesis module, an emotion migration module and a lip shape synchronization module. By constructing a few-sample speaker encoder, a cross-language rhythm migration module, a fine-grained emotion control module and a video lip shape synchronization module, end-to-end automatic generation from original dubbing audio to multi-language target dubbing is realized, and tone consistency, emotion authenticity and picture synchronism are kept.
Owner:JIANGSU HOPERUN SOFTWARE CO LTD

Speech synthesis method and device

The invention relates to a speech synthesis method and device, and the method comprises the steps: obtaining a to-be-synthesized text; performing phoneme conversion on the to-be-synthesized text to obtain a first phoneme sequence, the first phoneme sequence comprising at least one phoneme and a tone of a syllable where the at least one phoneme is located; decoupling phonemes and tones in the first phoneme sequence, and extracting phoneme features of the to-be-synthesized text based on a decoupled second phoneme sequence; wherein the second phoneme sequence comprises the at least one phoneme; performing text coding on the to-be-synthesized text in a syllable dimension, and extracting semantic features of the to-be-synthesized text; and performing fusion processing on the semantic features and the phoneme features, and generating voice based on the fused features. According to the invention, the synthesized voice is accurate in pronunciation and natural in rhythm, and it is ensured that the synthesized voice achieves expected effects in the aspects of naturalness, rhythm and semantic consistency.
Owner:YOUKU CULTURE TECH (BEIJING) CO LTD

Children language narrative ability evaluation method and tool based on multi-modal analysis

The invention discloses a children's language narrative ability assessment method and tool based on multi-modal analysis, and relates to the technical field of natural language processing, and the technical scheme is characterized in that audio data and video data of a testee during a narrative process are collected, and voice rhythm features, text language features and visual behavior features are extracted for fusion; the method comprises the steps of inputting a multi-modal scoring model to generate multi-dimensional quantitative scores of a macrostructure, a microstructure, a language organization and a pragmatic function of the language narrative ability of a testee child, generating a structured evaluation report for the testee child according to the scores, and giving personalized intervention suggestions. According to the method, multi-modal data such as voice, language texts and visual behaviors are collected, so that the narrative ability of the children can be comprehensively analyzed from multiple dimensions such as macrostructures, microstructures, language organizations and pragmatic functions, and the method is closer to the real natural expression state of the children; therefore, the comprehensive language ability and communication performance of children can be reflected more comprehensively and truly.
Owner:TONGJI HOSPITAL ATTACHED TO TONGJI MEDICAL COLLEGE HUAZHONG SCI TECH

ASV system risk assessment method and system based on multi-dimensional pronunciation characterization decoupling and fusion

The invention provides an ASV system risk assessment method and system based on multi-dimensional pronunciation characterization decoupling and fusion, and the method comprises the steps: obtaining a reference audio of a target user, and carrying out the decoupling extraction of a volume feature vector, a pitch feature vector, and a speaking style feature vector from the reference audio through a multi-dimensional pronunciation feature extraction network; a text to be verified is converted into a phoneme sequence, the phoneme sequence is input into a pronunciation feature prediction network for predicting a frame-level dynamic pronunciation feature sequence based on the phoneme sequence, and three pronunciation feature vectors are injected into the network through an adaptive instance normalization mechanism to dynamically modulate a prediction process. Generating a frame-level dynamic pronunciation feature sequence containing time sequence rhythm change; inputting a VAE-GAN synthesis trunk, re-injecting the three pronunciation feature vectors through an adaptive instance normalization mechanism, and generating a test voice sample consistent with the voiceprint characteristics of the target user; and initiating an identity verification query for the ASV system, and calculating safety indexes of different user groups according to a verification result so as to evaluate the risk of the ASV system.
Owner:FUJIAN NORMAL UNIV

TTS audio super-division method and device based on text semantics, equipment and medium

The invention discloses a TTS audio super-division method and device based on text semantics, equipment and a medium, and relates to the technical field of audio super-division, and the method comprises the steps: obtaining an initial TTS audio and initial text data, and carrying out the preprocessing of the initial TTS audio and initial text data, so as to obtain a preprocessed audio and a preprocessed text; performing feature extraction on the preprocessed text to obtain semantic features and rhythm features, and performing feature extraction on the preprocessed audio to obtain audio features; and fusing the semantic feature, the rhythm feature and the audio feature to obtain a fused feature, and generating a target TTS audio corresponding to the initial TTS audio by using the fused feature to realize audio super-division of the initial TTS audio. By introducing text features and rhythm features into audio super-division, the problem of semantic expression distortion caused by a pure audio driving method is solved.
Owner:MALANSHAN AUDIO & VIDEO LABORATORY

Anonymization privacy protection method and system for voice information retention

The embodiment of the invention provides an anonymization privacy protection method and system for voice information retention. The method comprises the following steps: extracting speaker embedding of an original audio, eliminating the tone of the speaker, and keeping semantic and rhythm speaker irrelevant features; embedding and inputting a speaker into a speaker anonymous module matched with a three-stage stream based on a U-Net architecture to obtain anonymous embedding; and combining irrelevant features of the speaker with anonymous embedding by using a pre-trained voice reconstruction model to generate anonymized voice with tone privacy. According to the embodiment of the invention, voice anonymization facing content privacy and tone privacy reserved by voice information is realized, the effectiveness of the anonymized voice generated by using the method in a downstream task is superior to that of a baseline model, meanwhile, the privacy of a speaker is also guaranteed, and safe use of data is realized.
Owner:SHANGHAI JIAOTONG UNIV

Text-to-voice method, system and device and storage medium

The invention discloses a text-to-speech method, system and device and a storage medium, and the method comprises the steps: carrying out the chapter processing of a novel text, and enabling each chapter text to comprise a front text, a middle text and a rear text; extracting phonemes, tones, rhythm information and text content from a middle text in each chapter text, and splicing the extracted phonemes, tones, rhythm information and text content to obtain a text vector; obtaining a timbre vector of a target speaker, and splicing the timbre vector and the text vector to obtain a target vector; and inputting the target vector into a text-to-speech model for speech synthesis, and outputting an audio corresponding to the novel text. According to the invention, the end-to-end speech synthesis is realized, and there is no need to label the white, emotion and role of the dialogue in advance, so that the speech synthesis efficiency is improved.
Owner:GUANGZHOU QUYAN NETWORK TECH CO LTD

English pronunciation error correction training method based on speech recognition

The invention relates to the technical field of speech recognition and processing, in particular to an English pronunciation error correction training method based on speech recognition, and the method comprises the following steps: S1, collecting a speech signal generated by a learner in a pronunciation training process, digitalizing the speech signal, associating the digitalized speech signal with a target standard text, and generating an original audio data record with a timestamp; according to the invention, phoneme-level decoding is carried out on the voice signal by using the recurrent neural network acoustic model, and accurate alignment of the pronunciation of the learner and the standard phoneme sequence is realized in combination with the dynamic time warping algorithm, so that pronunciation errors such as misreading, missed reading and increased reading can be accurately identified; meanwhile, acoustic features such as Mel frequency cepstrum coefficient, pitch and fundamental frequency are extracted to be quantitatively compared with a standard native language pronunciation database, multi-dimensional evaluation covering accuracy, integrity, fluency and rhythm is generated, and the accuracy and systematicness of oral English pronunciation error correction are remarkably improved.
Owner:吕丽沙

Voice conversion method and device, equipment and storage medium

The invention relates to the technical field of voice semantics, can be applied to the fields of medical health, financial science and technology and the like, and discloses a voice conversion method, device and equipment and a storage medium, and the method comprises the steps: obtaining an input voice signal, and converting the input voice signal into a Mel spectrogram; respectively extracting basic content information, global timbre information and rhythm feature information from the Mel spectrogram through a timbre encoder, a content encoder and a rhythm encoder; performing quantization processing on the global timbre information, the basic content information and the rhythm feature information to respectively generate timbre quantization information, content quantization information and rhythm quantization information; inputting the timbre quantization information, the content quantization information and the rhythm quantization information into a neural network model to obtain voice feature information; and inputting the voice feature information into a decoder to obtain a target voice signal. According to the method, the voice signals are decoupled into three independent attributes of timbre, content and rhythm, features are extracted through the special encoders respectively, and the encoding efficiency is improved.
Owner:PING AN TECH (SHENZHEN) CO LTD