Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

27 results about "Prosody" patented technology

In linguistics, prosody is concerned with those elements of speech that are not individual phonetic segments (vowels and consonants) but are properties of syllables and larger units of speech, including linguistic functions such as intonation, tone, stress, and rhythm. Such elements are known as suprasegmentals.

Voice dictation with audio large language model

A method comprising receiving audio data 102, generating a transcription 151 comprising a sequence of terms 152, such as “Buy some tomatoes and bananas. Change tomatoes to potatoes”, parallel processing the audio data and the transcription using a multimodal large language model (LLM) 150 to identify one or more revision terms 152R, for example “Change”, specifying a revision action to perform on at least one other term in the sequence, in this instance “tomatoes”, and modifying the transcription 151M accordingly – “Buy some potatoes and bananas”. Identifying the revision term(s) may be based on a corresponding user intent 154 determined for each respective term in the sequence, for example, the user 10 does not intend the final transcription to include “tomatoes”. For each term in the sequence, parallel processing may comprise correlating its speech characteristics 156 such as pitch, tone or prosody information determined from the audio data with its corresponding linguistic context 158 determined from the transcription. Transcription correction may be based on a revision token inserted into the sequence, the token indicating an N number of terms for replacement and their corresponding replacement terms. User context data 104 may be obtained to tailor the LLM to a particular user. [Figure 1A]
Owner:GOOGLE LLC

A method and system for prosody perception calculation based on curvature sequence

The application discloses a kind of based on fixed dimension curvature sequence's sense of rhythm perception and computing method and system, belong to data processing, signal analysis technical field.The method includes: obtaining sequence data, and it is discretized into 60-dimensional normalized curvature sequence;Preset reference curvature value 2π and total potential parameter 120 / π, construct reference curvature sequence;The point-by-point absolute deviation of both is calculated and weighted sum is obtained to get repulsion R, the compatibility P is calculated by formula P=1-R / Ψ_total, according to P value output rhythm evaluation result.The system corresponds to the above-mentioned method sets each functional module.The present application solves the problem that rhythm evaluation cannot be quantified in the prior art, and the calculation is complicated, the calculation is simple and efficient, and the rhythm evaluation of a variety of conventional signals can be adapted, and the practicality is strong.
Owner:成罡

Data processing method and apparatus

This specification provides a data processing method and apparatus, wherein the data processing method includes: preprocessing initial audio data to obtain audio data, and inputting the audio data into an audio recognition model to obtain initial text containing prosody identifiers; determining at least one text unit corresponding to the initial text, and determining time information corresponding to each of the at least one text unit based on the audio data; updating the initial text in the prosody identifier dimension based on the time information corresponding to each of the at least one text unit to obtain target text corresponding to the audio data, wherein the target text and the audio data are used to train an audio generation model.
Owner:BEIJING YUANLI WEILAI SCI & TECH CO LTD

A dialect generation method and system based on voiceprint features

PendingCN122116872ASpeech synthesisPattern recognitionVoice analysis
The present application relates to the field of artificial intelligence, and discloses a dialect generation method and system based on voiceprint features, comprising: extracting individual voiceprint identification and recording content information of a target area; using a pre-trained deep learning model to extract a voiceprint acoustic feature vector of the recording content information to construct a dialect feature vector corresponding to individuals in the target area; receiving target dialect data to be processed, analyzing voiceprint reference features corresponding to the target dialect data, and using a dialect classifier to analyze the dialect type corresponding to the target dialect data; performing voiceprint feature fusion processing on the target dialect data to obtain an adapted initial speech, and analyzing regional prosody features of the target area; and performing prosody optimization processing on the adapted initial speech to generate a dialect generation result corresponding to the target area. The present application can improve the accuracy of dialect generation.
Owner:SIMAI INTELLIGENT TECHNOLOGY (SHENZHEN) CO LTD

A national vocal dialect prosody intelligent correction method and system

This invention proposes an intelligent error correction method and system for ethnic vocal music dialect prosody. The method includes: collecting multimodal data, extracting features to form a training dataset, constructing and updating a dynamic dialect prosody map, and performing cross-modal comparative analysis to obtain a joint representation of cross-modal error features. A neural vocoder direct-connect correction model is constructed, inputting the joint representation to generate a corrected Mel spectrum. A cross-modal generative adversarial network is constructed, including a generator and three discriminators. The generator generates a joint latent representation based on high-fidelity corrected audio and three features. The three discriminators respectively judge the naturalness of the audio, the compliance of the text tone, and the rhythmic coordination of the musical score. This invention solves the problem of manual error correction by constructing a prosody map, overcomes the bottleneck of manual detection by utilizing cross-modal comparative analysis, improves the generalization ability of the model by combining meta-learning algorithms, and optimizes the correction results with the help of adversarial networks, thereby reducing manual costs and improving the error correction efficiency of ethnic vocal music works.
Owner:GUIZHOU RADIO & TV UNIV

A data processing method, device and electronic equipment

PendingCN122454955ASpeech rateSpeech sound
Embodiments of the present application provide a data processing method, device and electronic equipment. The method comprises: constructing a first text, and performing word segmentation processing on the first text to obtain a first word segmentation sequence; the first text contains a target keyword; a prosody control marker is randomly inserted after each word segmentation in the first word segmentation sequence to obtain a second word segmentation sequence; the prosody control marker is used to control any one of the following prosodic features of speech: pause, speech rate, pitch, stress, duration, volume, and continuity; the second word segmentation sequence is input into a trained text-to-speech model, and the text-to-speech model generates first speech data with corresponding prosodic features according to the prosody control markers in the second word segmentation sequence. Embodiments of the present application can improve the efficiency of obtaining training data, and can improve the accuracy and robustness of training a KWS model.
Owner:SHENZHEN MICROBT ELECTRONICS TECH CO LTD

A digital human incremental editing method and device based on natural language interaction

The application discloses a digital person editing method and device based on natural language interaction. The method proposes a state-aware digital person incremental editing architecture, first establishes and maintains a persistent multi-modal system state containing visual, voice and identity feature libraries; secondly, the language model is used for context-aware analysis and modal intention routing of multi-round dialogue instructions; then, local incremental update generation is performed for target modalities such as appearance, voice content or prosody; finally, an identity-aware gating mechanism is used to fuse and update the historical identity features and the candidate update state. The application effectively avoids the large amount of calculation overhead of repeated full generation in digital person editing, and can maintain the visual and voice identity consistency of the digital person in multi-round editing.
Owner:SOUTHEAST UNIV

Voice modification

ActiveUS12670918B2TimbreSpeech sound
A computing system that receives an audio waveform representing speech from an individual and produces as output a modified version of the audio waveform that maintains the speaker's speech characteristics as well as prosody for specific utterances (e.g., voice timbre, intonation, timing, intensity). The system uses a bottleneck-based autoencoder with speech spectrograms as input and output. To produce the output audio waveform, the system includes a reconstruction error-based loss function with two additional loss functions. The second loss function is speaker “real vs fake” discriminator that penalizes for the output not sounding like the speaker. The third loss function is a speech intelligibility scorer that penalizes the output for speech that is difficult for the target population to understand. The produced modified audio waveform is an enhanced speech output that delivers speech m a target accent without sacrificing the personality of the speaker.
Owner:SRI INTERNATIONAL

An AI audio synthesis optimization method based on deep reinforcement learning

PendingCN122157635ABiological modelsSpeech recognitionAudio synthesisFeature data
The application discloses an AI audio synthesis optimization method based on deep reinforcement learning, comprising the following steps: collecting training corpus data; using the training corpus data to perform supervised training on an improved VITS model to obtain a baseline improved VITS model; processing text data and extracting target feature data related to sentiment annotation data and style annotation data; setting a controllable parameter interface; calling the controllable parameter interface to generate synthesized speech data; evaluating the synthesized speech data to obtain reward signal data including sentiment consistency reward, style similarity reward and prosody naturalness reward; and using an improved GRPO algorithm to update the strategy of the baseline improved VITS model. The application realizes the effect of enhancing the expression of synthesized speech in emotion and style control while maintaining the naturalness of speech synthesis, and can be applied to intelligent voice interaction, virtual person broadcasting and multi-scene voice generation tasks.
Owner:XIAMEN RENZHI YOUXUE EDUCATION TECHNOLOGY CO LTD

A speech synthesis model method capable of synthesizing multi-emotional audio.

ActiveCN116798403BData setSpeech technology
This invention discloses a speech synthesis model method capable of synthesizing multi-emotion audio, relating to the field of intelligent speech technology. The method includes the following steps: processing raw data, distinguishing between training and validation sets, adding annotation files to each set, and simultaneously delivering the raw dataset to an emotion recognition module for processing; calling the emotion recognition module to preprocess the dataset, decomposing the audio into phonemes and emotion feature files; the complete multi-emotion text-to-speech model and dataset processing are specifically divided into dataset collection, unsupervised preprocessing, encoder training, and online inference. The final output includes a multi-emotion encoder with intermediate outputs and a final online synthesized independent WAV file, capable of achieving multi-emotion output and simulating prosody, making the effect close to that of a real person. No emotion annotation is required during data processing, and the method of constructing a continuous feature value spectrum greatly avoids the problem of inaccurate machine annotation.
Owner:UNICOM WOYUEDU TECH CULTURE CO LTD +1

Speech extraction method and device, electronic equipment and storage medium

The present application relates to the technical field of artificial intelligence, and provides a voice extraction method and device, electronic equipment and storage medium, the method comprising: obtaining a mixed voice segment at a current time and identity representation information of a target speaker; obtaining a historical voice segment of the target speaker at at least one historical time; fusing mixed voice features extracted from the mixed voice segment, historical context features extracted from the historical voice segment, and the identity representation information to obtain target fusion features; and extracting a target voice segment of the target speaker at the current time from the mixed voice segment according to the target fusion features. The present application introduces the historical voice segment containing rich phonemes, prosody and other sound states as a dynamic context reference, breaking the limitations of traditional stateless models, not only giving the model a memory ability for voice content, significantly improving the continuity of the output, thereby effectively avoiding the generation of stuttering, jumps and artifacts at the block boundary.
Owner:IFLYTEK CO LTD

Audio labeling method, speech synthesis model training method, device, and medium

The application relates to an audio labeling method, a speech synthesis model training method, equipment and a medium, and belongs to the computer technology field.The method comprises the following steps: obtaining a pre-labeling file of a text sample; obtaining speech data corresponding to the text sample uttered by a target speaker; performing word-level recognition on the speech data to obtain speech labeling information corresponding to the speech data; modifying the pre-labeling file based on the speech labeling information to obtain a labeling file consistent with the speech data, so as to train a speech synthesis model by using the labeling file; the problem that the speech synthesis effect of a speech synthesis model trained based on a labeling file obtained based on text sequence labeling is poor can be solved; the labeling file can have the speaking characteristics of the target speaker, and the speech synthesis effect can be improved. Meanwhile, the problem that the labeling results of different labeling personnel are different can be solved, and the differences between labeling files of different prosody levels are reduced.
Owner:AISPEECH CO LTD

Intelligent dubbing generation method and device fusing voiceprint features and semantic understanding

The application discloses a method and device for generating intelligent dubbing by fusing voiceprint features and semantic understanding, and relates to the technical field of artificial intelligence and speech synthesis. The application firstly acquires semantic features of a target text and voiceprint features of a reference audio; secondly, a feature alignment module based on a multi-head self-attention mechanism is constructed to map the voiceprint features to a semantic hidden space and generate semantic vectors with speaker identity attributes; subsequently, a sentiment reasoning engine is introduced to dynamically predict sentiment labels and prosody parameters according to semantic contexts; finally, an improved neural vocoder is used to generate high-fidelity audio waveforms. The application solves the problem of disconnection between tone color and sentiment and harsh prosody in existing speech synthesis technology, realizes deep fusion of tone color cloning and text sentiment expression, and significantly improves the realism, expressiveness and degree of personalization of the generated dubbing.
Owner:CHENGDU UNIVERSITY OF TECHNOLOGY

Speech synthesis method, apparatus and electronic device

The application discloses a speech synthesis method and device, electronic equipment and a computer readable storage medium. The method comprises: obtaining a target text to be converted into speech, a target emotion category and a target emotion intensity; determining start end information and end information of an emotion intensity interval to which the target emotion intensity belongs from a plurality of emotion intensities preset for the target emotion category; generating emotion feature data corresponding to the target text according to the start end information and the end information; processing the target text into each phoneme vector of a phoneme level with emotion features according to the emotion feature data; performing prosody prediction and prediction of the number of speech frames occupied on each phoneme vector respectively to obtain first prosody information and a speech content vector at a frame level; and processing the speech content vector into a time-domain speech signal according to the first prosody information. The scheme provided by the application enables the synthesized speech to correctly reflect the emotion type and emotion intensity, thereby improving the human-computer interaction experience of the user.
Owner:NETEASE (HANGZHOU) NETWORK CO LTD

Lightweight Intent-Assisted Input System Based on Speech Prosody Features

PendingCN122135693ASpeech recognitionSpeech synthesisDialog systemEngineering
This invention discloses a lightweight intent-assisted input scheme based on speech prosodic features, belonging to the field of artificial intelligence and voice interaction technology. The scheme utilizes a local speech synthesis or recognition front-end module on the client side to extract prosodic features such as tone, intonation, pauses, and rhythm from the user's speech, quantifying them into lightweight parameter labels with extremely small volumes, such as question level, exclamation level, pause length, and tone intensity. These labels are then encapsulated with the speech-to-text and sent to the server. The server-side intent recognition model uses these labels as prior information to dynamically adjust the inference path, pruning invalid intent branches, reducing computational consumption, and improving recognition efficiency and accuracy. This invention does not change the core model structure, does not significantly increase the transmission burden, and can reuse existing speech module logic, achieving front-end computation and system cost reduction and efficiency improvement. This invention is applicable to various voice assistants and intelligent dialogue systems, and has broad application prospects.
Owner:陈起洋

Voice synthesis model training method, voice synthesis method, device, electronic equipment, computer readable storage medium and computer program product

ActiveCN121862082BSpeech synthesisSpeech rateSynthesis methods
The application provides a speech synthesis model training method, a speech synthesis method, a device, an electronic device, a computer readable storage medium and a computer program product. The method comprises: synthesizing a first target speech based on a first text feature of a text sample and a first reference speech feature of a reference speech through a speech synthesis model; determining a speech length reward value based on a deviation between a first speech speed of the first target speech and a second speech speed of the reference speech; determining a prosody alignment reward value based on a matching degree between a first pause structure of the first target speech and a second pause structure pre-constructed for the text sample; determining an entropy regular reward value based on a difference between a distribution entropy when the first target speech is synthesized and a target entropy value labeled for the text sample; and training the speech synthesis model based on the speech length reward value, the entropy regular reward value and the prosody alignment reward value. The application can improve the quality of the speech synthesized by the speech synthesis model.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Sound replication method and related apparatus

The application discloses a sound replication method and related device, and relates to the technical field of speech synthesis. In one aspect, a hybrid expert sound replication model is adopted, each expert sub-model corresponds to only one acoustic category, and all the expert sub-models adopt a lightweight structure. Therefore, the end-side device only needs to activate a single expert sub-model matched with target acoustic category information for reasoning, which greatly reduces the model parameter quantity and computing cost, and thus successfully runs on the low-resource end-side device. In another aspect, the target speaker personalized acoustic features pre-delivered by the cloud and locally stored are directly used in the replication process. The features accurately depict the details of the speaker's timbre, accent, prosody, etc., and make up for the insufficient expression ability of the lightweight model. Therefore, the scheme can generate a highly similar replicated voice to the target speaker on the low-resource end-side device without relying on real-time calculation of the cloud large model, and takes into account the feasibility of end-side deployment and the high fidelity of the synthesized sound quality.
Owner:IFLYTEK CO LTD

Multi-modal based companion robot dialogue quality evaluation method and computer device

This application provides a multimodal dialogue quality assessment method and computer device for companion robots, relating to the field of robot control technology. The method includes: collecting multimodal raw data of the current dialogue turn during the interaction process; extracting semantic features, speech prosody features, facial expression features, and physiological emotion features based on dialogue speech, facial expression images, and user physiological signals; obtaining semantic relevance scores, emotional consistency scores, and contextual coherence scores based on facial expression images, semantic features, speech prosody features, facial expression features, physiological emotion features, response text, and historical text scores; and obtaining the current dialogue quality assessment result of the companion robot based on the semantic relevance score, emotional consistency score, and contextual coherence score. This application improves the accuracy of the current dialogue quality assessment result of the companion robot, thereby improving the service quality of the companion robot.
Owner:CHONGQING PHOENIX TECHNOLOGY CO LTD

Mongolian multi-modal sentiment analysis method based on cross-modal information enhancement and fusion

A Mongolian multi-modal sentiment analysis method based on cross-modal information enhancement and fusion. To solve the problems of scattered sentiment expression caused by the agglutinative characteristics of Mongolian and the difficulty of capturing implicit clues in multi-modal sentiment information, a three-level innovative solution is proposed: (1) Cross-modal explicit enhancement layer: Use multi-modal large language models to convert implicit sentiment clues such as prosody and expression in audio and video into text descriptions of sentiment, enriching the information of each modality through the "implicit→explicit→semantic fusion" path; (2) Semantic alignment layer based on Gram matrix: Extract the second-order statistical structure of the text modality as the semantic reference, and map heterogeneous audio and video features to a unified semantic space to solve the cross-modal alignment difficulty caused by the morphological changes of Mongolian; (3) Gating displacement fusion layer: Use an adaptive displacement mechanism to fine-tune the semantic position of text features to achieve precise injection of non-text information, avoiding the modality submersion problem in long morpheme sequences of Mongolian.
Owner:INNER MONGOLIA UNIV OF TECH

Cross-border e-commerce live interactive analysis and optimization system based on multi-modal big data

ActiveCN122053872BText streamBroadcast data
The application relates to the technical field of cross-border e-commerce live broadcast data processing, and particularly discloses a cross-border e-commerce live broadcast interaction analysis and optimization system based on multi-modal big data, which collects anchor audio stream, barrage text stream, audience behavior log and video frame sequence; cross-cultural sensitive word units and speech prosody fluctuation sections are extracted from the audio stream; emotional mutation points in the barrage text stream, churn rate jump points in the behavior log and picture heat attenuation areas in the video frame sequence are respectively located, and a multi-modal feature disturbance atlas is generated; the disturbance atlas is quantized to obtain a cross-cultural cognitive friction coefficient; when the cross-cultural cognitive friction coefficient exceeds a threshold value, a cultural buffer element is matched and inserted into the live broadcast stream; the application can quantize the cross-cultural cognitive load degree in real time, intervene in the cultural buffer element before the audience loses, and improve the conversion efficiency of cross-border live broadcast.
Owner:JIANGXI NORMAL UNIV

A multilingual automatic recognition method and system

The application relates to a multilingual automatic recognition method and system, and belongs to the technical field of language recognition. The recognition method comprises the following steps: receiving an original voice signal and performing pretreatment to obtain a pretreated voice frame sequence; simultaneously extracting an acoustic feature vector, a vocal organ movement feature matrix and a prosody feature vector from the voice frame sequence to form a feature triple; performing language family classification according to the acoustic feature vector and the vocal organ movement feature matrix to output a candidate language family set; inputting the candidate language family set and the prosody feature vector into a dialect clustering model to output a refined dialect cluster label; calculating acoustic feature weight values, vocal organ movement feature weight values and prosody feature weight values, performing weighted operation on the feature triple to generate a weighted feature vector; and inputting the weighted feature vector and the refined dialect cluster label into a language decision model to output a language recognition result containing a language label and a confidence value. The application improves the accuracy and robustness of multilingual recognition in a complex environment.
Owner:BEIJING HIZHI TECH CO LTD

Audio generation method, apparatus, and electronic device

PendingCN122116867ASpeech synthesisAcoustic modelLabeled data
The application provides an audio generation method, device and electronic equipment, and belongs to the technical field of audio processing. The audio generation method comprises: performing phoneme conversion processing on a text to be processed to obtain phoneme data of the text; performing prosody recognition processing on the text to obtain prosody data of the text; performing emotion recognition processing on the text to obtain emotion label data of the text; inputting the phoneme data, the prosody data and the emotion label data into an acoustic model to obtain acoustic feature data of the text and length data of each phoneme in the phoneme data; inputting the acoustic feature data into a vocoder for speech conversion processing to generate synthesized audio of the text; and adjusting a starting pronunciation interval duration between phonemes in the synthesized audio according to the length data of each phoneme to obtain adjusted synthesized audio. The synthesized audio can improve the degree of personification in terms of tone, pause and emotion, improve the naturalness of expression of the synthesized audio, and improve the listening experience of a user.
Owner:BEIJING CO WHEELS TECH CO LTD

Training text-to-speech model, text-to-speech method, device and equipment

ActiveCN119007706BAcousticsSpeech sound
The embodiment of the specification discloses a method and device for training a text-to-speech model and converting text into speech. The composition of the input data of the text-to-speech model is redefined. The input data includes not only the phoneme sequence corresponding to the text with inserted prosodic symbols, but also structure annotation information capable of representing the structural division of the text at at least one granularity level. Thus, in the process of predicting speech features, the text-to-speech model can not only refer to the prosody of the text at the phoneme level, but also refer to the prosody of the text at the granularity level of single word, phrase, sentence, etc. This can make the predicted speech features have the coherence of the pronunciation on the text structure, and the prosody is more natural. It should be noted that the present disclosure belongs to the technical solution in the field of artificial intelligence. In the implementation of the scheme, the privacy data used has been authorized by all parties.
Owner:ALIPAY (HANGZHOU) INFORMATION TECH CO LTD

A sentiment controllable poem generation method based on an improved loss function transformer model

ActiveCN117807991BData setAlgorithm
The application discloses a kind of sentiment controllable poem generation methods based on improved loss function Transformer model, collection public data set, data set screening refining is carried out, and label is preprocessed;Data is input into Transformer encoder, and is decoded using Linear layer neural network, to obtain decoding vector;In the training process, decoding vector is converted into prosody vector by word rhyme matrix, and gradient descent loss function is combined with two kinds of vectors to learn weighting;In the inference process, pruning search and heuristic method are used to improve the quality of the poem;Input prompt sentiment words to the deployment model and infer, obtain new generated poem corresponding to input sentiment.The method proposes word rhyme matrix, converts decoding vector to rhyme vector to obtain rhyme-related loss function;By merging the gradient descent of two kinds of loss functions, the poem generated by the model has high rhyme rate while ensuring quality;Combined with sentiment label, the literary quality of the poem is improved while ensuring the format of the poem.
Owner:NANJING UNIV OF SCI & TECH