Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

48 results about "Speaking style" patented technology

Text-to-speech synthesis using generative artificial intelligence models

A method and a system for generating human speech audio in a conversation using a trained generative AI model are provided. The method includes receiving a text input representing a portion of the conversation, receiving dialog context associated with the conversation, receiving information representing at least one voice and speaking style of at least one speaker in the conversation, generating the at least one voice and speaking style based on the received information, and generating at least one emotional audio response for the at least one speaker using the at least one voice and speaking style and without retraining the trained generative AI model.
Owner:PHEON INC

Voice synthesis method and system based on VITS improvement

The invention provides a voice synthesis method and system based on VITS improvement, and the method comprises the steps: optimizing a text encoder of a VITS model, introducing a large language model, and enabling the emotion, intention and speaking style of an input text to be captured when the text is encoded; a random disturbance item is introduced when the Q value is dynamically planned and solved, the alignment flexibility in the initial training stage is improved, meanwhile, monotonicity constraint is strictly kept, and it is avoided that a suboptimal solution is obtained through convergence too early; a ConvNeXt module is used as a basic backbone network of a decoder, and ISTFT is utilized to efficiently reconstruct a time domain signal, so that waveform up-sampling is realized, redundant calculation of traditional transpose convolution is avoided, and reasoning is accelerated. According to the method, the reasoning speed, the emotion expression ability and the style control flexibility of speech synthesis can be effectively improved, a new solution is provided for cross-language diversified speech synthesis, and a reference is provided for the more efficient and more intelligent development of the speech synthesis technology.
Owner:豫章师范学院

Method, electronic device, and computer program product for generating video

A method includes obtaining a reference image and a reference speech, the reference image specifying a head of a target object in the video, and the reference speech specifying a voice of the target object; and generating, based on the reference image and the reference speech, a fusion vector by combining a feature of the head and a feature of the voice. The method further includes generating, based on the fusion vector, a plurality of video frames in a video that represents the target object speaking in a timbre of the reference speech by denoising a plurality of initial frames including noise; and generating the video based on the plurality of video frames. In embodiments of the present disclosure, a video in which a semantic feature and a speaking style of the target object are merged can be generated, and the resolution and quality of the generated video are enhanced.
Owner:DELL PROD LP

ASV system risk assessment method and system based on multi-dimensional pronunciation characterization decoupling and fusion

PendingCN121354597ASpeech synthesisFeature extractionSpeaking style
The invention provides an ASV system risk assessment method and system based on multi-dimensional pronunciation characterization decoupling and fusion, and the method comprises the steps: obtaining a reference audio of a target user, and carrying out the decoupling extraction of a volume feature vector, a pitch feature vector, and a speaking style feature vector from the reference audio through a multi-dimensional pronunciation feature extraction network; a text to be verified is converted into a phoneme sequence, the phoneme sequence is input into a pronunciation feature prediction network for predicting a frame-level dynamic pronunciation feature sequence based on the phoneme sequence, and three pronunciation feature vectors are injected into the network through an adaptive instance normalization mechanism to dynamically modulate a prediction process. Generating a frame-level dynamic pronunciation feature sequence containing time sequence rhythm change; inputting a VAE-GAN synthesis trunk, re-injecting the three pronunciation feature vectors through an adaptive instance normalization mechanism, and generating a test voice sample consistent with the voiceprint characteristics of the target user; and initiating an identity verification query for the ASV system, and calculating safety indexes of different user groups according to a verification result so as to evaluate the risk of the ASV system.
Owner:FUJIAN NORMAL UNIV

Real-time live distribution system

PendingJP2025186961ASelective content distributionDistribution systemSpeaking style
To provide a real-time distribution system that is capable of enhancing communication between a distributor and a viewer at low costs and enables the distributor to obtain a benefit.SOLUTION: A real-time live distribution system for performing real-time live distribution of an idle talk dialog content toward a plurality of viewer terminals comprises a distributor terminal 1 including: an extraction section 11 having a virtual human for extracting viewer comments input from the plurality of viewer terminals in real time; an analysis section 12 for performing text analysis on the extracted viewer comments; a reading section 13 for reading a text using voice synthesis on the basis of a text analysis result; a feeling parameter setting section 14 for setting a virtual human feeling parameter for categorizing the viewer comments, which are at least a positive comment or negative comment, to determine the positive comment as a plus comment and the negative comment as a minus comment; and a change section 15 for changing a way of speaking according to the set feeling parameter.SELECTED DRAWING: Figure 2
Owner:CRYSTAL METHOD CO LTD

Video generation method, device, equipment and storage medium

The present disclosure provides a video generation method, apparatus, device, and storage medium, which relate to the fields of artificial intelligence technology, particularly computer vision, deep learning, and other technical fields, and can be applied to scenarios such as artificial intelligence-based content generation. A specific implementation scheme is as follows: a voice file and a face video to be driven are input into a pre-trained first model, and the first model outputs a three-dimensional face mesh sequence; wherein the three-dimensional face mesh sequence corresponds to the voice features of the voice file, and corresponds to the facial features and speaking style features of the face video to be driven; based on the three-dimensional face mesh sequence and the face video to be driven, a voice-driven face video matching the voice file is generated.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Audio-driven facial animation supporting varying identity and talking styles

The invention discloses an audio driven facial animation supporting varying identity and speaking styles. In various examples, systems and methods are disclosed for animating a virtual or digital actor or image using audio-driven animation. The system may identify an animation of the grid corresponding to the audio data and the indication of the speaking style. The system may generate a plurality of vertex increments using animations and neutral gestures of the grid. The system may update the machine learning model using the plurality of vertex increments, the audio data, and the indication of the speaking style to generate an output vertex increment for the grid given an input speaking style and input audio data.
Owner:NVIDIA CORP

System

PendingJP2026029847AData processing applicationsMoodSpeaking style
An object of a system according to an embodiment is to detect an emotion of a user and provide appropriate advice on the basis of the detected emotion.SOLUTION: A system according to an embodiment includes a smart emotion detection speaker, an emotion detection AI, and a mentoring generation AI. The smart speaker detects emotions from the user's voice and speech. The emotion detection AI generates words based on emotions. The mentoring generation AI listens to the user's distress and provides advice.SELECTED DRAWING: Figure 1
Owner:SOFTBANK GROUP CORP

System

A system is provided.SOLUTION: A system, comprising: means for selecting a person with whom a user wishes to interact; means for generating a AI model to mimic the person's thought patterns and speech; means for receiving a user question and generating a response based on the AI model; and means for displaying the generated response to the user.SELECTED DRAWING: Figure 1
Owner:SOFTBANK GROUP CORP

Speech style generation method and device, electronic equipment and storage medium

The present disclosure relates to a method and device for generating a speaking style, an electronic device and a storage medium. The method comprises: fitting target style feature attributes based on multiple style feature attributes, determining fitting coefficients of each style feature attribute; determining a target style feature vector according to the fitting coefficients of each style feature attribute and multiple style feature vectors, the multiple style feature vectors corresponding to the multiple style feature attributes one by one; inputting the target style feature vector into a speaking style model to output target speaking style parameters, the speaking style model being obtained based on a framework of training a speaking style model based on the multiple style feature vectors; and generating a target speaking style based on the target speaking style parameters. The method can realize fast migration of the speaking style and improve the generation efficiency of the speaking style.
Owner:HISENSE VISUAL TECH CO LTD

AUDIO-CONTROLLED FACIAL ANIMATION TO SUPPORT VARYING IDENTITIES AND SPEAKING STYLES

PendingDE102025141313A1Speech analysisAnimationAnimationMesh animation
Several examples disclose systems and methods for animating virtual or digital performers or avatars using audio-driven animation. A system can identify an animation for a mesh corresponding to audio data and a specified speech style. Using the animation and a neutral pose for the mesh, the system can generate a variety of vertex deltas. Based on this variety of vertex deltas, the audio data, and the specified speech style, the system can update a machine learning model to generate output vertex deltas for the mesh, taking into account an input speech style and input audio data.
Owner:NVIDIA CORP

System

PendingJP2026018011AInstrumentsProcessingServer
A system is provided.SOLUTION: Means for recording video and audio data of a face-to-face contact, means for transmitting the recorded video and audio data to a server, means for performing noise removal and text conversion on the transmitted video and audio data, means for analyzing wording and conversation content using a natural language processing technique on text data, means for analyzing a manner of speaking and a tone of voice from the audio data, means for analyzing facial expression and body language from the video data, and means for digitizing each analyzed feature, A system comprising: means for scoring; means for training a learning model using interview data of existing employees; means for evaluating new interview data using the learning model; and means for displaying evaluation results and providing feedback to a user.SELECTED DRAWING: Figure 1
Owner:SOFTBANK GROUP CORP

A voice-driven method for generating facial expression animations for 3D virtual avatars

This invention relates to a method for generating facial expression animations of a voice-driven 3D virtual avatar, comprising: inputting a speech segment and converting it into audio features; an encoder module encoding the audio features extracted by the audio processing module, a static 3D head mesh template, and the input one-hot encoding of the speaking style, encoding them into the latent space, and concatenating them; processing the concatenated feature vector sequence through a bidirectional Mamba module, capturing contextual information for each feature vector position, and outputting a feature vector sequence after information exchange of the same dimension; a decoder decoding the feature vector sequence output by the bidirectional Mamba module to obtain the vertex position offsets of each frame of the face mesh, adding them to the static 3D head mesh template to obtain a 3D head mesh sequence; and using the audio-head mesh sequence dataset to perform end-to-end training of the entire network model.
Owner:TIANJIN UNIV

Voice conversion method and device, electronic equipment and storage medium

PendingCN120727020ASpeech analysisSpeaking styleLanguage speech
The invention provides a voice conversion method and device, electronic equipment and a storage medium, and the method comprises the steps: carrying out the extraction of content features and fundamental frequency features of a source audio, and carrying out the extraction of speaker features of a target audio; inputting the content features, the fundamental frequency features and the speaker features into a voice conversion model for joint modeling processing, linear mapping processing and waveform reconstruction processing to generate a voice waveform; performing one-dimensional deep separation convolution processing and multi-receptive-field fusion processing on the voice waveform based on a vocoder in the voice conversion model to generate converted voice audio; wherein the converted voice audio shows that the speaking style of the source audio is replaced by the speaker style of the target audio. The voice conversion model is used for voice conversion, so that the timbre consistency and semantic retention capability of cross-language voice conversion are improved, and the naturalness and definition of reconstructed voice are improved.
Owner:BEIJING YUANJIAN INFORMATION TECH CO LTD

system

We provide the system. [Solution] A means for a person with specialized knowledge to input that knowledge and speaking style as information, A method for training an AI model using natural language processing technology based on that input information, A means of storing and making available trained AI models on the platform, A means for users to search for AI models through the platform and begin interacting with them, A means by which an AI model generates a response to that AC input and sends it back to the user, A means of providing responses as an interactive learning experience and enhancing immersion through visual and auditory means, A system that includes this.
Owner:SOFTBANK GROUP CORP

Virtual image construction method and device, computer equipment, storage medium and program product

The invention discloses a virtual image construction method and device, computer equipment, a storage medium and a program product. The method comprises the following steps: acquiring target text data and first video data; performing emotion recognition processing on the target text data to obtain a corresponding expression tag; performing voice conversion processing on the target text data based on the expression label to obtain target voice data; driving the lip of the target user in the first video data to act by taking the target voice data as a driving signal to obtain second video data; extracting a first mouth opening and closing feature in the first video data, and determining a first speaking style feature index; adjusting the general expression feature data based on the first speaking style feature index and the general speaking style feature index to obtain expression feature data corresponding to the target user; and performing expression editing processing on the second video data by using the expression feature data to obtain the target video data, thereby enhancing the reality sense of the virtual image and the immersion sense of the user.
Owner:MIGU CO LTD +1

System

An object of the system according to the embodiment is to reproduce the voice and the way of speaking of the deceased and to reestablish emotional connection.SOLUTION: A system includes a collection unit, a learning unit, a reception unit, a generation unit, and a reproduction unit. The collection unit collects video and audio data of the deceased. The learning unit learns the data collected by the collection unit and analyzes the voice and the way of speaking of the deceased. The receiving unit receives an input from a user. The generation unit generates new words in the voice of the deceased based on the input accepted by the acceptance unit. The reproduction unit reproduces the words generated by the generation unit.SELECTED DRAWING: Figure 1
Owner:SOFTBANK GROUP CORP

System

A system is provided.SOLUTION: A system comprising: a default setting unit configured to allow a user to register; a unit configured to collect and store conversation data with the user; a unit configured to analyze the collected conversation data and learn a manner of speaking and a context of the user; a unit configured to train a user-specific model based on the learned data; a unit configured to generate a reply of the conversation using the generated user-specific model; and a unit configured to transmit the generated reply to a user terminal.SELECTED DRAWING: Figure 1
Owner:SOFTBANK GROUP CORP

Method, electronic device, and computer program product for generating video utilizing a fusion vector based on reference image and reference speech

A method includes obtaining a reference image and a reference speech, the reference image specifying a head of a target object in the video, and the reference speech specifying a voice of the target object; and generating, based on the reference image and the reference speech, a fusion vector by combining a feature of the head and a feature of the voice. The method further includes generating, based on the fusion vector, a plurality of video frames in a video that represents the target object speaking in a timbre of the reference speech by denoising a plurality of initial frames including noise; and generating the video based on the plurality of video frames. In embodiments of the present disclosure, a video in which a semantic feature and a speaking style of the target object are merged can be generated, and the resolution and quality of the generated video are enhanced.
Owner:DELL PROD LP

system

PendingJP2026045369AData processing applicationsEngineeringSpeaking style
The system according to the embodiment aims to reproduce the speaking style and knowledge of great people, and to provide the user with the feeling that they are having a conversation with a great person. [Solution] A system according to an embodiment includes a collection unit, an analysis unit, a reproduction unit, a reception unit, and an answering unit. The collection unit collects voice data. The analysis unit analyzes the data collected by the collection unit. The reproduction unit reproduces the speaking style and knowledge of the great person based on the data analyzed by the analysis unit. The reception unit accepts input from a user. The answering unit answers like the great person based on the input accepted by the reception unit.
Owner:SOFTBANK GROUP CORP

AI interview method, electronic equipment, storage medium and program product

The invention relates to the technical field of artificial intelligence, and discloses an AI interview method, electronic equipment, a storage medium and a program product, and the method comprises the steps: obtaining an interview text; determining a first speech speed and a first pause frequency of the candidate according to the interview text; the first pause frequency is used for describing the pause frequency in the process that the candidate inputs the interview audio corresponding to the interview text; determining the semantic integrity probability of the interview text; determining the probability that the candidate does not continue speaking according to the first speech speed, the first pause frequency and the semantic integrity probability; and determining whether to continuously ask questions according to the probability that the candidates do not continuously speak. Therefore, the opportunity of continuing to ask questions by the AI interviewer can be adaptively adjusted according to the speaking style of the candidate, so that the AI interviewer takes over a conversation at a reasonable opportunity, and the interview experience feeling of the candidate is improved.
Owner:BEIJING NIUKE TECH CO LTD

Model updating method and device, voice conversion method, equipment and storage medium

ActiveCN116543780BEasy to updateimprove accuracySpeaking styleNetwork model
This application provides a model update method and apparatus, a speech conversion method, device, and storage medium, belonging to the field of financial technology. The method includes: acquiring sample speech data; inputting the sample speech data into a neural network model; encoding the sample speech data through an encoding network to obtain an initial audio feature vector; indexing the initial audio feature vector based on a preset codebook to obtain an audio frame index; extracting phoneme features from the initial audio feature vector based on the audio frame index to obtain an initial phoneme feature vector; performing speech alignment on the initial phoneme feature vector to obtain a sample audio embedding vector; decoding the sample audio embedding vector and the speech style embedding vector through a decoding network to obtain synthesized speech data; and updating the parameters of the neural network model based on the synthesized speech data and the sample speech data to obtain a speech conversion model. This application can improve the accuracy of the model in speech conversion.
Owner:PING AN TECH (SHENZHEN) CO LTD

Information processing system, information processing method, and information processing program

PCT designated stageWO2026079118A1Speech synthesisInformation processingSpeaking style
An information processing system according to the present disclosure comprises: a reception unit that receives, from a user, subjective evaluation information relating to speech evaluation, and designation of a speaking style; a generation unit that, on the basis of the subjective evaluation information and the designation of the speaking style, converts a low-dimensional embedding vector corresponding to the speaking style adjusted by the user to generate a high-dimensional speaking style embedding vector; and a speech synthesis unit that synthesizes speech from the speaking style embedding vector and text information defining speech to be synthesized.
Owner:SONY GROUP CORP

Controllable text-to-speech method based on decoupled multi-modal prompt and chain guide

The application provides a controllable text-to-speech method based on decoupled multi-modal cues and chain-based guidance, which comprises the following steps: constructing a unified multi-modal style encoder for mapping style cues from reference audio or descriptive text to a shared embedding space to generate style-conditioned embeddings, and training the multi-modal style encoder; constructing a text-to-speech model based on a latent diffusion transformer, which takes the latent representation of mel-spectrogram as the processing object and receives content text, speaker prosody and speaking style as input conditions; in the inference stage of the text-to-speech model, a chain-based classifier-free guidance mechanism is adopted to realize independent and continuous control of the content, prosody and style of the generated speech by independently adjusting the guidance strength of the content, prosody and style-conditioned embeddings. The independent and continuous control of the content, prosody and style of the generated speech is realized. The training stability and convergence speed of the model are significantly improved.
Owner:UNIV OF SCI & TECH OF CHINA

AUDIO-CONTROLLED FACIAL ANIMATION TO SUPPORT VARYING IDENTITIES AND SPEAKING STYLES

PendingDE102025141314A1Speech analysisAnimationAnimationMesh animation
Several examples disclose systems and methods for animating virtual or digital performers or avatars using audio-driven animation. A system can identify an animation for a mesh corresponding to audio data and a specified speech style. Using the animation and a neutral pose for the mesh, the system can generate a variety of vertex deltas. Based on this variety of vertex deltas, the audio data, and the specified speech style, the system can update a machine learning model to generate output vertex deltas for the mesh, taking into account an input speech style and input audio data.
Owner:NVIDIA CORP

Cross-language speech synthesis method and system based on dual speaker embedding

The embodiment of the application provides a cross-language speech synthesis method based on double speaker embedding. The method comprises the following steps: inputting text and native language speaker embedding into a txt2vec acoustic model, determining the phoneme sequence code of the text through a text encoder, and determining the vector quantization acoustic feature and the auxiliary feature from the phoneme sequence code and the native language speaker embedding through a decoder; inputting the target language speaker embedding, the vector quantization acoustic feature and the auxiliary feature into a vec2wav vocoder, extracting the X-vector feature of the target language speaker embedding, inputting the X-vector feature, the vector quantization acoustic feature and the auxiliary feature into a feature encoder, and obtaining a cross-language acoustic feature; and determining the cross-language synthesized speech of the cross-language acoustic feature by using a generator. The embodiment of the application constructs a cross-language TTS model based on VQTTS, models the language speaking style and the speaker timbre respectively, and thus realizes cross-language speech synthesis with high nativeness and similar timbre to the target speaker.
Owner:AISPEECH CO LTD

AI interviewing methods, electronic devices, storage media, and software products

ActiveCN121212135BThe comprehensiveness of the answer is obtainedSemantic analysisSpeech rateEngineering
This application relates to the field of artificial intelligence technology, disclosing an AI interview method, electronic device, storage medium, and program product. The method includes: acquiring interview text; determining the candidate's first speaking speed and first pause frequency based on the interview text; the first pause frequency describing the number of pauses during the candidate's input of the interview audio corresponding to the interview text; determining the semantic integrity probability of the interview text; determining the probability that the candidate will not continue speaking based on the first speaking speed, the first pause frequency, and the semantic integrity probability; and determining whether to continue asking questions based on the probability that the candidate will not continue speaking. This allows the AI ​​interviewer to adaptively adjust the timing of continuing to ask questions according to the candidate's speaking style, enabling the AI ​​interviewer to take over the conversation at an appropriate time, thereby improving the candidate's interview experience.
Owner:BEIJING NIUKE TECH CO LTD

Large model-based multi-feature voice identification method, apparatus and device, and medium

The invention relates to the technical field of voice semantics, and discloses a multi-feature voice identification method based on a large model, and the method comprises the steps: generating a voice fusion feature according to an initial semantic feature and an initial voice feature; and generating a voice identification result of the initial voice information based on a preset voice identification decision model and the voice fusion feature. Through the above mode, the semantic feature and the voice feature are fused to generate the voice fusion feature, and the semantic information and the acoustic information of the voice are comprehensively considered. The voice identification result generated based on the preset voice identification decision model and the voice fusion feature can adapt to voice data of different languages, different speaking styles and different background environments, and the accuracy of identifying the voice information by the voice identification system is improved in the business fields of financial science and technology, medical treatment, health, pension and the like.
Owner:PING AN TECH (SHENZHEN) CO LTD

Speech recognition methods, devices, equipment and storage media

ActiveCN115841813BSpeech recognitionSpeaking styleFeature parameter
This disclosure provides a speech recognition method, apparatus, device, and storage medium, relating to the field of computer technology, specifically to speech technology and artificial intelligence technologies such as deep learning. The specific implementation involves: preprocessing the speech data to be recognized to determine the speech feature parameters corresponding to the speech data; using a pre-trained style recognition model to recognize the speech feature parameters to determine the style feature vector corresponding to the speech data; and based on the style feature vector, using the pre-trained speech recognition model to recognize the speech data to generate a recognition result corresponding to the speech data. Therefore, in the speech recognition process, by recognizing the speech data based on the style feature vector corresponding to the speech recognition data, the influence of speaking style on speech recognition is avoided, improving the accuracy and reliability of the speech recognition results.
Owner:BEIJING YUANLI WEILAI SCI & TECH CO LTD

System

PendingJP2026018740AData processing applicationsEngineeringSpeaking style
An object of a system according to an embodiment is to provide specific advice for a user to act with more confidence.SOLUTION: A system according to an embodiment includes a speaking style analysis unit, a behavior analysis unit, an advice generation unit, and a successful experience recording unit. The speaking style analysis unit analyzes the speaking style of the user. The behavior analysis unit analyzes a behavior of a user. The advice generation unit generates specific advice for the user to act with more confidence on the basis of the data obtained by the speaking style analysis unit and the action analysis unit. The successful experience recording unit records a successful experience of a user.SELECTED DRAWING: Figure 1
Owner:SOFTBANK GROUP CORP