Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

39 results about "Speaking style" patented technology

Voice synthesis method and system based on VITS improvement

The invention provides a voice synthesis method and system based on VITS improvement, and the method comprises the steps: optimizing a text encoder of a VITS model, introducing a large language model, and enabling the emotion, intention and speaking style of an input text to be captured when the text is encoded; a random disturbance item is introduced when the Q value is dynamically planned and solved, the alignment flexibility in the initial training stage is improved, meanwhile, monotonicity constraint is strictly kept, and it is avoided that a suboptimal solution is obtained through convergence too early; a ConvNeXt module is used as a basic backbone network of a decoder, and ISTFT is utilized to efficiently reconstruct a time domain signal, so that waveform up-sampling is realized, redundant calculation of traditional transpose convolution is avoided, and reasoning is accelerated. According to the method, the reasoning speed, the emotion expression ability and the style control flexibility of speech synthesis can be effectively improved, a new solution is provided for cross-language diversified speech synthesis, and a reference is provided for the more efficient and more intelligent development of the speech synthesis technology.
Owner:豫章师范学院

Method, electronic device, and computer program product for generating video

A method includes obtaining a reference image and a reference speech, the reference image specifying a head of a target object in the video, and the reference speech specifying a voice of the target object; and generating, based on the reference image and the reference speech, a fusion vector by combining a feature of the head and a feature of the voice. The method further includes generating, based on the fusion vector, a plurality of video frames in a video that represents the target object speaking in a timbre of the reference speech by denoising a plurality of initial frames including noise; and generating the video based on the plurality of video frames. In embodiments of the present disclosure, a video in which a semantic feature and a speaking style of the target object are merged can be generated, and the resolution and quality of the generated video are enhanced.
Owner:DELL PROD LP

ASV system risk assessment method and system based on multi-dimensional pronunciation characterization decoupling and fusion

PendingCN121354597ASpeech synthesisFeature extractionSpeaking style
The invention provides an ASV system risk assessment method and system based on multi-dimensional pronunciation characterization decoupling and fusion, and the method comprises the steps: obtaining a reference audio of a target user, and carrying out the decoupling extraction of a volume feature vector, a pitch feature vector, and a speaking style feature vector from the reference audio through a multi-dimensional pronunciation feature extraction network; a text to be verified is converted into a phoneme sequence, the phoneme sequence is input into a pronunciation feature prediction network for predicting a frame-level dynamic pronunciation feature sequence based on the phoneme sequence, and three pronunciation feature vectors are injected into the network through an adaptive instance normalization mechanism to dynamically modulate a prediction process. Generating a frame-level dynamic pronunciation feature sequence containing time sequence rhythm change; inputting a VAE-GAN synthesis trunk, re-injecting the three pronunciation feature vectors through an adaptive instance normalization mechanism, and generating a test voice sample consistent with the voiceprint characteristics of the target user; and initiating an identity verification query for the ASV system, and calculating safety indexes of different user groups according to a verification result so as to evaluate the risk of the ASV system.
Owner:FUJIAN NORMAL UNIV

Real-time live distribution system

PendingJP2025186961ASelective content distributionDistribution systemSpeaking style
To provide a real-time distribution system that is capable of enhancing communication between a distributor and a viewer at low costs and enables the distributor to obtain a benefit.SOLUTION: A real-time live distribution system for performing real-time live distribution of an idle talk dialog content toward a plurality of viewer terminals comprises a distributor terminal 1 including: an extraction section 11 having a virtual human for extracting viewer comments input from the plurality of viewer terminals in real time; an analysis section 12 for performing text analysis on the extracted viewer comments; a reading section 13 for reading a text using voice synthesis on the basis of a text analysis result; a feeling parameter setting section 14 for setting a virtual human feeling parameter for categorizing the viewer comments, which are at least a positive comment or negative comment, to determine the positive comment as a plus comment and the negative comment as a minus comment; and a change section 15 for changing a way of speaking according to the set feeling parameter.SELECTED DRAWING: Figure 2
Owner:CRYSTAL METHOD CO LTD

Audio-driven facial animation supporting varying identity and talking styles

The invention discloses an audio driven facial animation supporting varying identity and speaking styles. In various examples, systems and methods are disclosed for animating a virtual or digital actor or image using audio-driven animation. The system may identify an animation of the grid corresponding to the audio data and the indication of the speaking style. The system may generate a plurality of vertex increments using animations and neutral gestures of the grid. The system may update the machine learning model using the plurality of vertex increments, the audio data, and the indication of the speaking style to generate an output vertex increment for the grid given an input speaking style and input audio data.
Owner:NVIDIA CORP

System

PendingJP2026029847AData processing applicationsMoodSpeaking style
An object of a system according to an embodiment is to detect an emotion of a user and provide appropriate advice on the basis of the detected emotion.SOLUTION: A system according to an embodiment includes a smart emotion detection speaker, an emotion detection AI, and a mentoring generation AI. The smart speaker detects emotions from the user's voice and speech. The emotion detection AI generates words based on emotions. The mentoring generation AI listens to the user's distress and provides advice.SELECTED DRAWING: Figure 1
Owner:SOFTBANK GROUP CORP

System

A system is provided.SOLUTION: A system, comprising: means for selecting a person with whom a user wishes to interact; means for generating a AI model to mimic the person's thought patterns and speech; means for receiving a user question and generating a response based on the AI model; and means for displaying the generated response to the user.SELECTED DRAWING: Figure 1
Owner:SOFTBANK GROUP CORP

Speech style generation method and device, electronic equipment and storage medium

The present disclosure relates to a method and device for generating a speaking style, an electronic device and a storage medium. The method comprises: fitting target style feature attributes based on multiple style feature attributes, determining fitting coefficients of each style feature attribute; determining a target style feature vector according to the fitting coefficients of each style feature attribute and multiple style feature vectors, the multiple style feature vectors corresponding to the multiple style feature attributes one by one; inputting the target style feature vector into a speaking style model to output target speaking style parameters, the speaking style model being obtained based on a framework of training a speaking style model based on the multiple style feature vectors; and generating a target speaking style based on the target speaking style parameters. The method can realize fast migration of the speaking style and improve the generation efficiency of the speaking style.
Owner:HISENSE VISUAL TECH CO LTD

AUDIO-CONTROLLED FACIAL ANIMATION TO SUPPORT VARYING IDENTITIES AND SPEAKING STYLES

PendingDE102025141313A1Speech analysisAnimationAnimationMesh animation
Several examples disclose systems and methods for animating virtual or digital performers or avatars using audio-driven animation. A system can identify an animation for a mesh corresponding to audio data and a specified speech style. Using the animation and a neutral pose for the mesh, the system can generate a variety of vertex deltas. Based on this variety of vertex deltas, the audio data, and the specified speech style, the system can update a machine learning model to generate output vertex deltas for the mesh, taking into account an input speech style and input audio data.
Owner:NVIDIA CORP

System

PendingJP2026018011AInstrumentsProcessingServer
A system is provided.SOLUTION: Means for recording video and audio data of a face-to-face contact, means for transmitting the recorded video and audio data to a server, means for performing noise removal and text conversion on the transmitted video and audio data, means for analyzing wording and conversation content using a natural language processing technique on text data, means for analyzing a manner of speaking and a tone of voice from the audio data, means for analyzing facial expression and body language from the video data, and means for digitizing each analyzed feature, A system comprising: means for scoring; means for training a learning model using interview data of existing employees; means for evaluating new interview data using the learning model; and means for displaying evaluation results and providing feedback to a user.SELECTED DRAWING: Figure 1
Owner:SOFTBANK GROUP CORP

system

We provide the system. [Solution] A means for a person with specialized knowledge to input that knowledge and speaking style as information, A method for training an AI model using natural language processing technology based on that input information, A means of storing and making available trained AI models on the platform, A means for users to search for AI models through the platform and begin interacting with them, A means by which an AI model generates a response to that AC input and sends it back to the user, A means of providing responses as an interactive learning experience and enhancing immersion through visual and auditory means, A system that includes this.
Owner:SOFTBANK GROUP CORP

Virtual image construction method and device, computer equipment, storage medium and program product

The invention discloses a virtual image construction method and device, computer equipment, a storage medium and a program product. The method comprises the following steps: acquiring target text data and first video data; performing emotion recognition processing on the target text data to obtain a corresponding expression tag; performing voice conversion processing on the target text data based on the expression label to obtain target voice data; driving the lip of the target user in the first video data to act by taking the target voice data as a driving signal to obtain second video data; extracting a first mouth opening and closing feature in the first video data, and determining a first speaking style feature index; adjusting the general expression feature data based on the first speaking style feature index and the general speaking style feature index to obtain expression feature data corresponding to the target user; and performing expression editing processing on the second video data by using the expression feature data to obtain the target video data, thereby enhancing the reality sense of the virtual image and the immersion sense of the user.
Owner:MIGU CO LTD +1

System

An object of the system according to the embodiment is to reproduce the voice and the way of speaking of the deceased and to reestablish emotional connection.SOLUTION: A system includes a collection unit, a learning unit, a reception unit, a generation unit, and a reproduction unit. The collection unit collects video and audio data of the deceased. The learning unit learns the data collected by the collection unit and analyzes the voice and the way of speaking of the deceased. The receiving unit receives an input from a user. The generation unit generates new words in the voice of the deceased based on the input accepted by the acceptance unit. The reproduction unit reproduces the words generated by the generation unit.SELECTED DRAWING: Figure 1
Owner:SOFTBANK GROUP CORP

System

A system is provided.SOLUTION: A system comprising: a default setting unit configured to allow a user to register; a unit configured to collect and store conversation data with the user; a unit configured to analyze the collected conversation data and learn a manner of speaking and a context of the user; a unit configured to train a user-specific model based on the learned data; a unit configured to generate a reply of the conversation using the generated user-specific model; and a unit configured to transmit the generated reply to a user terminal.SELECTED DRAWING: Figure 1
Owner:SOFTBANK GROUP CORP

Method, electronic device, and computer program product for generating video utilizing a fusion vector based on reference image and reference speech

A method includes obtaining a reference image and a reference speech, the reference image specifying a head of a target object in the video, and the reference speech specifying a voice of the target object; and generating, based on the reference image and the reference speech, a fusion vector by combining a feature of the head and a feature of the voice. The method further includes generating, based on the fusion vector, a plurality of video frames in a video that represents the target object speaking in a timbre of the reference speech by denoising a plurality of initial frames including noise; and generating the video based on the plurality of video frames. In embodiments of the present disclosure, a video in which a semantic feature and a speaking style of the target object are merged can be generated, and the resolution and quality of the generated video are enhanced.
Owner:DELL PROD LP

system

PendingJP2026045369AData processing applicationsEngineeringSpeaking style
The system according to the embodiment aims to reproduce the speaking style and knowledge of great people, and to provide the user with the feeling that they are having a conversation with a great person. [Solution] A system according to an embodiment includes a collection unit, an analysis unit, a reproduction unit, a reception unit, and an answering unit. The collection unit collects voice data. The analysis unit analyzes the data collected by the collection unit. The reproduction unit reproduces the speaking style and knowledge of the great person based on the data analyzed by the analysis unit. The reception unit accepts input from a user. The answering unit answers like the great person based on the input accepted by the reception unit.
Owner:SOFTBANK GROUP CORP

AI interview method, electronic equipment, storage medium and program product

The invention relates to the technical field of artificial intelligence, and discloses an AI interview method, electronic equipment, a storage medium and a program product, and the method comprises the steps: obtaining an interview text; determining a first speech speed and a first pause frequency of the candidate according to the interview text; the first pause frequency is used for describing the pause frequency in the process that the candidate inputs the interview audio corresponding to the interview text; determining the semantic integrity probability of the interview text; determining the probability that the candidate does not continue speaking according to the first speech speed, the first pause frequency and the semantic integrity probability; and determining whether to continuously ask questions according to the probability that the candidates do not continuously speak. Therefore, the opportunity of continuing to ask questions by the AI interviewer can be adaptively adjusted according to the speaking style of the candidate, so that the AI interviewer takes over a conversation at a reasonable opportunity, and the interview experience feeling of the candidate is improved.
Owner:BEIJING NIUKE TECH CO LTD

Model updating method and device, voice conversion method, equipment and storage medium

ActiveCN116543780BEasy to updateimprove accuracySpeaking styleNetwork model
This application provides a model update method and apparatus, a speech conversion method, device, and storage medium, belonging to the field of financial technology. The method includes: acquiring sample speech data; inputting the sample speech data into a neural network model; encoding the sample speech data through an encoding network to obtain an initial audio feature vector; indexing the initial audio feature vector based on a preset codebook to obtain an audio frame index; extracting phoneme features from the initial audio feature vector based on the audio frame index to obtain an initial phoneme feature vector; performing speech alignment on the initial phoneme feature vector to obtain a sample audio embedding vector; decoding the sample audio embedding vector and the speech style embedding vector through a decoding network to obtain synthesized speech data; and updating the parameters of the neural network model based on the synthesized speech data and the sample speech data to obtain a speech conversion model. This application can improve the accuracy of the model in speech conversion.
Owner:PING AN TECH (SHENZHEN) CO LTD

Information processing system, information processing method, and information processing program

PCT designated stageWO2026079118A1Speech synthesisInformation processingSpeaking style
An information processing system according to the present disclosure comprises: a reception unit that receives, from a user, subjective evaluation information relating to speech evaluation, and designation of a speaking style; a generation unit that, on the basis of the subjective evaluation information and the designation of the speaking style, converts a low-dimensional embedding vector corresponding to the speaking style adjusted by the user to generate a high-dimensional speaking style embedding vector; and a speech synthesis unit that synthesizes speech from the speaking style embedding vector and text information defining speech to be synthesized.
Owner:SONY GROUP CORP

Controllable text-to-speech method based on decoupled multi-modal prompt and chain guide

The application provides a controllable text-to-speech method based on decoupled multi-modal cues and chain-based guidance, which comprises the following steps: constructing a unified multi-modal style encoder for mapping style cues from reference audio or descriptive text to a shared embedding space to generate style-conditioned embeddings, and training the multi-modal style encoder; constructing a text-to-speech model based on a latent diffusion transformer, which takes the latent representation of mel-spectrogram as the processing object and receives content text, speaker prosody and speaking style as input conditions; in the inference stage of the text-to-speech model, a chain-based classifier-free guidance mechanism is adopted to realize independent and continuous control of the content, prosody and style of the generated speech by independently adjusting the guidance strength of the content, prosody and style-conditioned embeddings. The independent and continuous control of the content, prosody and style of the generated speech is realized. The training stability and convergence speed of the model are significantly improved.
Owner:UNIV OF SCI & TECH OF CHINA

AUDIO-CONTROLLED FACIAL ANIMATION TO SUPPORT VARYING IDENTITIES AND SPEAKING STYLES

PendingDE102025141314A1Speech analysisAnimationAnimationMesh animation
Several examples disclose systems and methods for animating virtual or digital performers or avatars using audio-driven animation. A system can identify an animation for a mesh corresponding to audio data and a specified speech style. Using the animation and a neutral pose for the mesh, the system can generate a variety of vertex deltas. Based on this variety of vertex deltas, the audio data, and the specified speech style, the system can update a machine learning model to generate output vertex deltas for the mesh, taking into account an input speech style and input audio data.
Owner:NVIDIA CORP

Cross-language speech synthesis method and system based on dual speaker embedding

The embodiment of the application provides a cross-language speech synthesis method based on double speaker embedding. The method comprises the following steps: inputting text and native language speaker embedding into a txt2vec acoustic model, determining the phoneme sequence code of the text through a text encoder, and determining the vector quantization acoustic feature and the auxiliary feature from the phoneme sequence code and the native language speaker embedding through a decoder; inputting the target language speaker embedding, the vector quantization acoustic feature and the auxiliary feature into a vec2wav vocoder, extracting the X-vector feature of the target language speaker embedding, inputting the X-vector feature, the vector quantization acoustic feature and the auxiliary feature into a feature encoder, and obtaining a cross-language acoustic feature; and determining the cross-language synthesized speech of the cross-language acoustic feature by using a generator. The embodiment of the application constructs a cross-language TTS model based on VQTTS, models the language speaking style and the speaker timbre respectively, and thus realizes cross-language speech synthesis with high nativeness and similar timbre to the target speaker.
Owner:AISPEECH CO LTD

AI interviewing methods, electronic devices, storage media, and software products

ActiveCN121212135BThe comprehensiveness of the answer is obtainedSemantic analysisSpeech rateEngineering
This application relates to the field of artificial intelligence technology, disclosing an AI interview method, electronic device, storage medium, and program product. The method includes: acquiring interview text; determining the candidate's first speaking speed and first pause frequency based on the interview text; the first pause frequency describing the number of pauses during the candidate's input of the interview audio corresponding to the interview text; determining the semantic integrity probability of the interview text; determining the probability that the candidate will not continue speaking based on the first speaking speed, the first pause frequency, and the semantic integrity probability; and determining whether to continue asking questions based on the probability that the candidate will not continue speaking. This allows the AI ​​interviewer to adaptively adjust the timing of continuing to ask questions according to the candidate's speaking style, enabling the AI ​​interviewer to take over the conversation at an appropriate time, thereby improving the candidate's interview experience.
Owner:BEIJING NIUKE TECH CO LTD

System

PendingJP2026018740AData processing applicationsEngineeringSpeaking style
An object of a system according to an embodiment is to provide specific advice for a user to act with more confidence.SOLUTION: A system according to an embodiment includes a speaking style analysis unit, a behavior analysis unit, an advice generation unit, and a successful experience recording unit. The speaking style analysis unit analyzes the speaking style of the user. The behavior analysis unit analyzes a behavior of a user. The advice generation unit generates specific advice for the user to act with more confidence on the basis of the data obtained by the speaking style analysis unit and the action analysis unit. The successful experience recording unit records a successful experience of a user.SELECTED DRAWING: Figure 1
Owner:SOFTBANK GROUP CORP

Speech synthesis method and system based on improved vits

The application provides a speech synthesis method and system based on an improved VITS, wherein a large language model is introduced by optimizing a text encoder of a VITS model, so that the input text can capture the emotion, intention and speaking style when the text is encoded; a random disturbance term is introduced when a Q value is solved by dynamic programming, so that the alignment flexibility in the early training stage is improved, and the monotonicity constraint is strictly maintained to avoid early convergence to a suboptimal solution; a ConvNeXt module is used as a basic backbone network of a decoder, and ISTFT is used to efficiently reconstruct a time domain signal, so that the upsampling of the waveform is realized, the redundant calculation of a traditional transpose convolution is avoided, and the inference is accelerated. The application can effectively improve the inference speed, emotion expression ability and flexibility of style control of speech synthesis, provides a new solution for cross-language diversified speech synthesis, and provides a reference for the development of speech synthesis technology in a more efficient and intelligent direction.
Owner:豫章师范学院

system

PendingJP2026045517AData processing applicationsAnimationEngineeringSpeaking style
The system according to the embodiment aims to reproduce the speaking style and facial expressions of the participant and generate realistic video. [Solution] A system according to an embodiment includes a collection unit, an analysis unit, a generation unit, and a video generation unit. The collection unit collects audio or video data of the participant. The analysis unit analyzes the data collected by the collection unit and reproduces the participant's speaking style or facial expression. The generation unit generates a model based on the data analyzed by the analysis unit. The video generation unit generates a video based on the model generated by the generation unit.
Owner:SOFTBANK GROUP CORP

Processing method, intelligent terminal and storage medium

The invention provides a processing method, an intelligent terminal and a storage medium. The method comprises the following steps: S11, acquiring a learning audio; and S12, outputting difference information between the voice style of the learning audio and the target voice style. Therefore, the user is allowed to send out the audio, namely the learning audio, according to the target voice style, the voice style of the learning audio is compared with the target voice style, and the difference information between the two styles is obtained, so that the user can clearly know the difference, and the user experience is improved. And correcting the pronunciation mode of the voice style according to the relationship of the mouth shape, expression and the like of the user, so that a processing method including new languages, different languages or imitation of specific speaking styles and the like is realized, and personalized learning experience of partner training and free practice can be realized.
Owner:SHENZHEN TRANSSION HLDG CO LTD

system

PendingJP2026105315AEngineeringSpeaking style
We provide the system. [Solution] A means of receiving presentation materials and audio information from users using an information terminal, A means of converting speech information into text data using speech recognition technology, A means for analyzing text data and presentation materials to evaluate speaking style characteristics and presentation material structure, A means for generating feedback based on analysis results using generative artificial intelligence technology, A means of providing improvement instructions to users, assuming store operations, A means for presenting the aforementioned feedback to the user, A system that includes this.
Owner:SOFTBANK GROUP CORP

AI-controlled voice optimization CTI

We provide a CTI (Computer Telephony Integration) system that combines AI-powered voice conversion control and customer information analysis to assist with making calls or to perform AI-powered automated calling and response. [Solution] In a cloud-based or on-premise CTI system, a customer information acquisition mechanism acquires industry, company size, job title, past call results, etc., and an analysis engine analyzes performance indicators for each voice conversion mode and recommends the optimal mode. The voice conversion engine can convert and output operator voice or AI-generated voice in real time and conducts calls according to the recommended mode. Furthermore, a learning mechanism analyzes the call results, and the analysis engine continuously improves the accuracy of voice mode recommendations, thereby achieving uniformity in call quality and improved call outcomes through optimal voice quality and speaking style according to customer attributes.
Owner:田中 芳明

Expression generation algorithm based on voice-driven robot

The invention relates to the technical field of artificial intelligence, in particular to an expression generation algorithm based on a voice-driven robot, and the algorithm comprises the steps: S1, obtaining the voice information of a user, and carrying out the preprocessing of the voice information, and obtaining the audio information; coding based on the audio information to generate phoneme logarithm features, time domain information, speaking style features and mixed shape coefficients; s2, based on the phoneme logarithm features, the time domain information, the speaking style features and mixed shape coefficient decoding, screening out a plurality of corresponding current basic shapes from a plurality of basic shapes of a pre-stored expression base space; s3, generating an actuator control vector instruction based on the plurality of current basic shapes and a preset basic model; and S4, controlling the face of the robot to generate an expression corresponding to the voice information based on the actuator control vector instruction. Control is achieved through the mapping relation between the voice information and the actuator, the encoding-decoding process is simple, and the method is suitable for a real-time interaction scene.
Owner:王晶