Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

3975results about "Speech synthesis" patented technology

Comprehensive AI-enabled systems for immersive voice, companion, and augmented / virtual reality interaction solutions

A computer-implemented method for operating an artificial intelligence voice agent system includes receiving voice input through communication channels; analyzing converted text through natural language processing (NLP) pipelines implementing intent recognition and sentiment analysis detecting emotional cues using a multimodal large language model (LLM); generating response content using machine learning models trained on domain-specific corpora; converting generated responses to synthetic speech through text-to-speech (TTS) engines; integrating with a customer relationship management (CRM) platforms or an enterprise resource planning (ERP) database; and implementing continuous learning by updating language understanding models using conversation logs, voice recognition parameters based on user feedback, and response generation patterns. One implementation is a computer-implemented system and method that operates a suite of intelligent interactive devices and platforms including an artificial intelligence voice agent, enhanced communication platforms, an intimacy companion system, and augmented / virtual reality eyeglasses. Further, one implementation includes AR / VR eyeglasses that project visual content onto interchangeable lenses or directly onto the user's retina via laser-based retinal projection, provide prescription adjustments, incorporate ear-mounted sensors for monitoring physiological parameters like heart rate, oxygen saturation, and blood pressure, and utilize wireless data transmission, onboard environmental sensing, and remote calibration, all designed to offer dynamically adaptive, secure, and context-aware interactions across communication, personal assistance, health monitoring, and immersive augmented or virtual reality environments.
Owner:TRAN BAO

Personalized and dynamic text to speech voice cloning using incompletely trained text to speech models

Systems and methods are provided for machine learning models configured as zero-shot personalized text-to-speech models which comprise a feature extractor, a speaker encoder, and a text-to-speech module. The feature extractor is configured to extract acoustic features and prosodic features from new target reference speech associated with the new target speaker. The speaker encoder is configured to generate a speaker embedding corresponding to the new target speaker based on the acoustic features extracted from the new target reference speech. The text-to-speech module is configured to generate the personalized voice corresponding for the new target speaker based on the speaker embedding and the prosodic features extracted from the new target reference speech without applying the text-to-speech module on new labeled training data associated with the new target speaker.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Precise international communication digital human real-time dialogue method fused with multi-modal technology

The invention discloses an accurate international communication digital human real-time dialogue method fused with a multi-modal technology. The method comprises the following steps: S1, constructing a digital human image and tone; s2, propagation content generation and problem guidance; s3, semantic analysis and intention clarification based on the real-time voice dialogue; s4, geographic preference modeling and path planning; s5, cross-context propagation content generated based on retrieval enhancement is generated; s6, visually displaying the propagation content; and S7, carrying out digital human-driven multi-language propagation content real-time output and feedback closed-loop optimization. Through accurate utterance expression analysis, accurate international propagation problem recommendation is provided, cross-context propagation content generation based on semantic understanding is realized, digital people with voice features and visual images are constructed, real-time dialogue interaction of users is realized, and user experience is improved. The system can carry out geographic modeling according to the region where the accurate problem is located, language preference and propagation object culture characteristics, and differential propagation path planning is achieved.
Owner:HUNAN NORMAL UNIVERSITY

Training and speech generation methods and apparatuses for speech generation model, electronic device, computer-readable storage medium, and computer program product

The present application provides training and speech generation methods and apparatuses for a speech generation model, an electronic device, a computer-readable storage medium, and a computer program product. The method comprises: obtaining a first speech generation model; obtaining sample data of a plurality of modalities; on the basis of a prompt image sequence and speech text, respectively calling a plurality of encoders to perform encoding, so as to obtain a multi-modal encoding vector sequence; on the basis of the multi-modal encoding vector sequence, calling a decoder to perform decoding, so as to obtain decoded text; determining a probability distribution for the decoded text and the sample data of the plurality of modalities, and determining a target loss on the basis of the probability distribution; and on the basis of the target loss, updating parameters of the decoder and at least one of the encoders, wherein the updated decoder and the plurality of updated encoders are configured to form a second speech generation model, and the second speech generation model is used to generate target speech text corresponding to a prompt image sequence of a target object.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

System and method for emotionally intelligent, personalized AI avatar-based health coaching using multi-domain data and adaptive behavioral intelligence

A programmatically generated AI avatar includes a customizable personality module, acting as the embodied interface for a powerful AI “mind” that delivers personalized coaching to improve user health, well-being, and longevity. The system uses machine learning, large language models, and biometric modeling to synthesize real-time, multi-modal health data—including sleep, nutrition, glucose, mood, and activity—and generate forward-prescribed KHAs. Unlike human coaches, it continuously adapts based on context and behavior, targeting the root cause: metabolic dysfunction—namely by restoring healthy, sustainable body composition through the preservation or building of lean muscle mass and reduction of excess fat. KHAs can also be shared with friends or programmatically generated AI avatars, allowing for coordinated action, emotional support, and accountability through social connection—further reinforcing positive behavior and adherence. The system's reinforcement learning engine incorporates both individual response data and anonymized population-level insights to optimize recommendations over time, learning which interventions are most effective for users with similar physiological and behavioral profiles. First validated with Olympic athletes—resulting in measurable improvements and medal-winning outcomes—this system offers a scalable, emotionally intelligent coaching engine that exceeds human capability, designed for the ultimate purpose of supporting sustainable health, resilience, and human thriving.
Owner:GOLD AI LLC

Automatic lecturer video generation method based on AI speech synthesis and animation driving

The invention discloses a lecturer video automatic generation method based on AI speech synthesis and animation driving. The method comprises the following steps: performing structured analysis on a PPT or a text script through an improved interior point method and an incremental shortest path algorithm; performing semantic grouping by applying a full-dynamic parallel single-link clustering algorithm and generating an enhanced script with an expressive mark; a CosyVoice technology is combined with a low-rank approximation method to generate a high-quality voice data stream; establishing a mapping relation between contents and action expressions through semantic analysis, and generating a complete action expression instruction set; and driving the digital human model by using the msueTalk technology, and generating a final lecturer teaching video through a parallel rendering algorithm. According to the invention, the method achieves the efficient and automatic generation of the education video, remarkably improves the content production efficiency, reduces the production cost, and guarantees the specialty and expressive force of the teaching video.
Owner:SHENZHEN XUEYOU TECHNOLOGY CO LTD

Bluetooth communication intelligent speech translation method and system based on multi-mode enhancement

The invention relates to the technical field of artificial intelligence, and discloses a Bluetooth communication intelligent speech translation method and system based on multi-mode enhancement, and the method comprises the steps: collecting a multi-channel audio signal through a built-in multi-microphone array of a Bluetooth device, carrying out the dynamic direction self-adaptive beam forming of the multi-channel audio signal, and carrying out the self-adaptive beam forming of the multi-channel audio signal; extracting a Mel spectrogram feature of the direction enhancement signal, identifying lip regions of a plurality of candidate speakers in each frame of real-time speaking video captured by a camera, performing time sequence convolution on the lip regions to obtain a lip movement time sequence embedded vector, calculating a correlation score with the Mel spectrogram feature, separating the direction enhancement signal, and obtaining a lip movement time sequence embedded vector; and performing text transcription and conversion on the high-confidence separation voice to obtain a translation language text, and sending the synthesized target translation voice to a preset mobile terminal through the Bluetooth device to obtain a target translation result. According to the method, the real-time performance and accuracy of speech translation are improved in a multi-person scene, far-field speech, noise interference and accent difference.
Owner:SHENZHEN DIE MICRO SEMICON CO LTD

System and method for multi-modal ai conversational interface improving website navigation and user interaction

The present invention relates to a system for transforming static websites into artificial intelligence (AI)-enabled interactive multi-modal conversational platforms. The system comprises a computing device having a processor for receiving user queries as text or speech input through an input module cooperating with a speech-to-text module. A natural language processing (NLP) module interprets intent, classifies user context, and retrieves grounded information from multiple webpages. A persona adaptation module dynamically modifies vocabulary, tone, and avatar representation across roles such as sales assistant, recruiter, educator, healthcare professional, etc. A response generator module produces structured natural language output, transmitted to a text-to-speech synthesis module and an avatar generation module to render synchronized lifelike video responses. An output rendering module displays multi-modal responses include text, audio, and video, thereby enabling direct navigation and escalation beyond limitations of conventional static websites.
Owner:NALLAM SREE RAMA CHANDRA MURTY

Large model cross-modal collaborative understanding method and device

The invention provides a large-model cross-modal collaborative understanding method and device. The fusion efficiency and the understanding capability of multi-modal information can be improved. The large-model cross-modal collaborative understanding method comprises the following steps: preprocessing visual data, language data and sound data to obtain preprocessed visual data, language data and sound data; wherein the preprocessing comprises the steps of performing adaptive size adjustment and normalization on visual data, performing word segmentation and dynamic truncation on language data, and performing band-pass filtering and spectral noise reduction on sound data; extracting visual features, language features and sound features based on the preprocessed visual data, language data and sound data; on the basis of an adaptive mapping network and a mixed granularity cross-modal attention mechanism, performing feature alignment on the visual features, the language features and the sound features to obtain aligned feature vectors; and based on a dynamic routing architecture, fusing the aligned feature vectors to generate a unified multi-modal representation.
Owner:SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD

Computer implemented system and method for automatically generating offer ranges for candidates in an interviewing process

A computer implemented system and method for generating offer ranges for candidates in an interviewing process is disclosed. The system generates an AI-based interviewer simulating human-based interactions for conducting an ongoing interview with candidates. The system analyzes data associated with candidates obtained during ongoing interview. The system process analyzed responses of candidates to determine contextual attributes associated with responses using ML models. The system automatically generates follow-up interview questions to be delivered to candidates during ongoing interview based on analyzed responses from candidates, by applying AI model to contextual attributes associated with responses. The system generates recruitment scores for candidates based on analyzed responses, contextual attributes, and interpreted non-verbal cues, associated with candidates, using AI model. The system generates offer ranges for candidates based on recruitment scores using AI model. The system provides information associated with selected candidates, and offer ranges generated for selected candidates, to users.
Owner:TALVIEW INC

Video translation method and system based on artificial intelligence

The invention discloses a video translation method and system based on artificial intelligence. The method relates to the technical field of video translation and comprises the following steps of original sound track extraction, target AI speaker adaptation, AI dubbing generation and mouth shape synchronization and video synthesis. According to the method, independent audio and video streams are obtained by adopting an audio and video separation technology, and multiple original sound tracks are extracted through a voice separation model; matching or generating an adaptive target AI speaker module in a preset tone library; converting the original language voice into a text, translating the text into a target language text, and synthesizing an AI dubbing audio track in combination with a target AI speaker module; and finally, the independent video stream and the multi-AI dubbing audio track are input into the mouth shape synchronization model to output a translated video, so that the timbre fitting degree, the voice quality and the voice consistency of the same speaker of AI dubbing are improved, and meanwhile, the resource utilization rate of video translation and the processing efficiency under batch tasks are improved. The problem that in the prior art, video translation is low in quality and efficiency is solved.
Owner:BEIJING DEEP LOGIC INTELLIGENT TECHNOLOGY CO LTD

Method and system for converting or encoding text

A publishing system with components, including:a system configured to receive at least one document including text that defines a base alphabet in one or more formats;a system configured to provide additional data for a reader to better understand the document which includes:a method of encoding or marking up non-phonetic words in the document to enable the reader to decode sounds of each non-phonetic word; anda system configured to output an encoded document with the text and the additional data in one or more formats,wherein the method of automatically encoding the non-phonetic words to make the encoded words phonetic:for at least one character (“spelling character”) in the non-phonetic word, using a compound character that includes the spelling character and a sound character, wherein the sound characters:are human-readable characters in the base alphabet and / or in one or more secondary alphabets,are added to the spelling characters to indicate that each spelling character makes the usual sound of the sound character,are added so that spelling characters can be visually discriminated from sound characters,are added such that a reader can recognize the non-phonetic word by sight because the spelling of the word is unchanged, andare added to the spelling characters such that the spelling characters and the sound characters remain human-readable such that the spelling character and the sound character of each compound character are within one visual field; andautomatically outputting the encoded words in a human-readable form / format such that the compound characters in the encoded word visually indicate which of the spelling characters have a sound other than their usual sound and what sound each character makes in the non-phonetic word when it does not make its usual sound.
Owner:STEPHEN CHRISTOPHER COLIN

Voice interaction processing method, device and system, intelligent door lock and cloud server

The invention is suitable for the technical field of intelligent door locks, and provides a voice interaction processing method, device and system, an intelligent door lock and a cloud server, and the method comprises the steps: receiving a voice instruction sent by a user, and carrying out the local recognition, so as to obtain a local recognition result and the confidence of the local recognition result; when the local recognition result is matched with the local instruction set and the confidence is high, executing a local operation corresponding to the local recognition result; otherwise, establishing a secure communication session with the cloud server based on voiceprint biological characteristics extracted from the voice instruction; sending the voice instruction to a cloud server through the secure communication session, and receiving a cloud operation instruction and / or a dynamic response text returned by the cloud server; and executing corresponding operation according to the cloud operation instruction, and / or converting the dynamic response text into voice for broadcasting through a text-to-voice engine. The intelligent door lock voice interaction system solves the problem that the real-time performance, the reliability, the safety and the interaction intellectualization are difficult to balance in the existing intelligent door lock voice interaction technology.
Owner:SHANGHAI ZHENGZHI INTELLIGENT TECH CO LTD

Business processing method and device based on multi-agent cooperation, equipment and medium

The invention provides a business processing method and device based on multi-agent collaboration, equipment, a medium and a program product, which can be applied to the technical field of digital human and artificial intelligence. The method comprises the following steps: in response to a service request initiated by a user through a first service channel, obtaining input information of the user, and creating or updating a session context object corresponding to the service request; analyzing the input information, generating a subtask sequence, and writing the subtask sequence into a session context object; according to the subtask sequence, a corresponding target agent is scheduled to execute operation, and an execution result is written back to the session context object; monitoring updating of the session context object, and generating a migration decision for migrating from the first service channel to the second service channel under the condition that the updated session context object meets a cross-channel migration condition; and based on the migration decision, synchronizing the session context object to the second service channel in a full amount, so that the business of the user is handled in the second service channel.
Owner:ANHUI BRANCH OF INDAL & COMML BANK OFCHINA

Method, system and device for voice interaction inside and outside vehicle and storage medium

The invention discloses a method, a system and equipment for voice interaction inside and outside a vehicle and a storage medium, and relates to the technical field of intelligent cabins and human-computer interaction. Identifying an interaction trigger type (in-vehicle active request, out-of-vehicle passive request, or system active trigger) based on the perceived data and / or the vehicle event; determining a corresponding response permission strategy according to the trigger type, and determining whether to allow the system to respond based on the response permission strategy; and if so, calling a large language model to generate a voice text, synthesizing voice by combining external perception characteristics, and broadcasting the voice through a loudspeaker outside the vehicle or a sound box in the vehicle. According to the scheme, three types of interaction intentions including the in-vehicle active request, the out-vehicle passive request or the system active triggering are recognized, the large language model is called to generate the scene-adaptive voice text, and personalized voice synthesis and broadcasting are performed in combination with the external perception characteristics, so that the response efficiency, the expression naturalness and the object adaptability of the in-vehicle and out-vehicle voice interaction are improved.
Owner:ZHEJIANG GEELY HLDG GRP CO LTD +1

Computing system, method, and medium for processing customer inquiries using speech-to-text, language model analysis, and text-to-speech services

An autonomous communication system includes a language model, a retrieval module, a caching mechanism, an autonomous agent, and a human interface for managing interactions and responses. A method and computer-readable medium for managing communication also include these components for processing, enhancing, summarizing, managing interactions, and allowing human intervention.
Owner:CDW LLC

Server, display device and digital human processing method

The embodiment of the invention provides a server, display equipment and a digital human processing method. The method comprises the following steps: receiving voice data input by a user and sent by the display equipment; broadcast voice is determined based on the voice data; extracting voice features of the broadcast voice; determining mouth shape parameters based on the voice features; determining emotion parameters and acquiring user image data; generating digital human image data based on the user image data, the emotion parameters and the mouth shape parameters; and sending the broadcast voice and the digital human image data to the display device, so that the display device plays the broadcast voice and displays a digital human image based on the digital human image data. According to the embodiment of the invention, the expression parameters and the mouth shape parameters are determined according to the voice data input by the user, the expression parameters and the mouth shape parameters are combined to generate the digital human image with better facial expression expression, and emotion customization and control are realized.
Owner:HISENSE VISUAL TECH CO LTD

Voice decoding method, system and equipment based on electroencephalogram signals and medium

The invention discloses a voice decoding method, system and device based on electroencephalogram signals and a medium, and relates to the technical field of electroencephalogram signal processing.The method comprises the steps that reading electroencephalogram signals and voice signals in the reading process of a to-be-tested person and imagination electroencephalogram signals in the imagination reading process of the to-be-tested person are collected; inputting the reading electroencephalogram signals and the voice signals into the DRCL to generate electroencephalogram characteristics containing voice information; training the DBM by using the electroencephalogram characteristics as input and using the voice signals as output, and adjusting the DBM by using imaginary electroencephalogram signals to construct a voice synthesizer; a mapping relation between the imaginary electroencephalogram signals and the voice signals is generated through the voice synthesizer, and decoding from the electroencephalogram signals to the voice signals is completed; according to the method, a deep representation correlation learning method is provided, potential correlation between electroencephalogram and voice signals can be deeply mined through a multi-layer network structure, a complex mode which is difficult to recognize by a traditional model is captured, and the voice decoding process is more accurate.
Owner:HARBIN INST OF TECH

Intelligent accompanying robot system

The invention discloses an intelligent accompanying robot system, and the system comprises the following modules: a multi-mode interaction module which is composed of a voice recognition unit, a voice synthesis unit, an emotion recognition unit, and a visual recognition unit, and is used for collecting the voice, facial expression, motion, and environment information of a user; the localized AI decision module is used for deploying a lightweight DeepSeekR1 large language model based on an ESP32-S3 edge computing chip, performing real-time processing on the multi-modal data and generating a social guidance strategy and an emotion intervention instruction; the data management module comprises a user behavior database and a privacy protection unit and is used for desensitizing the data and realizing sensitive data isolation through local storage; according to the method, the single-person intervention cost is reduced, and the time consumption of manual scene simulation is reduced.
Owner:NANTONG UNIV

Rhythm migration method and device, electronic equipment and storage medium

The invention relates to the technical field of voice processing, and provides a rhythm migration method and device, electronic equipment and a storage medium, and the method comprises the steps: obtaining a decoupled rhythm feature based on a source rhythm voice, and a decoupled tone feature based on the voice of a target speaker, the decoupled rhythm feature represents the rhythm of the source rhythm voice, and the decoupled tone feature represents the tone of the target speaker; the decoupled timbre features represent the timbre of the voice of the target speaker; generating a target voice vector sequence based on the text features of the target text, the voice features of the voice of the target speaker, the decoupled rhythm features and the decoupled timbre features; and synthesizing a target audio based on the target voice vector sequence. According to the method and the device, the decoupled rhythm features and the decoupled timbre features are acquired, and the target voice is generated based on the features, so that the problem of feature mixing is effectively relieved, the timbre purity of the target speaker in cross-person rhythm migration is ensured, the expressive force of rhythm migration is improved, and the synthesized audio is more natural and vivid.
Owner:IFLYTEK CO LTD

Real-time sound duplicating method and system based on end-cloud fusion

The invention provides a real-time sound copying method and system based on end-cloud fusion. The method comprises the following steps: a cloud end carries out real-time tone copying and voice synthesis on a small amount of voice data of a user based on an AI large model; timbre samples are collected when a user registers voice audio data, and user timbre voice data of a preset text are synchronously generated by a large model and are used as fine tuning training data of an end-side voice synthesis model; user tone voice data of a preset text and voice audio data registered by a user are utilized to carry out migration fine tuning training on an end-side voice synthesis model to adapt to the personalized tone of the user, so that high-quality output of the end-side voice synthesis model is ensured, and personalized voice replication is realized; and issuing the trained end-side speech synthesis model to the user equipment, and independently completing speech replication in a network-free or weak network environment. According to the invention, through automatic generation of the user tone data and adaptive fine tuning of the model, the tone of the user is deployed to the end side after fine tuning, and high-quality and high-adaptability sound replication of end-cloud collaboration is realized.
Owner:PACHIRA TIMES (ZHUHAI HENGQIN) INFORMATION TECH CO LTD

Providing private answers to non-vocal questions

Systems, methods, and non-transitory computer readable media including instructions for providing private answers to silent questions are described. Providing private answers to silent questions includes receiving signals indicative of particular facial micromovements in an absence of perceptible vocalization; accessing a data structure correlating facial micromovements with words; using the received signals to perform a lookup in the data structure of particular words associated with the particular facial micromovements; determining a query from the particular words; accessing at least one data structure to perform a look up for an answer to the query; and generating a discreet output that includes the answer to the query.
Owner:APPLE INC

Speech synthesis method and device, vehicle and storage medium

The embodiment of the invention provides a voice synthesis method and device, a vehicle and a storage medium, and the method comprises the steps: obtaining multi-modal data which comprises a voice signal, a user instruction, a historical interaction log and bus data; determining model input parameters according to the multi-modal data; according to the model input parameters, the static corpus and a preset large model, generating a personalized script, the preset large model being used for dynamically generating the personalized script in combination with the model input parameters and the static corpus; and synthesizing the broadcast audio according to the personalized script. According to the invention, the technical problem that the voice synthesis technology in the related technology cannot adapt to the personalized demands of the user is solved.
Owner:GUANGZHOU AUTOMOBILE GROUP CO LTD

Digital picture frame with life chronology storytelling

A method and system for automated routing of pictures taken on mobile electronic devices to a digital picture frame including a camera, microphone, and speaker integrated with the frame, and a network connection module allowing the frame for direct contact and upload of photos from electronic devices or from photo collections of community members. Clustering photos by content is used to improve display and to respond to photo viewer desires. Trends or patterns can be detected from the photo collections and that information is used for various purposes beyond photo display. The frame includes a conversational intelligence that provides a verbal communication with a viewer, such as for determining an identity or preferences of the frame viewer, determining photos to display for the viewer, discussing displayed photos with the viewer, or telling stories and life histories to the viewer based upon photo content.
Owner:PUSHD INC

Enabling user-centered and contextually relevant interaction

An approach is disclosed for enabling contextually relevant conversational interaction. Environment data is received by an AI System which detects a plurality of physical objects in a physical environment and forms a contextual understanding of the plurality of physical objects and the physical environment and identifies a user relevant to the contextual understanding. A most relevant contextual information to the user is predicted by the AI system and transformed into a textual form. A set of intents and objectives is predicted by the AI system for user-centered interaction. The AI system and the user interact iteratively through the user-centered interaction to determine an understanding of a most relevant intent and a most relevant objective which is validated by the AI system with the user until the user agrees. The validated most relevant intent and the most relevant objective is utilized to facilitate the user-centered and contextually relevant conversational interaction.
Owner:POLYPIE INC

Natural language prompt generation

Techniques for determining a follow-up natural language prompt, to continue a user-system dialog, are described. The system determines ASR output data representing a user input and / or a system-generated responsive thereto. The system determines one or more entities represented in the ASR output data and / or the system-generated response, and identifies one or more natural language prompts associated with the one or more entities in storage. The system filters out prompts classified as likely to result in an unsatisfactory user experience, using dialog history data including a previous user input(s) and / or a previous system-generated response(s). The system determines context(s) associated with the instant user and / or device, and uses this context(s), the ASR output data, and / or the system-generated response to determine which of the follow-up prompts is to be presented to the user.
Owner:AMAZON TECH INC

Text-to-speech synthesis using generative artificial intelligence models

A method and a system for generating human speech audio in a conversation using a trained generative AI model are provided. The method includes receiving a text input representing a portion of the conversation, receiving dialog context associated with the conversation, receiving information representing at least one voice and speaking style of at least one speaker in the conversation, generating the at least one voice and speaking style based on the received information, and generating at least one emotional audio response for the at least one speaker using the at least one voice and speaking style and without retraining the trained generative AI model.
Owner:PHEON INC

Text-to-speech processing

A speech-processing system may be configured to generate expressive synthesized speech. The system may include a prosody prediction model that generates a combination of durations and acoustic representations that may be based on the content of the text as well as additional context information. The model may be trained to predict a joint probability between linguistic representations (e.g., derived from text) and combined duration / acoustic representations (e.g., derived from audio). At inference, the model can process linguistic representations derived from text to predict combined duration / acoustic representations. In some implementations, the model may process additional information; for example, semantic embeddings output by a language model based on the text. In another example, the model may receive a speaker embedding representing voice characteristics of a particular speaker. A decoder may process the durations and acoustic representations output by the model to generate audio data representing the synthesized speech and representing expressive prosodic variation.
Owner:AMAZON TECH INC

Voice synthesis method and system based on VITS improvement

The invention provides a voice synthesis method and system based on VITS improvement, and the method comprises the steps: optimizing a text encoder of a VITS model, introducing a large language model, and enabling the emotion, intention and speaking style of an input text to be captured when the text is encoded; a random disturbance item is introduced when the Q value is dynamically planned and solved, the alignment flexibility in the initial training stage is improved, meanwhile, monotonicity constraint is strictly kept, and it is avoided that a suboptimal solution is obtained through convergence too early; a ConvNeXt module is used as a basic backbone network of a decoder, and ISTFT is utilized to efficiently reconstruct a time domain signal, so that waveform up-sampling is realized, redundant calculation of traditional transpose convolution is avoided, and reasoning is accelerated. According to the method, the reasoning speed, the emotion expression ability and the style control flexibility of speech synthesis can be effectively improved, a new solution is provided for cross-language diversified speech synthesis, and a reference is provided for the more efficient and more intelligent development of the speech synthesis technology.
Owner:豫章师范学院

Virtual Meeting Coaching

In one embodiment, a system receives a set of coaching items including a number of questions each associated with an expected answer; connects to a coaching session including one or more participants and a virtual coaching agent; for each question and for at least a subset of the participants: transmitting the question, by the virtual coaching agent, to the client device used by the participant; receiving an answer to the question by the participant, the answer including media of the participant; receiving text of utterances spoken by the participant during the answer; generating one or more evaluation scores for the answer based on evaluating at least the content of the answer to the question; and transmitting an overall evaluation score for each of the subset of participants based on the generated evaluation scores for that participant.
Owner:ZOOM COMMUNICATIONS INC