Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

4375results about "Speech synthesis" patented technology

Rich-Media Document Auxiliary Generation Apparatus

Disclosed in the present disclosure is a rich-media document auxiliary generation apparatus. The apparatus comprises a material extraction module, a theme sorting module, a semantic retrieval module, a structured data text generation module, an illustration recommendation module and a video composition module. The present disclosure uses intelligent means to assist a user to efficiently generate a high-quality rich-media composite document, thereby quickly and accurately describing a theme event in an all-round way.
Owner:10TH RES INST OF CETC

Face-translator: end-to-end system for speech-translated lip-synchronized and voice preserving video generation

A neural end-to-end system is provided for the face and voice preserving translation of videos. The system is a pipeline of multiple models that produces a video of the original speaker speaking in the target language with modified lip movement to match the target speech, while preserving emphases and prosody of the original speech, and voice characteristics of the original speaker. The pipeline starts with automatic speech recognition including emphasis detection, followed by the translation model. The translated text is then synthesized by a Text-to-Speech model that recreates the original emphases in the target sentence. The resulting synthetic speech is then converted back to the original speakers' voice using a voice conversion model. Finally, to synchronize the lips of the speaker with the translated audio, a generative model generates frames of adapted lip movements which are combined with the audio to produce the final output. The disclosure further describes several use-cases and configurations that apply these techniques to video conferencing, dubbing, low-bandwidth transmission, speech enhancement and assistive technology for the hearing impaired.
Owner:WAIBEL ALEXANDER

Comprehensive AI-enabled systems for immersive voice, companion, and augmented / virtual reality interaction solutions

A computer-implemented method for operating an artificial intelligence voice agent system includes receiving voice input through communication channels; analyzing converted text through natural language processing (NLP) pipelines implementing intent recognition and sentiment analysis detecting emotional cues using a multimodal large language model (LLM); generating response content using machine learning models trained on domain-specific corpora; converting generated responses to synthetic speech through text-to-speech (TTS) engines; integrating with a customer relationship management (CRM) platforms or an enterprise resource planning (ERP) database; and implementing continuous learning by updating language understanding models using conversation logs, voice recognition parameters based on user feedback, and response generation patterns. One implementation is a computer-implemented system and method that operates a suite of intelligent interactive devices and platforms including an artificial intelligence voice agent, enhanced communication platforms, an intimacy companion system, and augmented / virtual reality eyeglasses. Further, one implementation includes AR / VR eyeglasses that project visual content onto interchangeable lenses or directly onto the user's retina via laser-based retinal projection, provide prescription adjustments, incorporate ear-mounted sensors for monitoring physiological parameters like heart rate, oxygen saturation, and blood pressure, and utilize wireless data transmission, onboard environmental sensing, and remote calibration, all designed to offer dynamically adaptive, secure, and context-aware interactions across communication, personal assistance, health monitoring, and immersive augmented or virtual reality environments.
Owner:TRAN BAO

Personalized and dynamic text to speech voice cloning using incompletely trained text to speech models

Systems and methods are provided for machine learning models configured as zero-shot personalized text-to-speech models which comprise a feature extractor, a speaker encoder, and a text-to-speech module. The feature extractor is configured to extract acoustic features and prosodic features from new target reference speech associated with the new target speaker. The speaker encoder is configured to generate a speaker embedding corresponding to the new target speaker based on the acoustic features extracted from the new target reference speech. The text-to-speech module is configured to generate the personalized voice corresponding for the new target speaker based on the speaker embedding and the prosodic features extracted from the new target reference speech without applying the text-to-speech module on new labeled training data associated with the new target speaker.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Precise international communication digital human real-time dialogue method fused with multi-modal technology

The invention discloses an accurate international communication digital human real-time dialogue method fused with a multi-modal technology. The method comprises the following steps: S1, constructing a digital human image and tone; s2, propagation content generation and problem guidance; s3, semantic analysis and intention clarification based on the real-time voice dialogue; s4, geographic preference modeling and path planning; s5, cross-context propagation content generated based on retrieval enhancement is generated; s6, visually displaying the propagation content; and S7, carrying out digital human-driven multi-language propagation content real-time output and feedback closed-loop optimization. Through accurate utterance expression analysis, accurate international propagation problem recommendation is provided, cross-context propagation content generation based on semantic understanding is realized, digital people with voice features and visual images are constructed, real-time dialogue interaction of users is realized, and user experience is improved. The system can carry out geographic modeling according to the region where the accurate problem is located, language preference and propagation object culture characteristics, and differential propagation path planning is achieved.
Owner:HUNAN NORMAL UNIVERSITY

Training and speech generation methods and apparatuses for speech generation model, electronic device, computer-readable storage medium, and computer program product

The present application provides training and speech generation methods and apparatuses for a speech generation model, an electronic device, a computer-readable storage medium, and a computer program product. The method comprises: obtaining a first speech generation model; obtaining sample data of a plurality of modalities; on the basis of a prompt image sequence and speech text, respectively calling a plurality of encoders to perform encoding, so as to obtain a multi-modal encoding vector sequence; on the basis of the multi-modal encoding vector sequence, calling a decoder to perform decoding, so as to obtain decoded text; determining a probability distribution for the decoded text and the sample data of the plurality of modalities, and determining a target loss on the basis of the probability distribution; and on the basis of the target loss, updating parameters of the decoder and at least one of the encoders, wherein the updated decoder and the plurality of updated encoders are configured to form a second speech generation model, and the second speech generation model is used to generate target speech text corresponding to a prompt image sequence of a target object.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Multiple results presentation

In some embodiments, a natural language understanding (NLU) hypothesis may be determined for a natural language input to a device and a first component may be used to obtain first visual content corresponding to a first skill associated with the first NLU hypothesis. A second component may also be used to obtain second visual content corresponding to a second skill. The device may be caused to output a first graphical user interface (GUI) element including the first visual content and a second GUI element including the second visual content. In response to a subsequent input to the device corresponding to the second GUI element, the device may be caused to output content corresponding to the second skill.
Owner:AMAZON TECH INC

System and method for emotionally intelligent, personalized AI avatar-based health coaching using multi-domain data and adaptive behavioral intelligence

A programmatically generated AI avatar includes a customizable personality module, acting as the embodied interface for a powerful AI “mind” that delivers personalized coaching to improve user health, well-being, and longevity. The system uses machine learning, large language models, and biometric modeling to synthesize real-time, multi-modal health data—including sleep, nutrition, glucose, mood, and activity—and generate forward-prescribed KHAs. Unlike human coaches, it continuously adapts based on context and behavior, targeting the root cause: metabolic dysfunction—namely by restoring healthy, sustainable body composition through the preservation or building of lean muscle mass and reduction of excess fat. KHAs can also be shared with friends or programmatically generated AI avatars, allowing for coordinated action, emotional support, and accountability through social connection—further reinforcing positive behavior and adherence. The system's reinforcement learning engine incorporates both individual response data and anonymized population-level insights to optimize recommendations over time, learning which interventions are most effective for users with similar physiological and behavioral profiles. First validated with Olympic athletes—resulting in measurable improvements and medal-winning outcomes—this system offers a scalable, emotionally intelligent coaching engine that exceeds human capability, designed for the ultimate purpose of supporting sustainable health, resilience, and human thriving.
Owner:GOLD AI LLC

Artificial intelligence speech recognition system

The invention discloses an artificial intelligence speech recognition system, and the system comprises a multi-modal feature extraction module which employs an improved Conformer architecture to synchronously extract the time-frequency features and text embedding vectors of speech signals; the joint training module is used for performing joint optimization on ASR and NMT loss functions through an adversarial training strategy, learning voice recognition and machine translation tasks at the same time through joint training, and completing direct mapping from voice features to a target language; the context perception translation engine is used for integrating an attention mechanism of a pre-training language model, carrying out deep coding on the extracted speech features and generating cross-language semantic representation; the self-adaptive post-processing module is used for dynamically optimizing an output result by adopting a reinforcement learning framework, dynamically adjusting the output result according to a reward function, and optimizing translation quality and a speech synthesis effect; the dynamic language recognition module is a real-time language classifier based on a Wave2Vec 2.0 framework and is used for recognizing the language of the input voice in real time; and the incremental field adaptation module is used for quickly updating a field term library by using a LoRA fine tuning technology.
Owner:ANKANG UNIV

Automatic lecturer video generation method based on AI speech synthesis and animation driving

The invention discloses a lecturer video automatic generation method based on AI speech synthesis and animation driving. The method comprises the following steps: performing structured analysis on a PPT or a text script through an improved interior point method and an incremental shortest path algorithm; performing semantic grouping by applying a full-dynamic parallel single-link clustering algorithm and generating an enhanced script with an expressive mark; a CosyVoice technology is combined with a low-rank approximation method to generate a high-quality voice data stream; establishing a mapping relation between contents and action expressions through semantic analysis, and generating a complete action expression instruction set; and driving the digital human model by using the msueTalk technology, and generating a final lecturer teaching video through a parallel rendering algorithm. According to the invention, the method achieves the efficient and automatic generation of the education video, remarkably improves the content production efficiency, reduces the production cost, and guarantees the specialty and expressive force of the teaching video.
Owner:SHENZHEN XUEYOU TECHNOLOGY CO LTD

Intelligent oral English training system based on Prompt engineering

The invention discloses an intelligent oral English training system and method based on Prompt engineering, and relates to the technical field of oral English training. According to the method, a user is allowed to freely describe training requirements by using a natural language, the system automatically constructs a training Prompt, and personalized customization is realized; content deviating from a training target is accurately recognized through a deviation question detection module, a user is guided to return to a theme through real-time voice or text, and effectiveness and pertinence of training are guaranteed; a clear template definition ensures systematicness and reproducibility of a training task, and clear control and evaluation of a teaching target are facilitated; in combination with voice recognition, voice synthesis and text generation, natural interaction in a real context is realized, and the training experience is improved; on the basis of multi-dimensional automatic evaluation of a keyword hit rate, pronunciation quality and the like, personalized learning suggestions are supported; teachers are supported to generate training tasks in batches, students are supported to customize scenes, and different education environment requirements are met.
Owner:CHUXIONG NORMAL UNIV

Bluetooth communication intelligent speech translation method and system based on multi-mode enhancement

The invention relates to the technical field of artificial intelligence, and discloses a Bluetooth communication intelligent speech translation method and system based on multi-mode enhancement, and the method comprises the steps: collecting a multi-channel audio signal through a built-in multi-microphone array of a Bluetooth device, carrying out the dynamic direction self-adaptive beam forming of the multi-channel audio signal, and carrying out the self-adaptive beam forming of the multi-channel audio signal; extracting a Mel spectrogram feature of the direction enhancement signal, identifying lip regions of a plurality of candidate speakers in each frame of real-time speaking video captured by a camera, performing time sequence convolution on the lip regions to obtain a lip movement time sequence embedded vector, calculating a correlation score with the Mel spectrogram feature, separating the direction enhancement signal, and obtaining a lip movement time sequence embedded vector; and performing text transcription and conversion on the high-confidence separation voice to obtain a translation language text, and sending the synthesized target translation voice to a preset mobile terminal through the Bluetooth device to obtain a target translation result. According to the method, the real-time performance and accuracy of speech translation are improved in a multi-person scene, far-field speech, noise interference and accent difference.
Owner:SHENZHEN DIE MICRO SEMICON CO LTD

System and method for multi-modal ai conversational interface improving website navigation and user interaction

The present invention relates to a system for transforming static websites into artificial intelligence (AI)-enabled interactive multi-modal conversational platforms. The system comprises a computing device having a processor for receiving user queries as text or speech input through an input module cooperating with a speech-to-text module. A natural language processing (NLP) module interprets intent, classifies user context, and retrieves grounded information from multiple webpages. A persona adaptation module dynamically modifies vocabulary, tone, and avatar representation across roles such as sales assistant, recruiter, educator, healthcare professional, etc. A response generator module produces structured natural language output, transmitted to a text-to-speech synthesis module and an avatar generation module to render synchronized lifelike video responses. An output rendering module displays multi-modal responses include text, audio, and video, thereby enabling direct navigation and escalation beyond limitations of conventional static websites.
Owner:NALLAM SREE RAMA CHANDRA MURTY

Large model cross-modal collaborative understanding method and device

The invention provides a large-model cross-modal collaborative understanding method and device. The fusion efficiency and the understanding capability of multi-modal information can be improved. The large-model cross-modal collaborative understanding method comprises the following steps: preprocessing visual data, language data and sound data to obtain preprocessed visual data, language data and sound data; wherein the preprocessing comprises the steps of performing adaptive size adjustment and normalization on visual data, performing word segmentation and dynamic truncation on language data, and performing band-pass filtering and spectral noise reduction on sound data; extracting visual features, language features and sound features based on the preprocessed visual data, language data and sound data; on the basis of an adaptive mapping network and a mixed granularity cross-modal attention mechanism, performing feature alignment on the visual features, the language features and the sound features to obtain aligned feature vectors; and based on a dynamic routing architecture, fusing the aligned feature vectors to generate a unified multi-modal representation.
Owner:SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD

Computer implemented system and method for automatically generating offer ranges for candidates in an interviewing process

A computer implemented system and method for generating offer ranges for candidates in an interviewing process is disclosed. The system generates an AI-based interviewer simulating human-based interactions for conducting an ongoing interview with candidates. The system analyzes data associated with candidates obtained during ongoing interview. The system process analyzed responses of candidates to determine contextual attributes associated with responses using ML models. The system automatically generates follow-up interview questions to be delivered to candidates during ongoing interview based on analyzed responses from candidates, by applying AI model to contextual attributes associated with responses. The system generates recruitment scores for candidates based on analyzed responses, contextual attributes, and interpreted non-verbal cues, associated with candidates, using AI model. The system generates offer ranges for candidates based on recruitment scores using AI model. The system provides information associated with selected candidates, and offer ranges generated for selected candidates, to users.
Owner:TALVIEW INC

Video translation method and system based on artificial intelligence

The invention discloses a video translation method and system based on artificial intelligence. The method relates to the technical field of video translation and comprises the following steps of original sound track extraction, target AI speaker adaptation, AI dubbing generation and mouth shape synchronization and video synthesis. According to the method, independent audio and video streams are obtained by adopting an audio and video separation technology, and multiple original sound tracks are extracted through a voice separation model; matching or generating an adaptive target AI speaker module in a preset tone library; converting the original language voice into a text, translating the text into a target language text, and synthesizing an AI dubbing audio track in combination with a target AI speaker module; and finally, the independent video stream and the multi-AI dubbing audio track are input into the mouth shape synchronization model to output a translated video, so that the timbre fitting degree, the voice quality and the voice consistency of the same speaker of AI dubbing are improved, and meanwhile, the resource utilization rate of video translation and the processing efficiency under batch tasks are improved. The problem that in the prior art, video translation is low in quality and efficiency is solved.
Owner:BEIJING DEEP LOGIC INTELLIGENT TECHNOLOGY CO LTD

Fine emotion control TTS method and device based on large model, equipment and medium

The invention discloses a fine emotion control TTS method and device based on a large model, equipment and a medium, relates to the technical field of artificial intelligence, and can generate a voice service with fine emotion expression and improve the user experience in the high-sensitivity emotion interaction fields of banks, financial customer service, insurance consultation, medical hospital guide and the like. Acquiring an emotional feature vector of the input text based on a preset large language model; encoding the emotion feature vector to obtain a voice encoding vector; determining an emotion curve vector of the input text based on a pre-trained neural network model; and generating target voice corresponding to the input text based on the voice coding vector and the emotion curve vector. According to the method, emotion feature vectors are extracted through a large language model, a dynamic emotion curve is generated in combination with a neural network, speech synthesis is cooperatively driven, emotion fineness and continuity are remarkably improved, and the problems of traditional TTS emotion expression roughening and fragmentation are solved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Method and system for converting or encoding text

A publishing system with components, including:a system configured to receive at least one document including text that defines a base alphabet in one or more formats;a system configured to provide additional data for a reader to better understand the document which includes:a method of encoding or marking up non-phonetic words in the document to enable the reader to decode sounds of each non-phonetic word; anda system configured to output an encoded document with the text and the additional data in one or more formats,wherein the method of automatically encoding the non-phonetic words to make the encoded words phonetic:for at least one character (“spelling character”) in the non-phonetic word, using a compound character that includes the spelling character and a sound character, wherein the sound characters:are human-readable characters in the base alphabet and / or in one or more secondary alphabets,are added to the spelling characters to indicate that each spelling character makes the usual sound of the sound character,are added so that spelling characters can be visually discriminated from sound characters,are added such that a reader can recognize the non-phonetic word by sight because the spelling of the word is unchanged, andare added to the spelling characters such that the spelling characters and the sound characters remain human-readable such that the spelling character and the sound character of each compound character are within one visual field; andautomatically outputting the encoded words in a human-readable form / format such that the compound characters in the encoded word visually indicate which of the spelling characters have a sound other than their usual sound and what sound each character makes in the non-phonetic word when it does not make its usual sound.
Owner:STEPHEN CHRISTOPHER COLIN

Sound duplicating and low-delay streaming speech synthesis method and system based on ultra-short sample

The invention provides a voice replication and low-delay streaming voice synthesis method and system based on an ultra-short sample, relates to the technical field of artificial intelligence, and is suitable for intelligent interaction, outbound service and multi-mode communication scenes. Deep personalized customization of intelligent voice interaction and real-time generation of ultra-low delay are realized, and bidirectional streaming interaction of a system level is supported, so that the fluency and response speed of dialogues are improved; an ultra-short sample sound duplicating module specially designed for processing ultra-short audio samples and a speech synthesis engine with ultra-low delay and bidirectional streaming output capability are integrated, and the capabilities of the ultra-short sample sound duplicating module and the speech synthesis engine are applied to a real-time and interactive intelligent speech interaction process to form a complete and efficient solution. The industrial pain point is directly solved in a targeted manner, and the method has important commercial application value.
Owner:GUANGDONG CHAOTENG INFORMATION TECHNOLOGY CO LTD

Voice interaction processing method, device and system, intelligent door lock and cloud server

The invention is suitable for the technical field of intelligent door locks, and provides a voice interaction processing method, device and system, an intelligent door lock and a cloud server, and the method comprises the steps: receiving a voice instruction sent by a user, and carrying out the local recognition, so as to obtain a local recognition result and the confidence of the local recognition result; when the local recognition result is matched with the local instruction set and the confidence is high, executing a local operation corresponding to the local recognition result; otherwise, establishing a secure communication session with the cloud server based on voiceprint biological characteristics extracted from the voice instruction; sending the voice instruction to a cloud server through the secure communication session, and receiving a cloud operation instruction and / or a dynamic response text returned by the cloud server; and executing corresponding operation according to the cloud operation instruction, and / or converting the dynamic response text into voice for broadcasting through a text-to-voice engine. The intelligent door lock voice interaction system solves the problem that the real-time performance, the reliability, the safety and the interaction intellectualization are difficult to balance in the existing intelligent door lock voice interaction technology.
Owner:SHANGHAI ZHENGZHI INTELLIGENT TECH CO LTD

Business processing method and device based on multi-agent cooperation, equipment and medium

The invention provides a business processing method and device based on multi-agent collaboration, equipment, a medium and a program product, which can be applied to the technical field of digital human and artificial intelligence. The method comprises the following steps: in response to a service request initiated by a user through a first service channel, obtaining input information of the user, and creating or updating a session context object corresponding to the service request; analyzing the input information, generating a subtask sequence, and writing the subtask sequence into a session context object; according to the subtask sequence, a corresponding target agent is scheduled to execute operation, and an execution result is written back to the session context object; monitoring updating of the session context object, and generating a migration decision for migrating from the first service channel to the second service channel under the condition that the updated session context object meets a cross-channel migration condition; and based on the migration decision, synchronizing the session context object to the second service channel in a full amount, so that the business of the user is handled in the second service channel.
Owner:ANHUI BRANCH OF INDAL & COMML BANK OFCHINA

Method, system and device for voice interaction inside and outside vehicle and storage medium

The invention discloses a method, a system and equipment for voice interaction inside and outside a vehicle and a storage medium, and relates to the technical field of intelligent cabins and human-computer interaction. Identifying an interaction trigger type (in-vehicle active request, out-of-vehicle passive request, or system active trigger) based on the perceived data and / or the vehicle event; determining a corresponding response permission strategy according to the trigger type, and determining whether to allow the system to respond based on the response permission strategy; and if so, calling a large language model to generate a voice text, synthesizing voice by combining external perception characteristics, and broadcasting the voice through a loudspeaker outside the vehicle or a sound box in the vehicle. According to the scheme, three types of interaction intentions including the in-vehicle active request, the out-vehicle passive request or the system active triggering are recognized, the large language model is called to generate the scene-adaptive voice text, and personalized voice synthesis and broadcasting are performed in combination with the external perception characteristics, so that the response efficiency, the expression naturalness and the object adaptability of the in-vehicle and out-vehicle voice interaction are improved.
Owner:ZHEJIANG GEELY HLDG GRP CO LTD +1

Computing system, method, and medium for processing customer inquiries using speech-to-text, language model analysis, and text-to-speech services

An autonomous communication system includes a language model, a retrieval module, a caching mechanism, an autonomous agent, and a human interface for managing interactions and responses. A method and computer-readable medium for managing communication also include these components for processing, enhancing, summarizing, managing interactions, and allowing human intervention.
Owner:CDW LLC

Method of performing task based on large model and electronic device

A method of performing a task based on a large model and an electronic device are provided, which relate to artificial intelligence technology, and in particular to fields of voice interaction, deep learning, large model, etc. The method includes: acquiring a demand feature characterizing a demand intention; performing a task by using the large model according to the demand feature, to obtain a response text, in which a target response word is determined based on: determining a query feature for each attention subtask in the task based on an associated response word feature; and performing, based on the demand feature read from a storage unit as a value feature and a key feature shared by the plurality of attention subtasks, the plurality of attention subtasks by using a computing unit according to a plurality of query.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Server, display device and digital human processing method

The embodiment of the invention provides a server, display equipment and a digital human processing method. The method comprises the following steps: receiving voice data input by a user and sent by the display equipment; broadcast voice is determined based on the voice data; extracting voice features of the broadcast voice; determining mouth shape parameters based on the voice features; determining emotion parameters and acquiring user image data; generating digital human image data based on the user image data, the emotion parameters and the mouth shape parameters; and sending the broadcast voice and the digital human image data to the display device, so that the display device plays the broadcast voice and displays a digital human image based on the digital human image data. According to the embodiment of the invention, the expression parameters and the mouth shape parameters are determined according to the voice data input by the user, the expression parameters and the mouth shape parameters are combined to generate the digital human image with better facial expression expression, and emotion customization and control are realized.
Owner:HISENSE VISUAL TECH CO LTD

Voice decoding method, system and equipment based on electroencephalogram signals and medium

The invention discloses a voice decoding method, system and device based on electroencephalogram signals and a medium, and relates to the technical field of electroencephalogram signal processing.The method comprises the steps that reading electroencephalogram signals and voice signals in the reading process of a to-be-tested person and imagination electroencephalogram signals in the imagination reading process of the to-be-tested person are collected; inputting the reading electroencephalogram signals and the voice signals into the DRCL to generate electroencephalogram characteristics containing voice information; training the DBM by using the electroencephalogram characteristics as input and using the voice signals as output, and adjusting the DBM by using imaginary electroencephalogram signals to construct a voice synthesizer; a mapping relation between the imaginary electroencephalogram signals and the voice signals is generated through the voice synthesizer, and decoding from the electroencephalogram signals to the voice signals is completed; according to the method, a deep representation correlation learning method is provided, potential correlation between electroencephalogram and voice signals can be deeply mined through a multi-layer network structure, a complex mode which is difficult to recognize by a traditional model is captured, and the voice decoding process is more accurate.
Owner:HARBIN INST OF TECH

Multi-modal fusion driven emotion perception enhanced TTS speech synthesis method

The invention provides a TTS speech synthesis method for emotion perception enhancement under multi-modal fusion driving, and the method comprises the following steps: S1, carrying out the collection and preprocessing of multi-modal data, the multi-modal data comprising text data, speech data and facial expression data; s2, extracting and analyzing emotion features; s3, emotion perception speech synthesis model training; s4, speech synthesis and post-processing; s5, performing model evaluation and optimization; by collecting and analyzing multi-mode data such as texts, voices and facial expressions, emotion features can be captured more comprehensively and accurately, complementary information among different modes is fully mined through application of a multi-mode fusion network and a collaborative attention mechanism, the emotion expression of the synthesized voices is closer to real emotions, and the user experience is improved. And the precision of emotion perception is greatly improved.
Owner:TIANJIN CHENGJIAN UNIV

Intelligent accompanying robot system

The invention discloses an intelligent accompanying robot system, and the system comprises the following modules: a multi-mode interaction module which is composed of a voice recognition unit, a voice synthesis unit, an emotion recognition unit, and a visual recognition unit, and is used for collecting the voice, facial expression, motion, and environment information of a user; the localized AI decision module is used for deploying a lightweight DeepSeekR1 large language model based on an ESP32-S3 edge computing chip, performing real-time processing on the multi-modal data and generating a social guidance strategy and an emotion intervention instruction; the data management module comprises a user behavior database and a privacy protection unit and is used for desensitizing the data and realizing sensitive data isolation through local storage; according to the method, the single-person intervention cost is reduced, and the time consumption of manual scene simulation is reduced.
Owner:NANTONG UNIV

Rhythm migration method and device, electronic equipment and storage medium

The invention relates to the technical field of voice processing, and provides a rhythm migration method and device, electronic equipment and a storage medium, and the method comprises the steps: obtaining a decoupled rhythm feature based on a source rhythm voice, and a decoupled tone feature based on the voice of a target speaker, the decoupled rhythm feature represents the rhythm of the source rhythm voice, and the decoupled tone feature represents the tone of the target speaker; the decoupled timbre features represent the timbre of the voice of the target speaker; generating a target voice vector sequence based on the text features of the target text, the voice features of the voice of the target speaker, the decoupled rhythm features and the decoupled timbre features; and synthesizing a target audio based on the target voice vector sequence. According to the method and the device, the decoupled rhythm features and the decoupled timbre features are acquired, and the target voice is generated based on the features, so that the problem of feature mixing is effectively relieved, the timbre purity of the target speaker in cross-person rhythm migration is ensured, the expressive force of rhythm migration is improved, and the synthesized audio is more natural and vivid.
Owner:IFLYTEK CO LTD

Determining device context

A system may be configured to receive and process various signals to generate a natural language description of a user's environment, called situational context data. The signals may include sensor data, device status, user activity, user input, and / or inferences made using such data. The situational context data may express a user-centric description of the user's environment; for example: “User is taking a walk in the park on a sunny afternoon” or “activity: driving location: highway”, etc. The system may send the situational context data to various system components that may, for example, process speech, select applications / skills for handling user inputs, and / or that implement those applications / skills. The applications / skills may use the situational context data to provide recommendations, generate responses, and / or perform actions that are more relevant to the user's current environment.
Owner:AMAZON TECH INC