Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

1346 results about "Speech synthesis" patented technology

Speech synthesis is the artificial production of human speech. A computer system used for this purpose is called a speech computer or speech synthesizer, and can be implemented in software or hardware products. A text-to-speech (TTS) system converts normal language text into speech; other systems render symbolic linguistic representations like phonetic transcriptions into speech.

Multi-mode-based AI digital human intelligent interaction method, system and equipment

The invention relates to the technical field of computer vision and human-computer interaction, and discloses an AI digital human intelligent interaction method, system and equipment based on multiple modalities, and the method comprises the steps: pre-awakening a digital human when a human face is detected, and further thoroughly awakening the digital human based on recognized preset voice information or preset gesture information; voice and video information of a user in the interaction process is obtained, a keyword extraction result, a gesture recognition result and an emotional state tag are generated, a pre-constructed knowledge base is utilized to retrieve related information, a big language generation model module is combined to generate an answer text, and the answer text is input into a preset voice synthesis model to generate emotional voice output. And based on the current emotional state label of the user, driving the digital human animation to be output in an emotional manner. According to the method and the system, the digital human for understanding the emotion of the user, generating personalized answers, providing voices with rich emotions and displaying natural expressions and actions can be created, better interaction with the user can be realized, and more humanized and effective services can be provided.
Owner:BEI JING WAN JIE SHU JU KE JI YOU XIAN ZE REN GONG SI WU HAN FEN GONG SI +1

Vehicle state monitoring and early warning method and system

The invention belongs to the technical field of vehicle state detection and early warning, and particularly relates to a vehicle state monitoring and early warning method and system.The monitoring and early warning method comprises the steps that a sensor obtains real-time data, an anomaly detection algorithm is applied, and an abnormal event is marked; calculating an information priority according to the abnormal severity and the driving scene, and distributing the information priority to a high-priority queue; multi-mode early warning is generated, and high-frequency sound and vibration are used during high-speed driving; extracting a voice prompt, and generating voice waveform data through a voice synthesis module; voice input of a driver is recognized, intention is analyzed, and abnormal information feedback is provided; according to the scene and the abnormal state, interactive output is optimized, and detailed information is displayed during low-speed congestion; dynamically adjusting the interface of the instrument panel, and amplifying the key area in case of abnormal severity; and integrating the image data and the voice data, generating a multi-mode signal, and transmitting the multi-mode signal after rendering processing. The vehicle abnormal information can be timely and accurately transmitted to a driver, and the driving safety is effectively improved.
Owner:XIAN HUODA NETWORK TECH CO LTD

Intelligent dialogue method and device combining voiceprint recognition and voice synthesis

The invention relates to the technical field of voiceprint recognition, and particularly provides a voiceprint recognition and voice synthesis combined intelligent dialogue method, which comprises the following steps: acquiring a user voice input signal in real time; voiceprint feature extraction is carried out on the voice input signal, and a composite voiceprint vector containing user identity features and rhythm features is obtained; user identity matching is carried out based on the composite voiceprint vector, and a dynamic user portrait database is associated and called; generating personalized semantic feedback content according to the weight-updated dynamic parameters in the dynamic user portrait database; generating a target timbre synthesis parameter set based on the timbre representation vector in the composite voiceprint vector in combination with the spatio-temporal characteristic parameters of the interaction scene; and performing voice synthesis on the personalized semantic feedback content by adopting the target timbre synthesis parameter set, adjusting rhythm expression parameters of the synthesized voice based on real-time rhythm characteristics in the composite voiceprint vector, and finally outputting personalized voice, thereby realizing natural interaction of one person, one voice and one strategy.
Owner:MINAMI ACOUSTICS LTD

Video and voice automatic translation method based on pre-training model

The invention belongs to the technical field of speech translation, and particularly relates to a video speech automatic translation method based on a pre-training model, and the method comprises the following steps: 1, preprocessing video and audio data; step 2, voice recognition and language detection; step 3, machine translation and text post-processing; step 4, speech synthesis and audio mixing; step 5, synchronizing video processing and subtitles; step 6, quality control and multi-dimensional evaluation; 7, carrying out model iteration and data closed loop; and step 8, system deployment and engineering implementation. Through deep fusion of efficient transfer learning of the pre-training model and the multi-modal technology, a high-precision, low-cost and easy-to-expand video speech translation solution is constructed, the time and labor cost of globalized content production is greatly reduced, the cross-language communication efficiency is improved, immersive multi-language experience is provided, and the method is suitable for popularization and application. And a data-driven continuous optimization mechanism is established, so that the system performance is improved along with the increase of the use scale.
Owner:ZHE JIANG YAN HUANG KE JI YOU XIAN GONG SI

Voice conversation method and system, electronic equipment, storage medium and program product

The embodiment of the invention provides a voice conversation method and system, electronic equipment, a storage medium and a program product. One method is implemented as follows: a voice dialogue stream (which can be an inquiry voice stream input by a user for a target commodity) of a user is subjected to streaming response, a plurality of dialogue text segments are generated in a segmented manner in the streaming response process, and the dialogue text segments are output in real time after being generated and stored in a first cache queue; and when it is detected that the reply text fragment is stored in the first cache queue, the stored reply text fragment is converted into a reply voice fragment in real time and played. The streaming response may be performed by an agent. Visibly, according to the response of the voice dialogue stream of the user, the dialogue is played while being generated, the waiting time from text to voice synthesis can be greatly shortened, and therefore the response time can be shortened.
Owner:ZHEJIANG TMALL TECH CO LTD

Artificial intelligence speech recognition system

The invention discloses an artificial intelligence speech recognition system, and the system comprises a multi-modal feature extraction module which employs an improved Conformer architecture to synchronously extract the time-frequency features and text embedding vectors of speech signals; the joint training module is used for performing joint optimization on ASR and NMT loss functions through an adversarial training strategy, learning voice recognition and machine translation tasks at the same time through joint training, and completing direct mapping from voice features to a target language; the context perception translation engine is used for integrating an attention mechanism of a pre-training language model, carrying out deep coding on the extracted speech features and generating cross-language semantic representation; the self-adaptive post-processing module is used for dynamically optimizing an output result by adopting a reinforcement learning framework, dynamically adjusting the output result according to a reward function, and optimizing translation quality and a speech synthesis effect; the dynamic language recognition module is a real-time language classifier based on a Wave2Vec 2.0 framework and is used for recognizing the language of the input voice in real time; and the incremental field adaptation module is used for quickly updating a field term library by using a LoRA fine tuning technology.
Owner:ANKANG UNIV

Automatic lecturer video generation method based on AI speech synthesis and animation driving

The invention discloses a lecturer video automatic generation method based on AI speech synthesis and animation driving. The method comprises the following steps: performing structured analysis on a PPT or a text script through an improved interior point method and an incremental shortest path algorithm; performing semantic grouping by applying a full-dynamic parallel single-link clustering algorithm and generating an enhanced script with an expressive mark; a CosyVoice technology is combined with a low-rank approximation method to generate a high-quality voice data stream; establishing a mapping relation between contents and action expressions through semantic analysis, and generating a complete action expression instruction set; and driving the digital human model by using the msueTalk technology, and generating a final lecturer teaching video through a parallel rendering algorithm. According to the invention, the method achieves the efficient and automatic generation of the education video, remarkably improves the content production efficiency, reduces the production cost, and guarantees the specialty and expressive force of the teaching video.
Owner:SHENZHEN XUEYOU TECHNOLOGY CO LTD

Immersive concert experience system based on virtual reality

The invention provides an immersive concert experience system based on virtual reality, and belongs to the technical field of virtual reality, and the system comprises a dynamic adjustment module which is used for predicting the preference of a user for stage arrangement, light effects and music types through a machine learning model according to the historical preference data of the user, and outputting the preference to the stage arrangement, light effects and music types; dynamically adjusting visual elements and sound field parameters of the virtual scene and a refresh rate and a field angle of the virtual scene in combination with the user physiological indexes and the user head motion trail which are acquired in real time; the action capture and feedback module is used for capturing an interaction instruction sent to the dynamically adjusted virtual scene by the limb action of the user in real time; and the interaction feedback module is used for enabling the virtual performer to respond to the interaction instruction of the user in real time based on the expression recognition and voice synthesis technology, and generating personalized interaction feedback content. A user constructs an immersive music interaction field domain, and experience logic of a virtual music meeting is completely reconstructed.
Owner:SHIJIAZHUANG UNIVERSITY

Intelligent oral English training system based on Prompt engineering

The invention discloses an intelligent oral English training system and method based on Prompt engineering, and relates to the technical field of oral English training. According to the method, a user is allowed to freely describe training requirements by using a natural language, the system automatically constructs a training Prompt, and personalized customization is realized; content deviating from a training target is accurately recognized through a deviation question detection module, a user is guided to return to a theme through real-time voice or text, and effectiveness and pertinence of training are guaranteed; a clear template definition ensures systematicness and reproducibility of a training task, and clear control and evaluation of a teaching target are facilitated; in combination with voice recognition, voice synthesis and text generation, natural interaction in a real context is realized, and the training experience is improved; on the basis of multi-dimensional automatic evaluation of a keyword hit rate, pronunciation quality and the like, personalized learning suggestions are supported; teachers are supported to generate training tasks in batches, students are supported to customize scenes, and different education environment requirements are met.
Owner:CHUXIONG NORMAL UNIV

System and method for multi-modal ai conversational interface improving website navigation and user interaction

The present invention relates to a system for transforming static websites into artificial intelligence (AI)-enabled interactive multi-modal conversational platforms. The system comprises a computing device having a processor for receiving user queries as text or speech input through an input module cooperating with a speech-to-text module. A natural language processing (NLP) module interprets intent, classifies user context, and retrieves grounded information from multiple webpages. A persona adaptation module dynamically modifies vocabulary, tone, and avatar representation across roles such as sales assistant, recruiter, educator, healthcare professional, etc. A response generator module produces structured natural language output, transmitted to a text-to-speech synthesis module and an avatar generation module to render synchronized lifelike video responses. An output rendering module displays multi-modal responses include text, audio, and video, thereby enabling direct navigation and escalation beyond limitations of conventional static websites.
Owner:NALLAM SREE RAMA CHANDRA MURTY

Voice text bidirectional conversion method and device, equipment and medium

The invention relates to the technical field of voice processing, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a voice text bidirectional conversion method, device, equipment and medium, and the method comprises the steps: respectively executing voice recognition or voice synthesis operation according to the type of input information; for the voice information, noise suppression parameters are generated in combination with the lip movement video data, noise reduction processing is executed, and the recognition accuracy is improved; for text information, a pre-generated speaker style vector is obtained, the vector is cited in the speech synthesis process to generate natural personalized speech, and lip movement information and tactile feedback which are synchronous with speech output are generated. According to the method, complex noise is suppressed by fusing lip movement data, personalized voice is generated by using the style vector, and lip movement and touch information is output, so that bidirectional real-time conversion of voice and text in a complex environment is realized, and recognition accuracy, voice naturalness and interaction synchronism are effectively improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Lightweight intelligent traditional Chinese medicine inquiry system and construction method thereof

The invention relates to the field of artificial intelligence medical application, and discloses a lightweight intelligent traditional Chinese medicine inquiry system and a construction method thereof, and the system comprises a multi-dialect adaptive speech recognition module, a traditional Chinese medicine intelligent dialogue large language model module, a natural speech synthesis module, and a continuous learning mechanism module. The multi-dialect adaptive speech recognition module is used for converting dialect speech input of a patient into a standard text; the traditional Chinese medicine intelligent dialogue big language model module is the core of the system and is used for carrying out natural language understanding, dialectical reasoning and inquiry dialogue generation, and the natural speech synthesis module is used for converting a text response generated by the system into speech output; and the continuous learning mechanism module realizes continuous optimization of the large language model through incremental learning architecture and clinical feedback integration. According to the method, while the professional traditional Chinese medicine diagnosis capability is maintained, the calculation complexity is remarkably reduced, and the universality and sustainable development capability of system application are improved.
Owner:SUZHOU ANGSHENG NETWORK TECHNOLOGY CO LTD

Real-time interaction 3D digital holographic cabin method based on deep learning and sound cloning

The invention discloses a real-time interaction 3D digital holographic cabin method based on deep learning and sound cloning. The method comprises the following steps that S1, user data are collected and preprocessed; s2, extracting feature vectors of facial expressions and limb actions; s3, generating speech synthesis data by using the improved GE2E network and a preset target speech text; s4, generating a synthesized voice audio based on the voice synthesis data; s5, generating a three-dimensional digital human motion sequence according to the facial expression feature vector and the body motion feature vector; s6, performing timestamp alignment on the three-dimensional digital human action sequence and the synthesized voice audio, and constructing a synchronous output stream; and S7, rendering the synchronous output stream, and performing three-dimensional visual output. According to the method, the improved GE2E network, deep learning and sound cloning methods are fused, three-dimensional virtual human voice action synchronous control is achieved, and the method has the advantages of being high in real-time performance, high in immersion and natural in interaction.
Owner:HANGZHOU SECOND LIFE TECH CO LTD

Fine emotion control TTS method and device based on large model, equipment and medium

The invention discloses a fine emotion control TTS method and device based on a large model, equipment and a medium, relates to the technical field of artificial intelligence, and can generate a voice service with fine emotion expression and improve the user experience in the high-sensitivity emotion interaction fields of banks, financial customer service, insurance consultation, medical hospital guide and the like. Acquiring an emotional feature vector of the input text based on a preset large language model; encoding the emotion feature vector to obtain a voice encoding vector; determining an emotion curve vector of the input text based on a pre-trained neural network model; and generating target voice corresponding to the input text based on the voice coding vector and the emotion curve vector. According to the method, emotion feature vectors are extracted through a large language model, a dynamic emotion curve is generated in combination with a neural network, speech synthesis is cooperatively driven, emotion fineness and continuity are remarkably improved, and the problems of traditional TTS emotion expression roughening and fragmentation are solved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Intelligent voice interaction method and device based on voice large model, terminal and medium

The invention discloses an intelligent voice interaction method and device based on a voice large model, a terminal and a medium, and belongs to the technical field of voice recognition and interaction, and the method comprises the steps: recognizing the voiceprint information of a speaker through a voiceprint recognition model when a wake-up instruction is received; when the identity of the current pronunciator is the specified user registered through voiceprint, acquiring question information of the current pronunciator, calling a prompt word of a character event text of the specified user corresponding to the current pronunciator from a cloud end, and combining with a question proposed by the current pronunciator to obtain the question information of the current pronunciator; generating event text information related to a specified user character corresponding to the current pronunciator; and calling the speech synthesis model, synthesizing the generated event text information into speech according to the tone characteristics of the speaker, and outputting and playing the speech. According to the method, the user is supported to upload the character as the cue word, and the cue word is combined with the large model, so that the dialogue content generation based on the personalized information of the user is realized, and convenience is provided for the use of the user.
Owner:NANJING KUKAI SMART SCREEN TECH CO LTD

User timbre cloning and speech synthesis method and device based on deep learning

The invention discloses a user tone cloning and speech synthesis method and device based on deep learning, and relates to the technical field of speech recognition application, and the method comprises the steps: collecting the speech data of a user; extracting timbre features corresponding to the user from the voice data; training the personalized timbre model by adopting a transfer learning technology to obtain a trained timbre model; obtaining picture book content needing to be read, and extracting emotion and rhythm key information contained in the picture book content; according to the extracted picture book content including emotion and rhythm key information, dynamically adjusting each playing parameter of speech synthesis; and inputting the content of the picture book needing to be read into the trained tone model for speech synthesis, adjusting the synthesized speech in combination with the playing parameters of the adjusted speech synthesis, and outputting the adjusted speech. According to the invention, personalized cloning can be carried out according to the timbre characteristics of a specific user, a more natural and vivid speech synthesis effect is provided, and convenience is provided for use of the user.
Owner:SHENZHEN KUKAI SOFTWARE TECH CO LTD

Simulation digital human real-time intelligent voice interaction system and method based on vision and large model

The invention relates to a simulation digital human real-time intelligent voice interaction system and a simulation digital human real-time intelligent voice interaction method based on vision and a large model, and aims to solve the problems of inaccurate target speaker recognition, high response delay and the like in digital human voice interaction in a complex scene. The system circles an effective recognition range through a camera, triggers audio collection in combination with face detection, locks a target speaker and reduces noise by using lip movement recognition and sound image fusion technologies, converts the target speaker into a text through voice wake-up, generates an answer by means of a large language model (LLM) and knowledge retrieval enhancement (RAG) technologies, generates low-delay voice through a voice synthesis technology accelerated by the vLLM, and performs voice recognition on the target speaker. And driving the preloaded digital human image to synthesize a video stream and pushing the video stream to a front end for rendering in real time. Accurate pickup, low-delay interaction and rapid digital human image switching in a complex environment are realized, the accuracy and real-time performance of intelligent voice question answering are improved, and the method is suitable for government affair halls, exhibition halls and other scenes.
Owner:UNICOM (HENAN) IND INTERNET CO LTD

Sound duplicating and low-delay streaming speech synthesis method and system based on ultra-short sample

The invention provides a voice replication and low-delay streaming voice synthesis method and system based on an ultra-short sample, relates to the technical field of artificial intelligence, and is suitable for intelligent interaction, outbound service and multi-mode communication scenes. Deep personalized customization of intelligent voice interaction and real-time generation of ultra-low delay are realized, and bidirectional streaming interaction of a system level is supported, so that the fluency and response speed of dialogues are improved; an ultra-short sample sound duplicating module specially designed for processing ultra-short audio samples and a speech synthesis engine with ultra-low delay and bidirectional streaming output capability are integrated, and the capabilities of the ultra-short sample sound duplicating module and the speech synthesis engine are applied to a real-time and interactive intelligent speech interaction process to form a complete and efficient solution. The industrial pain point is directly solved in a targeted manner, and the method has important commercial application value.
Owner:GUANGDONG CHAOTENG INFORMATION TECHNOLOGY CO LTD

Voice cloning method and device based on emotion enhancement and related medium

The invention discloses a speech cloning method and device based on emotion enhancement, and a related medium. The method comprises the following steps: respectively obtaining a reference audio and a prediction text corresponding to the reference audio; respectively preprocessing the reference audio and the prediction text to generate standardized audio data for feature extraction; inputting the standardized audio data into an emotion enhancement module so as to generate acoustic features matched with a target voice style through multi-round feature fusion and an autoregression generation mechanism; and inputting the acoustic features into a speech synthesis module for decoding processing, and outputting predicted speech. According to the invention, by introducing the emotion enhancement module, accurate modeling of the emotion style of the speaker in the acoustic feature generation process is realized, so that the emotion simulation degree and the personality restoration capability of the synthesized speech are significantly improved.
Owner:AFIRSTSOFT CO LTD

Medical intelligent question-answering method and system based on multi-modal data, medium and product

The invention discloses a medical intelligent question and answer method and system based on multi-modal data, a medium and a product, and relates to the field of medical intelligent question and answer. Performing knowledge retrieval in a preset vector database according to the query content to obtain a knowledge retrieval result; performing query intention recognition on the query content to obtain a query intention recognition result; generating a query answer according to the knowledge retrieval result and the query intention recognition result; wherein the query answer comprises one or more of text, image annotation, voice synthesis and video interpretation. The system supports various output forms such as text, image annotation, voice synthesis and video interpretation, explains and analyzes the illness state or medical diagnosis and treatment results of the patient from multiple dimensions, provides rich and visual diagnosis and treatment reports and interactive interfaces for doctors and patients, and effectively improves the user experience and decision-making efficiency.
Owner:DATA SPACE RES INST

Voice compression method and system based on multi-scale back projection feature fusion

The invention discloses a voice compression method and system based on multi-scale back projection feature fusion, and belongs to the technical field of voice synthesis, and the method comprises the steps: calling a plurality of multi-scale back projection feature fusion layers in an encoder to encode a to-be-synthesized voice signal, and obtaining voice features; calling a plurality of multi-scale back projection feature fusion layers in a decoder to decode the voice features to obtain synthetic voice; wherein the process of encoding or decoding the input features by the multi-scale back projection feature fusion layer comprises the following steps: carrying out cross learning on the input features by using convolution kernels of different scales to obtain multi-scale features, carrying out back projection on the multi-scale features respectively to obtain back projection features, the back projection features and the multi-scale features are fused and then fused with features input into the multi-scale back projection feature fusion layer, and output features of the multi-scale back projection feature fusion layer are obtained. The speech synthesis quality is improved, and the technical problem that the speech synthesis quality of a current method is limited is solved.
Owner:QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +1

Voice decoding method, system and equipment based on electroencephalogram signals and medium

The invention discloses a voice decoding method, system and device based on electroencephalogram signals and a medium, and relates to the technical field of electroencephalogram signal processing.The method comprises the steps that reading electroencephalogram signals and voice signals in the reading process of a to-be-tested person and imagination electroencephalogram signals in the imagination reading process of the to-be-tested person are collected; inputting the reading electroencephalogram signals and the voice signals into the DRCL to generate electroencephalogram characteristics containing voice information; training the DBM by using the electroencephalogram characteristics as input and using the voice signals as output, and adjusting the DBM by using imaginary electroencephalogram signals to construct a voice synthesizer; a mapping relation between the imaginary electroencephalogram signals and the voice signals is generated through the voice synthesizer, and decoding from the electroencephalogram signals to the voice signals is completed; according to the method, a deep representation correlation learning method is provided, potential correlation between electroencephalogram and voice signals can be deeply mined through a multi-layer network structure, a complex mode which is difficult to recognize by a traditional model is captured, and the voice decoding process is more accurate.
Owner:HARBIN INST OF TECH

Multi-modal fusion driven emotion perception enhanced TTS speech synthesis method

The invention provides a TTS speech synthesis method for emotion perception enhancement under multi-modal fusion driving, and the method comprises the following steps: S1, carrying out the collection and preprocessing of multi-modal data, the multi-modal data comprising text data, speech data and facial expression data; s2, extracting and analyzing emotion features; s3, emotion perception speech synthesis model training; s4, speech synthesis and post-processing; s5, performing model evaluation and optimization; by collecting and analyzing multi-mode data such as texts, voices and facial expressions, emotion features can be captured more comprehensively and accurately, complementary information among different modes is fully mined through application of a multi-mode fusion network and a collaborative attention mechanism, the emotion expression of the synthesized voices is closer to real emotions, and the user experience is improved. And the precision of emotion perception is greatly improved.
Owner:TIANJIN CHENGJIAN UNIV

Hierarchical coding and decoding speech synthesis method, device, equipment and medium

The invention relates to the technical field of speech synthesis, can be applied to business system platforms such as financial science and technology, medical health and the like, and discloses a layered coding and decoding speech synthesis method, device, equipment and medium, and the method comprises the following steps: carrying out semantic feature extraction on a pre-acquired original text to obtain text semantic features; performing prosodic feature extraction on the original text to obtain a text prosodic feature; performing feature fusion compression on the text semantic features and the text rhythm features to obtain fusion compression features; performing low-frequency decoding on the fusion compression features to obtain low-frequency decoding features; performing high-frequency decoding on the fusion compression feature to obtain a high-frequency decoding feature; and combining the low-frequency decoding features and the high-frequency decoding features, and converting the combined features into target synthetic speech. Through high and low frequency decoding, richer and more accurate voice details are restored, so that the synthesized voice is listened more naturally and smoothly, and the credibility and fluency of the voice are improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Intelligent accompanying robot system

The invention discloses an intelligent accompanying robot system, and the system comprises the following modules: a multi-mode interaction module which is composed of a voice recognition unit, a voice synthesis unit, an emotion recognition unit, and a visual recognition unit, and is used for collecting the voice, facial expression, motion, and environment information of a user; the localized AI decision module is used for deploying a lightweight DeepSeekR1 large language model based on an ESP32-S3 edge computing chip, performing real-time processing on the multi-modal data and generating a social guidance strategy and an emotion intervention instruction; the data management module comprises a user behavior database and a privacy protection unit and is used for desensitizing the data and realizing sensitive data isolation through local storage; according to the method, the single-person intervention cost is reduced, and the time consumption of manual scene simulation is reduced.
Owner:NANTONG UNIV

Real-time sound duplicating method and system based on end-cloud fusion

The invention provides a real-time sound copying method and system based on end-cloud fusion. The method comprises the following steps: a cloud end carries out real-time tone copying and voice synthesis on a small amount of voice data of a user based on an AI large model; timbre samples are collected when a user registers voice audio data, and user timbre voice data of a preset text are synchronously generated by a large model and are used as fine tuning training data of an end-side voice synthesis model; user tone voice data of a preset text and voice audio data registered by a user are utilized to carry out migration fine tuning training on an end-side voice synthesis model to adapt to the personalized tone of the user, so that high-quality output of the end-side voice synthesis model is ensured, and personalized voice replication is realized; and issuing the trained end-side speech synthesis model to the user equipment, and independently completing speech replication in a network-free or weak network environment. According to the invention, through automatic generation of the user tone data and adaptive fine tuning of the model, the tone of the user is deployed to the end side after fine tuning, and high-quality and high-adaptability sound replication of end-cloud collaboration is realized.
Owner:PACHIRA TIMES (ZHUHAI HENGQIN) INFORMATION TECH CO LTD

Speech synthesis method and device, vehicle and storage medium

The embodiment of the invention provides a voice synthesis method and device, a vehicle and a storage medium, and the method comprises the steps: obtaining multi-modal data which comprises a voice signal, a user instruction, a historical interaction log and bus data; determining model input parameters according to the multi-modal data; according to the model input parameters, the static corpus and a preset large model, generating a personalized script, the preset large model being used for dynamically generating the personalized script in combination with the model input parameters and the static corpus; and synthesizing the broadcast audio according to the personalized script. According to the invention, the technical problem that the voice synthesis technology in the related technology cannot adapt to the personalized demands of the user is solved.
Owner:GUANGZHOU AUTOMOBILE GROUP CO LTD

Real-time emotion perception and voice interaction system for intelligent cockpit

The invention discloses a real-time emotion perception and voice interaction system for an intelligent cabin. The system comprises a multi-modal data acquisition module used for synchronously acquiring a facial image, a voice signal and text input of a driver; the visual feature enhancement unit is used for carrying out restoration and emotion distribution extraction on the low-quality image; the audio noise reduction and feature extraction unit is used for extracting voice emotion features; the text emotion coding unit is used for fusing relative position coding and context semantic information; the cross-modal fusion module is used for outputting an emotion classification result by integrating visual, audio and text features; the personalized emotion database is used for storing historical emotion data of the user and performing emotion trend prediction and early warning judgment; the large language model feedback module is used for generating structured cue words according to the emotion recognition result and the driving situation and generating natural language feedback; the voice synthesis and output module is used for adjusting voice parameters and performing feedback output through a vehicle-mounted multi-channel; according to the invention, the emotion expression of human-vehicle interaction is enhanced.
Owner:SUZHOU UNIV

Text-to-speech synthesis using generative artificial intelligence models

A method and a system for generating human speech audio in a conversation using a trained generative AI model are provided. The method includes receiving a text input representing a portion of the conversation, receiving dialog context associated with the conversation, receiving information representing at least one voice and speaking style of at least one speaker in the conversation, generating the at least one voice and speaking style based on the received information, and generating at least one emotional audio response for the at least one speaker using the at least one voice and speaking style and without retraining the trained generative AI model.
Owner:PHEON INC

Intelligent cockpit implementation method and system based on single-model multi-task reasoning

The invention discloses an intelligent cockpit implementation method and system based on single-model multi-task reasoning, and belongs to the field of automotive electronics, and the method comprises the steps: inputting shared depth features into a detection head and a first classification head after original image data in an intelligent cockpit is preprocessed; the output of the detection head is subjected to detection post-processing to obtain a target detection anchor frame and position information, and the output of the first classification head is subjected to classification post-processing to obtain distraction identification classification; after the target detection anchor frame and the position information are preprocessed, human body area sub-images are input into a safety belt recognition and classification model; inputting the face region sub-images into a fatigue recognition classification model, a face ID matching model, a fixation point estimation model, a sight line estimation model and a head posture estimation model in parallel; and a face image sequence and a voice sequence in the continuous video stream are collected, the face image sequence and the voice sequence are respectively input into the lip language recognition model and the voice recognition model, a recognition result is input into the large language model to be processed to generate an interaction instruction, and a response is output through the voice synthesis module.
Owner:北极雄芯智驾科技(无锡)有限公司