Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

866 results about "Speech synthesis" patented technology

Speech synthesis is the artificial production of human speech. A computer system used for this purpose is called a speech computer or speech synthesizer, and can be implemented in software or hardware products. A text-to-speech (TTS) system converts normal language text into speech; other systems render symbolic linguistic representations like phonetic transcriptions into speech.

System and method for multi-modal ai conversational interface improving website navigation and user interaction

The present invention relates to a system for transforming static websites into artificial intelligence (AI)-enabled interactive multi-modal conversational platforms. The system comprises a computing device having a processor for receiving user queries as text or speech input through an input module cooperating with a speech-to-text module. A natural language processing (NLP) module interprets intent, classifies user context, and retrieves grounded information from multiple webpages. A persona adaptation module dynamically modifies vocabulary, tone, and avatar representation across roles such as sales assistant, recruiter, educator, healthcare professional, etc. A response generator module produces structured natural language output, transmitted to a text-to-speech synthesis module and an avatar generation module to render synchronized lifelike video responses. An output rendering module displays multi-modal responses include text, audio, and video, thereby enabling direct navigation and escalation beyond limitations of conventional static websites.
Owner:NALLAM SREE RAMA CHANDRA MURTY

Lightweight intelligent traditional Chinese medicine inquiry system and construction method thereof

The invention relates to the field of artificial intelligence medical application, and discloses a lightweight intelligent traditional Chinese medicine inquiry system and a construction method thereof, and the system comprises a multi-dialect adaptive speech recognition module, a traditional Chinese medicine intelligent dialogue large language model module, a natural speech synthesis module, and a continuous learning mechanism module. The multi-dialect adaptive speech recognition module is used for converting dialect speech input of a patient into a standard text; the traditional Chinese medicine intelligent dialogue big language model module is the core of the system and is used for carrying out natural language understanding, dialectical reasoning and inquiry dialogue generation, and the natural speech synthesis module is used for converting a text response generated by the system into speech output; and the continuous learning mechanism module realizes continuous optimization of the large language model through incremental learning architecture and clinical feedback integration. According to the method, while the professional traditional Chinese medicine diagnosis capability is maintained, the calculation complexity is remarkably reduced, and the universality and sustainable development capability of system application are improved.
Owner:SUZHOU ANGSHENG NETWORK TECHNOLOGY CO LTD

Simulation digital human real-time intelligent voice interaction system and method based on vision and large model

The invention relates to a simulation digital human real-time intelligent voice interaction system and a simulation digital human real-time intelligent voice interaction method based on vision and a large model, and aims to solve the problems of inaccurate target speaker recognition, high response delay and the like in digital human voice interaction in a complex scene. The system circles an effective recognition range through a camera, triggers audio collection in combination with face detection, locks a target speaker and reduces noise by using lip movement recognition and sound image fusion technologies, converts the target speaker into a text through voice wake-up, generates an answer by means of a large language model (LLM) and knowledge retrieval enhancement (RAG) technologies, generates low-delay voice through a voice synthesis technology accelerated by the vLLM, and performs voice recognition on the target speaker. And driving the preloaded digital human image to synthesize a video stream and pushing the video stream to a front end for rendering in real time. Accurate pickup, low-delay interaction and rapid digital human image switching in a complex environment are realized, the accuracy and real-time performance of intelligent voice question answering are improved, and the method is suitable for government affair halls, exhibition halls and other scenes.
Owner:UNICOM (HENAN) IND INTERNET CO LTD

Medical intelligent question-answering method and system based on multi-modal data, medium and product

The invention discloses a medical intelligent question and answer method and system based on multi-modal data, a medium and a product, and relates to the field of medical intelligent question and answer. Performing knowledge retrieval in a preset vector database according to the query content to obtain a knowledge retrieval result; performing query intention recognition on the query content to obtain a query intention recognition result; generating a query answer according to the knowledge retrieval result and the query intention recognition result; wherein the query answer comprises one or more of text, image annotation, voice synthesis and video interpretation. The system supports various output forms such as text, image annotation, voice synthesis and video interpretation, explains and analyzes the illness state or medical diagnosis and treatment results of the patient from multiple dimensions, provides rich and visual diagnosis and treatment reports and interactive interfaces for doctors and patients, and effectively improves the user experience and decision-making efficiency.
Owner:DATA SPACE RES INST

Voice decoding method, system and equipment based on electroencephalogram signals and medium

The invention discloses a voice decoding method, system and device based on electroencephalogram signals and a medium, and relates to the technical field of electroencephalogram signal processing.The method comprises the steps that reading electroencephalogram signals and voice signals in the reading process of a to-be-tested person and imagination electroencephalogram signals in the imagination reading process of the to-be-tested person are collected; inputting the reading electroencephalogram signals and the voice signals into the DRCL to generate electroencephalogram characteristics containing voice information; training the DBM by using the electroencephalogram characteristics as input and using the voice signals as output, and adjusting the DBM by using imaginary electroencephalogram signals to construct a voice synthesizer; a mapping relation between the imaginary electroencephalogram signals and the voice signals is generated through the voice synthesizer, and decoding from the electroencephalogram signals to the voice signals is completed; according to the method, a deep representation correlation learning method is provided, potential correlation between electroencephalogram and voice signals can be deeply mined through a multi-layer network structure, a complex mode which is difficult to recognize by a traditional model is captured, and the voice decoding process is more accurate.
Owner:HARBIN INST OF TECH

Intelligent accompanying robot system

The invention discloses an intelligent accompanying robot system, and the system comprises the following modules: a multi-mode interaction module which is composed of a voice recognition unit, a voice synthesis unit, an emotion recognition unit, and a visual recognition unit, and is used for collecting the voice, facial expression, motion, and environment information of a user; the localized AI decision module is used for deploying a lightweight DeepSeekR1 large language model based on an ESP32-S3 edge computing chip, performing real-time processing on the multi-modal data and generating a social guidance strategy and an emotion intervention instruction; the data management module comprises a user behavior database and a privacy protection unit and is used for desensitizing the data and realizing sensitive data isolation through local storage; according to the method, the single-person intervention cost is reduced, and the time consumption of manual scene simulation is reduced.
Owner:NANTONG UNIV

Speech synthesis method and device, vehicle and storage medium

The embodiment of the invention provides a voice synthesis method and device, a vehicle and a storage medium, and the method comprises the steps: obtaining multi-modal data which comprises a voice signal, a user instruction, a historical interaction log and bus data; determining model input parameters according to the multi-modal data; according to the model input parameters, the static corpus and a preset large model, generating a personalized script, the preset large model being used for dynamically generating the personalized script in combination with the model input parameters and the static corpus; and synthesizing the broadcast audio according to the personalized script. According to the invention, the technical problem that the voice synthesis technology in the related technology cannot adapt to the personalized demands of the user is solved.
Owner:GUANGZHOU AUTOMOBILE GROUP CO LTD

Real-time emotion perception and voice interaction system for intelligent cockpit

The invention discloses a real-time emotion perception and voice interaction system for an intelligent cabin. The system comprises a multi-modal data acquisition module used for synchronously acquiring a facial image, a voice signal and text input of a driver; the visual feature enhancement unit is used for carrying out restoration and emotion distribution extraction on the low-quality image; the audio noise reduction and feature extraction unit is used for extracting voice emotion features; the text emotion coding unit is used for fusing relative position coding and context semantic information; the cross-modal fusion module is used for outputting an emotion classification result by integrating visual, audio and text features; the personalized emotion database is used for storing historical emotion data of the user and performing emotion trend prediction and early warning judgment; the large language model feedback module is used for generating structured cue words according to the emotion recognition result and the driving situation and generating natural language feedback; the voice synthesis and output module is used for adjusting voice parameters and performing feedback output through a vehicle-mounted multi-channel; according to the invention, the emotion expression of human-vehicle interaction is enhanced.
Owner:SUZHOU UNIV

Voice synthesis method and system based on VITS improvement

The invention provides a voice synthesis method and system based on VITS improvement, and the method comprises the steps: optimizing a text encoder of a VITS model, introducing a large language model, and enabling the emotion, intention and speaking style of an input text to be captured when the text is encoded; a random disturbance item is introduced when the Q value is dynamically planned and solved, the alignment flexibility in the initial training stage is improved, meanwhile, monotonicity constraint is strictly kept, and it is avoided that a suboptimal solution is obtained through convergence too early; a ConvNeXt module is used as a basic backbone network of a decoder, and ISTFT is utilized to efficiently reconstruct a time domain signal, so that waveform up-sampling is realized, redundant calculation of traditional transpose convolution is avoided, and reasoning is accelerated. According to the method, the reasoning speed, the emotion expression ability and the style control flexibility of speech synthesis can be effectively improved, a new solution is provided for cross-language diversified speech synthesis, and a reference is provided for the more efficient and more intelligent development of the speech synthesis technology.
Owner:豫章师范学院

Automatic operation backtracking and digital person continuous talk control method and system oriented to multi-modal interaction

The invention relates to a multi-modal interaction-oriented automatic operation backtracking and digital person continuous talking control method and system, and the system comprises an operation snapshot module which records operation nodes through a semantic anchor point marking technology, such as an interface XPath page number and a PPT page number; the interruption detection module is used for capturing a user interruption instruction in real time; the progress prediction engine dynamically adjusts a follow-up demonstration path based on a program; and the continuous talk controller is used for coordinating RPA operation execution, digital human rendering and voice synthesis through a three-level priority thread pool. The problems of inaccurate operation flow backtracking and multi-modal thread conflict after demonstration interruption are solved, the breakpoint continuous talk error is obviously lower than that of a traditional scheme, and the method is suitable for PPT demonstration, cockpit explanation, business system training, course explanation and 3D modeling guidance scenes.
Owner:WUXI GANGWAN NETWORK TECH

Intelligent tactile stick interaction method and system based on multi-modal perception and edge calculation

The invention provides an intelligent tactile stick interaction method and system based on multi-modal perception and edge calculation, and relates to the technical field of intelligent auxiliary equipment. The method comprises the following steps: acquiring environment and position data through a camera, a GPS and an inertial measurement unit; performing multi-modal data fusion and deep learning identification by using an edge calculation module to generate a dynamic environment map; performing real-time path planning and obstacle avoidance based on a map and a destination; analyzing a user instruction through voice recognition and interactively adjusting a path; haptic and auditory bimodal feedback is provided through a vibration module and a voice synthesis module according to a navigation result; and when an emergency obstacle is detected or a user actively alarms, a local and remote alarm mechanism is triggered. The system correspondingly comprises an intelligent tactile stick terminal and a cloud server. According to the invention, accurate and real-time perception and intelligent navigation of the environment are realized, visual and natural interaction experience is provided, an effective emergency help channel is established, and the safety, independence and convenience of travel of visually impaired people are significantly improved.
Owner:GUANGXI NORMAL UNIV

Cross-platform calling method and system for corpus training library

The invention relates to the technical field of speech synthesis, in particular to a corpus training library cross-platform calling method and system, and the method comprises the steps: obtaining the text content of a to-be-synthesized audio, disassembling the text content to obtain a pause position sequence, selecting a corresponding corpus training library according to a preset language identifier, extracting the pronunciation duration of character phonemes, and calculating the duration of basic phonemes. Text emotion features are recognized through an emotion classification model, a three-dimensional emotion feature vector containing emotion density, emotion variance and distribution entropy is constructed, and a first phoneme adjustment coefficient is determined according to the three-dimensional emotion feature vector to correct basic phoneme duration. And when a difference value exists between the corrected phoneme duration and the target audio duration, carrying out compression processing on the phonemes when the difference value is a positive number, and generating an optimal second phoneme adjustment coefficient sequence and a pause duration sequence by using a genetic algorithm when the difference value is a negative number. According to the invention, the problem of speech speed and pause duration optimization under the preset audio length is solved.
Owner:SICHUAN NORMAL UNIV

Normalizing flows with neural splines for high-quality speech synthesis

Disclosed are apparatuses, systems, and techniques that may use machine learning for implementing generative text-to-speech models. The techniques include identifying a mapping of speech characteristics (SC) on a target distribution of a latent variable using a non-linear transformation for at least a subset of the SC. Parameters of the non-linear transformation are determined using a neural network that approximates a statistics of the SC with a statistics predicted for the SC based on the identified mapping and the target distribution of the latent variable.
Owner:NVIDIA CORP

Voice interaction large model family health assistant dialogue method, device and equipment and medium

The invention relates to a voice interaction large model family health assistant dialogue method and device, equipment and a medium. The method comprises the following steps: carrying out fragmentation processing according to an original voice stream of a user to generate an audio fragment with a medical mark, and carrying out voice recognition and entity extraction on the audio fragment to generate a dynamic entity map; generating an evidence-based decision prompt based on the map, inputting the prompt into a preset medical big model for processing, and outputting a result containing an essential symptom list; and according to the symptom matching degree of the necessary symptom list and the dynamic entity map, generating a diagnosis report or a question-asking list, if the diagnosis report is output, performing medical rule chain verification operation on the diagnosis report to generate a quality control report, and based on the question-asking list or the quality control report, generating a synthetic voice stream. According to the method, through medical intention directional screening, map entity analysis, large model diagnosis, voice synthesis and the like, the voice recognition accuracy, the diagnosis suggestion reliability and the inquiry interaction efficiency of the family health assistant in the medical scene are improved.
Owner:SHANGHAI LOHAS YUAN MEDICAL TECHNOLOGY CO LTD

Intelligent voice interaction system and method based on streaming multi-mode fusion and equipment control protocol

PendingCN121260156ASpeech recognitionSpeech synthesisSpeech comprehensionEngineering
The embodiment of the invention discloses an intelligent voice interaction system and method based on streaming multi-mode fusion and an equipment control protocol, the system comprises a voice input processing module, a voice understanding and generating module and a voice synthesis module, the voice input processing module is used for converting an audio signal into a first token sequence, and the first token sequence is used for converting the audio signal into a second token sequence; the voice understanding and generating module is used for determining a response token sequence according to the first token sequence on the basis of a multi-modal Transform architecture so as to realize voice understanding and generation; and the voice synthesis module is used for synthesizing the response token sequence into an output audio so as to carry out at least one of the following adjustments on the converted audio of the response token sequence: emotion parameter adjustment, tone adjustment and rhythm adjustment. By adopting the embodiment of the invention, low-delay and high-naturalness intelligent voice interaction can be realized, multi-modal fusion and equipment control are supported, and the user experience is remarkably improved.
Owner:SHENZHEN HUANZHI TECHNOLOGY CO LTD

Asynchronous anthropomorphic communication method and system based on voice transfer

The invention provides an asynchronous anthropomorphic communication method and system based on voice transfer, which are suitable for various terminals such as wearable equipment, earphones, dolls, bolsters and the like. The method comprises the steps of user voice input, edge or cloud recognition, playback confirmation, content translation and voice synthesis, asynchronous transmission and broadcast and the like. The system does not depend on a specific hardware form, emphasizes a user confirmation mechanism and anthropomorphic voice broadcast and supports multilingual translation and personalized voice styles, the communication process is bound based on a device ID or nickname identity, social account login is not needed, and interaction privacy security is guaranteed. The method is widely applicable to various asynchronous social application scenes such as children, old people, lovers, autism rehabilitation and the like. The system supports nickname binding and friend relationship establishment, users can complete social connection through voice instructions or two-dimensional codes, and controllability and interestingness of communication interaction are enhanced.
Owner:GUANGDONG OPERATOR WIRE INTELLIGENT TECHNOLOGY CO LTD

Intelligent communication content adaptation system and method based on user intention recognition

The invention discloses an intelligent communication content adaptation system and method based on user intention recognition. The system integrates and processes texts, voices, emoticons and unstructured behavior data through a multi-modal input analysis module; the deep learning intention recognition engine adopts a triple attention mechanism, current input and weighted historical interaction features are fused, and multi-level intention classification including basic operation, semantic targets and emotion driving is output; the real-time emotion state analysis module fuses acoustics, semantics and physiological indexes to generate a dynamic emotion matrix; the content adaptation decision engine is based on the intention confidence and the emotional state; and the multi-channel output optimization module performs cooperative adjustment on the speech synthesis rhythm, the text abstract and the visual interface according to the speech synthesis rhythm and the text abstract. According to the method, the perception precision of the deep intention and emotion of the user in a complex scene is remarkably improved, the individuation and multi-channel adaptive optimization of the response content are realized, and the communication efficiency and the user experience are effectively improved.
Owner:CHINA INFOMRAITON CONSULTING & DESIGNING INST CO LTD

Customer service information generation method and system based on multi-modal intention recognition

The invention discloses a customer service information generation method and system based on multi-modal intention recognition, and the method comprises the steps: obtaining customer service session data, and carrying out the heterogeneous data classification operation of the session data; performing feature extraction operation on each piece of parting modal information; establishing a semantic bridging matrix among the modal features, performing hierarchical attention routing, and outputting an intention label; a knowledge graph is injected, the knowledge graph comprises a static knowledge base, a real-time service flow and a user historical portrait, and three-source dynamic knowledge is output; constructing a decision tree based on the intention label, the three-source dynamic knowledge and the space-time identifier; and executing a corresponding text generation operation, a visual mark generation operation or a voice synthesis operation based on a response action corresponding to the leaf node of the decision tree, and generating customer service information. According to the method and the system, the buyer intention recognition accuracy of the multi-modal session data can be greatly improved, and the response delay time of generating the customer service information is shortened.
Owner:深圳乐搏科技有限公司

Video generation method and system based on AI voice cloning and mouth shape synchronization

The embodiment of the invention provides a video generation method and system based on AI voice cloning and mouth shape synchronization, and the method comprises the steps: carrying out the fusion of voiceprint features of an input video and an input text through a voice synthesis model after the input video and the input text are obtained, so as to generate a natural voice; analyzing lip key points of the input video by using a lip shape displacement model, and matching lip shape change data according to the lip key points and natural voice; and generating an output video according to the input video and the lip shape change data. According to the method, the phoneme duration prediction of the speech synthesis model and the lip displacement model can be coupled through time sequence convolution, so that the mouth shape of the output video is matched with the speech content, lightweight video restoration is realized by redrawing the lip region, the dynamic response to the input text modified by the user in real time is supported, and the response efficiency is improved.
Owner:成都安易迅科技有限公司

Real-time voice stream dialogue interaction method and system based on large language model

The invention discloses a real-time voice stream dialogue interaction method and system based on a large language model, and belongs to the technical field of voice recognition, natural language processing and human-computer interaction. The method comprises the following steps: S1, segmenting a voice signal according to a 320 millisecond time window; s2, through double-channel parallel processing, generating a preliminary text hypothesis locally and uploading a frame to a server side to accurately reconstruct a text; s3, fusing to generate a final text with a confidence label; s4, inputting the text into a large language model with a context management mechanism to generate a natural language response; s5, performing speech synthesis output of rhythm perception; s6, updating a dialogue cache and synchronizing a language model state; s7, predicting the semantic trend in advance by adopting a delayed slow release mechanism; and S8, introducing round-level intonation / rhythm parameters for semantic adjustment. The method has the beneficial effects that the response delay is reduced, the recognition accuracy is improved by fusing the confidence label, and meanwhile, the continuity and naturalness of multi-round interaction are enhanced through context and intonation modeling.
Owner:MINIMALIST INTERNET (BEIJING) INFORMATION TECHNOLOGY CO LTD

Speech synthesis method and device, computer equipment and storage medium

The invention discloses a speech synthesis method and device, computer equipment and a storage medium. The method comprises the following steps: acquiring multi-mode background sound condition input data; performing modal integrity detection on the multi-modal background sound condition input data to obtain a detection result; generating an environment background sound feature embedding vector according to a detection result; obtaining to-be-synthesized text data and speaker reference audio data, and performing feature extraction to obtain text semantic features and speaker timbre features; inputting the environment background sound feature embedded vector, the text semantic feature and the speaker timbre feature into an acoustic model to generate a Mel spectrum; and converting the Mel spectrum into a target voice waveform to obtain synthetic voice data. By implementing the method, scene requirements can be deeply matched, diversified scene types can be covered, accurate matching of background sounds and voice semantics is realized, and the technical scheme can be applied to the fields of finance and medical health.
Owner:PING AN TECH (SHENZHEN) CO LTD

News broadcasting method based on artificial intelligence and related device

The invention provides a news broadcasting method based on artificial intelligence and a related device. The method comprises the following steps: acquiring text information and chart information corresponding to target news; performing information processing on the character information and the chart information to obtain a target broadcast text and corresponding target feature information; analyzing the target feature information to obtain a reference content attribute and a reference broadcast emotion; determining a target broadcast style according to the reference content attribute and the reference broadcast emotion; obtaining a target broadcast demand of a target user; performing voice synthesis on the target broadcast text based on the target broadcast style and the target broadcast demand through a preset AI broadcast model to obtain a target broadcast voice; and in response to a voice broadcasting operation of the target user, broadcasting with the target broadcasting voice. The picture and text information of the news is processed, and the adaptive voice is synthesized through the AI broadcasting model in combination with the personalized requirements of the user, so that the news broadcasting quality is improved.
Owner:JIANGXI RONG MEDIA BRAIN TECHNOLOGY CO LTD

Large model visual impairment auxiliary robot system based on edge cloud collaboration

The invention discloses a large-model visual impairment assisting robot system based on edge cloud collaboration, and is applied to the technical field of intelligent robots. Comprising a walking assembly, a framework support, a sitting and lying structure and an electrical control system. The walking assembly is arranged below the framework support and used for achieving movement and steering of the auxiliary robot. The sitting-lying structure is arranged in the middle of the framework support and used for achieving a seat function. And the electrical control system is used for realizing sensing and control functions. According to the invention, a large language model, a visual identification model, a voice synthesis model and a mobile robot platform are fused, multiple auxiliary services such as autonomous navigation, environment description, voice interaction and the like can be provided for visually impaired people in public culture or living places, and the travel independence, the information acquisition capability and the culture participation degree of the visually impaired people are remarkably improved.
Owner:ZHEJIANG UNIV OF TECH

Voice synthesis method and device based on voice token fusion, equipment and medium

The invention relates to the technical field of speech semantics, can be applied to business system platforms of financial science and technology, medical health and the like, and discloses a speech synthesis method, device, equipment and medium based on speech token fusion. Performing text coding on the initial potential representation to obtain a target text feature; generating a semantic token corresponding to the initial text according to the target text feature, and performing time sequence alignment on the semantic token and the target text feature to obtain a target semantic token; obtaining user voice of a reference user, extracting timbre characteristics of the user voice, and generating a Mel spectrogram frame by frame according to the timbre characteristics and the target semantic token; and performing speech synthesis according to the Mel spectrogram to obtain target speech. The speech synthesis efficiency and quality can be improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Live voice synthetization

PendingUS20260097308A1Video gamesSpeech synthesisLive voiceTelecommunications
Systems, devices, methods, and machine-readable media configured to provide voice synthetization in a multiplayer video game are provided. A system can include a multiplayer video game including a character selection interface through which a player selects a character to represent them in playing the video game, and a voice model trained to convert audio from the player directly into audio in a voice of the character and provide an output that includes the audio in the voice of the character.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Method for training speech synthesis model, speech synthesis method, and electronic device

A method for training a speech synthesis model includes obtaining training data; obtaining an initial speech synthesis model; training a semantic encoding network and a semantic decoding network in the speech synthesis model respectively based on a style sample speech, a timbre sample speech, an input sample text, and an output sample speech in training samples of the training data, to obtain a trained speech synthesis model.
Owner:BAIDU INT TECH (SHENZHEN) CO LTD

Paper text processing method based on intelligent glasses

The invention relates to the technical field of artificial intelligence, and discloses a paper text processing method based on intelligent glasses. The method comprises the steps that intelligent glasses are awakened through an awakening word, a user can conduct image recognition through a voice instruction, and whether a textbox is complete or not is judged. If the textbox is complete, character recognition is carried out; if not, the intelligent glasses help the user to adjust through voice guidance until the textbox is completely displayed. If the recognized characters are not the language set by the user, the intelligent glasses automatically translate the recognized characters into the user language, and voice synthesis is carried out to generate a voice file. And after generation, inquiring the user whether to play the voice, and if not, encrypting and storing the voice file. The problems that current text detection precision is not high, the reading process is not natural, user control is complex, and a feedback mechanism is single are solved.
Owner:SHENZHEN SENSING FUTURE TECHNOLOGY CO LTD

False wakeup corpus acquisition method, related method, device, equipment and storage medium

The invention discloses a false wake-up corpus acquisition method, a related method, device and equipment and a storage medium, and the false wake-up corpus acquisition method comprises the steps: constructing a model instruction based on a product demand of a wake-up model; performing voice synthesis on a false wake-up text output by responding to the model instruction based on the generative model to obtain false wake-up voice; based on the response result of the wake-up model to the false wake-up voice, the recognition text of the false wake-up voice and the phoneme sequence of the false wake-up text, obtaining a reward value generated by the generative model at this time; and adjusting network parameters of the generative model based on the reward value, determining whether to select a false wake-up text and a false wake-up voice as sample corpora for training of the wake-up model based on the reward value, and returning to the step of constructing the model instruction based on the product demand of the wake-up model until an end condition is met. According to the scheme, the acquisition efficiency and the data quality of the false wake-up corpus can be improved.
Owner:IFLYTEK CO LTD

Corpus expansion method, system and equipment based on speech synthesis and medium

The invention relates to the technical field of speech synthesis, in particular to a speech synthesis-based corpus expansion method, system and device and a medium, and the method comprises the steps: carrying out the preprocessing including data annotation based on the collected audio and corresponding text of a target speaker; extracting acoustic features from the preprocessed audio; on the basis of a pre-trained acoustic model, performing personalized fine tuning by using the annotation data and the acoustic features, and training personalized acoustic models of a plurality of speakers at the same time through multi-thread parallel computing; calling the trained personalized acoustic model, and synthesizing a voice corpus of the target text in combination with a vocoder; and based on the trained personalized acoustic model, continuously expanding the corpus by changing the text. The personalized voice corpus is quickly generated through a small number of voice samples, the data acquisition cost is remarkably reduced, and the corpus construction efficiency is improved.
Owner:深圳市友杰智新科技有限公司