Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

44 results about "Speech technology" patented technology

Speech technology relates to the technologies designed to duplicate and respond to the human voice. They have many uses. These include aid to the voice-disabled, the hearing-disabled, and the blind, along with communication with computers without a keyboard. They enhance game software and aid in marketing goods or services by telephone.

Industrial internet software intelligent customer service interaction method based on multi-modal deep fusion

The invention relates to the technical field of industrial internet, artificial intelligence and human-computer interaction, in particular to an industrial internet software intelligent customer service interaction method based on multi-modal deep fusion. According to the method, an intelligent customer service system integrating natural language processing, a voice technology, computer vision and augmented reality (AR) is constructed aiming at the characteristics of complex operation, multiple user groups, high problem speciality and the like in an industrial scene. The core of the method is that operation questions, interface states and potential faults of engineers in software use are accurately understood through multi-modal input; an accurate solution is generated in combination with the knowledge graph of the industrial software and the real-time data of the equipment; a 3D virtual expert image is driven, and immersive and scene-based guidance is provided by fusing voice explanation, interface marking guidance, AR superposition demonstration and operation demonstration videos. The method has the working condition self-adaptive capability, can identify the professional level (such as a green hand / expert mode) of a user, and adjusts the explanation depth; and cross-region collaboration is supported, a plurality of users are allowed to synchronously share a guidance picture of a virtual expert, and efficient remote collaboration troubleshooting is realized. The method solves the technical problems that traditional industrial software is slow in customer service response, abstract in guidance, difficult to process complex problems on site and the like, and operation and maintenance efficiency and user experience are remarkably improved.
Owner:SHANGHAI CAIJIANG INTELLIGENT TECH CO LTD

Interaction method and device, intelligent agent, equipment, medium and program product

The invention provides an interaction method, an interaction device, an intelligent agent, equipment, a medium and a program product, and relates to the technical field of artificial intelligence, in particular to the technical field of man-machine interaction, computer vision and voice. According to the specific implementation scheme, in response to received multi-modal information input by a target object to a virtual object, intention recognition is conducted on the multi-modal information, and text information representing the intention of the current round is determined; based on the text information and historical visual perception information crossing time domains with the text information, response analysis is carried out, response information is obtained, and the historical visual perception information is obtained by carrying out visual perception on historical multi-modal information input by the target object in historical rounds; and broadcasting the response information through the virtual object.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

High-performance zero-sample text-to-speech conversion method and system based on vLLM acceleration

The invention discloses a high-performance zero-sample text-to-speech method and system based on vLLM acceleration, and belongs to the technical field of intelligent speech, and the system comprises a dynamic streaming sentence segmentation module which carries out the intelligent segmentation of an input text, and obtains a to-be-converted text; the multi-reference audio fusion module is used for receiving a plurality of reference audios of a speaker and combining weight fusion to obtain a final audio coding feature; the session management and control module is used for judging whether the to-be-converted text is a new session request according to the session ID of the to-be-converted text, and if the to-be-converted text is the new session request, distributing a speaker ID and an associated audio signaling feature to the to-be-converted text; if not, searching a speaker ID (Identity) and an audio conditioning feature; and the text-to-voice module is used for performing voice synthesis on the text to be converted by adopting the IndexTTS model accelerated by the vLLM. According to the method, the reasoning performance can be remarkably improved, the audio quality and the response time delay stability in a multi-round dialogue scene are ensured, and the natural continuity of long text synthesis is ensured.
Owner:JIANGSU HAOBAI INFORMATION SERVICE CO LTD

Bluetooth connection voice prompter

This utility model relates to the field of Bluetooth voice technology, specifically a Bluetooth-connected voice prompt device, including a speaker, a sound hole, and a protective component. The sound hole is located on the outer arc surface of the speaker. A light guide ring is mounted on the upper surface of the speaker, and the light guide ring is arc-shaped. A button is fixedly mounted on the upper surface of the speaker. The protective component is located on the outer arc surface of the speaker and includes a guide rail. The guide rail is fixedly connected to the left side of the speaker near the sound hole and is rectangular in shape. A guide block is slidably connected to the inner wall of the guide rail, and a baffle is fixedly connected to the side of the guide block away from the guide rail. This utility model, by incorporating the protective component, can effectively prevent dust, hair, debris, and other foreign objects from entering the sound hole. If these foreign objects enter the interior, they may adhere to the speaker diaphragm and other sound-producing components, affecting sound propagation and sound quality, and may even damage the diaphragm, shortening the lifespan of the voice prompt device.
Owner:SHENZHEN LONGQIANSHU INVESTMENT SERVICES CO LTD

3D digital human lip shape driving method and device, electronic equipment and storage medium

The present disclosure provides a 3D digital human lip shape driving method and device, electronic equipment and storage medium, which relates to the technical field of computer vision. The method comprises: obtaining input text information; converting the text information into a phoneme sequence, audio data and timestamp information based on text-to-speech (TTS) technology; deleting corresponding silent phonemes in the phoneme sequence according to the timestamp information; performing a preset multiple sampling on the phoneme sequence after the deletion processing to obtain a bs animation coefficient sequence; and generating a 3D digital human lip shape animation according to the bs animation coefficient sequence, the audio data, a preset phoneme lip shape mapping table and a preset optimization of special phonemes. The present disclosure improves the robustness and fluency of 3D digital human lip shape driving.
Owner:CHINA TELECOM CORP LTD

Audio training dataset screening method based on phoneme matching pronunciation table

This invention belongs to the field of speech technology, specifically relating to a method for selecting audio training datasets based on phoneme matching pronunciation tables. Addressing the problems of high redundancy, high acquisition and annotation costs, and difficulty in verifying content integrity in existing technologies, the following solution is proposed: A target phoneme set is constructed according to the needs of the target task or domain; a pronunciation table is constructed or obtained and stored in key-value pair format; the original audio-text pair dataset is obtained, and the text is preprocessed; the preprocessed text is converted into corresponding phoneme sequences in the pronunciation table, and each phoneme sequence is analyzed; based on preset selection criteria, sample pairs in the original audio-text pair dataset are selected, retaining those that meet the selection criteria to form a reduced dataset; the reduced dataset is output. This solution reduces annotation costs while ensuring data comprehensiveness, and is suitable for processing datasets for lightweight edge speech synthesis models and fine-tuning of speech models.
Owner:HANGZHOU JUNTONG FUTURE TECHNOLOGY CO LTD

A controller and chip

The present disclosure provides a kind of controller and chip, it is related to integrated circuit technical field, more particularly to chip technology and voice technology.The specific implementation scheme is: a kind of controller, configuration is in chip, the controller includes multiple storage and interconnection interface and configuration interface;Wherein, the configuration interface is configured to the working mode of the multiple storage and interconnection interface, wherein the working mode includes storage interface mode and interconnection interface mode;Each storage and interconnection interface is configured based on the configuration of the configuration interface, and it is externally connected memory in the storage interface mode, or it is externally connected other chip in the interconnection interface mode.The present disclosure can realize multi-port storage and multi-chip interconnection in the same interface in the controller of chip by configuration, on the premise of controllable cost, not only can expand the required storage capacity, but also can improve overall computing power through inter-chip interconnection, meet different product requirements.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Audio recall method, model training method, device and electronic equipment

The present disclosure provides an audio recall method, a model training method, a device and an electronic device, relates to the technical field of cloud computing, in particular to the technical field of deep learning, intelligent search and voice technology, and the audio recall method comprises: acquiring a first audio; segmenting the first audio to obtain N audio segments, any two adjacent audio segments in the N audio segments partially coincide, and N is an integer greater than 1; recalling a second audio corresponding to each of the N audio segments from a sample pool.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Voice wake-up method and device

The invention relates to the technical field of voice, and discloses a voice wake-up method and device, and the method comprises the steps: obtaining the living body indication information and spatial information, detected by voice interaction equipment based on target perception, of a target user relative to the voice interaction equipment, and carrying out the voice wake-up according to the target information of the target user relative to each piece of voice interaction equipment; and determining at least one target activation device in each voice interaction device, and controlling each target activation device to respond to a voice instruction sent by the target user. It can be seen that the probability that the voice interaction device is awakened by non-user sound sources such as environment sound can be reduced, the nearby awakening and nearby interaction performance of the voice interaction device and the user response accuracy are improved, and the voice interaction performance between the user and the device and the use experience of the user on the voice interaction device are improved.
Owner:FOSHAN VIOMI ELECTRICAL TECH

system

Provide a system. 【Solution means】 Means for receiving travel-related information from a user, Means for generating a travel plan based on the received information, Means for presenting the generated travel plan to the user and receiving feedback, Means for readjusting the travel plan by reflecting the user's feedback, Means for visualizing and displaying the travel plan on a map, Means for supporting reservations for accommodation facilities and transportation, Means for dynamically adjusting the plan in consideration of real-time information during the trip, Means for providing information in multiple languages, Means for accepting the user's travel request in natural language using voice recognition technology, Means for presenting the generated travel plan in voice using text-to-speech technology, A system including the above.
Owner:SOFTBANK GROUP CORP

Audio processing method and device and electronic equipment

The invention provides an audio processing method and device and electronic equipment, and relates to the technical field of artificial intelligence, in particular to the technical fields of deep learning, natural language processing, computer vision, voice technology large models and the like. The specific implementation scheme is as follows: acquiring an audio generation request of a video application; the video application is in an audio mode, and the audio generation request comprises an audio identifier of a target audio in a playing state in the video application; obtaining a target video corresponding to the target audio according to the audio identifier; generating a reference audio corresponding to the target audio according to the comment information of the target video; playing the reference audio through the video application after the target audio is played; wherein the reference audio generated according to the comment information can reflect the climax fragment information or the comment information and the like in the target video, so that the object can obtain the information amount matched with the watched video, and the matching degree between the played audio and the object demand is further improved.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Text-to-voice conversion method and device

The invention discloses a text-to-speech method and device, and relates to the technical field of artificial intelligence speech. The method comprises the following steps: training an audio sample, and generating a timbre feature model weight and a model vocoder; receiving a new timbre registration request containing user metadata, a path of timbre feature model weight and audio data, distributing a unique identifier for a new timbre, and setting an initial state as an inactive state; selecting a target timbre from the plurality of timbres, and setting the state of the target timbre as an activated state; receiving a synthesis request, and segmenting the text into ordered text segments; and loading the target timbre in the activated state, processing the text segments in sequence by using the timbre feature model weight corresponding to the target timbre, generating audio features corresponding to the text segments, converting the audio features into an audio data stream by using a model vocoder, and transmitting the audio data stream to a user. The problems that an existing speech synthesis method is long in generation time and does not support low-delay streaming playing are solved.
Owner:CETC XINGHE BEIDOU TECH (XIAN) CO LTD

Voice technology engine pool dynamic scheduling method and system

PendingCN122369466ASpeech technologySpeech sound
This invention provides a dynamic scheduling method for a voice technology engine pool, belonging to the field of automotive technology. The method involves acquiring voice signals from users or devices; selecting an engine to process the voice signals; initiating a monitoring program to monitor the processing; and aggregating the output after the voice signals are processed. This invention integrates the advantages of multiple engines to form a heterogeneous engine pool, compensating for weaknesses in recognition rate, dialect support, and noise resistance, resulting in an overall service quality exceeding that of a single engine. Through real-time monitoring and a rapid failover mechanism, uninterrupted service is achieved, single-point failures are transparent to users, and the system exhibits extremely high robustness.
Owner:DONGFENG MOTOR GRP

Speech recognition method, apparatus, device, storage medium and product

ActiveCN115831125Beasy to understandSpeech analysisSpeech technologySpeech sound
The present disclosure provides a speech recognition method, device and equipment, and a storage medium, relates to the technical field of artificial intelligence, in particular to the technical field of deep learning and speech. The speech recognition method comprises: obtaining speech data; extracting acoustic feature vectors of the speech data according to at least two frame lengths; clustering the acoustic feature vectors to obtain a speaker label and a time label corresponding to the speaker label; and identifying the speech data according to the speaker label and the time label to generate a speech recognition result with the speaker label and the time label. The speech recognition method can improve the accuracy of distinguishing different speakers in the speech data, improve the accuracy and reliability of the speech recognition result, and can be applied to real-time speech recognition scenarios such as meetings and interviews.
Owner:BAIDU COM TIMES TECH (BEIJING) CO LTD

Target recognition-oriented multi-dimensional collaborative visual prosthesis electrical stimulation coding method

The invention relates to a target recognition-oriented multi-dimensional collaborative visual prosthesis electrical stimulation coding method and system, electronic equipment and a storage medium. The method comprises the following steps: acquiring a two-dimensional target instance; saliency map quantization and edge and key point feature extraction are carried out on the two-dimensional target instance to obtain a linkage quantization grading atlas comprising multi-dimensional features; converting the linkage quantization grading atlas into a multi-dimensional feature quantization atlas adaptive to the size of the target electrode array; acquiring a multi-dimensional feature electrical stimulation parameter sub-atlas sequence corresponding to the multi-dimensional feature quantization atlas; and modulating an electrode array stimulation strategy according to a visual prosthesis communication protocol by using the multi-dimensional characteristic electrical stimulation parameter sub-atlas sequence, and generating a voice feedback signal accurately synchronized with a visual electrical stimulation time sequence by using a text-to-voice technology. According to the method, accurate locking of a semantic effectiveness and visual saliency target is realized, the electrical stimulation coding identification degree of multi-dimensional feature linkage quantitative classification is also remarkably improved, the electrical stimulation power consumption under a visual persistence effect is effectively reduced, and the accuracy of target identification is greatly improved through cross-modal collaboration.
Owner:THE FIRST MEDICAL CENT CHINESE PLA GENERAL HOSPITAL

Video generation method and device and electronic equipment

The invention provides a video generation method and device and electronic equipment, and relates to the technical field of artificial intelligence, in particular to the technical fields of deep learning, natural language processing, computer vision, voice technologies, large models and the like. The specific implementation scheme is as follows: determining a knowledge text, a corresponding display related text and a candidate video clip set; performing text rendering processing on candidate video clips in the candidate video clip set according to the display related text, and / or generating video clips according to the display related text and adding the video clips into the candidate video clip set to obtain a processed candidate video clip set; performing splicing processing on each candidate video clip to obtain a knowledge video; according to the technical scheme, the knowledge video can be obtained by performing text rendering and other processing on the candidate video clips in the candidate video clip set corresponding to the knowledge text according to the display related text corresponding to the knowledge text, so that batch production of the knowledge video is supported, the video production efficiency is improved, and the video production cost is reduced.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Voice interaction method and device, equipment, storage medium and program product

The invention discloses a voice interaction method and device, equipment, a storage medium and a program product, and relates to the technical field of intelligent voice. The method comprises the following steps: acquiring first voice content automatically acquired by voice equipment; under the condition that the voice device fails to recognize the first voice content, in response to a first interaction operation received within a preset duration range, recording an association relationship between a first interaction function and a first voiceprint feature, the first voiceprint feature and a voiceprint feature of the first voice content conforming to a voiceprint matching relationship, the first interaction operation is a manual operation used for triggering execution of the first interaction function under the condition that the voice equipment fails to recognize; under the condition that the association relationship meets the preset effective condition, the association relationship is updated to be in an effective state, the effective state is used for indicating the voice equipment to execute the first interaction function based on the second voice content, the second voice content is the voice content of which the voiceprint feature and the first voiceprint feature meet the voiceprint matching relationship, and the voice interaction efficiency is improved.
Owner:GUANGDONG MURORA INTELLIGENT LIGHTING CO LTD

An external voice broadcasting device and vehicle

ActiveCN224439018UImprove experienceSpeech is beneficial toTelecommunicationsAcquisition apparatus
This application discloses an external voice broadcasting device and vehicle, belonging to the field of vehicle-mounted voice technology. The device includes at least a voice acquisition device, a vehicle control circuit, a power amplifier circuit, a voice playback device, a voice wake-up button, and button hardware circuitry. The button hardware circuitry is electrically connected to the vehicle control circuit. When the voice wake-up button is pressed for more than a preset time, it uploads a long-press signal to the vehicle control circuit. The vehicle control circuit is electrically connected to both the power amplifier circuit and each voice acquisition device. Upon receiving the long-press signal, it uses the voice acquisition device to acquire user voice information and then sends the user voice information to the power amplifier circuit. The voice playback device is electrically connected to the power amplifier circuit and installed externally in the vehicle to play the amplified user voice information. This application can at least enable in-vehicle personnel to speak to people outside the vehicle, improving the user experience.
Owner:CHINA FAW CO LTD +1

A speech synthesis model method capable of synthesizing multi-emotional audio.

ActiveCN116798403BData setSpeech technology
This invention discloses a speech synthesis model method capable of synthesizing multi-emotion audio, relating to the field of intelligent speech technology. The method includes the following steps: processing raw data, distinguishing between training and validation sets, adding annotation files to each set, and simultaneously delivering the raw dataset to an emotion recognition module for processing; calling the emotion recognition module to preprocess the dataset, decomposing the audio into phonemes and emotion feature files; the complete multi-emotion text-to-speech model and dataset processing are specifically divided into dataset collection, unsupervised preprocessing, encoder training, and online inference. The final output includes a multi-emotion encoder with intermediate outputs and a final online synthesized independent WAV file, capable of achieving multi-emotion output and simulating prosody, making the effect close to that of a real person. No emotion annotation is required during data processing, and the method of constructing a continuous feature value spectrum greatly avoids the problem of inaccurate machine annotation.
Owner:UNICOM WOYUEDU TECH CULTURE CO LTD +1

Eternity Chat: AI-Driven System Configured to Provide User Simulation Conversation with Deceased Individuals

An AI-driven and machine learning system provides users with simulated conversations with deceased individuals. Users may interact with an AI module that reproduces both the voice and persona of a deceased loved one. The system integrates an AI algorithm, a machine learning module, a sentiment analysis component, a voice cloning module, a natural language processing and generation engine, a multimodal interaction module, a data input interface, a user interface, cloud storage, a real-time processing unit, multiple security protocols, and API integrations to external third-party services and databases. Unlike existing solutions, the system functions as a hub that combines proprietary and third-party technologies, continuously learning from user-provided data and prior interactions to deliver emotionally resonant, contextually accurate responses in real time. The architecture supports flexible component substitution, enabling integration with evolving external AI, sentiment, or voice technologies while maintaining personalized, secure, and seamless user experiences.
Owner:MANESH CYRUS VAHABZADEH

A method, system and terminal for intelligent voice real-time communication in a fund transaction room

The application belongs to the technical field of intelligent voice, and discloses a fund transaction room intelligent voice real-time communication method, system and terminal. The communication method comprises the following steps: first, acquiring fund transaction room multi-source voice data through the terminal, and identifying voice instructions related to the user and time domain, frequency domain, nonlinear dynamics and other emotional expression characteristics; the system background normalizes various emotional characteristics of the user to obtain characteristic scores, and obtains a comprehensive emotional score by weighted summation and determines a comprehensive emotional category; judging the risk level of the voice instruction according to a preset risk control threshold, and sending the comprehensive emotional category and the risk level to the risk control or compliance department; through accurate separation of the voice instruction and the emotional information, the emotional state of the trader is comprehensively quantitatively evaluated, so that the transaction decision is more scientific and reasonable; the risk level is effectively identified, and the warning instruction is accurately determined, so that the transaction risk is reduced.
Owner:SHANGHAI SHOUEE TECH CO LTD

Streaming AI speaker (S2502)

1. The name of the design product: streaming AI sound box (S2502). 2. The use of the design product: AI streaming sound box for whole-house intelligent electrical control through far-field voice technology. 3. The design points of the design product: in shape. 4. The picture or photo that best indicates the design points: perspective view 1.
Owner:SHENZHEN SDMC TECH CO LTD

Microphone array, sound velocity estimation method, sound pickup method and related device and apparatus

The microphone array, the sound velocity estimation method, the sound pickup method and the related device and equipment provided by the embodiment of the application are applied to the technical field of voice, the microphone array comprises: N sound vector microphones, which are used to pick up sound signals of a target sound source; wherein the distance between at least two sound vector microphones in the N sound vector microphones is within a distance range of 0.02 meters to 5 meters, and N is an integer greater than 1. The accuracy and real-time performance of sound velocity estimation are improved.
Owner:HUAWEI TECH CO LTD

A method and system for dynamic emotion-based speech generation based on stream matching and multimodal context

This invention discloses a dynamic emotion-based speech generation method and system based on stream matching and multimodal context, belonging to the field of intelligent speech technology. The method includes: acquiring the text to be synthesized in the current round and context information from N previous rounds of dialogue; obtaining semantic representation vectors and emotion representation vectors using a causal cross-attention mechanism; obtaining the discrete coordinates of each phoneme in the text to be synthesized in the current round in a three-dimensional continuous space based on the emotion representation vectors; generating a phoneme-level VAD emotion trajectory sequence based on the discrete coordinates of adjacent phonemes; inputting the phoneme-level VAD emotion trajectory sequence and the semantic representation vector into a constructed vector field to obtain an emotional acoustic representation; and inputting the emotional acoustic representation into a residual vector quantization decoder to output a speech waveform. This invention eliminates acoustic artifacts, achieves self-consistent emotional logic in multi-turn dialogues, and restores highly realistic micro-emotional details such as breathing and vibrato, making it suitable for real-time interaction scenarios such as intelligent assistants and virtual humans.
Owner:YUANYU HUANYU ARTIFICIAL INTELLIGENCE TECHNOLOGY (SUZHOU) CO LTD

A COMMUNICATION AID FOR THE SPEAKINGLY IMPAIRED WITH AUDIO OUTPUT VIA ANALOG CONTROL INPUT BASED ON TEXT-TO-VOICE TECHNOLOGY AND CLOUD COMPUTING

This invention discloses a communication aid for individuals with speech impairments that enables the user to generate audio output based on text-to-speech technology and cloud computing. The device utilizes an analog joystick as the primary input to select operational modes including "Sentence," "Word," and "Letter." In the "Sentence" and "Word" modes, the user can select from various predefined activity categories to generate relevant sentences or words. The "Letter" mode allows the selection of alphabets from A to Z to form unique words. The text selected by the user is sent to a cloud computing system for processing into audio data, which is then output through a speaker. This invention is designed to provide an intuitive, effective, and flexible communication experience, and to support the independence of individuals with speech impairments in daily communication.This tool also supports vocabulary updates and can be used universally for various communication needs.
Owner:UNIVS GADJAH MADA

An audio recording method, device, electronic equipment and readable storage medium

The application provides a recording method and device, electronic equipment and readable storage medium, the method is applied to the voice technical field, the method comprises the following steps: in the case that a recording request sent by a first application is received, if a microphone is opened by a second application, it is determined whether the first application is a first type of application which needs to occupy an audio focus; if the first application is the first type of application, it is determined whether the second application is a second type of application which does not need to occupy the audio focus; if the second application is the second type of application, the recording permission of the first application is opened, so that the first application and the second application jointly acquire the audio data currently collected by the microphone. The application can make multiple applications share the microphone resource, thereby improving the user experience.
Owner:GREAT WALL MOTOR CO LTD

A method and apparatus for communicating over a VoIP and CT network

The embodiment of the present specification provides a method and device for communication based on VoIP and CT network, wherein the method comprises: a calling terminal initiates a call request on an Internet side, wherein the call request contains a calling identification; an Internet server accesses a landing gateway; the Internet server or the landing gateway converts the calling identification into a calling number recognizable by a CT network; according to the call request, an Internet side signaling is generated, wherein the calling number is encapsulated in the Internet side signaling; the landing gateway performs authentication and verification on the Internet side signaling, and converts the Internet side signaling into a telecommunication side signaling after the verification; the telecommunication side signaling is sent to the CT network; the CT network receives the telecommunication side signaling, and connects a called terminal to realize communication with full-process traceability. The embodiment of the present specification can integrate traditional voice technology and emerging VoIP technology, and realize efficient and stable voice communication.
Owner:SHANGHAI SHIJI TECHNOLOGY CO LTD

Intelligent broadcast generation method and device, electronic equipment and storage medium

The invention discloses an intelligent broadcast generation method and device, electronic equipment and a storage medium, and relates to the technical field of computers, in particular to the artificial intelligence fields of natural language processing, voice technologies, large models, agents and the like. According to the specific implementation scheme, multi-dimensional semantic feature analysis is conducted on an original text, and an analysis result is obtained; segmenting the original text according to the analysis result to obtain a segmentation result of the original text; according to the segmentation result, generating an initial broadcast script; wherein the initial broadcaster script is a structured text fusing semantic content of the original text and audio expression control information; generating a target audio according to the initial broadcasting script; and generating a target broadcaster according to the initial broadcaster script and the target audio.
Owner:BAIDU COM TIMES TECH (BEIJING) CO LTD

Video generation method and device of digital human and electronic equipment

The invention provides a video generation method and device of a digital human and electronic equipment, and relates to the technical fields of artificial intelligence, voice technologies, natural language processing and the like. Comprising the following steps: acquiring enhanced prompt information, and encoding the enhanced prompt information to obtain time sequence constraint information; acquiring a reference image sequence including a digital human image, and encoding the reference image sequence to obtain an identity style vector of the digital human; obtaining a time domain audio according to the enhancement prompt information and the reference image sequence; and generating a target digital human video according to the time sequence constraint information, the identity style vector and the time domain audio. According to the method and the device, coherent digital human actions, expressions and mouth shapes can be quickly generated according to voice, texts and instructions coming in real time, the actions and expressions of the digital human can change in real time along with line content, real dialogue interaction is realized, high-cost video shooting is not depended on, and deployment in a real-time scene can be realized. The method is suitable for the application fields of intelligent e-commerce, agents and the like.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD