Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

220 results about "Mouth shape" patented technology

Real-time video translation and audio and picture synchronization method and system based on multi-modal large model

The invention provides a real-time video translation and audio and picture synchronization method and system based on a multi-modal large model, and relates to the technical field of video translations, and the method comprises the steps: obtaining a source video; extracting the source video based on the multi-modal large model to obtain a multi-modal feature; fusing the multi-modal features through a cross-modal attention mechanism to generate a context semantic vector; translating into a target language text in real time based on the context semantic vector, and processing the translated language text based on the multi-modal features to obtain a translated language sound source; and performing mouth shape adjustment on the source video based on the translation language sound source, and merging the translation language sound source and the mouth shape animation video to obtain a real-time translation video with synchronous sound and picture. According to the method, the limitation of traditional single-modal translation is broken through, and the semantic accuracy of translation is remarkably improved by dynamically aligning the context information through the multi-modal features in combination with a cross-modal attention mechanism.
Owner:SHANGHAI YINGZHUO INFORMATION TECH CO LTD

Multi-modal time sequence alignment AI video translation method and system

The invention relates to the technical field of subtitle translation, in particular to a multi-modal time sequence alignment AI video translation method and system, and the method comprises the steps: 1, carrying out the multi-modal analysis of a to-be-translated video, and obtaining audio separation data, voiceprint feature data and visual time sequence data; 2, performing cross-language translation and context optimization on the basis of the voice of the audio separation data to generate a target language text, and synthesizing target language voice retaining the original voice color in combination with the voiceprint feature data and the target language text; generating a mouth shape animation matched with the target language voice based on the lip key point data and the limb action time sequence data; and step 3, performing four-dimensional alignment on the target language voice, the translated text, the mouth shape animation and the limb action sequence through a cross-modal time sequence encoder, and dynamically adjusting the layout of the bilingual subtitles to adapt to a video picture. According to the method and the device, multi-mode synchronization can be taken into consideration during video translation, so that the body actions such as voice, subtitles and mouth shapes are kept aligned.
Owner:HANGZHOU BAOMIHUA TECH CO LTD

Video translation method and system based on artificial intelligence

The invention discloses a video translation method and system based on artificial intelligence. The method relates to the technical field of video translation and comprises the following steps of original sound track extraction, target AI speaker adaptation, AI dubbing generation and mouth shape synchronization and video synthesis. According to the method, independent audio and video streams are obtained by adopting an audio and video separation technology, and multiple original sound tracks are extracted through a voice separation model; matching or generating an adaptive target AI speaker module in a preset tone library; converting the original language voice into a text, translating the text into a target language text, and synthesizing an AI dubbing audio track in combination with a target AI speaker module; and finally, the independent video stream and the multi-AI dubbing audio track are input into the mouth shape synchronization model to output a translated video, so that the timbre fitting degree, the voice quality and the voice consistency of the same speaker of AI dubbing are improved, and meanwhile, the resource utilization rate of video translation and the processing efficiency under batch tasks are improved. The problem that in the prior art, video translation is low in quality and efficiency is solved.
Owner:BEIJING DEEP LOGIC INTELLIGENT TECHNOLOGY CO LTD

Intelligent dialogue method and system based on digital human

The invention relates to the technical field of artificial intelligence and multimedia processing, and particularly provides an intelligent dialogue method and system based on a digital human, and the method comprises the following steps: S1, an input stage: a user inputs a question or an instruction through a text or voice; s2, the processing stage comprises large model generation reply, TTS model conversion and audio driving mouth shape; and S3, the output stage comprises video generation, stream pushing and display. Compared with the prior art, the method has the advantages that the reply can be generated through the large model, and the TTS model and the audio-driven mouth shape model are combined, so that more natural and anthropomorphic dialogue interaction is realized, and the user experience is improved.
Owner:浪潮智慧城市科技有限公司

Digital human rendering method and device, storage medium and program product

One or more embodiments of the invention provide a digital human rendering method and device, a storage medium and a program product. The digital human rendering method comprises the following steps: analyzing a target voice stream matched with a to-be-rendered digital human to obtain a phoneme sequence and voice rhythm characteristics; a mouth shape control parameter sequence corresponding to the phoneme sequence is determined, all mouth shape control parameters in the mouth shape control parameter sequence are in one-to-one correspondence with all phonemes in the phoneme sequence, and a mouth shape control curve is generated based on the mouth shape control parameter sequence; performing rhythm alignment processing on the mouth shape control curve according to the voice rhythm characteristics, executing rasterization conversion, and generating a mouth shape image block sequence synchronized with the target voice stream; and synthesizing each mouth shape image block in the mouth shape image block sequence with other image contents of the digital human to generate a digital human image frame sequence.
Owner:HANGZHOU ANT KUAI TECHNOLOGY CO LTD

Intelligent cabin multi-mode voice interaction system and method

The invention belongs to the technical field of voice processing, and discloses an intelligent cockpit multi-mode voice interaction system and method, and the system comprises a voice triggering unit which collects the environment audio and video information in a cockpit, judges whether to enter a voice interaction mode or not through combining with the environment perception parameters in a vehicle, and sends the voice interaction mode to the vehicle; when the voice interaction triggering condition is satisfied, generating a voice interaction input signal matched with the current environment; the mouth shape analysis unit is used for carrying out acoustic feature extraction on the voice interaction input signal, synchronously analyzing the lip motion trail of the driver in the video information, establishing a corresponding relation between voice phonemes and mouth shape motion, and forming a joint analysis feature; the candidate generation unit is used for carrying out segmented alignment on the joint analysis features and constructing a continuous multi-modal fragment sequence; performing time synchronization on the multi-modal fragment sequence, and projecting the multi-modal fragment sequence to a predefined intention space to obtain a candidate intention set containing different candidate intentions; and the man-machine interaction experience of the intelligent cabin is improved.
Owner:SHENZHEN SHENHANG HUACHUANG AUTOMOBILE TECH CO LTD

Device for Creating Digital Persona

A device for creating digital persona, which includes a data collection module responsible for collecting personality data of a target object, a personality training module utilizing a large language model and the personality data to train and generate a virtual personality model with personality characteristics of the target object, thereby generating a virtual personality consistent with the personality characteristics of the target object, an appearance video generation module and voice generation module respectively utilizing face replacement and voice cloning technologies to extract the pictures and sounds from the personality data to generate a virtual personality model with the appearance and voice characteristics of the target object, a lip synchronization module, utilizing lip synchronization technology to ensure that the mouth shape and voice of the digital persona are synchronized, and an interactive module, providing an interactive interface that allows users to interact with the virtual personality and receive responses from it.
Owner:MORPHUSAI CO LTD

Bionic robot control method, device and equipment and storage medium

The invention discloses a bionic robot control method and device, equipment and a storage medium, and relates to the technical field of bionic robots, and the method comprises the steps: obtaining a target character through a speech synthesis model, determining a phoneme sequence based on the target character, coding the phoneme sequence, and carrying out the preset convolution operation of the obtained coded sequence, a corresponding phoneme feature map is obtained; determining voice data based on the phoneme feature map through a voice synthesis model, and determining a voice time sequence based on the voice data; determining a target mouth shape based on the voice time sequence through a voice synthesis model, and determining a mouth shape time sequence based on the voice time sequence and the target mouth shape; and when the voice data is played, a motor of the mouth is controlled to perform corresponding motion based on the mouth shape time sequence, so that the bionic robot generates a corresponding mouth shape. Therefore, the action of the mouth of the bionic robot can be accurately controlled while the bionic robot plays the voice.
Owner:DIGITAL HUAXIA (SHENZHEN) TECHNOLOGY CO LTD

Geometric constraint-based face mouth shape coordination generation method

The invention discloses a human face mouth shape coordination generation method based on geometric constraints, and relates to the technical field of image generation, and the method comprises the steps: 1, synchronously generating face vertex index tables corresponding to mouth opening and closing spatial form matrixes one by one; 2, based on the facial vertex index table, mapping the lip and gum combined boundary constraint set to a unified coordinate system to form a geometric prior constraint tensor; 3, performing multi-modal topological cascade reversible mixed transformation on the geometric prior constraint tensor to obtain a high-dimensional tooth texture candidate tensor and constraint consistency confidence spectrum; and step 4, using the high-dimensional tooth texture candidate tensor and the constraint consistency confidence spectrum as input, completing anti-aliasing, illumination compensation and gamma correction on a unified rendering pipeline, and finally obtaining a mouth shape coordination synthetic frame sequence. According to the method, the sense of reality, geometric consistency and time sequence coherence of tooth textures in the mouth shape synthesis process are effectively improved, and the problem that lips and teeth are not coordinated in an existing method is solved.
Owner:CLOUD ATTACK NETWORK TECH HEBEI CO LTD

Video generation method and system based on AI voice cloning and mouth shape synchronization

The embodiment of the invention provides a video generation method and system based on AI voice cloning and mouth shape synchronization, and the method comprises the steps: carrying out the fusion of voiceprint features of an input video and an input text through a voice synthesis model after the input video and the input text are obtained, so as to generate a natural voice; analyzing lip key points of the input video by using a lip shape displacement model, and matching lip shape change data according to the lip key points and natural voice; and generating an output video according to the input video and the lip shape change data. According to the method, the phoneme duration prediction of the speech synthesis model and the lip displacement model can be coupled through time sequence convolution, so that the mouth shape of the output video is matched with the speech content, lightweight video restoration is realized by redrawing the lip region, the dynamic response to the input text modified by the user in real time is supported, and the response efficiency is improved.
Owner:成都安易迅科技有限公司

Virtual digital human lip synchronization optimization method, device, equipment and storage medium

The present invention relates to the field of computer vision technology, and discloses a virtual digital human lip synchronization optimization method, device, equipment and storage medium. The virtual digital human lip synchronization optimization method includes: obtaining the target audio segment to be output by the virtual digital human at the next moment; judging whether the target audio segment belongs to the audio type to be processed; if the target audio segment belongs to the audio type to be processed, then based on a preset lip synchronization optimization strategy, generating a 3D human face mouth shape parameter frame sequence corresponding to the target audio segment; based on the 3D human face mouth shape parameter frame sequence, generating a corresponding 3D human face mouth shape image frame sequence and rendering it into the virtual digital human. The present invention can adapt to various audio types and improve the fluency and naturalness of the virtual digital human's mouth shape under different audio types.
Owner:GUANGZHOU HUYA TECH CO LTD

Method and device for synchronizing voice and mouth shape of digital human and electronic equipment

The invention relates to the field of digital people, and discloses a method and a device for synchronizing voice and mouth shape of a digital person, and electronic equipment. The method comprises the steps of obtaining natural language information, converting the natural language information into character information, determining a video frame and an audio frame based on the character information, and controlling the digital person to speak when the first frame time of the audio frame is the same as the first frame time of the video frame, so that the voice and mouth shape of the digital person can be kept consistent.
Owner:WUHAN AOTUO INTELLIGENT TECH CO LTD

Instrument of intelligent breathing and motion capture system based on six-character table health maintenance

The invention discloses an instrument of an intelligent breathing and motion capture system based on six-character table health maintenance, and relates to the technical field of intelligent health maintenance equipment, the instrument comprises a hardware part and a software part; the hardware part comprises a mouth shape visual capture assembly, an action force and angle sensing assembly, an audio acquisition and feedback assembly and a main control and power supply assembly; the software part comprises a data acquisition module, a signal preprocessing module, a parameter analysis and judgment module, a real-time feedback module and a data storage and management module; according to the invention, through integration of the mouth-shaped visual capture assembly, the action force and angle sensing assembly, the audio acquisition and feedback assembly and the master control and power supply assembly, synchronous acquisition and processing of multi-source data are realized; the mouth shape change, the body movement and the audio pronunciation of the user can be accurately captured in real time, and powerful technical support is provided for standardized practice of six-character table health maintenance.
Owner:SHENZHEN TRADITIONAL CHINESE MEDICINE HOSPITAL

Education robot voice signal processing method

The invention discloses a voice signal processing method for an education robot, relates to the technical field of voice signal processing, and aims to solve the problems of multi-person overlapping language and environmental noise interference in a classroom, multi-channel data is acquired by relying on a microphone array and a camera, and an environmental model is constructed in the step 1 to determine a noise baseline and student distribution; in the second step, overlapped voices are detected, and sound source localization is carried out in combination with the time difference of arrival and mouth shape data; in the third step, directional gain is executed in the target direction, and a deep network is used for separating aliasing voice; and in the step 4, the separated voice is input into a children customized recognition engine to complete high-precision recognition and interaction in combination with confidence evaluation. The recognition accuracy and the interaction efficiency can be remarkably improved under the complex scenes of classroom reverberation and simultaneous speaking of multiple persons, meanwhile, the noise change is tracked through the global environment model so that the education robot can keep stable recognition performance in diversified teaching interaction, and the teaching effect is remarkably enhanced.
Owner:北京爱宾果科技有限公司

Voice action synchronization method and device of virtual character, equipment and storage medium

The embodiment of the invention discloses a voice action synchronization method and device for a virtual character, equipment and a storage medium, and relates to the technical field of artificial intelligence, and the method comprises the steps: carrying out the voice generation of an answer text, obtaining an answer voice sequence and the attribute information of each phoneme, determining a standard mouth shape identifier of the corresponding phoneme based on the phoneme identifier of each phoneme, and obtaining a voice action synchronization result; generating a mouth shape animation sequence based on the standard mouth shape identifier of each phoneme according to the time range of each phoneme; determining an emotional action sequence and an emotional action starting timestamp corresponding to the emotional label of the answer text, and generating a supplementary animation sequence based on the emotional action sequence and the emotional action starting timestamp; and synchronizing the answer voice sequence, the mouth shape animation sequence and the supplementary animation sequence according to the time sequence to obtain a synchronization relationship, and driving the virtual response character of the user based on the synchronization relationship, the mouth shape animation sequence and the supplementary animation sequence when the answer voice sequence is played, so that the voice action synchronization function of the virtual character is realized.
Owner:SHANGHAI JIACHE INFORMATION TECH CO LTD

Voice data extraction method and system adopting artificial intelligence

The invention relates to a voice data extraction method and system adopting artificial intelligence, and relates to the technical field of voice intelligent extraction, and the method comprises the steps: monitoring and collecting video and audio data generated by video communication, and obtaining a corresponding data sequence; carrying out speech speed and volume identification on the audio sequence, obtaining characteristic parameters, and respectively configuring mouth shape and audio weight to obtain a first group of weight and a second group of weight; performing mouth shape image proportion analysis on the video sequence, and obtaining a third group of weights and a fourth group of weights in combination with vocabulary analysis of historical voice data of the user; and fusing the four groups of mouth shapes and audio weights, and performing voice content recognition extraction on the audio and video sequence to obtain final voice data. The problems that in the prior art, dependence on voice data extraction is single, the voice data extraction quality is poor, and personalized adaptation is lacked are solved.
Owner:HANGZHOU ZHILIAO INFORMATION TECH CO LTD

Control method and device of voice air conditioner, voice air conditioner and medium

The invention discloses a voice air conditioner control method and device, a voice air conditioner and a medium. The invention relates to the technical field of air conditioners. The method comprises the steps that dialect voice and a mouth shape video corresponding to a dialect instruction are obtained; processing the dialect voice to obtain acoustic features, and processing the mouth shape video to obtain visual features; fusing the acoustic features and the visual features to obtain multi-modal features; acquiring environment data, and determining an activity scene according to the environment data; and according to the multi-mode characteristics, the activity scene and a preset dialect instruction library, whether the voice air conditioner is controlled or not according to the dialect instruction is judged. According to the method, the dialect voice corresponding to the dialect instruction and the feature of the mouth shape video are fused to obtain the multi-mode feature, then the voice air conditioner is controlled according to the multi-mode feature, the activity scene determined by the environment data and the preset dialect instruction library, and the dialect recognition accuracy of the voice air conditioner is improved.
Owner:GREE ELECTRIC APPLIANCE INC OF ZHUHAI

User emotion monitoring intervention method and device based on multiple models and computer program product

According to the multi-model-based user emotion monitoring intervention method and device and the computer program product provided by the invention, the communication style of the target object is simulated by training multiple types of models, a virtual image is synchronously generated to match mouth shapes, expressions and actions, and interaction with more emotion support can be provided for the user. Meanwhile, an emotion analysis model is introduced to carry out emotion recognition, early warning and risk level analysis, and when emotion abnormity is recognized, a pacific dialogue is carried out based on a second model obtained by learning a psychological intervention strategy library and a psychological knowledge graph. And sudden emotional abnormal conditions can be detected and responded in real time, and professional pacifying dialogues can be intervened.
Owner:BEIJING INFORMATION SCI & TECH UNIV

Video generation method and device, computer equipment and storage medium

The embodiment of the invention relates to a video generation method and device, computer equipment and a storage medium, and the method comprises the steps: carrying out the filling processing of each frame of image of an initial video, and obtaining a first image; obtaining a second image from the initial video, and obtaining a target audio; extracting audio features from the target audio, extracting first image features from the first image, and extracting second image features from the second image; performing alignment operation on the first image feature and the second image feature; performing spatial deformation on the second image feature according to the audio feature and the aligned second image feature; generating a mouth shape image according to the deformed second image feature and the first image feature; and generating a target video according to each mouth shape image and the target audio. Therefore, the target video with highly synchronous mouth shape and voice content can be generated while the character identity features are kept, and the naturalness and visual reality of mouth shape and voice matching generation are improved.
Owner:SHANGHAI IQIYI NEW MEDIA TECH CO LTD

Digital human question and answer system

The invention provides a digital human question-answering system, and the system comprises a voice receiving and processing module which is used for receiving the voice input of a user in real time, splitting the input voice, extracting the acoustic features of each frame of voice in real time, and caching the acoustic features of continuous frames to form a streaming data queue; the voice feature reasoning module is used for inputting the streaming data queue into a pre-training question and answer model and outputting a question and answer text result and a voice synthesis instruction in real time; and the digital human synchronous driving module is used for generating a synchronous voice signal based on the question and answer text result and the voice synthesis instruction, mapping the synchronous voice signal to a mouth shape mapping library and an action template library, generating a mouth shape sequence and a limb action sequence matched with the voice time sequence, and completing digital human broadcasting. According to the method, user text or voice input is received, the question and answer result is returned after model reasoning, and the digital person is driven to complete voice broadcasting.
Owner:ZHONGCHUANG (WUHAN) TECH CO LTD

Screen operation method, system and equipment based on air blowing and medium

The invention relates to a blowing-based screen operation method, system and equipment and a medium, and belongs to the field of blowing detection. The method comprises the following steps: acquiring a blowing audio and a face image of a user, and performing face contour detection on the face image based on a face contour recognition algorithm; when the face contour of the user meets a preset detection condition, whether the mouth shape of the user in the face image is a blowing mouth shape or not is judged based on a blowing mouth shape recognition algorithm; when the user is in the blowing nozzle type, the identity of the user is recognized according to the blowing audio of the user based on a blowing voiceprint recognition algorithm; wherein the identity of the user comprises a registered user and a non-registered user; when the identity of the user is a registered user, acquiring a sight focus from the face image based on an eyeball focus recognition algorithm; and executing an operation of the corresponding screen area according to the sight focus. The accuracy of blowing behavior detection is greatly improved, and effective blowing interaction between the screen and the user is achieved.
Owner:ZHEJIANG ZEEKR INTELLIGENT TECH CO LTD +1

Speech recognition and speech synthesis optimization method and system based on large model

The invention provides a voice recognition and voice synthesis optimization method and system based on a large model, and the method comprises the steps: extracting lip motion features, lip feature timestamps and audio feature timestamps through obtaining a real-time voice input signal and an image frame sequence of a user face region, so as to generate an initial space-time offset sequence; generating a dynamic offset compensation parameter sequence based on the initial space-time offset sequence in combination with a reference alignment template in the voice and vision synchronization data set; and performing joint processing on the dynamic offset compensation parameter sequence, the lip motion characteristics and the real-time voice input signal by using a large model to reconstruct a target voice segment, generating a corrected phoneme sequence in combination with the lip motion characteristics, retrieving a mouth shape parameter group corresponding to the corrected phoneme sequence from a phoneme mouth shape mapping rule base, and performing mouth shape correction on the mouth shape parameter group. To generate a voice waveform in phase synchronization with the lip motion; according to the invention, the naturalness and immersion of man-machine interaction and the robustness in a voice missing or delay scene are improved.
Owner:LUSTER LIGHTWAVE CO LTD

Digital human rendering method and device, storage medium and program product

One or more embodiments of the invention provide a digital human rendering method and device, a storage medium and a program product. The digital human rendering method comprises the following steps: analyzing a target voice stream matched with a to-be-rendered digital human to obtain a phoneme sequence and a voice feature vector of each phoneme in the phoneme sequence; obtaining a pre-stored mouth shape vector library, wherein the mouth shape vector library is used for storing a mapping relationship between the voice feature vectors of different phonemes and mouth shape control parameters; based on the voice feature vector of each phoneme in the phoneme sequence, a mouth shape control parameter sequence is retrieved from a mouth shape vector library, and each mouth shape control parameter in the mouth shape control parameter sequence is in one-to-one correspondence with each phoneme in the phoneme sequence; and executing a digital human rendering task based on the mouth shape control parameter sequence.
Owner:HANGZHOU ANT KUAI TECHNOLOGY CO LTD

Multi-language interactive learning system based on speech recognition

The invention relates to the field of voice signal processing, in particular to a multi-language interactive learning system based on voice recognition, which is characterized in that sound waves and mouth shape images are synchronously acquired and discretized by the system, and cross-modal coding is formed after alignment; secondly, the code is injected into a micro-ring photon reserve network through phase modulation to unfold time sequence characteristics, the code is mapped into a quaternion graph to be embedded, and segment boundaries are extracted through a diffusion-pulse coupling method; then, according to a fixed field sequence, encapsulating the quaternion graph embedding and segmentation data into an object mark prompt, inputting the object mark prompt into a low-rank adaptive language model, and generating semantic segmentation data and text transcription; and finally, the adaptive learning module implements sparse gradient updating on the low-rank weight and pulse network by using an integer fractal hash exclusive-or difference mask, and adopts support-query element learning for synchronous iteration after user clarification. According to the system, low-power-consumption, high-precision and second-level accent self-adaptive multi-language voice interaction is realized on the end side.
Owner:SICHUAN COLLEGE OF ARCHITECTURAL TECH

Video figure mouth shape synchronization method based on deep learning

The invention belongs to the technical field of computer vision and artificial intelligence, and particularly relates to a video figure mouth shape synchronization method based on deep learning. According to the method, high-precision and low-delay lip action generation is achieved through multi-modal feature fusion, a generative adversarial network (GAN) and a differentiable rendering technology, the method is suitable for scenes such as film and television post-production, virtual reality (VR) real-time interaction, voice-driven animation generation and multi-language video translation, the synchronization error of the method on a standard data set is reduced by 62.5%, and the real-time performance of the method is improved. The method supports 30fps real-time processing, has strong noise robustness and multilingual adaptability, and can be widely applied to film and television production, virtual reality and real-time interaction scenes.
Owner:ZHE JIANG YAN HUANG KE JI YOU XIAN GONG SI

Video parameter adjustment method and apparatus

The present application provides a video parameter adjustment method and apparatus. The method comprises: acquiring a first video segment, the first video segment comprising a digital person, and the digital person being generated by a neural network model; determining a video quality monitoring index on the basis of the first video segment, the video quality monitoring index comprising at least one of the following: the accuracy of the mouth shape of the digital person, the degree of synchronization between the picture and sound of the digital person, and the degree of consistency of the figure of the digital person; and determining a parameter of the neural network model on the basis of the video quality monitoring index, the parameter being used for generating the digital person in a second video segment. In the described method, a quality monitoring index related to a digital person in a livestreaming video can be monitored and corresponding feedback adjustment settings are executed, thereby improving the quality of the livestreaming video of the digital person.
Owner:HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD

Real-time digital human audio and video synchronization method and device based on TRTC

The invention is suitable for the technical field of real-time audio and video communication and digital human interaction, and provides a real-time digital human audio and video synchronization method and device based on TRTC. In the embodiment of the invention, for the voice data needing to be synchronized, the device firstly converts the voice data into the audio stream and uploads the audio stream to the cloud server, and then the cloud server generates the lip-shaped video stream according to the lip-shaped action video clip pre-generated based on the digital human image; the method comprises the following steps: receiving a lip-shaped video stream, generating a second audio stream of which the timestamp is synchronous with that of the lip-shaped video stream, returning the lip-shaped video stream and the second audio stream to equipment, and finally playing by the equipment according to the timestamps in the lip-shaped video stream and the second audio stream, thereby realizing audio and video synchronization. Therefore, the system delay of the real-time dialogue of the digital person is greatly reduced, the phenomenon that the sound and the mouth shape are not synchronous is effectively avoided, and the problem that low delay and high-precision mouth shape synchronization are difficult to consider when the digital person type and the voice are synchronous is solved.
Owner:HONGYING GROWTH (HANGZHOU) TECHNOLOGY CO LTD

Device and method for generating avatar lip-sync animation based on multimodal biosignals

The present disclosure relates to a device and method for generating avatar lip-sync animation based on multimodal biosignals, The device comprises a multimodal data collection unit configured to collect data including biosignal data including brain waves when a user imagines speaking and image data; a preprocessing unit configured to preprocess the multimodal data; a feature extraction unit configured to extract feature vectors including the user's biosignal feature and facial feature from the preprocessed multimodal data; an avatar generation unit configured to generate an avatar; a lip-sync reconstruction unit configured to predict the mouth shape and facial movement when the user imagines speaking by inputting the extracted feature vectors to a pre-prepared lip-sync reconstruction model; and a lip-sync animation implementation unit for implementing an avatar lip-sync animation by applying the mouth shape and facial movement predicted by the lip-sync reconstruction unit to the avatar generated by the avatar generation unit.
Owner:KOREA UNIV RES & BUSINESS FOUND

Orthodontic adhesive

ActiveCN309697184SIngrown nailDentistry
1. The name of the design product: the mouth shape of the nail paste. 2. The use of the design product: paste on the surface of the nail with adhesive, correct the ingrown nail with its own rebound tension, and treat paronychia. 3. The design points of the design product: in shape. 4. The picture or photo that best shows the design points: perspective drawing.
Owner:何浩强

Mandarin pronunciation evaluation system based on deep learning

The invention discloses a mandarin pronunciation evaluation system based on deep learning, and relates to the technical field of voice scoring, and the system comprises a data input port which is used for obtaining a copy, a standard pitch curve and a standard mouth shape video, shooting the face mouth shape motion and sound of a user during reading into a video, and extracting an audio from the video; the data processor is used for processing videos and audios; the tone evaluator is used for evaluating the audio to obtain a tone score; the mouth shape evaluator is used for obtaining a mouth shape score; the pronunciation evaluator is used for obtaining a pronunciation score; and the score output end is used for combining the tone score, the mouth shape score and the pronunciation score to generate a final score. The mandarin pronunciation scoring method and device have the effect of improving the accuracy and adaptability of mandarin pronunciation scoring.
Owner:AI TUER