Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

8205results about "Animation" patented technology

AI visual special effect dynamic generation system fused with multi-modal perception

The invention belongs to the technical field of visual special effects, and discloses an AI visual special effect dynamic generation system fusing multi-modal perception. Through cooperative work of six core modules of multi-modal input processing, importance analysis, parameter configuration, parameterized rendering, parameter optimization and rendering and output, deep fusion of special effects and contents is realized. An image sequence, audio data and scene parameters can be processed at the same time, a multi-modal feature data set is constructed, a scene key visual area is scientifically recognized, a visual importance distribution map is generated, and the special effect parameter configuration and rendering process is guided. An improved neural radiation field technology and a physical constraint model are adopted to ensure the consistency and reality of the special effect at different visual angles; and a multi-dimensional quality evaluation and parameter automatic optimization mechanism is introduced, so that the visual expressive force and the artistic value of the special effect are guaranteed. The creation threshold is remarkably reduced, the cooperative expression ability of the special effect and the content is improved, and the special effect becomes a powerful tool for enhancing the narration and enhancing the emotion.
Owner:SHENZHEN XINGHUO MUTUAL ENTERTAINMENT DIGITAL TECH CO LTD

Intelligent real-time interactive question-answering system based on virtual digital human

The invention provides an intelligent real-time interactive question-answering system based on a virtual digital human, and belongs to the technical field of voice signal processing and voice recognition, and the system comprises a data acquisition module which receives a voice or text interaction request input by a user, collects the expression dynamic parameter sequence and limb movement sequence data of the user in real time, and transmits the data to a user interaction module; obtaining a standardized voice feature vector and structured text data; the cross-modal fusion module is used for constructing an interactive feature matrix; the behavior decision module outputs a decision instruction set; the knowledge retrieval module is used for generating an answer text with emotional adaptability and voice features; and the voice generation module is used for generating a mouth shape animation key frame, a micro expression parameter sequence and a limb action track of the virtual digital human, generating a voice response in combination with the answer text and the voice characteristics, and pushing the voice response to the user terminal. According to the method, the interaction experience and adaptability of the virtual digital human are remarkably improved.
Owner:XIAMEN DUOXIANG ANIMATION CO LTD

Digital twin processing method and system, and cloud platform

The present invention relates to a digital twin processing method and system, and a cloud platform. The method comprises: acquiring production system elements, carrying out abstraction definition and parameterization description on the production system elements by means of digital-to-analog knowledge fusion of a mathematical model and a mechanistic model, so as to construct a digital twin ontology model; analyzing and reconstructing model data to obtain a mapping model of which object variables can be directly accessed and operated by a collective motion control method, so that the model is visualized at the cloud; and using an external data source to drive parameter update and operation matching of the model by means of a motion control method, so as to complete cooperative deployment and synchronous evolution of an actual physical device and the model in the production process on a cloud server. According to the present invention, a model is constructed by means of digital-to-analog knowledge fusion of a mathematical model and a mechanistic model, the model is mapped to achieve motion visualization, model parameter update and operation matching on the cloud are achieved, and then cooperative deployment and synchronous evolution of a physical device and the model are completed.
Owner:HAINAN UNIV

Intelligent multi-mode virtual digital human interaction system based on AI language large model, interaction method and application

The invention discloses an intelligent multi-modal virtual digital human interaction system based on an AI language large model. The system comprises a high-authenticity face generation module; the high-authenticity face generation module uses an AdaAN network, based on adaptive feature fusion and voice driving and time sequence modeling of voice features, feature information related to voice is extracted, the extracted voice features are processed through a deep neural network, it is ensured that the voice and facial expressions are highly aligned in time and space, and the face recognition accuracy is improved. Collecting a bio-electricity signal, mapping the signal to facial muscle movement, generating a final facial expression, and interacting with a user; the system further comprises an intelligent interaction module, a training optimization and efficient generation module, an efficient integration module, a multi-modal data acquisition module, an AI large model core processing module, a digital human image generation and driving module, an interaction scene adaptation module and a feedback optimization module. The invention further discloses a multi-mode digital human interaction method which has wide application value.
Owner:EAST CHINA NORMAL UNIV

Dynamic animation based on waiting period

ActiveUS20250232503A1Character and pattern recognitionAnimationAnimationWaiting period
An example operation may include one or more of receiving context of a user during an inquiry of a feature via a software application, executing a waiting period via the software application, during the waiting period, selecting an animation to display via the software application based on the context of the user and the feature inquiry wherein the animation provides contextual data associated with the feature, wherein the contextual data is based on a determined need of the user, displaying the animation via the software application during the waiting period, and determining if the user has accepted the feature via the software application. At least one portion of the example operation: integrates with an artificial intelligence (AI) chatbot, interacts with the AI chatbot, is performed by the AI chatbot, and / or is associated with an AI model.
Owner:THE TORONTO DOMINION BANK

Digital human video generation method based on multi-modal large model

The invention belongs to the technical field of virtual person generation, and particularly relates to a digital person video generation method based on a multi-modal large model, and the method comprises the following steps: 1, constructing a multi-modal data system; 2, multi-modal large model training and adaptation are carried out; 3, constructing a digital human three-dimensional model; step 4, performing semantic analysis and modal mapping; 5, generating a time sequence action and a mouth shape; step 6, building and rendering a virtual scene; step 7, audio and video synchronous rendering and synthesis; step 8, quality optimization and defect repair; and step 9, performing user interaction and iterative optimization. Through technical innovation and engineering, the core pain point in digital human video generation is solved, efficient, vivid and customizable content production capacity is provided for virtual anchors, intelligent customer service, enterprise training and other scenes, and the AI digital human technology is promoted to be applied to large-scale business from experiments.
Owner:ZHE JIANG YAN HUANG KE JI YOU XIAN GONG SI

Rehabilitation training action evaluation method and device based on multi-view vision

The invention provides a rehabilitation training action evaluation method and device based on multi-view vision. According to the method, a multi-view camera is adopted to synchronously collect a rehabilitation training image sequence, two-dimensional coordinates of key points of a human body are extracted by improving an HRNet deep learning model, three-dimensional reconstruction is performed in combination with a Gaussian process, motion feature data are extracted, and rehabilitation training quality is evaluated. Accurate motion capture without wearing mark points by the patient is realized, the training constraint feeling of the patient is effectively reduced, the rehabilitation evaluation accuracy is improved, and a personalized rehabilitation scheme can be generated.
Owner:JILIN UNIVERSITY

Method of generating fullbody animatable person avatar from single image of person, computing device and computer-readable medium implementing the same

PendingUS20250209712A1Image enhancementImage analysis
A computer-implemented method of generating fullbody animatable avatar of a person from a single image of the person includes: obtaining an image of a person body and a parametric body model defined by pose parameters and shape parameters of the person body in the image, and by camera parameters used when capturing the image; defining, based on the parametric body model, a texturing function including a mapping between each pixel corresponding to a part of the person body shown in the image and corresponding texture coordinates in a texture space, and corresponding texture coordinates in the texture space for a part of the person body not shown in the image; sampling RGB texture of the person body based on the mapping and obtaining a map of sampled pixels.
Owner:SAMSUNG ELECTRONICS CO LTD

Avatar JSON interchange file format

Some embodiments of a method may include: obtaining Avatar JSON Interchange File (AJIF) data; decoding the AJIF data to generate avatar data; obtaining animation parameters from the avatar data; obtaining user movement data corresponding to a movement of a user; generating updated avatar data based on the user movement data; and sending the updated avatar data to a client device. For some embodiments, a method may include: generating Avatar JSON Interchange File (AJIF) data corresponding to an avatar; sending, to an application server, the AJIF data; obtaining user movement data corresponding to a movement of a user; sending, to the application server, the user movement data corresponding to the movement of the user; receiving, from the application server, scene update data, wherein the scene update data includes updated avatar data corresponding to the movement of the user; and rendering a scene update based on the scene update data.
Owner:INTERDIGITAL CE PATENT HOLDINGS SAS

Method for real-time generation of empathy expression of virtual human based on multimodal emotion recognition and artificial intelligence system using the method

Provided are a conversational artificial intelligence (AI) system and method based on real-time multimodal emotion recognition. The system includes a model server configured to provide a machine learning-based conversational model, a terminal configured to perform a conversation with the machine learning-based conversational model through the model server, display a virtual human responding to a user during a conversation with the user, and capture a facial image of the user during the conversation, and a multimodal empathetic conversation-generation system configured to access the model server and receive a response to a question of the user from the terminal, and assess an emotion of the user from the facial image of the user and control, based on the assessed emotion, an expression of the virtual human displayed on the terminal.
Owner:SANGMYUNG UNIV IND ACAD COOP FOUND

Face-translator: end-to-end system for speech-translated lip-synchronized and voice preserving video generation

A neural end-to-end system is provided for the face and voice preserving translation of videos. The system is a pipeline of multiple models that produces a video of the original speaker speaking in the target language with modified lip movement to match the target speech, while preserving emphases and prosody of the original speech, and voice characteristics of the original speaker. The pipeline starts with automatic speech recognition including emphasis detection, followed by the translation model. The translated text is then synthesized by a Text-to-Speech model that recreates the original emphases in the target sentence. The resulting synthetic speech is then converted back to the original speakers' voice using a voice conversion model. Finally, to synchronize the lips of the speaker with the translated audio, a generative model generates frames of adapted lip movements which are combined with the audio to produce the final output. The disclosure further describes several use-cases and configurations that apply these techniques to video conferencing, dubbing, low-bandwidth transmission, speech enhancement and assistive technology for the hearing impaired.
Owner:WAIBEL ALEXANDER

Video generation method and interaction method based on digital human, and device, storage medium and program product

Provided in the embodiments of the present application are a video generation method and interaction method based on a digital human, and a device, a storage medium and a program product. In the embodiments of the present application, text-to-speech processing is performed on the basis of voice features of a user and an emotion label, speech-to-expression processing is performed on the basis of a mapping relationship between the voice features of the user and expression coefficients, and a digital human model is rendered on the basis of speech signals and the expression coefficients, so as to obtain video data of the digital human model. Thus, voice features of a user are accurately simulated, so as to ensure that a speech output of a digital human sounds natural and is also highly personalized, thereby realizing personalized driving of the digital human, and improving the realism of the digital human in terms of voice and dynamic images. Thus, the user experience is improved, and the interactivity of the digital human and the authenticity and immersion are enhanced.
Owner:TAOBAO CHINA SOFTWARE

AI-based animation sub-mirror script automatic generation and visual preview method and system

The invention discloses an AI-based animation split script automatic generation and visual preview method and system, and the method comprises the following steps: 1, receiving a natural language script text inputted by a user, the natural language script text comprising scene description, role action, dialogue and shot indication information; step 2, performing semantic analysis and structured analysis on the script text based on a natural language processing technology, and identifying and extracting key narrative elements; by introducing an artificial intelligence technology, end-to-end automatic generation and interactive optimization from a character script to a dynamic split rehearsal video are realized, the system can deeply understand scenes, actions, role emotions and shot languages in the script, corresponding visual elements are automatically matched and generated, and the dynamic split rehearsal effect is improved. And the timeline and the rhythm conforming to the film and television grammar are constructed, so that the efficiency and the consistency of the split creation are greatly improved, and the professional threshold and the manufacturing cost are reduced.
Owner:NEW AXIS ANIMATION TECHNOLOGY DEVELOPMENT (BEIJING) CO LTD

Intelligent surveying and mapping method and system based on AI and BIM fusion

The embodiment of the invention discloses an intelligent surveying and mapping method and system based on AI and BIM fusion. The method comprises the steps that an unmanned aerial vehicle platform carrying a laser radar, an RGB camera and a positioning system is used for scanning ancient building cultural relics and surroundings in a multi-angle flight mode, point cloud data, multi-view image data and position and attitude data are synchronously collected, and the three are associated through timestamps; after the point cloud data and the multi-view image data are preprocessed, cross-modal registration is completed through feature matching and pose estimation in combination with the position and pose data, and a registration data set is obtained; semantic segmentation is carried out on the point cloud data and the image data in the registration data set, and semantic segmentation results are fused based on the incidence relation; classifying and aggregating the original point cloud components according to category labels, constructing a topological relation reasoning assembly relation, calling corresponding BIM template instantiation model components based on the assembly relation, and hooking a segmentation result to generate a semantic enhanced BIM model; and integrating the BIM model and the GIS base map to form a fusion model so as to plot the historic building cultural relics.
Owner:XIAN UNVERSITY OF ARTS & SCI

Game animation character display method based on virtual reality technology

The invention provides a game cartoon character display method based on a virtual reality technology. The method comprises the following steps: acquiring three-dimensional model data of a game cartoon character in a target display area through a user interaction terminal; the virtual reality content management platform determines a dynamic rendering precision level according to the model data, and generates a real-time rendering strategy including a model patch reduction coefficient, a texture compression rate and a skeleton animation updating frequency in combination with terminal performance parameters; the strategy is sent to a virtual reality supervision platform and a user interaction terminal, and a virtual reality rendering engine platform is instructed to execute real-time rendering; and the virtual reality supervision platform monitors the frame rate fluctuation data of the head-mounted display device, calculates the scene rendering stability, sends a rendering optimization instruction if the scene rendering stability is lower than a threshold value, and adjusts the model data acquisition frequency and the video memory cleaning period to optimize the performance. The rendering efficiency and the system stability can be improved, and the hardware load and the frame rate fluctuation are reduced.
Owner:JIANGSU JIUQU INTERACTIVE ENTERTAINMENT NETWORK TECHNOLOGY CO LTD

Digital delivery topology mapping method and system for multi-source real-time data fusion

The invention belongs to the field of digital delivery, and particularly relates to a digital delivery topology mapping method and system for multi-source real-time data fusion, and the method comprises the steps: obtaining factory building distribution, equipment distribution and operation control logic and preset function block operation logic, and constructing a hierarchical clustering function mapping space through combining an association analysis and clustering algorithm; in response to a target function demand, obtaining a layered response mapping path in combination with a deep search algorithm; layered synchronous response and distributed node anomaly monitoring are realized based on the path, the three-dimensional simulation model and the display system equipment performance and the network state. Tracing abnormities based on a monitoring result in combination with a hidden Markov algorithm and a forward reasoning model, performing iterative verification after conflict resolution until the function is free of abnormities, and updating a mapping space; and adjusting the demand repeating steps to obtain a complete and updated mapping space, and realizing accurate function and picture collaboration under multi-source data fusion.
Owner:NANJING CHANCE ENG TECH SERVICES INC

Methods and systems of text-conditioned audio-visual speech generation with multi-modal latent diffusion models

Methods, systems, and computer programs are presented for audio-visual speech generation with multi-modal latent diffusion models. One method includes encoding raw audio signals and video frames into respective latent spaces using audio and visual autoencoders. A text transcript is processed into phoneme sequences using a text transcript processor. The audio and visual latent spaces are conditioned using the text transcript and a conditioning variable. Joint distributions of the visual and audio latent spaces, text transcript, and conditioning variable are learned using a multi-modal latent diffusion model. The model adds noise to the latent audio-visual representations and predicts the noise through denoising neural networks. An inverted diffusion process is utilized to generate diverse speech content and speaker characteristics, resulting in realistic audio-visual speech. The technology presented provides a novel approach to conditional speech generation with potential applications in speech synthesis, voice conversion, and speech recognition.
Owner:TENSORTYPE INC

Virtual human design and application platform and method based on artificial intelligence, equipment and medium

The invention provides a virtual human design and application platform, method and device based on artificial intelligence, and relates to the technical field of virtual digital humans. The method comprises the steps of performing local anonymization on multi-modal input data on user equipment, encoding generated anonymized multi-modal features to obtain a multi-modal feature vector, and inputting the multi-modal feature vector into an emotion calculation model to obtain a user emotion intensity quantized value; inputting the multi-modal feature vector into a context sensing model, and generating a user intention vector after context correction in combination with a knowledge graph; generating an updated personality parameter matrix according to the user emotion intensity quantized value and the user intention vector; and outputting voice waveform data, facial muscle motion parameters and skeleton joint coordinate data based on the personality parameter matrix, and driving the virtual digital human three-dimensional model to perform real-time rendering. According to the scheme, the naturalness, emotional resonance and long-term user retention rate of virtual digital human interaction can be improved, and user privacy data security is protected.
Owner:郑雯月

Devices, systems, and methods for machine consciousness

Certain aspects of the disclosure generally relate to devices, systems, applications, and / or objects of applications, and may be generally directed to artificial learning and / or use of artificial knowledge. Other aspects of the disclosure generally relate to consciousness, and may be generally directed to learning and / or implementing one or more purposes. One or more purposes may drive the use of artificial knowledge in implementing the one or more purposes. Therefore, in some aspects, a conscious device, system, application, and / or object of application may include one or more purposes and artificial knowledge so that the device, system, application, and / or object of application can act upon a world in implementing the one or more purposes. The disclosure also describes other functionalities.
Owner:COSIC JASMIN

Commodity display interaction visualization method and device

The invention relates to the field of commodity visualization, in particular to a commodity display interaction visualization method and device. The method comprises the following steps: collecting a multi-azimuth image of a commodity, carrying out three-dimensional texture modeling, and constructing a three-dimensional texture mapping model; performing material light rendering on the three-dimensional texture mapping model to generate a material rendering result; collecting an environment detection image of a commodity display environment, and performing environment illumination adaptation compensation on a material rendering result to obtain an illumination compensation rendering commodity; carrying out attribute information visual layout on the illumination compensation rendering commodity to obtain a commodity visual space; and carrying out interaction response animation analysis according to the commodity visualization space, carrying out multi-target parallel rendering, and executing commodity interaction visualization operation. The form and surface details of the commodity in the real world are accurately restored, the visual reality sense is improved, and the interactive experience feeling of browsing the commodity by a user is enhanced.
Owner:SHENZHEN XIAOYI SHUZHI TECH CO LTD

System for generating conversational content by utilizing generative ai and method thereof

The invention discloses a system (100) for generating conversational content using a generative artificial intelligence (AI), said system (100) comprising: a user (101), an administrator (102), an application programming interface (API) server (103), a generative artificial intelligence (AI) server (104), a plurality of databases (105), a generative artificial intelligence (AI) processor (106), an audio generate processor (107), a text-to-speech processor / service provider (108), a video generation service (109), a video generation processor (110), and a memory communicatively coupled to the processor, wherein the memory stores processors instructions, which, on execution, causes the processor to generate at least one of conversational script, audio, video, or combination thereof. The system (100) allows users to create and customize various aspects of conversational content, including characters / personas / speakers, groups (of personas / characters / speakers), tones, content types, topics, conversation formats, and tone.
Owner:SINGH HEMENDRA +1

Controllers in MPEG avatar representation format

Some embodiments of a method may include: obtaining a MPEG Avatar Representation Format (MARF)-based file; parsing the MARF-based file into a data structure; selecting an asset from the data structure; verifying the asset comprises a controller set property; obtaining controller set data corresponding to the controller set property; obtaining container data using the controller set data; determining a mime type using the controller set data; decoding the container data using the mime type; obtaining a mesh and a skeleton using the container data; and animating an avatar corresponding to the mesh and the skeleton.
Owner:INTERDIGITAL CE PATENT HOLDINGS SAS

Virtual-real fusion exhibition display interaction system and multi-mode perception method

The invention discloses a virtual-real fusion exhibition display interaction system and a multi-mode perception method, and belongs to the technical field of exhibition display interaction. The system collects audience eyeball fixation points, gesture actions and ambient light data through AR / VR equipment, analyzes coordinates of a region of interest through an eyeball fixation point attention mechanism, a gesture space-time encoder and a multi-modal fusion unit, triggers holographic projection explanation and virtual exhibition stand light and shadow dynamic adjustment (including illumination intensity, color, Gaussian blur and the like) based on a threshold value, and performs real-time display on the virtual exhibition stand. And multi-user collaborative interaction is realized through federal learning. According to the method, reinforcement learning is adopted to optimize an event-driven threshold value, and virtual and real visual splitting is eliminated in combination with ambient light adaptive mapping. The problems of low participation degree, insufficient single-mode interaction information and multi-user cooperation of traditional exhibition are solved, interest analysis accuracy is improved through multi-mode fusion, personalized experience is enhanced through dynamic interaction, the method is suitable for multiple scenes such as museums and science and technology museums, and exhibition intellectualization, immersion and group interaction efficiency are effectively improved.
Owner:SUZHOU ART & DESIGN TECH INST

Personalized digital human generation method based on single video

The invention discloses a personalized digital human generation method based on a single video, and relates to the field of virtual digital human modeling and driving, and the method aims at the video and voice data of a target person, through introducing a multi-modal alignment constrained voice driving synchronization mechanism and combining semantic understanding and an emotion label expression generation model, a personalized digital human model is generated. High synchronization, nature and vividness of digital human facial expressions and voice contents are realized. According to the method, a self-supervised style consistency constraint is added in model training to ensure that the style of a generated character image is stable, and a 3D semantic mask is added in an image fusion stage to improve the fusion precision and realistic effect of a synthetic facial expression and a reference face. According to the method, the digital human can be rapidly cloned and driven to synthesize the expression through the voice only through a single video sample, the generation process is efficient, and the obtained digital human video has excellent sense of reality and interactivity.
Owner:LIANGSHENG DIGITAL CREATIVE DESIGN (HANGZHOU) CO LTD

Three-dimensional attitude high-precision acquisition and reconstruction system for exercise training

The invention discloses a three-dimensional attitude high-precision acquisition and reconstruction system for exercise training, which relates to the technical field of three-dimensional attitude acquisition and reconstruction and comprises a camera calibration module, an image acquisition module, a key point extraction module, a three-dimensional reconstruction module, an attitude optimization module, a data fusion module and a feedback generation module. The camera calibration module is used for performing internal and external parameter calibration on a multi-view camera assembly by adopting a self-adaptive joint calibration method to obtain a projection matrix parameter and a transformation matrix of each camera; the image acquisition module is used for synchronously acquiring multi-view video images of the athlete in the training process based on the calibration parameters, and performing timestamp alignment processing to obtain an original image sequence; and the key point extraction module is used for detecting and identifying a human body image by adopting a deep learning model, extracting two-dimensional coordinates of key human body joint points of each frame from an original image sequence, and obtaining a two-dimensional key point set under each view angle.
Owner:洪永帅

Non-perpetual culture immersive interactive experience device and system based on virtual digital human

The invention discloses a non-perpetual culture immersive interactive experience device and system based on a virtual digital human, and relates to the technical field of intangible cultural heritage intelligent interaction, and the device comprises a multi-mode sensing module which collects data such as user actions and voices through a depth camera and the like; the non-perpetual knowledge base module stores a non-perpetual knowledge graph and a case library; the digital human modeling engine fuses the inheritor characteristics and the user data to generate a virtual digital human; the immersion interaction engine constructs a cross-platform rendering environment to realize five-sense fusion experience; and the adaptive learning module analyzes the interaction sequence to optimize the digital human behavior, and all the modules cooperate to realize the intelligent interaction experience of non-genetic culture. Through the multi-modal perception module, the digital human modeling module and the like, the immersive interactive experience of the non-perpetual culture is realized, the user behavior can be accurately captured, the characteristic virtual digital human is generated, the five-sense fusion experience is provided, the content can be optimized based on the user interaction, the inheritance and propagation of the non-perpetual culture are promoted, and the participation degree and the sense of identity of the user are improved.
Owner:HUNAN INSTITUTE OF ENGINEERING

Avatar creation user interface

The present disclosure generally relates to creating and editing avatars, and navigating avatar selection interfaces. In some examples, an avatar feature user interface includes a plurality of feature options that can be customized to create an avatar. In some examples, different types of avatars can be managed for use in different applications. In some examples, an interface is provided for navigating types of avatars for an application.
Owner:APPLE INC