Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

4315results about "Animation" patented technology

AI-based animation sub-mirror script automatic generation and visual preview method and system

The invention discloses an AI-based animation split script automatic generation and visual preview method and system, and the method comprises the following steps: 1, receiving a natural language script text inputted by a user, the natural language script text comprising scene description, role action, dialogue and shot indication information; step 2, performing semantic analysis and structured analysis on the script text based on a natural language processing technology, and identifying and extracting key narrative elements; by introducing an artificial intelligence technology, end-to-end automatic generation and interactive optimization from a character script to a dynamic split rehearsal video are realized, the system can deeply understand scenes, actions, role emotions and shot languages in the script, corresponding visual elements are automatically matched and generated, and the dynamic split rehearsal effect is improved. And the timeline and the rhythm conforming to the film and television grammar are constructed, so that the efficiency and the consistency of the split creation are greatly improved, and the professional threshold and the manufacturing cost are reduced.
Owner:NEW AXIS ANIMATION TECHNOLOGY DEVELOPMENT (BEIJING) CO LTD

Controllers in MPEG avatar representation format

Some embodiments of a method may include: obtaining a MPEG Avatar Representation Format (MARF)-based file; parsing the MARF-based file into a data structure; selecting an asset from the data structure; verifying the asset comprises a controller set property; obtaining controller set data corresponding to the controller set property; obtaining container data using the controller set data; determining a mime type using the controller set data; decoding the container data using the mime type; obtaining a mesh and a skeleton using the container data; and animating an avatar corresponding to the mesh and the skeleton.
Owner:INTERDIGITAL CE PATENT HOLDINGS SAS

Virtual historical character dialogue method and system with role knowledge and context awareness

The invention discloses a virtual historical character dialogue method and system with role knowledge and context awareness, and relates to the technical field of man-machine interaction, and the method comprises the steps: constructing a multi-level role depth model; when a question of a current user is received, identifying information of a virtual scene where the current user is located, analyzing micro-expressions of the face of the user and voice rhythm characteristics of speech of the user, analyzing an emotional state and an interaction intention of the user based on a multi-modal fusion algorithm, and generating a user state vector; executing a dynamic Prompt construction program, extracting related information from the multi-level role depth model and the user state vector, and generating a structured Prompt; and inputting the structured Prompt into a large language model, generating a reply text conforming to role features based on questions of the current user, and driving a virtual character model. The method solves the problem that in the prior art, virtual historical figures cannot provide real immersion and credible interactive experience with emotional connection.
Owner:BEIJING GROWLIB TECH CO LTD

System and method for emotionally intelligent, personalized AI avatar-based health coaching using multi-domain data and adaptive behavioral intelligence

A programmatically generated AI avatar includes a customizable personality module, acting as the embodied interface for a powerful AI “mind” that delivers personalized coaching to improve user health, well-being, and longevity. The system uses machine learning, large language models, and biometric modeling to synthesize real-time, multi-modal health data—including sleep, nutrition, glucose, mood, and activity—and generate forward-prescribed KHAs. Unlike human coaches, it continuously adapts based on context and behavior, targeting the root cause: metabolic dysfunction—namely by restoring healthy, sustainable body composition through the preservation or building of lean muscle mass and reduction of excess fat. KHAs can also be shared with friends or programmatically generated AI avatars, allowing for coordinated action, emotional support, and accountability through social connection—further reinforcing positive behavior and adherence. The system's reinforcement learning engine incorporates both individual response data and anonymized population-level insights to optimize recommendations over time, learning which interventions are most effective for users with similar physiological and behavioral profiles. First validated with Olympic athletes—resulting in measurable improvements and medal-winning outcomes—this system offers a scalable, emotionally intelligent coaching engine that exceeds human capability, designed for the ultimate purpose of supporting sustainable health, resilience, and human thriving.
Owner:GOLD AI LLC

System and method for multi-modal ai conversational interface improving website navigation and user interaction

The present invention relates to a system for transforming static websites into artificial intelligence (AI)-enabled interactive multi-modal conversational platforms. The system comprises a computing device having a processor for receiving user queries as text or speech input through an input module cooperating with a speech-to-text module. A natural language processing (NLP) module interprets intent, classifies user context, and retrieves grounded information from multiple webpages. A persona adaptation module dynamically modifies vocabulary, tone, and avatar representation across roles such as sales assistant, recruiter, educator, healthcare professional, etc. A response generator module produces structured natural language output, transmitted to a text-to-speech synthesis module and an avatar generation module to render synchronized lifelike video responses. An output rendering module displays multi-modal responses include text, audio, and video, thereby enabling direct navigation and escalation beyond limitations of conventional static websites.
Owner:NALLAM SREE RAMA CHANDRA MURTY

AI digital human interactive response method based on large language model

The invention discloses an AI digital human interactive response method based on a large language model, and relates to the technical field of digital human interaction, and the method comprises the steps: analyzing collected user voice data and visual data through a natural language processing method, generating a cross-modal feature vector, carrying out the cross-modal association analysis of the cross-modal feature vector, and carrying out the cross-modal association analysis of the cross-modal feature vector. Generating a semantic association topological graph; calculating a vertex coordinate and a joint activity threshold value of the semantic association topological graph through high-digital human correlation, inputting the vertex coordinate and the joint activity threshold value into a constructed coordinate index database to execute attention weight calibration, and outputting a multi-dimensional association graph; and performing information density analysis based on the multi-dimensional association map, generating an information density gradient vector field, and dividing a high-density core region and a low-density edge region, the high-density core region generating a semantic core coding tensor, and the low-density edge region generating an edge feature package. According to the method, the cross-modal fusion vector is converted into the cross-modal feature vector, so that the modeling of the cross-modal association relationship is realized.
Owner:BEI JING XIN ZHI YUAN LANG WANG LUO KE JI YOU XIAN GONG SI

Computer implemented system and method for automatically generating offer ranges for candidates in an interviewing process

A computer implemented system and method for generating offer ranges for candidates in an interviewing process is disclosed. The system generates an AI-based interviewer simulating human-based interactions for conducting an ongoing interview with candidates. The system analyzes data associated with candidates obtained during ongoing interview. The system process analyzed responses of candidates to determine contextual attributes associated with responses using ML models. The system automatically generates follow-up interview questions to be delivered to candidates during ongoing interview based on analyzed responses from candidates, by applying AI model to contextual attributes associated with responses. The system generates recruitment scores for candidates based on analyzed responses, contextual attributes, and interpreted non-verbal cues, associated with candidates, using AI model. The system generates offer ranges for candidates based on recruitment scores using AI model. The system provides information associated with selected candidates, and offer ranges generated for selected candidates, to users.
Owner:TALVIEW INC

Interactive automatic explanation method for converting traditional video into artificial intelligence digital human

The invention provides an interactive automatic explanation method for converting a traditional video into an artificial intelligence digital human, and relates to the technical field of artificial intelligence, and the method comprises the steps: obtaining original video data and an audio track, and carrying out the semantic analysis of the audio track, and obtaining multi-mode deconstruction data; generating an explanation script for each time period of the video based on the explanation text, and performing timestamp labeling on the visual elements to form a time sequence synchronization data structure; in the playing process, a virtual image generator is driven to synthesize digital human dynamic expression output in real time according to the current playing time point; after a user interruption request is received, semantic matching is carried out on a query intention in the explanation script, a target explanation fragment and visual elements are positioned, and complementary explanation content is generated; and driving the virtual image generator to synthesize dynamic output synchronized with the supplementary explanation, and after interaction is completed, recovering playing or skipping to a specified time point according to a user instruction. According to the invention, the conversion from the traditional video to the interactive intelligent explanation video is realized, and the watching experience and learning efficiency of the user are improved.
Owner:BEIJING MENGKE TECH CO LTD

Intelligent agent digital image interaction generation method based on multi-modal perception

The invention discloses an intelligent agent digital image interaction generation method based on multi-modal perception, which comprises the following steps: collecting multi-modal input data of a user, and respectively carrying out preprocessing and feature extraction on the multi-modal input data; inputting to an improved efficient modal cross learning network, and carrying out multi-modal feature fusion processing; constructing a semantic intention map, introducing a time index edge weight and an emotion driving edge weight, and encoding the map by using a structure perception map neural network; a modal style vector is extracted through a cross-modal style contrast learning mechanism, and a personalized style coding vector is generated through a hierarchical nested structure; inputting a personalized regulation and control gating mechanism, and regulating and controlling the middle layer representation in the interaction strategy generation process by adopting a feature channel linear modulation method; inputting the representation vector into a behavior strategy generation module to generate a multi-modal behavior output sequence; and the sequence is output to drive the digital image to perform synchronous response, and natural response generation in the user interaction process is completed.
Owner:JIANGSU ELECTRIC POWER INFORMATION TECH

Virtual human real-time generation method and system based on expression control embedding space

The invention relates to a multi-modal virtual human real-time generation method based on an expression control embedding space, and belongs to the field of artificial intelligence. According to the method, an expression control embedding space is constructed and used for fusing voice semantics, a rhythm structure and multi-dimensional emotion information, and continuous and controllable multi-modal driving vectors are generated. The whole system has an end-to-end linkage mechanism from audio input to expression and action output. Semantic features, rhythm structures and emotional states jointly act on generation paths of lip and upper body postures and expression modalities, and all modal features are fused and expressed in a unified control space through a collaborative coding and time sequence alignment mechanism. And finally, a high-consistency and high-fidelity virtual human video is generated in real time through an output scheduling mechanism. The method has remarkable advantages in the aspects of modal fusion consistency, generation expression naturalness and emotion control flexibility, and can be widely applied to key scenes such as virtual human broadcasting, voice interaction agency and meta-universe digital identity construction.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Digital face image generation method and device and electronic equipment

The invention provides a digital face image generation method and apparatus, and an electronic device. The method comprises the steps of extracting multi-modal data from a to-be-processed face video; performing feature modulation on the multi-modal data through a coarse fusion network in the progressive fusion model to obtain multi-modal features; and performing feature modulation on the multi-modal features through a fine fusion network in the progressive fusion model to obtain a digital face image, the feature resolution of the coarse fusion network being lower than the feature resolution of the fine fusion network. In the implementation process of the scheme, the feature fusion process is divided into two stages of coarse fusion and fine fusion, and gradual feature optimization from global to local is realized by adopting different feature granularities, so that the quality fineness of the finally generated digital face can be effectively improved under the condition of ensuring that the feature extraction efficiency is not reduced.
Owner:北京天数智芯半导体科技有限公司

System and Method for Event-Driven Video Synthesis Using Textual Descriptions

A video generation framework that is controllable, unsupervised and based on events (CUBE) includes an event camera, which captures changes in light intensity at each pixel of a scene asynchronously and generates event camera data. A text-to-image diffusion model that is conditioned on textual descriptions integrates the event camera data to control video synthesis. Further, an edge extraction module translates event data into a format usable by the text-to-image diffusion model, whereby the diffusion model synthesizes detailed and contextually accurate videos based on textual prompts. Further, an improved system (CUBE Plus) includes a content frame identification module which selectively identifies and uses only the most information-rich event segments of the event camera data to drive cross-frame attention, and an event driven attention mechanism that allows the framework to focus on event-dense moments.
Owner:THE UNIVERSITY OF HONG KONG

Techniques for automated generation and rigging of objects for animation

One embodiment of a method for generating animations includes generating one or more images of an object based on user input, generating textured geometry based on the one or more images, generating a weight map based on at least one image included in the one or more images, and generating an animation of the object based on the textured geometry, the weight map, and a skeleton.
Owner:DISNEY ENTERPRISES INC

Multi-modal portrait video editing method, electronic device, and storage medium

The present invention provides a multi-modal portrait video editing method, which can be applied to the technical field of video editing. The method comprises: given a portrait video, preprocessing same to obtain camera parameters, a human identity coefficient, a human expression coefficient, a human pose coefficient, a human semantic segmentation map, and a two-dimensional portrait mask; on the basis of a neural Gaussian texture mechanism, embedding a learnable three-dimensional Gaussian feature into a parameterized human geometric surface, using a neural renderer to convert a three-dimensional Gaussian splatting feature map into an image, and optimizing the reconstruction of a three-dimensional portrait on the basis of RGB and segmentation map information of the video; and using an iterative dataset update technique to distill knowledge of a multi-modal two-dimensional image generation model into three-dimensional portrait editing, and using expression similarity guidance and a face-aware portrait editing model to improve editing quality. The method elevates a two-dimensional editing task to three-dimensional space, and ensures good three-dimensional consistency and temporal consistency. By means of the knowledge of the multi-modal generation model, high-quality portrait video editing functions can be achieved.
Owner:UNIV OF SCI & TECH OF CHINA

Remote maintenance auxiliary method integrating video monitoring and three-dimensional modeling

The invention relates to the technical field of industrial internet of things operation and maintenance, and particularly provides a remote maintenance auxiliary method integrating video monitoring and three-dimensional modeling. The method comprises the following steps: acquiring engineering graphic data and point cloud scanning data of maintenance equipment, and collecting video stream data of a maintenance equipment site; the video stream data is used for describing the operation state of maintenance equipment; the video stream data comprises a plurality of video frames; matching the point cloud scanning data with the engineering graphic data, and constructing a watertight three-dimensional grid model according to a matching result; mapping texture features of the maintenance equipment in a target video frame to the surface of the watertight three-dimensional grid model to obtain a target three-dimensional model; and receiving a first maintenance instruction marked in the target three-dimensional model by a remote expert, and sending the first maintenance instruction to a video picture of a client of an on-site maintainer. According to the technical scheme provided by the invention, the time consumption for positioning the overhaul part can be reduced.
Owner:STATE GRID JIANGSU ELECTRIC POWER CO LIANYUNGANG POWER SUPPLY CO

Method and device for generating an animation graph

ActiveUS12499599B1AnimationAlgorithmAnimation
In some implementations, the method includes: obtaining a plurality of animations; determining initial and end motion states for each of the plurality of animations; generating an animation graph including nodes for each of the plurality of animations by connecting, with a directional edge, a first node with an end motion state to a second node with an initial motion state that matches the end motion state of the first node; generating a transitional animation that is not included among the plurality of animations from an initial reference motion state to a target motion state that corresponds to a path that traverses the animation graph from a third node associated with the initial reference motion state to a fourth node associated with the target motion state; and updating the animation graph by removing one or more nodes from the animation graph based at least in part on the transitional animation.
Owner:APPLE INC

Generative ai pet avatar generation

Described is a system for virtual pet generation by receiving, by a computing device of a first user, a real life image of a pet; identifying a prompt corresponding to desired characteristics of a first virtual pet avatar for the first user; processing the real life image of the pet and the prompt by a first generative Artificial Intelligence (AI) model, the first generative AI model being trained to receive images and prompts and to generate virtual pet avatars based on the received images and prompts; and receiving a first virtual pet avatar from the first generative AI model.
Owner:SNAP INC

Rendering Video Of A Scene Using Three-Dimensional Gaussians

A set of images of a scene re received. Each image includes temporal data and spatial data relating to the scene. Based on the spatial data of each image, three-dimensional (3D) Gaussian splatting data is generated. The temporal data of each image and the 3D Gaussian splatting data are inputted to a neural network to generate spatial-temporal 3D Gaussian embeddings. Offset data based on the spatial-temporal 3D Gaussian embeddings is generated. The video of the scene is rendered based on the 3D Gaussian splatting data and the offset data, allowing for improved rendering of video of the scene.
Owner:YINWANG INTELLIGENT TECHNOLOGIES CO LTD

Dynamic generation method of digital animation character based on generative model

The invention discloses a digital animation role dynamic generation method based on a generative model. The method comprises the following steps: inputting a role original image and scene background information, extracting skeleton key points and scene features, analyzing an action sequence, predicting a motion track, calculating an optimal position of a role in a picture, and carrying out dynamic adjustment. And according to the role position, obtaining morphological feature data, evaluating the quality grade, and optimizing the coordination of the role image. And finally, comprehensively scoring by adopting a multi-dimensional quality evaluation system, and determining the visual presentation quality of the role. According to the method, deep fusion of role actions, positions and scenes is realized, the dynamic expressive force and visual coordination of animation pictures are improved, and an efficient solution is provided for generating high-quality animation contents.
Owner:HEBEI XIONGAN PEPSI HENGXING NETWORK TECHNOLOGY CO LTD

Five-dimensional motion vector real-time construction and calibration method based on multi-sensor fusion

The invention discloses a five-dimensional motion vector real-time construction and calibration method based on multi-sensor fusion, and relates to the technical field of multi-sensor information fusion and dynamic state estimation, and the method comprises the steps: collecting and preprocessing motion carrier data, obtaining a preprocessing data set and a feature data set, inputting the feature data set into a long short-term memory network model, and obtaining a multi-sensor fusion model; and outputting a sensor error offset prediction vector to the extended Kalman filtering model, outputting a preliminary five-dimensional motion vector, and performing consistency verification and correction on the preliminary five-dimensional motion vector through a kinematics constraint equation to obtain a corrected five-dimensional motion vector so as to drive a virtual model corresponding to a motion carrier to perform synchronous position and attitude updating. According to the method, the sensor error is predicted through the long-short-term memory network model, and the dynamic motion model constraint module is additionally arranged to perform physical constraint correction, so that the problems of inaccurate error compensation and lack of physical authenticity of the calculation result under the dynamic working condition are solved, and the construction precision and reliability of the five-dimensional motion vector are improved.
Owner:SHANDONG PRECISION INTELLIGENT MEDICAL EQUIPMENT CO LTD

Fire scene spreading time sequence reconstruction method and system based on Doppler weather radar mountain fire echo

The invention relates to the technical field of forest fire monitoring and power grid safety prevention and control, and discloses a fire scene spreading time sequence reconstruction method based on Doppler weather radar mountain fire echoes, which comprises the following steps: acquiring original data of the Doppler weather radar mountain fire echoes and carrying out quality control processing to obtain radar echo data; performing automatic segmentation of a smoke plume area on radar echo data by adopting an optimized Otsu-Unet network model, and extracting a centroid coordinate of a mountain fire echo area; a Kalman filtering algorithm is used and combined with external three-dimensional wind field data to predict a future motion trend of centroid coordinates of a forest fire echo region so as to reconstruct a continuous fire scene spreading time sequence voxel sequence; performing 3DTiles slicing processing on the fire scene spreading time sequence voxel sequence, and performing real-time rendering display on a power grid monitoring platform by using a WebGL technology; and obtaining an external expansion boundary of fire scene spreading based on the fire scene spreading time sequence voxel sequence, and calculating the distance between the external expansion boundary and the power transmission line GIS data in real time. The fire prevention and control capability of the power transmission line of the power grid is improved.
Owner:STATE GRID SICHUAN ELECTRIC POWER CORP ELECTRIC POWER RES INST

Scene generation method and device based on multi-source GIS data fusion

The invention discloses a scene generation method and device based on multi-source GIS data fusion, and relates to the technical field of digital twinning and programmed generation. The method comprises the following steps: acquiring GIS data, preprocessing the GIS data, and storing the preprocessed GIS data in a geographic information resource library; a structured resource library is constructed, semantic parameters are added to the three-dimensional model in the structured resource library through the configuration file, and three-dimensional model resources with structured semantics are constructed; on the basis of the three-dimensional model resources with structured semantics and GIS data in a geographic information resource library, building and road generation and terrain processing are carried out in a programmed modeling engine through a configuration file, and scene data are generated; and importing the generated scene data into a real-time rendering engine, carrying out dynamic environment interaction and biocenosis simulation, and generating a city scene. The problems that in the prior art, an urban three-dimensional modeling method is low in efficiency and insufficient in environment interaction reality sense are solved.
Owner:TUDOU DATA (HANGZHOU) HOLDINGS CO LTD

Backboard video generation method for real-time interactive digital human and related device

The invention provides a backplane video generation method for real-time interactive digital humans and a related device, and relates to the technical field of image generation, in particular to the technical field of artificial intelligence such as human-computer interaction, digital humans, end-cloud integration and large models. The method comprises the following steps: acquiring a plurality of original images presented by the same target person at different angles; generating a digital human image taking a pure color as a background based on the plurality of original images; generating prompt information based on a preset action expression demand and the digital human image; and inputting the prompt information into a preset video generation large model, and generating a digital person bottom plate video which enables a digital person corresponding to the target person to show a target action corresponding to the action expression demand. According to the method, efficient and low-cost generation of the digital human bottom plate video is realized.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Using Game Metadata to Animate User-Generated Object in Video Game

A technique for generating, from a video from a computer game, a three-dimensional (3D) representation of space in which Gaussians represent objects in the video. Metadata from the game can be used in creating the 3D representation. User-input content such as a hand-drawn game path is inserted into the 3D representation of space and aligned and scaled. The opacity of the Gaussians in the 3D representation of space is then set to zero such that Gaussians representing objects in the video are transparent and only one or more portions of the user-input content are not transparent. The 3D representation of space is then combined with the video so that the user-input content is presented with the video and animated according to the metadata.
Owner:SONY INTERACTIVE ENTERTAINMENT LLC

Multi-modal large model-based AI recruitment candidate comprehensive ability evaluation method

The invention discloses an AI recruitment candidate comprehensive ability evaluation method based on a multi-modal large model, and relates to the technical field of artificial intelligence recruitment and multi-modal data processing. The method comprises the following steps: S1, constructing a post digital twinborn environment; S2, collecting candidate multi-modal interaction data; S3, training and optimizing a multi-modal large model; and S4, evaluating the comprehensive ability of the candidate and outputting a result. According to the method, a high-simulation post digital twinborn environment is constructed, multi-modal interaction data acquisition is combined, and deep analysis and evaluation are performed by using a multi-modal large model, so that the method not only considers the traditional resume and interview performance of candidates, but also improves the evaluation efficiency by simulating a real working scene. The method comprehensively evaluates the actual operation capability, the communication cooperation capability, the problem solving capability and other multi-dimensional capabilities of the candidates, and compared with a traditional recruitment mode, the method can more accurately predict the scene adaptation speed and the comprehensive capability performance of the candidates after the candidates enter the job.
Owner:SHANGHAI DAOAN INFORMATION TECHNOLOGY CO LTD

Three-dimensional grounded video generation

Systems and methods are disclosed related to a 3D grounded video foundation model. A video generation method and system provide 3D conditioning information to a video diffusion model to improve generated video quality (object and temporal consistency) that is grounded in three dimensions (3D). The video generation method and system also enable precise camera control, cinematic effects, and scene editing. Video output corresponding to a set of camera specifications is generated for a scene from input image(s) including one or more images of a static scene or a sequence of images (video) for a dynamic scene. The input image(s) are used to calculate a 3D cache representing the scene. The 3D cache is rendered according to the set of camera specifications to produce a frame sequence and a mask sequence that identifies missing pixels in each frame. The frame sequence is encoded and masked to generate the output video.
Owner:NVIDIA CORP