Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

1051 results about "Digital human" patented technology

Precise international communication digital human real-time dialogue method fused with multi-modal technology

The invention discloses an accurate international communication digital human real-time dialogue method fused with a multi-modal technology. The method comprises the following steps: S1, constructing a digital human image and tone; s2, propagation content generation and problem guidance; s3, semantic analysis and intention clarification based on the real-time voice dialogue; s4, geographic preference modeling and path planning; s5, cross-context propagation content generated based on retrieval enhancement is generated; s6, visually displaying the propagation content; and S7, carrying out digital human-driven multi-language propagation content real-time output and feedback closed-loop optimization. Through accurate utterance expression analysis, accurate international propagation problem recommendation is provided, cross-context propagation content generation based on semantic understanding is realized, digital people with voice features and visual images are constructed, real-time dialogue interaction of users is realized, and user experience is improved. The system can carry out geographic modeling according to the region where the accurate problem is located, language preference and propagation object culture characteristics, and differential propagation path planning is achieved.
Owner:HUNAN NORMAL UNIVERSITY

Digital human interaction system and method based on multi-modal emotion recognition

ActiveCN121116129ASemantic analysisSpeech analysisInteractive modelingData stream
The embodiment of the invention provides a digital human interaction system and method based on multi-modal emotion recognition, and belongs to the technical field of digital human interaction. The system comprises a multi-modal sensing module used for collecting multi-modal data and preprocessing the multi-modal data to generate a standardized data stream; the cross-modal fusion and emotion recognition module is used for carrying out interactive modeling on the multi-modal features and outputting a current emotion label and emotion intensity; the reaction planning module is used for generating a composite reaction strategy; and the digital human rendering module is used for mapping the composite reaction strategy into control signals corresponding to the voice, the facial expression and the action respectively, and driving a digital human to execute corresponding voice output, facial expression change and limb action through the control signals so as to realize interaction. According to the method, multi-modal data are deeply fused through the cross-modal graph neural network and comparative learning, the weight is dynamically adjusted in combination with the modal confidence, and the emotion recognition accuracy and robustness are improved.
Owner:XIAODUO INTELLIGENT TECH (BEIJING) CO LTD

Intelligent digital human training method and system based on multi-modal interaction

The invention discloses an intelligent digital human training method and system based on multi-modal interaction, and belongs to the technical field of semantic indexing.The method specifically comprises the steps that voice, vision and text data are analyzed and converted into high-dimensional feature vectors through a modal exclusive encoder, the high-dimensional feature vectors are projected to a unified semantic space through a cross-modal semantic mapping model, and the high-dimensional feature vectors are obtained; generating a semantic primitive containing a modal identifier, a core semantic tag and a feature weight; semantic primitives are used as nodes, directed edges and edge weight table association strength are established based on semantic similarity, typical scene node connection weights are strengthened, and a mesh map containing intra-modal hierarchy and inter-modal cross association is formed; constructing a double-layer index on the basis of the mesh map; semantic primitives are extracted from newly added data, the position of a new node in an association graph is determined through a graph matching algorithm, an association edge with an existing node is automatically established, and a lower-layer modal exclusive index is synchronously updated.
Owner:JIANGXI INST OF FASHION TECH

Digital human interaction method and device based on multi-modal sentiment analysis and medium

The invention discloses a digital human interaction method and device based on multi-modal sentiment analysis and a medium, and relates to the field of artificial intelligence, and the method comprises the steps: collecting multi-modal data of a user in real time through a multi-source sensor device; the multi-modal data comprises face video stream data, voice audio stream data and text dialogue data; calling data analysis engines corresponding to different modalities, and extracting corresponding modal feature sequences; according to the current interaction scene, the modal feature sequence and the historical dialogue context features are fused, and a comprehensive emotion evaluation result is generated; outputting a corresponding multi-modal response data packet based on the modal feature sequence through an interactive response engine corresponding to a comprehensive emotion evaluation result; and executing the multi-modal response data packet. And after feature fusion is carried out in combination with the current interaction scene, the generated response can more accurately fit the current emotion demand and communication context of the user, so that the digital human can be more easily fused into various scenes needing emotion interaction.
Owner:INSPUR ZHUOSHU BIG DATA IND DEV CO LTD

AI digital human interactive response method based on large language model

The invention discloses an AI digital human interactive response method based on a large language model, and relates to the technical field of digital human interaction, and the method comprises the steps: analyzing collected user voice data and visual data through a natural language processing method, generating a cross-modal feature vector, carrying out the cross-modal association analysis of the cross-modal feature vector, and carrying out the cross-modal association analysis of the cross-modal feature vector. Generating a semantic association topological graph; calculating a vertex coordinate and a joint activity threshold value of the semantic association topological graph through high-digital human correlation, inputting the vertex coordinate and the joint activity threshold value into a constructed coordinate index database to execute attention weight calibration, and outputting a multi-dimensional association graph; and performing information density analysis based on the multi-dimensional association map, generating an information density gradient vector field, and dividing a high-density core region and a low-density edge region, the high-density core region generating a semantic core coding tensor, and the low-density edge region generating an edge feature package. According to the method, the cross-modal fusion vector is converted into the cross-modal feature vector, so that the modeling of the cross-modal association relationship is realized.
Owner:BEI JING XIN ZHI YUAN LANG WANG LUO KE JI YOU XIAN GONG SI

Intelligent dancing garment generation method and system based on multi-modal action analysis

The invention provides an intelligent dancing garment generation method and system based on multi-modal action analysis. Dance movement biomechanical data are collected through an inertial sensor array and a multi-view visual system, features are extracted through a space-time diagram convolutional network, style semantics are analyzed in combination with a CLIP model, and a design drawing is generated through a diffusion model and fused into physical constraint optimization. And binding the 3D model with the virtual digital human to simulate a dynamic effect, and outputting a production instruction containing the fabric, the model and the process parameters. The system comprises a data acquisition module, an action analysis module and the like, and the whole-process intelligentization is realized. According to the method, traditional limitation is broken through, by fusing biomechanical data and artistic style semantics, the tear strength and style matching accuracy of the clothes are improved, the design period is shortened, automatic design and production of the dancing clothes are achieved, the dynamic adaptability and artistic expressive force of the clothes are improved, and an efficient scheme is provided for customization of the dancing clothes.
Owner:XIAMEN UNIV OF TECH

Digital human interaction system based on web terminal

The invention discloses a digital human interaction system based on a web end, relates to the technical field of digital human interaction, and aims to solve the problem of accumulated dislocation of browser end digital population animation and actual audio playback caused by multiple clocks and buffer scheduling. The system comprises a visual position cooperative control module, a session initialization module, an audio track and rendering canvas binding module, a multi-domain alignment time base cluster establishment module, construction of a time base cluster containing a system reference sub-time base and a content logic sub-time base, potential candidate anchor point generation module, an optimal anchor point selection module, synchronization error calculation, and judgment of a synchronization steady state, a fine adjustment state or a lost state. An optimal anchor point is screened to adjust the animation, and a visual effect studio dynamic maintenance module and an enhancement generation module assist in out-of-step processing and parameter optimization; through cooperation of multiple modules, accurate synchronization of audio and digital human animation is realized.
Owner:NANJING SUPERMIND INFORMATION TECHNOLOGY CO LTD

Interactive automatic explanation method for converting traditional video into artificial intelligence digital human

The invention provides an interactive automatic explanation method for converting a traditional video into an artificial intelligence digital human, and relates to the technical field of artificial intelligence, and the method comprises the steps: obtaining original video data and an audio track, and carrying out the semantic analysis of the audio track, and obtaining multi-mode deconstruction data; generating an explanation script for each time period of the video based on the explanation text, and performing timestamp labeling on the visual elements to form a time sequence synchronization data structure; in the playing process, a virtual image generator is driven to synthesize digital human dynamic expression output in real time according to the current playing time point; after a user interruption request is received, semantic matching is carried out on a query intention in the explanation script, a target explanation fragment and visual elements are positioned, and complementary explanation content is generated; and driving the virtual image generator to synthesize dynamic output synchronized with the supplementary explanation, and after interaction is completed, recovering playing or skipping to a specified time point according to a user instruction. According to the invention, the conversion from the traditional video to the interactive intelligent explanation video is realized, and the watching experience and learning efficiency of the user are improved.
Owner:BEIJING MENGKE TECH CO LTD

Digital human AGI dialogue system based on cloud side-end collaborative architecture

The invention provides a digital human AGI dialogue system based on a cloud side-end collaborative architecture, and relates to the technical field of digital humans, the system is characterized in that a sensing module, a processing module, a decision module, a rendering module and an output module are deployed at a side end, and a decision module, a driving module and a rendering module are deployed at a cloud end; the sensing module collects and preprocesses an input signal of a user; the processing module is connected with the sensing module and is used for extracting features of the input signals and generating context vectors; the decision-making module is connected with the processing module, and generates a decision-making result containing an answer text and an emotion label according to the context vector; the driving module is connected with the decision module, generates an audio stream and a phoneme sequence according to the answer text, and calculates skeleton driving parameters and mouth shape driving parameters of the digital human; the rendering module is connected with the driving module to generate a rendered picture; and the interaction module is connected with the driving module and the rendering module, and aligns the rendered picture and the audio stream to obtain an output result. And low time delay and high performance are realized by adopting cloud edge collaboration.
Owner:SUZHOU PENGYU ZHISHENG NETWORK TECHNOLOGY CO LTD

Digital Humanoid Robots with Dynamical Models for Robot Guidance and Control System Design

This patent discloses a computer system for humanoid robot control system design and implementation, featuring a digital humanoid robot with dynamical models and a set of single-input-single-output (SISO) and multi-input-multi-output (MIMO) controllers. The system comprises a main software program, a generative Al humanoid robot intelligence engine, a robot motion path planner module, and a control system simulation engine. It enables efficient design, testing, validation, and implementation of robot control systems, significantly reducing time to market. The system supports seamless upgrades to accommodate new designs and components, enhancing applications in industrial automation, healthcare, public safety, and more, aligning with the goals of the 4th Industrial Revolution.
Owner:GEN CYBERNATION GROUP

Business processing method and device based on multi-agent cooperation, equipment and medium

The invention provides a business processing method and device based on multi-agent collaboration, equipment, a medium and a program product, which can be applied to the technical field of digital human and artificial intelligence. The method comprises the following steps: in response to a service request initiated by a user through a first service channel, obtaining input information of the user, and creating or updating a session context object corresponding to the service request; analyzing the input information, generating a subtask sequence, and writing the subtask sequence into a session context object; according to the subtask sequence, a corresponding target agent is scheduled to execute operation, and an execution result is written back to the session context object; monitoring updating of the session context object, and generating a migration decision for migrating from the first service channel to the second service channel under the condition that the updated session context object meets a cross-channel migration condition; and based on the migration decision, synchronizing the session context object to the second service channel in a full amount, so that the business of the user is handled in the second service channel.
Owner:ANHUI BRANCH OF INDAL & COMML BANK OFCHINA

Digital human construction method and device based on heterogeneous emotion semantic graph and long sequence emotion modeling

The invention discloses a digital human construction method and device based on a heterogeneous emotion semantic graph and long-sequence emotion modeling, and the method comprises the steps: obtaining multi-modal emotion input data of a text, voice and a visual image, extracting features, and constructing a multi-modal emotion feature set with a timestamp; constructing a heterogeneous emotion semantic graph which comprises user entity nodes, modal feature nodes and emotion concept nodes, modeling a semantic association, state transition and conflict suppression relationship through a multi-type edge structure, and introducing a dynamic evolution and conflict discrimination mechanism; performing time sequence modeling on the emotional state sequence by utilizing a local-global double-layer emotional modeling mechanism, and respectively capturing short-time fluctuation and long-time trend; performing cross-modal fusion on the emotional state and the modal features, and decoding the emotional state and the modal features into behavior parameters for controlling expressions, voices and actions of the digital human; and multi-modal emotion expression of the digital human is driven. Compared with the prior art, the emotion recognition accuracy and expression continuity and naturalness can be effectively improved.
Owner:HUAIYIN INSTITUTE OF TECHNOLOGY

Server, display device and digital human processing method

The embodiment of the invention provides a server, display equipment and a digital human processing method. The method comprises the following steps: receiving voice data input by a user and sent by the display equipment; broadcast voice is determined based on the voice data; extracting voice features of the broadcast voice; determining mouth shape parameters based on the voice features; determining emotion parameters and acquiring user image data; generating digital human image data based on the user image data, the emotion parameters and the mouth shape parameters; and sending the broadcast voice and the digital human image data to the display device, so that the display device plays the broadcast voice and displays a digital human image based on the digital human image data. According to the embodiment of the invention, the expression parameters and the mouth shape parameters are determined according to the voice data input by the user, the expression parameters and the mouth shape parameters are combined to generate the digital human image with better facial expression expression, and emotion customization and control are realized.
Owner:HISENSE VISUAL TECH CO LTD

Digital human rendering method based on Gaussian splashing and multi-scale characteristic field distillation

The invention discloses a digital human rendering method based on Gaussian splashing and multi-scale characteristic field distillation, and belongs to the field of three-dimensional human body digital reconstruction. According to the method, feature extraction is carried out through a Vision Transformer encoder, based on an SMPL model, through a cross-modal parameter estimation module and dynamic human body modeling of three-dimensional Gaussian splashing and semantic feature rendering, three-dimensional Gaussian is projected to a two-dimensional image plane to calculate a covariance matrix and color mixing, and after a rendered color image and an initial feature field image are output, a three-dimensional image is obtained. A student feature map is obtained through a convolution acceleration module, a teacher feature map is obtained after feature extraction is carried out through a two-dimensional basic model, the constructed model is trained, and optimization training is completed through comprehensive total loss function calculation; according to the method, the problems of fuzzy semantics, detail missing, low rendering efficiency and inaccurate human body-scene separation in the existing method are effectively solved, so that more efficient, fine and robust monocular or multi-view human body three-dimensional reconstruction is realized.
Owner:YUNNAN UNIV

Digital human interaction method and system based on large model

The invention discloses a digital human interaction method and system based on a large model. The method comprises the following steps: acquiring a multi-mode instruction; generating contextual information based on the basic settings and historical information of the digital human; inputting the multi-mode instruction and the context information into a core model to obtain an initial text response; adjusting the initial text response based on the emotional state and character traits of the digital person to obtain a text reply; generating an expression reply according to the text reply, and controlling digital human output; receiving feedback information input by a user, and generating interaction data; and optimizing model parameters of the core model based on the interaction data. According to the method and the device, the reply which better fits the user demand and the emotional state is generated by acquiring the multi-mode instruction and combining the basic setting and the historical information of the digital human. The digital human continuously accumulates knowledge from the interaction of the user, and the behavior mode and evolution character are optimized, so that the continuous learning and growth of the digital human are realized.
Owner:ZHEJIANG FENGWO IOT TECH CO LTD

Bionic robot control system based on digital human action mapping

The invention relates to a biomimetic robot control system based on digital human action mapping, in particular to the technical field of robot control, solves the problem that a biomimetic robot is difficult to perceive physical attributes only by means of vision, achieves accurate pre-judgment of contact force and material attributes by analyzing digital human fingertip microscopic color change and object texture deformation, and improves the accuracy of the biomimetic robot control system. A robot can establish an adaptive mechanical model before contact, the system dynamically adjusts the flexibility and damping of a mechanical arm according to the environment rigidity and friction characteristics, soft grabbing of hard objects and stable grabbing of sliding objects are achieved, the hysteresis quality of traditional feedback control is effectively overcome, meanwhile, the time sequence feed-forward and energy monitoring mechanism is combined, and the control precision is improved. Collision protection is provided while zero-delay synchronization of actions is ensured, and the operation stability and safety of the robot for fragile and flexible objects in a complex environment are remarkably improved.
Owner:SHANGHAI SECOND POLYTECHNIC UNIVERSITY

Secure authentication of digital humans

A video stream that depicts at least the face of an individual, and information identifying a known individual is received. Predetermined validation data derived from the known individual is accessed. An analysis of a segment of the video stream based on the predetermined validation data is performed. Based on the analysis, an output signal indicative of a confidence level that the video stream is a video stream generated by the known individual is provided.
Owner:CHARTER COMM OPERATING LLC

Backboard video generation method for real-time interactive digital human and related device

The invention provides a backplane video generation method for real-time interactive digital humans and a related device, and relates to the technical field of image generation, in particular to the technical field of artificial intelligence such as human-computer interaction, digital humans, end-cloud integration and large models. The method comprises the following steps: acquiring a plurality of original images presented by the same target person at different angles; generating a digital human image taking a pure color as a background based on the plurality of original images; generating prompt information based on a preset action expression demand and the digital human image; and inputting the prompt information into a preset video generation large model, and generating a digital person bottom plate video which enables a digital person corresponding to the target person to show a target action corresponding to the action expression demand. According to the method, efficient and low-cost generation of the digital human bottom plate video is realized.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Dynamic Gaussian digital human image rendering method and device, equipment and storage medium

The invention discloses a dynamic Gaussian digital human image rendering method and device, equipment and a storage medium, and relates to the technical field of artificial intelligence. The method comprises the following steps: determining a target multi-view image frame from a multi-view RGB video image sequence according to a preset attitude, and constructing a deformable parameterized model according to the target multi-view image frame; constructing an attitude space driving attitude corresponding to the multi-view RGB video image sequence based on a three-dimensional attitude estimation technology, and generating a two-dimensional position map according to the attitude space driving attitude and the deformable parameterized model; performing three-dimensional Gaussian binding on the deformable parameterized model to obtain a local attribute of the three-dimensional Gaussian; training is carried out according to the two-dimensional position map, and a target StyleUNet neural network is obtained; and predicting the new attitude through the target StyleUNet neural network to obtain a dynamic Gaussian digital human image. In this way, the high-fidelity drivable high-frequency detail digital human image can be automatically rendered.
Owner:MALANSHAN AUDIO & VIDEO LABORATORY

Intelligent item selection recommendation system of live broadcast room based on AI digital human

The invention discloses an intelligent item selection recommendation system of a live broadcast room based on AI digital people, relates to the technical field of intelligent item selection recommendation, and solves the technical problems of generating a dynamic and high-confidence matching scheme based on highly matched commodities and adjusting a combination strategy in combination with a real-time scene and user feedback. Through the multi-dimensional weighted calculation of the feature matching score, the crowd matching score and the scene association score, the hidden integrating degree of the commodity and the user demand is quantified, the highly matched commodity is screened out, the invalid recommendation is reduced, the comprehensive confidence degree of the commodity combination is calculated based on the association rule algorithm, the personalized matching scheme conforming to the real-time scene is generated, and the user experience is improved. A high-potential scheme is screened through a self-adaptive threshold value, the matching rationality and transformation potential are improved, a comprehensive feedback processing module is introduced, a user is supported to customize combined commodities, and optimization suggestions are provided through system evaluation; and for user change feedback, generating hidden collocation based on a user portrait, and realizing real-time iteration of a recommendation scheme.
Owner:BEIJING BTG HUILIAN TECH CO LTD

Automatic generation method and system for AI digital person recommendation strategy

The invention relates to the technical field of data processing, and discloses an AI digital person recommendation strategy automatic generation method and system. The method comprises the steps of collecting target market data, extracting a behavior feature vector and constructing a strategy vector, predicting a promotion effect through a deep neural network, generating an optimal strategy vector based on a multi-agent game framework and a genetic algorithm, generating digital human expression content, action and scene configuration according to a strategy coefficient, and cross-market strategy reuse is realized by using transfer learning. According to the method, the culture adaptability and the generation efficiency of the AI digital person recommendation strategy are improved.
Owner:TIANJIN BAIMA PLANET INTELLIGENT TECHNOLOGY CO LTD

Real-time digital human-oriented multi-process decoupling and double-state self-adaptive flow pushing method

The invention provides a multi-process decoupling and double-state self-adaptive flow pushing method for a real-time digital human, and relates to the technical field of real-time digital humans. The method comprises the steps that a server initializes a digital human instance and starts two decoupling processes by responding to a client session request; the method comprises the following steps: processing input into an audio frame sequence, forming an audio buffer area through a sliding window, extracting structured features by a first process, and transmitting the structured features with corresponding original audio through cross-process communication; in the second process, a dynamic face area is generated based on the structural features and the visual parameters, and a digital human image frame is synthesized; packaging the image frame, the original audio and the state information into a multimedia unit, and stabilizing frame rate plug flow; based on state and audio input trigger event notification, sessions are monitored, processes are terminated and resources are recycled when the sessions are terminated or abnormal, multi-process decoupling processing and double-state self-adaptive stream pushing of a real-time digital human can be achieved, stream pushing stability is guaranteed, and dynamic monitoring and resource recycling can be achieved.
Owner:XIAODUO INTELLIGENT TECH (BEIJING) CO LTD

Digital human knowledge graph dynamic iteration system and method based on user feedback

The invention relates to the technical field of artificial intelligence and knowledge engineering, and provides a digital human knowledge graph dynamic iteration system and method based on user feedback. According to the method, the timeliness and automation level of knowledge updating are improved; the configuration efficiency of technical resources is optimized; a data-driven iterative verification closed loop is constructed; according to the knowledge service system, the maintainability and robustness of the system are enhanced, the system has the capabilities of tracking, querying and rollback any change, the risk of system service degradation caused by misoperation or invalid updating is greatly reduced, and the stability of the whole knowledge service system is improved.
Owner:SUPER SENSE DIGITAL TECHNOLOGY (DONGGUAN) CO LTD

AI digital human conference proxy method and device under off-line local area network and medium

The invention discloses an AI digital human conference proxy method and device under an offline local area network and a medium, and the method comprises the steps: collecting conference voice in real time, converting the conference voice into a real-time text, and pre-judging a subject set in combination with a localized industry knowledge base and a user historical conference track; if it is detected that the user or the to-be-decided item is mentioned, reply voice is generated in combination with historical corpora of the user; and if the conference enters the pre-judgment topic set, calling the pre-loaded user feature packet to generate reply voice, and controlling the digital person to generate a corresponding audio and video stream. The invention provides an AI digital human conference proxy method and device under an offline local area network and a medium, and aims to realize topic pre-judgment in combination with a local knowledge base and a historical conference track of a user and provide preparation for real-time reply; meanwhile, for different scenes, the user historical corpus and the user feature packet are called respectively to generate the reply, so that the problem that the real-time personalized reply generation of the AI digital person conference agency and the conference issue pre-judgment are difficult to collaboratively realize in an offline local area network environment can be solved.
Owner:GUANGZHOU BAOLUN ELECTRONICS CO LTD

Dynamic human body reconstruction method and system based on double-motion embedding and point cloud fusion

The invention discloses a dynamic human body reconstruction method and system based on double-motion embedding and point cloud fusion, and relates to the field of digital human reconstruction processing. According to the method, the input frame sequence composed of three adjacent frames is intercepted from the monocular video and preprocessed to obtain necessary input parameters, then the dynamic human body reconstruction model is designed to process the input frame sequence and the input parameters to obtain the human body reconstruction image, the overall operation is fast and convenient, and the use effect is good. According to the dynamic human body reconstruction model designed by the invention, on one hand, a DMPF network part is adopted, dual-motion embedding is adopted to extract multi-modal motion features, and efficient fusion of 2D and 3D motion features is realized, so that richer motion information supervision is obtained, and on the other hand, a mixed point cloud encoder is adopted to fuse isolated point and overall point cloud features, so that the dynamic human body reconstruction model is more efficient in motion information supervision. Therefore, the dependency relationship of the local change of the human body surface on the global change is captured, and correct modeling of geometry and texture of the moving human body is further enhanced.
Owner:ANHUI UNIV

AI digital human real-time rendering method based on GPU acceleration

The invention discloses an AI digital human real-time rendering method based on GPU acceleration, and relates to the technical field of digital human rendering, and the method comprises the steps: S1, constructing a multi-GPU hardware feature and load monitoring module, and collecting hardware feature parameters and current load states of each GPU participating in collaborative rendering in real time; according to the method, by constructing the multi-GPU hardware feature and load monitoring module and combining a dynamic task allocation algorithm, dynamic allocation of the rendering sub-tasks to the optimal GPU is achieved, the problems of GPU performance bottleneck and resource idleness caused by traditional fixed task allocation are solved, the multi-GPU collaborative rendering efficiency is improved, and the multi-GPU collaborative rendering efficiency is improved. A scene complexity analysis module and a self-adaptive resource allocation algorithm are built, the proportion of GPU resources between an AI driving module and a rendering module is dynamically adjusted according to scene complexity, the AI digital human interaction naturalness is improved in a simple scene, the rendering frame rate and the picture quality are guaranteed in a complex scene, and the method is suitable for being used in a large-scale scene. And finally, dual optimization of rendering efficiency and picture quality in AI digital human real-time rendering is realized.
Owner:GUANGZHOU PERANG IND CHAIN DEVELOPMENT CO LTD

Robust multi-mode emotion understanding method for intelligent customer service digital human

The invention relates to an intelligent customer service digital human-oriented robust multi-mode emotion understanding method, and belongs to the field of artificial intelligence and human-computer interaction. The method is implemented by an intelligent customer service system, and comprises the following steps: S1, acquiring user data in real time; s2, extracting modal features by using a deep learning network; s3, for the deficiency of modal features, adopting a noise condition fractional network based on a diffusion model for recovery; s4, using the reinforcement learning network to optimize the fused weight corresponding to the modal features; s5, combining the weight of the fusion, and achieving the fusion of the modal features through an attention mechanism; and S6, taking the fused modal features as input, establishing an emotion recognition model by using a deep learning network, and completing an emotion recognition task. The method can effectively solve the problem of data missing in real environment interaction of the intelligent customer service digital person, has high emotion recognition accuracy, provides stable and high-quality service for the user, and has good robustness.
Owner:CHONGQING INST OF GREEN & INTELLIGENT TECH CHINESE ACAD OF SCI

Personal health service system based on smart watch digital human

The invention relates to the technical field of intelligent wearable equipment, and provides a personal health service system based on an intelligent watch digital human, the system comprises an intelligent data analysis platform module comprising an intelligent watch end and a cloud end, the intelligent watch end comprises a digital human interaction module and a multi-source data acquisition module, the digital human interaction module constructs a multi-mode interaction mechanism based on a virtual digital human, obtains user health related information, receives a user instruction and outputs a health analysis result; the multi-source data acquisition module acquires physiological feature data and behavior activity data of a user in real time, and performs dynamic alignment and consistency verification of multi-dimensional data through a space-time association algorithm; and the intelligent data analysis platform module adopts a hierarchical collaborative architecture of edge computing and cloud deep mining to intelligently analyze and process the collected data and generate a personalized health assessment report and a dynamic intervention suggestion, so that the digital human interaction experience can be optimized, and the personalization, intelligence and naturalization of health management can be realized.
Owner:HUNAN UNIV OF SCI & TECH SANYA RES INST

Lip shape synchronization model training method, digital human video generation method and device

The invention provides a training method of a lip shape synchronization model and a generation method and device of a digital human video. The invention relates to the technical field of artificial intelligence, in particular to the technical fields of computer vision, virtual reality, large models, digital human broadcast, digital human live broadcast and the like. According to the specific scheme, the method comprises the steps of determining a two-dimensional face feature truth value and a three-dimensional face feature truth value based on a video of a target object; extracting a three-dimensional face feature prediction value from the sample voice data through the first network sub-model; generating a two-dimensional face feature predicted value from the three-dimensional face feature predicted value through a second network sub-model; based on the two-dimensional face feature truth value, the three-dimensional face feature truth value, the two-dimensional face feature predicted value and the three-dimensional face feature predicted value, a first network sub-model and a second network sub-model are trained, a lip shape synchronization model is obtained, and the lip shape synchronization model comprises the first network sub-model and the second network sub-model.
Owner:BEIJING DUSHANG SOFTWARE TECH CO LTD

Conditional diffusion model-based digital human posture action generation method

The invention discloses a digital human posture action generation method based on a conditional diffusion model, and relates to the technical field of posture action generation. In the training process, a posture sequence is coded into potential representation by using an encoder; performing random masking on the potential representation to obtain a mask potential representation; performing forward diffusion on the potential representation to gradually add noise to obtain a noise representation; inputting the noise representation into a de-noising device, taking the mask potential representation as condition information, and de-noising step by step to obtain a de-noised potential representation; and decoding the denoised potential representation by using a decoder to obtain a recovered attitude sequence. In the reasoning process, filling a missing attitude sequence by using an interpolation mode to obtain an initialized attitude sequence; encoding the initialized attitude sequence by using an encoder as condition information; and representing the randomly initialized noise as noise. The method is used for synthesizing the connection frame in the coherent attitude video.
Owner:HEFEI UNIV OF TECH