Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

1595 results about "Digital human" patented technology

Digital human video generation method based on multi-modal large model

The invention belongs to the technical field of virtual person generation, and particularly relates to a digital person video generation method based on a multi-modal large model, and the method comprises the following steps: 1, constructing a multi-modal data system; 2, multi-modal large model training and adaptation are carried out; 3, constructing a digital human three-dimensional model; step 4, performing semantic analysis and modal mapping; 5, generating a time sequence action and a mouth shape; step 6, building and rendering a virtual scene; step 7, audio and video synchronous rendering and synthesis; step 8, quality optimization and defect repair; and step 9, performing user interaction and iterative optimization. Through technical innovation and engineering, the core pain point in digital human video generation is solved, efficient, vivid and customizable content production capacity is provided for virtual anchors, intelligent customer service, enterprise training and other scenes, and the AI digital human technology is promoted to be applied to large-scale business from experiments.
Owner:ZHE JIANG YAN HUANG KE JI YOU XIAN GONG SI

Lightweight digital human lesson preparation system based on intelligent agent

The invention provides a lightweight digital human lesson preparation system based on an intelligent agent, and belongs to the field of intelligent teaching. Through collaborative operation of four core modules of knowledge graph construction and reasoning, multi-modal cognitive agent, lightweight digital human generation and intelligent teaching plan assistance, the problems of low efficiency of resource integration, teaching content homogenization, insufficient digital human interaction experience and the like in traditional lesson preparation are solved. The knowledge graph construction and reasoning module is used for constructing a structured knowledge graph and realizing knowledge point association mining and teaching logic reasoning; the multi-modal cognitive agent module is used for generating personalized explanation content according with a teaching target by fusing multi-modal courseware analysis, semantic understanding and lecture style dynamic adaptation functions; the lightweight digital human generation module is combined with model pruning and emotion modeling technologies to synchronously output natural voice and a high-simulation digital human image; the intelligent teaching plan auxiliary module helps the teacher to intelligently generate a teaching plan and a teaching outline according to the courseware content based on the knowledge graph.
Owner:TAIYUAN UNIVERSITY OF TECHNOLOGY

Government affair digital human dynamic interaction method and system based on multi-modal large model

The embodiment of the invention provides a government affair digital human dynamic interaction method and system based on a multi-mode large model. The method is applied to the technical field of government affair intelligent services, and comprises the following steps: acquiring a policy announcement text, and performing cleaning and structuring processing to obtain a structured policy data set; and extracting old and new policy data, performing difference comparison, marking key change fields, and generating policy change data. Abstracting and element extraction are carried out on the change data to form structured semantic fragments, and the structured semantic fragments are incrementally embedded into the policy knowledge graph. And generating a question and answer pair sample based on the updated knowledge graph, and carrying out self-supervised fine tuning training on the multi-modal large model. A user inputs multi-modal data, and the model generates a government affair response and feeds back the government affair response. According to the scheme, the multi-modal large model can continuously keep the latest policy knowledge; the model is enabled to generate accurate government affair response with consistent context while understanding multi-modal input such as text, voice and image, and timeliness, accuracy and interactive experience of policy interpretation are improved.
Owner:JIANGSU FENGYUN TECH SERVICE CO LTD

Policy knowledge graph construction method and system based on digital human interaction data analysis

The invention relates to the field of knowledge graph construction, in particular to a policy knowledge graph construction method and system based on digital human interaction data analysis. The policy knowledge graph construction system based on digital human interaction data analysis comprises a knowledge graph preliminary construction module, a first graph updating module, a second graph updating module and a digital human interaction module. According to the method, the policy source webpage content change is monitored in real time, and the question information in the digital human interaction data is deeply mined, so that the full-life-cycle dynamic maintenance of the policy knowledge graph is realized; on one hand, a webpage policy updating event is accurately captured, and knowledge injection is automatically completed in combination with a semantic unit matching algorithm; and on the other hand, the conflict characteristics questioned by the user in the interaction data are extracted through dialogue semantic analysis, atlas correction is triggered after confidence assessment and multi-stage verification, it is ensured that the end-to-end timeliness from policy release to user perception is controlled within a reasonable period, and the policy response accuracy of the digital human service is remarkably improved.
Owner:SGSG SCI & TECH CO LTD

Multi-modal interaction method and system of digital human intelligent agent

The invention relates to the field of multi-modal interaction analysis, in particular to a multi-modal interaction method and system of a digital human agent. The method comprises the following steps: acquiring a real-time face image and a voice signal input stream of an interactive user based on an intelligent agent; performing real-time micro-expression recognition and deep emotion analysis based on the real-time facial image to obtain real-time emotion features of the user; performing time sequence evolution analysis on the real-time emotion characteristics of the user, performing holographic user emotion deep mining, and constructing a user emotion holographic characteristic spectrum; carrying out adaptive acoustic gain processing on the voice signal input stream, and carrying out voice-emotion association analysis based on the user emotion holographic characteristic spectrum to generate a voice-emotion linkage mapping spectrum; and carrying out eyeball fixation point migration tracking based on the user emotion holographic feature map and the real-time face image, and generating a user interaction depth intention signal. Through the real-time deep semantic understanding and emotion perception ability, the intelligent agent interaction intelligence and response accuracy are improved.
Owner:GUANGDONG HUITONG INFORMATION TECH CO LTD

Multi-mode-based AI digital human intelligent interaction method, system and equipment

The invention relates to the technical field of computer vision and human-computer interaction, and discloses an AI digital human intelligent interaction method, system and equipment based on multiple modalities, and the method comprises the steps: pre-awakening a digital human when a human face is detected, and further thoroughly awakening the digital human based on recognized preset voice information or preset gesture information; voice and video information of a user in the interaction process is obtained, a keyword extraction result, a gesture recognition result and an emotional state tag are generated, a pre-constructed knowledge base is utilized to retrieve related information, a big language generation model module is combined to generate an answer text, and the answer text is input into a preset voice synthesis model to generate emotional voice output. And based on the current emotional state label of the user, driving the digital human animation to be output in an emotional manner. According to the method and the system, the digital human for understanding the emotion of the user, generating personalized answers, providing voices with rich emotions and displaying natural expressions and actions can be created, better interaction with the user can be realized, and more humanized and effective services can be provided.
Owner:BEI JING WAN JIE SHU JU KE JI YOU XIAN ZE REN GONG SI WU HAN FEN GONG SI +1

Precise international communication digital human real-time dialogue method fused with multi-modal technology

The invention discloses an accurate international communication digital human real-time dialogue method fused with a multi-modal technology. The method comprises the following steps: S1, constructing a digital human image and tone; s2, propagation content generation and problem guidance; s3, semantic analysis and intention clarification based on the real-time voice dialogue; s4, geographic preference modeling and path planning; s5, cross-context propagation content generated based on retrieval enhancement is generated; s6, visually displaying the propagation content; and S7, carrying out digital human-driven multi-language propagation content real-time output and feedback closed-loop optimization. Through accurate utterance expression analysis, accurate international propagation problem recommendation is provided, cross-context propagation content generation based on semantic understanding is realized, digital people with voice features and visual images are constructed, real-time dialogue interaction of users is realized, and user experience is improved. The system can carry out geographic modeling according to the region where the accurate problem is located, language preference and propagation object culture characteristics, and differential propagation path planning is achieved.
Owner:HUNAN NORMAL UNIVERSITY

Multi-modal collaborative digital human interaction method based on thinking chain and related equipment

The invention discloses a multi-mode collaborative digital human interaction method and related equipment based on a thinking chain, and the method comprises the steps: receiving language, action and text signals inputted by a user through a multi-mode perception module, and generating a multi-mode thinking chain driven by the thinking chain in combination with a historical dialogue scene; the thinking chain dynamically binds time sequence nodes of voices, actions and expressions and defines semantic association of the time sequence nodes; emotional consistency, time sequence coherence and intention matching degree of multi-modal output are verified through a real-time cooperative verification mechanism, and local backtracking and dynamic correction of a thinking chain are triggered; and finally, driving the digital person to output collaborative voice, actions and expressions according to the optimized thinking chain. According to the method, a multi-modal interaction thinking chain is generated by adopting a large language model, semantic deep collaboration and dynamic adaptability of digital human output are realized through a sequential binding and real-time verification mechanism, the problems of multi-modal splitting, intention matching deviation and response stiffness in the traditional technology are solved, and the interaction simulation degree and scene robustness are remarkably improved.
Owner:SOUTH CHINA UNIV OF TECH

Non-perpetual culture immersive interactive experience device and system based on virtual digital human

The invention discloses a non-perpetual culture immersive interactive experience device and system based on a virtual digital human, and relates to the technical field of intangible cultural heritage intelligent interaction, and the device comprises a multi-mode sensing module which collects data such as user actions and voices through a depth camera and the like; the non-perpetual knowledge base module stores a non-perpetual knowledge graph and a case library; the digital human modeling engine fuses the inheritor characteristics and the user data to generate a virtual digital human; the immersion interaction engine constructs a cross-platform rendering environment to realize five-sense fusion experience; and the adaptive learning module analyzes the interaction sequence to optimize the digital human behavior, and all the modules cooperate to realize the intelligent interaction experience of non-genetic culture. Through the multi-modal perception module, the digital human modeling module and the like, the immersive interactive experience of the non-perpetual culture is realized, the user behavior can be accurately captured, the characteristic virtual digital human is generated, the five-sense fusion experience is provided, the content can be optimized based on the user interaction, the inheritance and propagation of the non-perpetual culture are promoted, and the participation degree and the sense of identity of the user are improved.
Owner:HUNAN INSTITUTE OF ENGINEERING

Three-dimensional Gaussian digital human generation system and method and electronic equipment

The invention provides a three-dimensional Gaussian digital human generation system and method and electronic equipment, and relates to the technical field of augmented reality, and the system comprises a data obtaining unit which is used for obtaining visual depth data, dynamic behavior data and environmental perception data, and generating point cloud data and multi-modal data according to the visual depth data, the dynamic behavior data and the environmental perception data; the rendering adjustment unit is used for monitoring the fixation area in real time and determining a rendering strategy; according to the real-time computing power parameter, determining a rendering specification; the driving unit is used for performing feature extraction on the multi-modal data to obtain multi-modal features and generating face driving parameters and limb driving parameters of the digital human model; and the modeling unit is used for generating a basic grid of the digital human model according to the point cloud data and the initial digital human model, and performing expression rendering and action rendering on the basic grid according to a rendering strategy and a rendering specification in combination with the face driving parameters and the limb driving parameters to obtain a rendered digital human model. The fluency, authenticity and naturalness of the display effect are improved.
Owner:CHINA UNICOM ONLINE INFORMATION TECHNOLOGY CO LTD +1

Digital human generation method based on multi-modal large model

The invention provides a digital human generation method based on a multi-modal large model. The method comprises the following steps: constructing a digital human basic model; generating a structured training set; generating a question and answer model supporting multi-channel interaction; semantic answers of the user questions are output, text emotional tendencies of the semantic answers are extracted, and emotional intensity parameters are output; generating facial muscle movement track data, and performing real-time rendering on the digital human basic model according to the facial muscle movement track data to output a digital human three-dimensional image with emotion expression. According to the embodiment of the invention, cross-modal alignment is carried out on text, image and audio data, and a multi-modal large model containing visual, voice and knowledge models is optimized by using a joint training method, so that more natural and smoother multi-channel interaction experience is realized; in addition, by introducing an emotion recognition model and a face interaction model, the emotion tendency contained in the semantic answer can be captured and reflected more accurately, so that a digital human three-dimensional image with real emotion expression is output.
Owner:CHINA NAT BUILDING MATERIALS TECH CO LTD +2

Controllable digital human real-time interaction system based on multi-modal sensor

The invention discloses a controllable digital human real-time interaction system based on a multi-modal sensor, and relates to the technical field of digital humans, and the system comprises a multi-modal sensor data collection module which captures user interaction videos, collects interaction voice signals of a user, and collects the temperature of the user and the distance between the user and equipment; the multi-modal feature extraction module comprises a static feature extraction unit and a dynamic feature extraction unit; the multi-modal fusion module is used for analyzing the trustworthiness degree from the multi-modal feature extraction module, judging the importance of each information source in the current task, comprehensively considering and dynamically adjusting weight distribution, and fusing multi-modal feature data; the intention understanding module is used for understanding the interaction intention of the user by utilizing a large language model (LLM) based on the fused features and contexts; and the digital human interaction output module is used for generating an open domain text reply based on the interaction intention of the user and finally realizing digital human multi-mode fusion interaction output.
Owner:CHENGDU MEIZHIDA INFORMATION TECHNOLOGY CO LTD

Interaction method for driving digital human language understanding and corresponding reaction through artificial intelligence algorithm

The invention relates to the technical field of electric digital data processing, in particular to an artificial intelligence algorithm-driven digital human language understanding and corresponding reaction interaction method. The method comprises the following steps: collecting historical language data, text data, action instruction data and physical environment information of a user to obtain structured training data; vectorizing the text data to obtain a text semantic vector; environment feature vectors are extracted from the physical environment information; splicing the text semantic vector and the environment feature vector to obtain an environment enhanced text vector; and constructing a pre-training language model based on the environment enhanced text vector, and performing secondary pre-training to generate an optimized language model parameter. According to the invention, by fusing the user language and the physical environment information and combining with the multi-module intelligent component, context perception, efficient understanding and multi-modal natural interaction of the digital human in a complex scene are realized, and the intelligence, adaptability and user experience of the system are remarkably improved.
Owner:SHENZHEN NEITWAY INFORMATION & TECH DEV CO LTD +1

Digital human interaction system and method based on multi-modal emotion recognition

ActiveCN121116129ASemantic analysisSpeech analysisInteractive modelingData stream
The embodiment of the invention provides a digital human interaction system and method based on multi-modal emotion recognition, and belongs to the technical field of digital human interaction. The system comprises a multi-modal sensing module used for collecting multi-modal data and preprocessing the multi-modal data to generate a standardized data stream; the cross-modal fusion and emotion recognition module is used for carrying out interactive modeling on the multi-modal features and outputting a current emotion label and emotion intensity; the reaction planning module is used for generating a composite reaction strategy; and the digital human rendering module is used for mapping the composite reaction strategy into control signals corresponding to the voice, the facial expression and the action respectively, and driving a digital human to execute corresponding voice output, facial expression change and limb action through the control signals so as to realize interaction. According to the method, multi-modal data are deeply fused through the cross-modal graph neural network and comparative learning, the weight is dynamically adjusted in combination with the modal confidence, and the emotion recognition accuracy and robustness are improved.
Owner:XIAODUO INTELLIGENT TECH (BEIJING) CO LTD

Intelligent digital human training method and system based on multi-modal interaction

The invention discloses an intelligent digital human training method and system based on multi-modal interaction, and belongs to the technical field of semantic indexing.The method specifically comprises the steps that voice, vision and text data are analyzed and converted into high-dimensional feature vectors through a modal exclusive encoder, the high-dimensional feature vectors are projected to a unified semantic space through a cross-modal semantic mapping model, and the high-dimensional feature vectors are obtained; generating a semantic primitive containing a modal identifier, a core semantic tag and a feature weight; semantic primitives are used as nodes, directed edges and edge weight table association strength are established based on semantic similarity, typical scene node connection weights are strengthened, and a mesh map containing intra-modal hierarchy and inter-modal cross association is formed; constructing a double-layer index on the basis of the mesh map; semantic primitives are extracted from newly added data, the position of a new node in an association graph is determined through a graph matching algorithm, an association edge with an existing node is automatically established, and a lower-layer modal exclusive index is synchronously updated.
Owner:JIANGXI INST OF FASHION TECH

Automatic lecturer video generation method based on AI speech synthesis and animation driving

The invention discloses a lecturer video automatic generation method based on AI speech synthesis and animation driving. The method comprises the following steps: performing structured analysis on a PPT or a text script through an improved interior point method and an incremental shortest path algorithm; performing semantic grouping by applying a full-dynamic parallel single-link clustering algorithm and generating an enhanced script with an expressive mark; a CosyVoice technology is combined with a low-rank approximation method to generate a high-quality voice data stream; establishing a mapping relation between contents and action expressions through semantic analysis, and generating a complete action expression instruction set; and driving the digital human model by using the msueTalk technology, and generating a final lecturer teaching video through a parallel rendering algorithm. According to the invention, the method achieves the efficient and automatic generation of the education video, remarkably improves the content production efficiency, reduces the production cost, and guarantees the specialty and expressive force of the teaching video.
Owner:SHENZHEN XUEYOU TECHNOLOGY CO LTD

Digital human interaction method and device based on multi-modal sentiment analysis and medium

The invention discloses a digital human interaction method and device based on multi-modal sentiment analysis and a medium, and relates to the field of artificial intelligence, and the method comprises the steps: collecting multi-modal data of a user in real time through a multi-source sensor device; the multi-modal data comprises face video stream data, voice audio stream data and text dialogue data; calling data analysis engines corresponding to different modalities, and extracting corresponding modal feature sequences; according to the current interaction scene, the modal feature sequence and the historical dialogue context features are fused, and a comprehensive emotion evaluation result is generated; outputting a corresponding multi-modal response data packet based on the modal feature sequence through an interactive response engine corresponding to a comprehensive emotion evaluation result; and executing the multi-modal response data packet. And after feature fusion is carried out in combination with the current interaction scene, the generated response can more accurately fit the current emotion demand and communication context of the user, so that the digital human can be more easily fused into various scenes needing emotion interaction.
Owner:INSPUR ZHUOSHU BIG DATA IND DEV CO LTD

Water affair decision-making system based on multi-modal model fusion

The invention relates to the technical field of computers, in particular to a multi-modal model fusion-based water decision-making system, which comprises a virtual digital human interaction module, an intelligent agent business middle table and a large model technology base, and is characterized in that the virtual digital human interaction module is used for monitoring a voice instruction of a user; the method comprises the following steps: acquiring a voice instruction of a user, acquiring biological characteristic data of the user by using a multi-modal sensor, identifying relevance characteristics between the voice instruction and the biological characteristic data through a water conservancy emergency scene emotion identification model, acquiring indication information, and sending the indication information to an intelligent agent service middle station; the agent business middle platform performs task chain analysis on the indication information based on a water conservancy professional term intention recognition model to obtain different task requirements, and calls a large model technology base to make a decision based on the different task requirements to obtain decision information; a water conservancy large model trained based on a domain mechanism model library is integrated on the large model technology base. According to the invention, more efficient and intelligent services can be provided.
Owner:ZHEJIANG KEEPSOFT INFORMATIONTECHNOLOGY CORP LTD

Three-dimensional digital human generation method and system capable of voice interaction

The invention belongs to the technical field of three-dimensional reconstruction, and discloses a three-dimensional digital human generation method and system capable of voice interaction. According to the invention, brand new speaking audios in different languages are automatically generated according to different languages of the input target text and the sampled human voice audios; the sequential stability and detail reduction capability of three-dimensional human motion are guaranteed by using multi-model joint estimation and a sequential loss function, and facial expression details and hand postures in the image can be accurately estimated. After the high-precision three-dimensional human body model is obtained through estimation, human body action and expression generation is carried out based on voice driving, accurate synchronization of actions and expressions generated through voice is achieved, and facial expression movement and body posture movement, namely a whole-body three-dimensional human body model, conforming to brand-new speaking audio are accurately generated; and finally, rendering the whole-body three-dimensional human body model into a real digital human capable of voice interaction by using a three-dimensional neural rendering model. According to the invention, the realization of single person picture input, high-precision three-dimensional digital person generation and voice interaction is facilitated.
Owner:NANJING UNIV OF SCI & TECH

Vocational ability training system based on artificial intelligence technology

The invention relates to the technical field of vocational education, in particular to a vocational ability training system based on an artificial intelligence technology, which comprises a user portrait modeling module for generating a dynamic personal ability map based on historical data and an ability evaluation model of a user; the knowledge graph engine is used for integrating industry capability standards, post demand data and a real-time updated vocational skill knowledge base to form a structured knowledge network; the AI training recommendation module is used for dynamically generating a personalized training path by adopting a reinforcement learning algorithm in combination with a user portrait and a knowledge graph; the digital human application module constructs a simulation working scene through virtual reality and NLP technologies, and supports a user to complete training through natural dialogue and operation; and the real-time feedback and correction module is used for analyzing user performance by using multi-modal data and generating instant guidance suggestions. The system has the following beneficial effects of training personalized depth improvement, learning resource and scene expansion, learning effect evaluation and feedback optimization, and technology fusion and interactive innovation.
Owner:SHANGHAI ZHIYUN ZHIXUN EDUCATION TECH CO LTD

AI digital human interactive response method based on large language model

The invention discloses an AI digital human interactive response method based on a large language model, and relates to the technical field of digital human interaction, and the method comprises the steps: analyzing collected user voice data and visual data through a natural language processing method, generating a cross-modal feature vector, carrying out the cross-modal association analysis of the cross-modal feature vector, and carrying out the cross-modal association analysis of the cross-modal feature vector. Generating a semantic association topological graph; calculating a vertex coordinate and a joint activity threshold value of the semantic association topological graph through high-digital human correlation, inputting the vertex coordinate and the joint activity threshold value into a constructed coordinate index database to execute attention weight calibration, and outputting a multi-dimensional association graph; and performing information density analysis based on the multi-dimensional association map, generating an information density gradient vector field, and dividing a high-density core region and a low-density edge region, the high-density core region generating a semantic core coding tensor, and the low-density edge region generating an edge feature package. According to the method, the cross-modal fusion vector is converted into the cross-modal feature vector, so that the modeling of the cross-modal association relationship is realized.
Owner:BEI JING XIN ZHI YUAN LANG WANG LUO KE JI YOU XIAN GONG SI

Intelligent dancing garment generation method and system based on multi-modal action analysis

The invention provides an intelligent dancing garment generation method and system based on multi-modal action analysis. Dance movement biomechanical data are collected through an inertial sensor array and a multi-view visual system, features are extracted through a space-time diagram convolutional network, style semantics are analyzed in combination with a CLIP model, and a design drawing is generated through a diffusion model and fused into physical constraint optimization. And binding the 3D model with the virtual digital human to simulate a dynamic effect, and outputting a production instruction containing the fabric, the model and the process parameters. The system comprises a data acquisition module, an action analysis module and the like, and the whole-process intelligentization is realized. According to the method, traditional limitation is broken through, by fusing biomechanical data and artistic style semantics, the tear strength and style matching accuracy of the clothes are improved, the design period is shortened, automatic design and production of the dancing clothes are achieved, the dynamic adaptability and artistic expressive force of the clothes are improved, and an efficient scheme is provided for customization of the dancing clothes.
Owner:XIAMEN UNIV OF TECH

Digital human interaction system based on web terminal

The invention discloses a digital human interaction system based on a web end, relates to the technical field of digital human interaction, and aims to solve the problem of accumulated dislocation of browser end digital population animation and actual audio playback caused by multiple clocks and buffer scheduling. The system comprises a visual position cooperative control module, a session initialization module, an audio track and rendering canvas binding module, a multi-domain alignment time base cluster establishment module, construction of a time base cluster containing a system reference sub-time base and a content logic sub-time base, potential candidate anchor point generation module, an optimal anchor point selection module, synchronization error calculation, and judgment of a synchronization steady state, a fine adjustment state or a lost state. An optimal anchor point is screened to adjust the animation, and a visual effect studio dynamic maintenance module and an enhancement generation module assist in out-of-step processing and parameter optimization; through cooperation of multiple modules, accurate synchronization of audio and digital human animation is realized.
Owner:NANJING SUPERMIND INFORMATION TECHNOLOGY CO LTD

Interactive automatic explanation method for converting traditional video into artificial intelligence digital human

The invention provides an interactive automatic explanation method for converting a traditional video into an artificial intelligence digital human, and relates to the technical field of artificial intelligence, and the method comprises the steps: obtaining original video data and an audio track, and carrying out the semantic analysis of the audio track, and obtaining multi-mode deconstruction data; generating an explanation script for each time period of the video based on the explanation text, and performing timestamp labeling on the visual elements to form a time sequence synchronization data structure; in the playing process, a virtual image generator is driven to synthesize digital human dynamic expression output in real time according to the current playing time point; after a user interruption request is received, semantic matching is carried out on a query intention in the explanation script, a target explanation fragment and visual elements are positioned, and complementary explanation content is generated; and driving the virtual image generator to synthesize dynamic output synchronized with the supplementary explanation, and after interaction is completed, recovering playing or skipping to a specified time point according to a user instruction. According to the invention, the conversion from the traditional video to the interactive intelligent explanation video is realized, and the watching experience and learning efficiency of the user are improved.
Owner:BEIJING MENGKE TECH CO LTD

Digital human AGI dialogue system based on cloud side-end collaborative architecture

The invention provides a digital human AGI dialogue system based on a cloud side-end collaborative architecture, and relates to the technical field of digital humans, the system is characterized in that a sensing module, a processing module, a decision module, a rendering module and an output module are deployed at a side end, and a decision module, a driving module and a rendering module are deployed at a cloud end; the sensing module collects and preprocesses an input signal of a user; the processing module is connected with the sensing module and is used for extracting features of the input signals and generating context vectors; the decision-making module is connected with the processing module, and generates a decision-making result containing an answer text and an emotion label according to the context vector; the driving module is connected with the decision module, generates an audio stream and a phoneme sequence according to the answer text, and calculates skeleton driving parameters and mouth shape driving parameters of the digital human; the rendering module is connected with the driving module to generate a rendered picture; and the interaction module is connected with the driving module and the rendering module, and aligns the rendered picture and the audio stream to obtain an output result. And low time delay and high performance are realized by adopting cloud edge collaboration.
Owner:SUZHOU PENGYU ZHISHENG NETWORK TECHNOLOGY CO LTD

Digital Humanoid Robots with Dynamical Models for Robot Guidance and Control System Design

This patent discloses a computer system for humanoid robot control system design and implementation, featuring a digital humanoid robot with dynamical models and a set of single-input-single-output (SISO) and multi-input-multi-output (MIMO) controllers. The system comprises a main software program, a generative Al humanoid robot intelligence engine, a robot motion path planner module, and a control system simulation engine. It enables efficient design, testing, validation, and implementation of robot control systems, significantly reducing time to market. The system supports seamless upgrades to accommodate new designs and components, enhancing applications in industrial automation, healthcare, public safety, and more, aligning with the goals of the 4th Industrial Revolution.
Owner:GEN CYBERNATION GROUP

Simulation digital human real-time intelligent voice interaction system and method based on vision and large model

The invention relates to a simulation digital human real-time intelligent voice interaction system and a simulation digital human real-time intelligent voice interaction method based on vision and a large model, and aims to solve the problems of inaccurate target speaker recognition, high response delay and the like in digital human voice interaction in a complex scene. The system circles an effective recognition range through a camera, triggers audio collection in combination with face detection, locks a target speaker and reduces noise by using lip movement recognition and sound image fusion technologies, converts the target speaker into a text through voice wake-up, generates an answer by means of a large language model (LLM) and knowledge retrieval enhancement (RAG) technologies, generates low-delay voice through a voice synthesis technology accelerated by the vLLM, and performs voice recognition on the target speaker. And driving the preloaded digital human image to synthesize a video stream and pushing the video stream to a front end for rendering in real time. Accurate pickup, low-delay interaction and rapid digital human image switching in a complex environment are realized, the accuracy and real-time performance of intelligent voice question answering are improved, and the method is suitable for government affair halls, exhibition halls and other scenes.
Owner:UNICOM (HENAN) IND INTERNET CO LTD

Business processing method and device based on multi-agent cooperation, equipment and medium

The invention provides a business processing method and device based on multi-agent collaboration, equipment, a medium and a program product, which can be applied to the technical field of digital human and artificial intelligence. The method comprises the following steps: in response to a service request initiated by a user through a first service channel, obtaining input information of the user, and creating or updating a session context object corresponding to the service request; analyzing the input information, generating a subtask sequence, and writing the subtask sequence into a session context object; according to the subtask sequence, a corresponding target agent is scheduled to execute operation, and an execution result is written back to the session context object; monitoring updating of the session context object, and generating a migration decision for migrating from the first service channel to the second service channel under the condition that the updated session context object meets a cross-channel migration condition; and based on the migration decision, synchronizing the session context object to the second service channel in a full amount, so that the business of the user is handled in the second service channel.
Owner:ANHUI BRANCH OF INDAL & COMML BANK OFCHINA

Method for Training Image Generation Model, Method for Generating Digital Human Image, Electronic Device and Storage Medium

A method for training an image generation model, a method for generating a digital human image, and related apparatuses are provided, relating to the fields of artificial intelligence, big model, big data and other technologies. The method for training an image generation model includes: obtaining N target facial images of a target face, wherein N is an integer greater than 1; inputting the N target facial images and at least one target background image into a preset image generation model to obtain a target digital human image after the target face is fused with each target background image; and training the preset image generation model based on a degree of difference between a first facial feature in the target digital human image and a second facial feature of the target face in the target facial images to obtain a target image generation model.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Digital human construction method and device based on heterogeneous emotion semantic graph and long sequence emotion modeling

The invention discloses a digital human construction method and device based on a heterogeneous emotion semantic graph and long-sequence emotion modeling, and the method comprises the steps: obtaining multi-modal emotion input data of a text, voice and a visual image, extracting features, and constructing a multi-modal emotion feature set with a timestamp; constructing a heterogeneous emotion semantic graph which comprises user entity nodes, modal feature nodes and emotion concept nodes, modeling a semantic association, state transition and conflict suppression relationship through a multi-type edge structure, and introducing a dynamic evolution and conflict discrimination mechanism; performing time sequence modeling on the emotional state sequence by utilizing a local-global double-layer emotional modeling mechanism, and respectively capturing short-time fluctuation and long-time trend; performing cross-modal fusion on the emotional state and the modal features, and decoding the emotional state and the modal features into behavior parameters for controlling expressions, voices and actions of the digital human; and multi-modal emotion expression of the digital human is driven. Compared with the prior art, the emotion recognition accuracy and expression continuity and naturalness can be effectively improved.
Owner:HUAIYIN INSTITUTE OF TECHNOLOGY