Multi-modal self-adaptive interactive digital human generation method and system
By constructing an intelligent decision-making framework, the digital human generation technology has been platformized and made more flexible, solving the problem that existing technologies cannot dynamically adapt to multi-source heterogeneous data, and generating high-quality, highly adaptable, and interactive digital humans.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-13
- Publication Date
- 2026-03-27
AI Technical Summary
Existing digital human generation solutions cannot be dynamically adjusted according to data characteristics and business needs, resulting in poor technology reusability, high development costs, a lack of unified processing capabilities for multi-source heterogeneous data, and insufficient matching of generation quality with scenario requirements.
Construct an intelligent decision-making framework that includes a data perception layer and a business configuration layer. Through multimodal data input and adaptive analysis and decision-making, select and execute appropriate data processing and business model strategies, generate configurable and verifiable core components, integrate and interact with digital humans, and continuously improve the model by using feedback optimization.
It enables the efficient generation of interactive digital humans that adapt to diverse business scenarios within a unified technical framework, ensuring the quality of generation and the fit of the scenario, reducing development costs and improving technology reusability.
Smart Images

Figure CN121745147A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of artificial intelligence-generated content and digital human technology, specifically relating to a highly adaptable general-purpose digital human generation technology. More specifically, this invention relates to a method and system capable of automatically analyzing and making decisions based on the modal characteristics, source attributes, and specific business needs of input data, and adapting to different technical processing and verification paths to generate interactive digital humans suitable for diverse scenarios. Background Technology
[0002] As a new form of interaction between the virtual and physical worlds, digital humans are gradually being applied to various fields such as education, entertainment, customer service, emotional companionship, and cultural heritage due to their anthropomorphic interactive capabilities. Currently, the construction and application solutions for digital humans often exhibit a "siloed" characteristic, meaning that different application scenarios, such as "historical figure recreation," "virtual idols," and "private digital twins," often require the design and development of independent and non-reusable technical architectures. For example, historical figure digital humans used in education emphasize strict historical accuracy, while private digital twins used for emotional companionship focus more on emotional resonance and privacy security. The underlying data preprocessing, model training, and content verification logic of these two types of systems are completely different.
[0003] This fragmented development model leads to significant technical shortcomings: First, it suffers from poor technology reusability and high development costs, requiring the entire system to be built from scratch for each new scenario, making it difficult to create a unified platform product. Second, it lacks a unified processing capability for multi-source heterogeneous data; public, structured historical materials and private, unstructured personal data require different processing strategies, and existing technologies cannot flexibly adapt within a single framework. Finally, the quality of the generated data does not adequately match the needs of the scenario, and the fixed generation process struggles to achieve dynamic balance and precise assurance in terms of multi-dimensional quality requirements such as "factual accuracy," "emotional appropriateness," and "logical consistency."
[0004] Therefore, there is an urgent need in this field for a universal, highly adaptive digital human generation framework. This framework should serve as an underlying technology platform, intelligently sensing the characteristics of input data and business objectives, dynamically allocating internal resources and strategies, thereby efficiently and effectively supporting diverse business applications on it, and achieving a fundamental shift from "one system per scenario" to "one platform for multiple scenarios". Summary of the Invention
[0005] The technical problem to be solved by this invention is to provide a multimodal adaptive interactive digital human generation method and system, which addresses the shortcomings of existing digital human generation schemes that are fixed and unable to be dynamically adjusted according to data characteristics and business needs, thus making it difficult to uniformly adapt to multiple business scenarios and heterogeneous data from multiple sources.
[0006] To solve the above-mentioned technical problems, the present invention adopts the following technical solution: In its first aspect, the present invention provides a multimodal adaptive interactive digital human generation method, characterized in that it constructs an intelligent decision-making framework comprising a data perception layer and a business configuration layer, the method comprising the following steps: S1: Multimodal data input and service configuration: Receives multimodal input data containing text, images and audio, and simultaneously receives service scenario configuration instructions, wherein the configuration instructions specify at least the data source type and the service mode type; S2: Adaptive analysis and decision-making: Extract features from the multimodal input data, and based on the extracted features and the business scenario configuration instructions, analyze and determine the data source attributes and the appropriate business mode of the data; S3: Adaptive Strategy Selection and Execution: Based on the data source attributes determined in step S2, select and execute the corresponding data processing strategy; based on the business model determined in step S2, select and execute the corresponding business model processing strategy. S4: Configurable verification core component generation: Based on the execution result of step S3, digital human core components including personality model, voice model and virtual avatar model are generated in parallel; wherein, the process of generating the personality model is embedded with a configurable consistency verification module, which activates the corresponding verification logic according to the business mode type. S5: Digital Human Integration and Interaction: Multimodal alignment and fusion of the generated core components are used to form a unified digital human instance to provide interaction with the user in line with their role positioning; S6: Feedback Optimization: Collect user interaction data with the digital human instance, and optimize the adaptive analysis decision model of step S2 and / or the core component generation model of step S4 based on the interaction data.
[0007] A second aspect of the present invention provides a multimodal adaptive interactive digital human generation system for implementing the above-described method, characterized in that it comprises: The multimodal data input and business configuration interface is used to receive multimodal input data and business scenario configuration instructions. An adaptive analysis and decision center, connected to the interface, is used to extract and analyze features from input data and output data source determination results and business model matching results. The strategy execution engine, connected to the adaptive analysis and decision center, includes a data processing strategy unit and a business mode processing strategy unit set in parallel, which are used to call the corresponding strategy unit to perform processing according to the judgment result and the matching result. A core component generator, connected to the strategy execution engine, is used to generate personality, voice, and virtual avatar components; the generation process of the personality component includes a configurable consistency verification module. The digital human integration platform, connected to the core component generator, is used to fuse and drive the alignment of the generated components, and output an interactive digital human instance. The feedback optimization loop is used to collect interactive data and optimize the parameters of the models in the adaptive analysis and decision center and the core component generator. Attached Figure Description
[0008] Figure 1 This is a schematic diagram of the overall architecture and workflow of a multimodal adaptive interactive digital human generation system provided in an embodiment of the present invention.
[0009] Figure 2 This is a schematic diagram of the workflow of the personality interaction component generation and business configurable verification mechanism provided in an embodiment of the present invention. Detailed Implementation
[0010] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of the embodiments of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this invention, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0011] The core idea of the method and system provided in this invention lies in constructing an intelligent framework with a closed loop of "perception-decision-execution-verification-optimization". The following will combine... Figure 1 and Figure 2 The technical solution of the present invention will be described in detail through two typical business scenario embodiments.
[0012] This embodiment demonstrates how to utilize publicly available historical materials to construct a digital persona of "Comrade Lei Feng" on the system of this invention, which can be used for patriotic education and moral character learning. This embodiment fully follows... Figure 1 The overall system flow shown is as follows Figure 2 The core component generation process is shown below.
[0013] Step S1: Multimodal data input and service configuration Users perform initial configuration through the system interface: setting the "Business Mode" to "Historical Figures" and the "Data Source" to "Public". Subsequently, users upload a digital resource package containing multimodal publicly available historical materials, primarily including: Text data: a full digital version of Lei Feng's diary, the official biography "The Story of Lei Feng," and historical reports from authoritative media outlets; Image data: a standard photograph of Lei Feng during his lifetime, digital photos of his work and life scenes, and related propaganda posters; Audio data: a very limited number of existing recordings of Lei Feng speaking.
[0014] Step S2: Adaptive Analysis and Decision Making The system's adaptive analysis and decision center is activated, performing multimodal feature scanning and analysis on the input resource package. The analysis module identifies that: the total amount of text data is sufficient and the structure is standardized; the image data quality is acceptable but the perspective is relatively singular; and the audio data is very sparse. Based on the comprehensive judgment, the system marks the data package as "public data - modal imbalance" and confirms the business model as "historical figures".
[0015] Step S3: Adaptive Strategy Selection and Execution Based on the analysis results of step S2, the system triggers two strategy selection and execution paths in parallel: Data Processing Strategy Selection and Execution: Since the data source is determined to be "public," the system automatically selects the "Public Data Processing" path. This path utilizes a pre-trained large-scale language model to perform deep semantic understanding and structuring of the text, and connects to public knowledge bases (such as encyclopedias and historical event databases) for information association, completion, and cross-validation. Business Model Strategy Selection and Execution: Since the business model is determined to be "historical figures," the system automatically selects the "Historical Figures Processing" path. This path will highlight all declarative information related to time, place, people, and events, preparing for subsequent rigorous verification.
[0016] Step S4: Generation of Configurable Verification Core Components The system processes the data along the selected path mentioned above, generating three core components of the digital human in parallel, and then aggregates them into the core component generator. Personality Interaction Component Generation: Knowledge Graph Construction: From texts such as "Lei Feng's Diary," using named entity recognition and relation extraction technologies, a dedicated dynamic knowledge graph centered on "Lei Feng" is automatically constructed, containing nodes such as "person," "event," "location," "time," and "quotes," along with their relationships. Adaptive Training: Given the "historical figure" pattern and abundant data, the system selects the "deep fine-tuning" path. Using processed high-quality text data, a basic large language model undergoes supervised deep fine-tuning to generate a "Lei Feng" personality model with deeply personalized features. Configurable Consistency Verification: This module receives the "historical figure" business pattern signal from the system. A verification strategy selector dynamically loads the "historical accuracy verification" rule set accordingly. Before outputting any content, the generated personality model must be checked by this rule set to ensure that its content is consistent with the knowledge graph and external authoritative historical materials. After evaluation and result judgment, the model is encapsulated as a usable component. If the verification fails, a report is generated and fed back to the training phase for adjustments. Virtual avatar component generation: Under the "Public Data Processing" path, a 3D face reconstruction algorithm based on multi-view photos is used to generate a basic 3D head model. High-precision texture mapping and rendering are then performed, incorporating elements such as military uniform styles and cap badges recorded in historical materials to ensure the historical authenticity of the avatar. Voice component generation: For sparse raw audio, few-sample timbre cloning and transfer learning techniques are used to extract effective timbre features. These features are then fused and modeled with standard Mandarin speech features from the same era to synthesize speech that conforms to the historical context and the characteristics of the person.
[0017] Step S5: Digital Human Integration and Interaction All generated and verified components are sent to the digital human integration platform for multimodal fusion. Here, the personality model drives the dialogue logic, the virtual avatar model is bound to corresponding lip movements, facial expressions, and gestures, and the voice model provides sound output. The three are aligned through timestamps to form a unified, real-time interactive "digital Lei Feng" instance. When a student asks, "Uncle Lei Feng, what would you do if you saw a classmate in trouble?" the digital human instance will drive the virtual avatar to answer with historically accurate language (derived from the personality model and knowledge base), an engaging voice, and sincere facial expressions.
[0018] Step S6: Feedback Optimization All interactions are anonymously recorded by the interaction data collection module. This data (such as follow-up questions from users and feedback on their satisfaction with the answers) is used to continuously optimize the system. For example, feedback data can be used to fine-tune the classification model of the adaptive analysis decision center, making its judgments on similar data more accurate; it can also be used to enable the personality model to learn online, making its answers more relevant to the needs of educational scenarios, forming a closed-loop improvement process.
[0019] This example demonstrates how to use user-provided private data to create a digital twin of a deceased loved one for emotional support. This example also follows... Figure 1 and Figure 2 The process is shown, but the strategy choices are quite different.
[0020] Step S1: Multimodal data input and service configuration The user configures the business model as "emotional companionship" and the data source as "private". The user uploads a private data package of their grandfather, which may include: a handwritten biography, scanned copies of several family letters, several personal photos, and a short phone recording.
[0021] Step S2: Adaptive Analysis and Decision Making System analysis revealed that all modalities of data exhibited highly sparse, unstructured, and highly private characteristics, classifying them as "private data - scarce" and confirming the business model as "emotional companionship".
[0022] Step S3: Adaptive Strategy Selection and Execution Data Processing Strategy Selection and Execution: Because the data source is "private," the system immediately activates the "Private Data Processing" path. This path first activates the privacy protection engine to automatically anonymize sensitive information such as names, addresses, and phone numbers in the text; all data processing and model training are performed in a user-specified local or secure isolated environment, eliminating the risk of raw data leakage. Business Model Strategy Selection and Execution: Because the business model is "emotional companionship," the system selects the "Emotional Companionship Mode Processing" path. This path focuses on analyzing the emotional tendencies, language habits, values, and ways of addressing and caring for users revealed in the data.
[0023] Step S4: Generation of Configurable Verification Core Components Personality interaction component generation (following) Figure 2 process): Knowledge Graph Construction: Key life events, family relationships, and personality keywords are extracted from biographical sketches and family letters to construct a lightweight knowledge graph centered on kinship. Adaptive Training: Addressing data scarcity and privacy requirements, the system employs a "hinted learning" path. Anonymized private data (such as excerpts from family letters) serves as high-quality hindrances, combined with a basic large language model to quickly generate a contextualized personality model that mimics the grandfather's tone and manner of care, without requiring large-scale parameter fine-tuning, balancing effectiveness and privacy. Configurable Consistency Verification: The verification module receives the "emotional companionship" mode signal and loads the "emotional appropriateness verification" rule set. Verification focuses on evaluating the emotional support, empathy, warmth of tone, and consistency with the grandfather's image in the user's private memory, rather than rigidly adhering to historical facts. This ensures that the digital twin's output is psychologically safe, appropriate, and comforting. Virtual Avatar Generation: Under the premise of privacy protection, few-sample image generation technology is used to generate 3D face models from a limited number of photos. Age degradation / evolution models can be applied to generate avatars with a smooth transition from youth to old age. Voice Generation: Zero-sample or few-sample voice cloning technology is used to extract timbre codes from short recordings to achieve high-quality personalized voice synthesis.
[0024] Step S5: After integrating the digital human and interactive components, users can converse with their "digital grandfather." When a user expresses feeling "very stressed at work lately," the system drives the digital human instance to respond with a cloned, friendly voice and an encouraging tone commonly found in family letters, providing emotional support.
[0025] Step S6: Feedback Optimization Interaction data is also used for optimization, particularly to enhance the system's ability to understand users' emotional states and generate personalized responses, making the digital twin more "understanding" of users over time.
[0026] The two contrasting examples above demonstrate that the system and method of this invention, through a unified adaptive analysis and decision-making center, intelligently perceives data characteristics, dynamically selects differentiated data processing strategies and business model strategies based on a clearly defined business scenario configuration, and ensures that the output content accurately matches the core requirements of the scenario (historical accuracy or emotional appropriateness) through a configurable consistency verification module driven by the business model. Ultimately, within a unified technical framework, an interactive digital human that meets both high-quality standards and high scenario adaptability is efficiently generated. This fully verifies the significant progress and beneficial effects of this invention in realizing the platformization, intelligence, and flexibility of digital human generation technology.
Claims
1. A multimodal adaptive interactive digital human generation method, characterized in that, Includes the following steps: S1: Multimodal data input and service configuration: Receives multimodal input data containing text, images and audio, and simultaneously receives service scenario configuration instructions, wherein the configuration instructions specify at least the data source type and the service mode type; S2: Adaptive analysis and decision-making: Extract features from the multimodal input data, and based on the extracted features and the business scenario configuration instructions, analyze and determine the data source attributes and the appropriate business mode of the data; S3: Adaptive Strategy Selection and Execution: Based on the data source attributes determined in step S2, select and execute the corresponding data processing strategy; based on the business model determined in step S2, select and execute the corresponding business model processing strategy. S4: Configurable verification core component generation: Based on the execution result of step S3, digital human core components including personality model, voice model and virtual avatar model are generated in parallel; wherein, the process of generating the personality model is embedded with a configurable consistency verification module, which activates the corresponding verification logic according to the business mode type. S5: Digital Human Integration and Interaction: Multimodal alignment and fusion of the generated core components are used to form a unified digital human instance to provide interaction with the user in line with their role positioning; S6: Feedback Optimization: Collect user interaction data with the digital human instance, and optimize the adaptive analysis and decision model of step S2 and / or the core component generation model of step S4 based on the interaction data.
2. The method according to claim 1, characterized in that, In step S2, the data source types include public data and private data; the business model types include historical figures model and emotional companionship model.
3. The method according to claim 2, characterized in that, In step S3, selecting and executing the corresponding data processing strategy based on the data source attribute includes: if the data source attribute is public data, then calling the pre-trained model and accessing the public knowledge base to perform information enhancement and structuring processing on the input data; if the data source attribute is private data, then initiating a privacy protection processing procedure, which includes at least data desensitization and localized model training.
4. The method according to claim 2 or 3, characterized in that, In step S3, selecting and executing the corresponding business mode processing strategy according to the business mode includes: if the business mode is the historical figure mode, then the processed data is marked with historical correlation; if the business mode is the emotional companionship mode, then the processed data is analyzed for emotional tendency and personalized characteristics.
5. The method according to claim 1, characterized in that, In step S4, the configurable consistency verification module activates the corresponding verification logic according to the business mode type, specifically including: when the business mode is the historical figure mode, activating and executing the historical accuracy verification; when the business mode is the emotional companionship mode, activating and executing the emotional appropriateness verification; wherein, the verification logic is used to evaluate the compliance of the content generated by the personality model, and only the content that passes the verification is used for component encapsulation.
6. The method according to claim 5, characterized in that, The generation process of the personality model also includes constructing a knowledge graph corresponding to the input data; the historical accuracy verification specifically involves comparing the content generated by the personality model with the knowledge graph and external authoritative data sources to verify the consistency of facts.
7. The method according to claim 1, characterized in that, In step S4, when generating the personality model, different model adaptation methods are selected through the adaptive training interface according to the data scale and business model, including: generating a deep personality model by using a deep fine-tuning method; or generating a contextualized personality model by using a prompting learning method.
8. The method according to claim 1, characterized in that, In step S4, when generating the virtual avatar model and the voice model, different generation techniques are selected according to the data source attributes: for public data, 3D modeling technology based on multi-view reconstruction and / or small sample timbre cloning technology are used; for private data, few sample image generation technology and / or zero sample voice cloning technology are used.
9. A multimodal adaptive interactive digital human generation system for implementing the method of any one of claims 1-8, characterized in that, include: The multimodal data input and business configuration interface is used to receive multimodal input data and business scenario configuration instructions. An adaptive analysis and decision center, connected to the interface, is used to extract and analyze features from input data and output data source determination results and business model matching results. The strategy execution engine, connected to the adaptive analysis and decision center, includes a data processing strategy unit and a business mode processing strategy unit set in parallel, which are used to call the corresponding strategy unit to perform processing according to the judgment result and the matching result. A core component generator, connected to the strategy execution engine, is used to generate personality, voice, and virtual avatar components; the generation process of the personality component includes a configurable consistency verification module. The digital human integration platform, connected to the core component generator, is used to fuse and drive the alignment of the generated components, and output an interactive digital human instance. The feedback optimization loop is used to collect interactive data and optimize the parameters of the models in the adaptive analysis and decision center and the core component generator.