A digital human generation method and system based on multi-dimensional matching
By using multi-dimensional matching algorithms and video synthesis technology, the problems of single input methods and rigid tag libraries in digital human generation technology have been solved, enabling flexible input and personalized digital human generation, and improving the intelligence and realism of digital human generation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENBI (NANJING) SOFTWARE CO LTD
- Filing Date
- 2026-03-12
- Publication Date
- 2026-06-05
AI Technical Summary
Existing digital human generation technologies lack flexibility in input methods, making it difficult to support diverse inputs such as text and images. The tag library system is rigid, the correlation between the voice library and the action library is low, and the level of intelligence is low, resulting in digital humans that lack personalization and realism.
Employing a multi-dimensional matching algorithm, basic materials are generated from text or image input. Combined with a dynamic anthropomorphic tag library, background recognition, age judgment, and voice/action library, accurate matching of scene, age, voice, and action is achieved. Digital human videos are then generated using a large video synthesis model.
It enables flexible input methods, supports text and image input, builds an extensible tag library, enhances the personalization and realism of digital human generation, and meets diverse user needs.
Smart Images

Figure FT_1 
Figure FT_2
Abstract
Description
Technical Field
[0001] This invention specifically relates to a digital human generation method and system based on multi-dimensional matching, belonging to the field of artificial intelligence generation technology. Background Technology
[0002] With the rapid development of the digital content industry, digital human technology is being applied more and more widely in fields such as virtual live streaming, film and television production, and online education. Currently, digital human generation technology mainly relies on fixed templates or a single input method, which has the following drawbacks: First, the input method lacks flexibility, making it difficult to simultaneously support diverse input needs such as text and images; second, the tag library system is rigid, with low correlation between the voice and action libraries, making dynamic expansion of tags impossible; third, the matching of voice, action, scene, and age relies on manual settings, resulting in low levels of intelligence and a "one-size-fits-all" problem for the generated digital humans, making it difficult to meet users' personalized needs.
[0003] Therefore, developing a digital human generation method and system that supports multiple input methods, multi-dimensional intelligent matching, and an expandable tag library has become an urgent technical problem to be solved in this field. Summary of the Invention
[0004] To address the shortcomings of existing technologies, the present invention aims to provide a digital human generation method and system based on multi-dimensional matching.
[0005] To achieve the above objectives, the present invention adopts the following technical solution:
[0006] A digital human generation method based on multi-dimensional matching includes the following steps:
[0007] S1. Generate a basic template image with human-shaped features based on text information, or use image information with human-shaped features as the basic material for generating digital humans;
[0008] S2. Based on the aforementioned basic materials, extract character feature information including at least the profession, and match target tags from the dynamic anthropomorphic tag library;
[0009] The dynamic anthropomorphic tag library contains several basic tags based on occupational categories; and each basic tag is associated with a corresponding voice library and action library.
[0010] S3. Perform image recognition on the background in the basic materials to determine the corresponding scene type; extract age features from the people in the basic materials to determine the corresponding age range;
[0011] Based on scene type, match the corresponding scene's sound in the sound library, and select the appropriate timbre based on age range;
[0012] Based on the scene type, match the corresponding scene behavior in the action library, and at the same time combine the age range to match the body language that matches the behavioral characteristics of that age group.
[0013] S4. Using a large video synthesis model, combined with the sound, timbre, movement, and basic human materials obtained from the above matching, a digital human video is generated.
[0014] The aforementioned behaviors are simple actions, including raising hands, walking, sitting down, picking up a cup, and kicking legs; the aforementioned body language is the use of body movements to convey information, including crossing arms (=defensive, resistant), nodding (=agree), leaning forward (=interested), avoiding eye contact (=nervous, guilty), and spreading hands (=helpless, helpless).
[0015] The aforementioned character feature information also includes subcategories of gender and occupation. Correspondingly, the basic tags are derived from differentiated tags corresponding to the above subcategories, and the differentiated tags are associated with differentiated voice and motion libraries.
[0016] The aforementioned large-scale video synthesis model includes the Doubao video model.
[0017] The above age ranges include infancy (1 month to 1 year), early childhood (1 to 3 years), preschool age (3 to 6 years), school age (6 to 12 years), adolescence (12 to 18 years), youth (18 to 35 years), middle age (36 to 59 years), and old age (65 years and above).
[0018] The aforementioned dynamic anthropomorphic tag library includes at least 100 basic tags.
[0019] The above-mentioned scene types are identified using a background recognition algorithm, the age features are extracted using an age judgment algorithm, the voices are matched using a voice matching algorithm, and the behavioral actions and body language are matched using an action matching algorithm.
[0020] Furthermore, the aforementioned background recognition algorithm, age judgment algorithm, voice matching algorithm, and action matching algorithm are constructed into a multi-dimensional intelligent matching algorithm library.
[0021] A digital human generation system based on multi-dimensional matching includes:
[0022] a. Input module: Used to receive text or image information input by the user; and when receiving text information, it generates a basic template image with human-like features as the base material;
[0023] b. Tag Library Module: Used to build and store a dynamic anthropomorphic tag library. The dynamic anthropomorphic tag library contains at least 100 basic tags. Each basic tag is associated with a corresponding sound library and action library, and tag expansion is supported.
[0024] c. Feature recognition module: used to perform background recognition and extract age features of people from basic materials, and output scene type and age range;
[0025] d. Multi-dimensional matching module: used to match suitable voices from the sound library and actions from the action library based on scene type and age range, supporting random matching and user-defined intervention;
[0026] e. Video generation module: Used to call up a large video synthesis model, combine basic materials, matched sound and behavioral actions to generate digital human videos.
[0027] The advantages of this invention are:
[0028] The present invention provides a digital human generation method and system based on multi-dimensional matching, which has the following advantages:
[0029] 1. Flexible input methods, supporting either text or image input. Text input can generate a basic template image with human-like features as a base material to meet the input habits of different users.
[0030] 2. Construct a dynamic and scalable anthropomorphic tag library that supports the detailed expansion of basic tags and combines it with related sound and action libraries to achieve accurate positioning of digital human characteristics.
[0031] 3. Employing a multi-dimensional intelligent matching algorithm that integrates background recognition and age judgment, it achieves accurate matching of voice, action, scene, and age, enhancing the personalization and realism of digital humans.
[0032] 4. It integrates multiple technologies such as image recognition, large model generation, and video synthesis to form a complete digital human generation technology chain. It has strong technical integration and can be widely applied to various digital content creation scenarios.
[0033] This invention effectively solves the problems of low flexibility, insufficient matching accuracy, and poor personalization in existing digital human generation technologies, and improves the intelligence and personalization level of digital human generation. It can be widely used in fields such as digital content creation and virtual interaction, and has strong practicality and wide applicability. Attached Figure Description
[0034] Figure 1 A flowchart of the digital human generation method.
[0035] Figure 2 A structural block diagram of a digital human generation system. Detailed Implementation
[0036] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments.
[0037] A digital human generation system based on multi-dimensional matching comprises an input module, a tag library module, a feature recognition module, a multi-dimensional matching module, and a video generation module. These modules work collaboratively.
[0038] a. Input module: Used to receive text or image information input by the user; and when receiving text information, it generates a basic template image with human-like features as the base material;
[0039] b. Tag Library Module: Used to build and store a dynamic anthropomorphic tag library. The dynamic anthropomorphic tag library contains at least 100 basic tags. Each basic tag is associated with a corresponding sound library and action library, and tag expansion is supported.
[0040] c. Feature recognition module: used to perform background recognition and extract age features of people from basic materials, and output scene type and age range;
[0041] d. Multi-dimensional matching module: used to match suitable voices from the sound library and actions from the action library based on scene type and age range, supporting random matching and user-defined intervention;
[0042] e. Video generation module: Used to call up a large video synthesis model, combine basic materials, matched sound and behavioral actions to generate digital human videos.
[0043] A digital human generation method based on multi-dimensional matching, based on the above system, includes the following steps:
[0044] S1, Input Receiver
[0045] Receive text information input by the user to generate a basic template image with human-shaped features, or receive image information with human-shaped features, one of the two options;
[0046] If it is a text description, a basic template image with human-like features will be used as the basic material for digital human generation;
[0047] If the image is of a person, it will be used directly as the base image to ensure flexibility in the input method.
[0048] S2, Tag Matching
[0049] A dynamic anthropomorphic tag library was constructed, containing at least 100 basic tags covering common occupational types such as teacher and programmer. The basic tags are categorized based on occupation, and each basic tag is associated with a corresponding voice and action library; ensuring consistency between digital human characteristics and tag attributes. Occupational classifications are based on the "Occupational Classification Dictionary".
[0050] Based on the extraction of basic materials (occupation), target tags (basic tags) are matched from the dynamic anthropomorphic tag library.
[0051] It also supports tag expansion. Sub-tags can be derived from basic tags to meet differentiated needs. Other character feature information, such as gender and occupational subcategories, can be extracted from basic materials to create corresponding sub-tags (differentiated tags), and these differentiated tags are associated with differentiated voice and motion libraries.
[0052] For example, the profession of teacher can be further subdivided into different labels such as male teacher - middle school, male teacher - primary school, female teacher - humanities, female teacher - science, etc., based on the teaching stage or subject.
[0053] S3, Multi-dimensional Feature Matching
[0054] S31. Background Recognition: Using a background recognition algorithm (image recognition algorithm), the background in the basic material is identified to determine the corresponding scene type, such as playground, classroom, office, etc.
[0055] S32. Age Determination: Using an age determination algorithm, age features are extracted from the people in the basic materials (e.g., using facial feature extraction technology and a CNN-based age detection model to analyze the age features of the people in the basic materials) to determine the corresponding age ranges; the age ranges include infants aged 1 month to 1 year, toddlers aged 1 to 3 years, preschoolers aged 3 to 6 years, school-age children aged 6 to 12 years, teenagers aged 12 to 18 years, young adults aged 18 to 35 years, middle-aged adults aged 36 to 59 years, and the elderly aged 65 and above.
[0056] S33. Voice Matching: Through a voice matching algorithm, the system matches the corresponding voice to the identified scene type, such as matching a physical education teacher's voice to a playground scene; at the same time, it selects the appropriate timbre based on the determined age range, such as matching a crisp timbre to teenagers and a steady timbre to the elderly; the voice parameters can be randomly selected by the system or customized by the user; the voice library contains at least 50 scene voices.
[0057] S34. Action Matching: Using an action matching algorithm, corresponding behavioral actions are generated based on the identified scene type, such as triggering physical education actions in a playground scene. Simultaneously, body language matching that matches the behavioral characteristics of a given age range is used; for example, teenagers are lively while the elderly move slowly. The action library covers 100+ basic behavioral actions, with a differentiation rate of ≥80% for action resources corresponding to differentiated tags.
[0058] Based on scene type, match the corresponding scene's sound in the sound library, and select the appropriate timbre based on age range;
[0059] Based on the scene type, the system matches the corresponding scene behavior in the action library. The behavior is a simple action, including raising a hand, walking, sitting down, picking up a cup, and kicking a leg. At the same time, it matches body language that matches the behavioral characteristics of the age group, which is to convey information through body movements, including crossing arms (=defensive, resistant), nodding (=agree), leaning forward (=interested), avoiding eye contact (=nervous, guilty), and spreading hands (=helpless, helpless).
[0060] Among them, background recognition algorithm, age judgment algorithm, voice matching algorithm and action matching algorithm are constructed into a multi-dimensional intelligent matching algorithm library.
[0061] In addition to traditional background modeling / recognition algorithms such as frame difference, Gaussian mixture model, KNN (K-Nearest Neighbors), SuBSENSE / PBAS, the background recognition algorithm in this invention can also use deep learning-based background / foreground recognition algorithms such as CN / U-Net / DeepLab, Mask R-CNN, etc.
[0062] In addition to traditional machine learning algorithms such as Active Appearance Model (AAM) and Local Binary Pattern (LBP), the age estimation algorithm in this invention can also use mainstream deep learning algorithms, such as basic CNN age estimation (AlexNet / VGG / ResNet series) and classic dedicated age estimation models (DEX (Deep Expectation), ORCNN and other sequential regression networks, etc.).
[0063] The voice matching algorithm in this invention can use MFCC+cosine similarity (simple and fast comparison), DTW / lightweight audio hash, x-vector / ECAPA-TDNN (human voice matching (voiceprint)) and other algorithms depending on the scenario.
[0064] The action matching algorithm in this invention can use human keypoints + DTW + cosine similarity according to the scene, or it can use deep learning-based Pose Embedding + cosine similarity, SimSiam / MoCo self-supervised features, etc.
[0065] S4. Using video synthesis models such as the Doubao video model, and combining the matched sound, timbre, movement, and basic character materials, a digital human video is generated, completing the rendering and generation of the digital human video, forming a complete digital human generation technology chain.
[0066] Example 1: Digital Human Generation Based on Human Images
[0067] S1. Input reception: The user uploads a picture of a middle-aged man on a playground.
[0068] S2. Tag Matching: The system extracts human features (middle-aged male, playground scene) through facial feature recognition algorithms (such as CNN-based age detection model), matches the basic tag "male physical education teacher - middle school" from the dynamic anthropomorphic tag library, and associates it with the corresponding physical education teaching sound library (containing 50+ scene voices) and action library (covering 100+ basic behavioral actions).
[0069] S3, Multi-dimensional Feature Matching:
[0070] Background recognition: The system identifies the background of the image as a playground scene.
[0071] Age determination: The system extracts facial features of the person and determines the age range to be middle-aged.
[0072] Voice matching: Based on the playground scene, the voice of the physical education teacher is matched, and a steady and powerful tone is selected in combination with the middle-aged age range.
[0073] Action matching: Based on the playground scene, physical education teaching actions (such as demonstrating running posture) are triggered, and body language with appropriate range is matched in combination with the characteristics of middle-aged people.
[0074] S4. Video Generation: By calling the Doubao video model and self-developed video synthesis technology, the system integrates images of people, matching audio and actions to generate a digital human video of a middle-aged male physical education teacher demonstrating running movements on the playground.
[0075] Example 2: Digital Human Generation Based on Text Description
[0076] S1. Input reception: The user inputs the text description "an elderly female teacher is teaching in the classroom", and the system automatically generates a basic template image of an elderly female teacher with human-like features as the basic material.
[0077] S2. Tag Matching: The system uses natural language processing algorithms to parse the character attributes and scene information in the text description, matches the basic tag "elderly female teacher - liberal arts" from the tag library, and associates it with the corresponding teacher's teaching voice library and action library. The resource differentiation rate of the voice library / action library corresponding to the sub-tag is ≥80%.
[0078] S3, Multi-dimensional Feature Matching:
[0079] Background recognition: The system determines the scene to be a classroom based on the text description.
[0080] Age determination: The age range is determined to be elderly based on the text description.
[0081] Voice matching: Matches the teacher's voice to the classroom setting, and selects a gentle and soothing tone based on the elderly age range.
[0082] Action matching: Trigger classroom teaching actions (such as writing on the blackboard, explaining gestures), and match them with soothing body language based on the age characteristics of the elderly.
[0083] S4. Video Generation: Using the Doubao video model and self-developed video synthesis technology, a digital human video of an elderly female teacher lecturing in the classroom is generated.
[0084] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the above embodiments do not limit the present invention in any way, and all technical solutions obtained by equivalent substitution or equivalent transformation fall within the protection scope of the present invention.
Claims
1. A digital human generation method based on multi-dimensional matching, characterized in that, Includes the following steps: S1. Generate a basic template image with human-shaped features based on text information, or use image information with human-shaped features as the basic material for generating digital humans; S2. Based on the aforementioned basic materials, extract character feature information including at least the profession, and match target tags from the dynamic anthropomorphic tag library; The dynamic anthropomorphic tag library contains several basic tags based on occupational categories; and each basic tag is associated with a corresponding voice library and action library. S3. Perform image recognition on the background in the basic materials to determine the corresponding scene type; extract age features from the people in the basic materials to determine the corresponding age range; Based on scene type, match the corresponding scene's sound in the sound library, and select the appropriate timbre based on age range; Based on the scene type, match the corresponding scene behavior in the action library, and at the same time combine the age range to match the body language that matches the behavioral characteristics of that age group. S4. Using a large video synthesis model, combined with the sound, timbre, movement, and basic human materials obtained from the above matching, a digital human video is generated.
2. The method according to claim 1, characterized in that, The behavioral actions are simple actions, including raising hands, walking, sitting down, picking up a cup, and kicking legs; the body language is the use of body movements to convey information, including crossing arms, nodding, leaning forward, avoiding eye contact, and spreading hands.
3. The method according to claim 1, characterized in that, The personal characteristics information also includes subcategories of gender and occupation. Correspondingly, the basic tags are derived from differentiated tags corresponding to the above subcategories, and the differentiated tags are associated with differentiated voice and motion libraries.
4. The method according to claim 1, characterized in that, The large-scale video synthesis model includes the Doubao video synthesis model.
5. The method according to claim 1, characterized in that, The age ranges include infancy (1 month to 1 year), early childhood (1 to 3 years), preschool age (3 to 6 years), school age (6 to 12 years), adolescence (12 to 18 years), youth (18 to 35 years), middle age (36 to 59 years), and old age (65 years and above).
6. The method according to claim 1, characterized in that, The dynamic anthropomorphic tag library includes at least 100 basic tags.
7. The method according to claim 1, characterized in that, The scene type is identified by a background recognition algorithm, the age feature is extracted by an age judgment algorithm, the voice is matched by a voice matching algorithm, and the behavior and body language are matched by an action matching algorithm.
8. The method according to claim 7, characterized in that, The background recognition algorithm, age judgment algorithm, voice matching algorithm, and action matching algorithm are constructed into a multi-dimensional intelligent matching algorithm library.
9. A digital human generation system based on multi-dimensional matching, characterized in that, include: a. Input module: Used to receive text or image information input by the user; Upon receiving text information, a basic template image with human-like features is generated as the base material. b. Tag Library Module: Used to build and store a dynamic anthropomorphic tag library. The dynamic anthropomorphic tag library contains at least 100 basic tags. Each basic tag is associated with a corresponding sound library and action library, and tag expansion is supported. c. Feature recognition module: used to perform background recognition and extract age features of people from basic materials, and output scene type and age range; d. Multi-dimensional matching module: used to match suitable voices from the sound library and actions from the action library based on scene type and age range, supporting random matching and user-defined intervention; e. Video generation module: Used to call up a large video synthesis model, combine basic materials, matched sound and behavioral actions to generate digital human videos.