A method, system, electronic device, and storage medium for rapid development of digital humans
By receiving the requirements of the person and the scene, determining the appearance and personality characteristics, constructing and training the digital human, the problem of the difficulty in integrating appearance characteristics, personality characteristics and preset functions in the existing technology is solved, and the personalized performance and efficient development of the digital human in complex scenes and emotional interactions are realized.
Patent Information
- Application Number
- CN202411494574.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-24
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2044-10-24
AI Technical Summary
Existing technologies struggle to effectively integrate physical features, personality traits, and preset functions when constructing digital humans, especially in complex scenarios and emotional interactions, making it difficult to meet personalized needs.
By receiving the requirements of the person and the scene, determining the appearance and personality characteristics, constructing and training the digital human, including acquiring real human action and expression data, processing and mapping it, and combining voice interaction, action performance and emotional communication functions, constructing activity scenarios and conducting testing and optimization.
It achieves a high degree of personalization in appearance, personality, and function of digital humans, enhances the naturalness and immersion of interaction with users, improves adaptability and expressiveness, reduces development costs, and increases efficiency.
Smart Images

Figure CN119478207B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of digital human development, specifically to a rapid digital human development method, system, electronic device, and storage medium. Background Technology
[0002] With the development of computer technology, digital humans, serving as a bridge between the virtual and real worlds, are widely used in various fields, such as virtual customer service, virtual anchors, and personalized entertainment. Digital technology has not only improved the efficiency of content creation in the entertainment industry but has also driven the application of artificial intelligence and virtual reality technologies.
[0003] In existing technologies, common methods for rapidly constructing digital humans that meet specific needs include, but are not limited to: using readily available 3D model libraries for rapid configuration; combining animation tools to customize the digital human's appearance and movement; and writing interaction scripts for specific scenarios based on development requirements to simulate interactions and feedback in a real environment, thereby realizing the digital human's basic functions. While these methods can meet the demand for efficiency in digital human development to some extent, they are still insufficient when dealing with situations involving complex personalities and emotional interactions. Specifically, in the traditional digital human creation process, the integration between appearance features, personality traits, and preset functions needs further improvement, especially since personalized preset functions often fail to fully adapt to the needs of specific application scenarios.
[0004] Therefore, there is a need for a method to develop digital humans capable of handling complex scenarios and emotional interactions. Summary of the Invention
[0005] This application provides a method, system, electronic device, and storage medium for rapid development of digital humans, which can quickly build and optimize digital humans according to needs.
[0006] The first aspect of this application provides a rapid development method for digital humans, applied to a digital human development platform, the rapid development method for digital humans comprising:
[0007] When the first character's requirements and the first scene requirements are received, the first appearance characteristics and personality characteristics are determined according to the first character's requirements, and the second appearance characteristics and preset functions are determined according to the first scene requirements. The first appearance characteristics include gender, age, face shape and skin color, the second appearance characteristics include facial features, hairstyle and clothing, the personality characteristics include humor, seriousness and gentleness, and the preset functions include voice interaction, action performance and emotional communication.
[0008] A first digital human is constructed based on the first appearance feature, the second appearance feature, the personality feature, and the preset functions;
[0009] Based on the requirements of the first scenario, an activity scenario and a test scenario are constructed, wherein the test scenario includes triggering conditions, interactive objects, and expected results;
[0010] The first digital human is trained according to the test scenario to obtain the second digital human.
[0011] By adopting the above technical solution, the requirements of the first person and the first scenario are received and analyzed, enabling the rapid determination of the digital human's appearance, personality traits, and required functions, thus achieving efficient customized development. The determination of the first appearance features (gender, age, face shape, and skin tone), the second appearance features (facial features, hairstyle, and clothing), as well as personality traits (humorous, meticulous, gentle) and preset functions (voice interaction, motor expression, emotional communication) ensures the comprehensiveness and personalization of the digital human. Combining the first and second appearance features, the constructed digital human can highly replicate the appearance of a real person, enhancing the user's visual realism. Based on the personality traits, the digital human can exhibit specific behavioral patterns and emotional responses, making it closer to a real person or meeting the needs of specific scenarios. Preset functions include voice interaction, motor expression, and emotional communication; the integration of these functions allows the digital human to interact with the user more naturally and comprehensively. By constructing activity scenarios and test cases according to scenario requirements and conducting targeted training on the digital human, the adaptability and expressiveness of the digital human in different scenarios are improved. Training the digital human through test cases allows for the timely identification and resolution of potential problems, thereby continuously optimizing the digital human's performance and behavior.
[0012] Optionally, training the first digital human to obtain the second digital human based on the test scenario includes:
[0013] Acquire first limb movement data of a real human in the test scenario, and perform a first process on the first limb movement data to obtain second limb movement data. The first process includes noise removal, smoothing of movement curves, and adjustment of movement speed.
[0014] The joint angle and position information in the second limb movement data are mapped to the corresponding position of the first digital human to obtain the second digital human.
[0015] By employing the above technical solution to acquire real human limb movement data, it is possible to ensure that the digital human's movements are based on real human motion data, thereby increasing the realism and credibility of the movements. The first processing of the first limb movement data—removing noise, smoothing the movement curve, and adjusting the movement speed—effectively reduces errors and jitter in the data, making the digital human's movements smoother and more natural, conforming to the laws of human movement. Precisely mapping each joint angle and position information in the processed second limb movement data to the corresponding position on the first digital human ensures that the digital human's movements maintain a high degree of consistency with real human movements in detail. This precise mapping not only improves the accuracy of the digital human's movements but also enhances its expressiveness and immersion. Since this method is trained based on real human movement data in test scenarios, it can flexibly respond to movement requirements under different scenarios and needs. Whether it's simple daily movements or complex professional movements, this method can be used for training and mapping. As test scenarios change or needs are updated, new limb movement data can be acquired and processed again, and the first digital human can be retrained to obtain a second digital human adapted to new scenarios or needs. This flexibility and adjustability give the embodiments of this application broad application prospects and the possibility of continuous optimization. Because the movements of the second digital human are more natural and realistic, it can significantly enhance the immersion and quality of the user experience when interacting with the digital human. This improvement not only helps to strengthen the user's sense of identification and trust in the digital human, but also promotes a more natural and fluid interaction process.
[0016] Optionally, training the first digital human to obtain the second digital human based on the test scenario includes:
[0017] The first facial expression data of real humans under various emotions in the test scenario is obtained, and the first facial expression data is processed to obtain the second facial expression data. The second processing includes noise removal, classification of expressions according to emotion category, definition of a unified expression description language, expression coding standard and expression evaluation standard.
[0018] Identify the current emotional state of the first digital human, select a matching facial expression based on the current emotional state, and map the matching facial expression onto the first digital human to obtain the second digital human.
[0019] By employing the aforementioned technical solutions, first facial expression data of real humans under various emotions in test scenarios is obtained, providing a rich foundation for the emotional expression of digital humans. This data ensures the authenticity and credibility of digital humans in expressing emotions. Secondary processing, such as noise removal and expression classification according to emotion categories, improves the purity and usability of the first facial expression data. Simultaneously, defining a unified expression description language, expression coding standard, and expression evaluation standard provides a standardized framework for the processing and application of expression data, further enhancing the accuracy and consistency of expressions. By identifying the current emotional state of the first digital human, the system can accurately determine the type of emotion it needs to express. This function enables the digital human to more intelligently adjust its expressions according to the current context, enhancing the authenticity and naturalness of the interaction. Selecting matching facial expressions based on the current emotional state and mapping these expressions onto the first digital human yields the second digital human. This process achieves precise matching of emotion and expression, making the digital human's emotional expression more nuanced and vivid. Because the second digital human can express various emotions more accurately, users experience a stronger sense of immersion and realism when interacting with it. This immersion helps improve user engagement and satisfaction. Through precise facial expression mapping, the second digital human can better understand the user's emotional needs and respond with corresponding facial expressions. This emotional exchange not only enhances the interactivity between the user and the digital human but also promotes emotional connection and trust between them. This process involves multiple technical fields, including emotion recognition, facial expression classification, and facial expression mapping, driving continuous innovation and development in these technologies. These advancements provide strong support for the further application of digital human technology.
[0020] Optionally, training the first digital human to obtain the second digital human based on the test scenario includes:
[0021] The speech rate of the first digital human is set according to the personality traits, the timbre of the first digital human is set according to the age and gender, and the tone of the first digital human is set according to the current emotional state and dialogue content.
[0022] The second digital human is obtained by training the first digital human to engage in dialogue according to a preset dialogue script.
[0023] By employing the aforementioned technical solutions, setting the speech rate based on personality traits, the timbre based on age and gender, and the tone of voice based on the current emotional state and dialogue content, these personalized settings make the digital human's expression more in line with human behavioral habits and emotional expression, thereby enhancing the naturalness and realism of the expression. By adjusting the tone of voice according to the current emotional state, the digital human can more accurately convey its emotions, enhancing its emotional communication ability with users. This emotional adaptability makes the digital human more vivid and interesting during the interaction process, increasing user participation and satisfaction. Training the first digital human with pre-set dialogue scripts allows it to become familiar with and master expression methods in various dialogue scenarios and contexts, thereby improving the fluency and coherence of the dialogue. This training also helps the digital human better understand and respond to user input, achieving more efficient interaction. Digital humans with rich expressive capabilities can be applied in multiple fields, such as virtual customer service, online education, and entertainment interaction. In these fields, digital humans can provide a more natural, fluent, and personalized interactive experience, meeting diverse user needs. By improving the expressive capabilities of digital humans, users can experience a more realistic and vivid interactive experience, thereby increasing their identification with and satisfaction with the digital human. This positive user experience helps enhance users' trust and acceptance of digital human technology. The continuous exploration and application of technologies such as setting digital human speech rate, timbre, and intonation, as well as dialogue training, have driven the continuous innovation and development of digital human technology. These technological achievements not only improve the expressive capabilities of digital humans but also provide valuable experience and references for further research and application of digital human technology in the future.
[0024] Optionally, constructing the first digital human based on the first appearance feature, the second appearance feature, the personality feature, and the preset function includes:
[0025] Build an initial 3D model and create a skeletal system, which includes multiple joints and bones;
[0026] The skeletal system is bound to the initial 3D model so that the initial 3D model can undergo corresponding changes when the skeleton moves.
[0027] The initial 3D model is adjusted based on the first appearance feature, the second appearance feature, and the personality feature to obtain an initial digital human;
[0028] The preset functions are added to the initial digital human to construct the first digital human.
[0029] By adopting the above technical solutions, an initial 3D model was constructed and a skeletal system was created, providing a solid foundation for the digital human. The introduction of the skeletal system makes the digital human's movements more natural and fluid, while also facilitating subsequent appearance adjustments. Adjustments were made to the initial 3D model based on primary appearance features (such as face shape and hairstyle), secondary appearance features (such as clothing and accessories), and personality traits, resulting in a highly personalized initial digital human. This personalization is not only reflected in appearance but also conveys the digital human's personality traits through fine-tuning model details. Binding the skeletal system to the initial 3D model allows the movement of the skeleton to directly drive corresponding changes in the model. This design greatly enhances the digital human's ability to express movement, enabling it to perform various complex actions and expressions, improving the realism and immersion of the interaction. Pre-set functions, such as speech recognition, natural language processing, and emotion computing, were added to the initial digital human, giving it not only a realistic appearance and movements but also the ability to interact deeply with the user. The integration of these functions greatly expands the application scenarios and value of the digital human. The entire construction process adopted a modular design approach, including modules for 3D modeling, skeletal binding, appearance adjustment, and function integration. This design allows development teams to flexibly adjust the development process according to actual needs, improving development efficiency. At the same time, the modular design also helps reduce development costs, as developers can reuse existing modules to build new digital humans.
[0030] Optionally, constructing the activity scenario and test scenario based on the requirements of the first scenario includes:
[0031] When the activity scenario is an interview, the interviewee is constructed as the interactive object in the test scenario, the interviewee's self-introduction is constructed as the trigger condition, and the preset questions and preset answers are determined based on the self-introduction;
[0032] A first emotional state is determined based on the similarity between the interviewee's answer and the preset answer, and a second emotional state is determined based on the interviewee's answer emotion. The response state of the first digital human is determined based on the first emotional state and the second emotional state.
[0033] By employing the aforementioned technical solutions, and through pre-setting questions and answers, as well as automatically judging the interviewee's emotional state and the quality of their responses, First Digital Human can intelligently generate replies. This intelligent question-and-answer method not only improves interview efficiency but also makes the interview process more objective and accurate. Based on the interviewee's emotional state and the content of their answers, First Digital Human can adjust its response state, thus providing more personalized feedback. This personalized feedback helps create a better interview atmosphere and enhances the interviewee's sense of participation and trust. By constructing interview activity scenarios and test situations to simulate a real interview environment, interviewees can demonstrate their abilities and qualities in near-realistic situations. This realistic scenario simulation helps improve the credibility of the interview results. The interviewee's answers and emotional state affect First Digital Human's response state, forming a dynamic interactive process. This dynamic interaction not only makes the interview process more vivid and interesting but also helps to uncover deeper abilities and qualities in the interviewee. Through pre-setting questions and answers and an automated emotion judgment mechanism, First Digital Human can quickly process interviewees' answers and generate corresponding replies. This automated processing method greatly improves interview efficiency. By comprehensively analyzing interviewees' answers and emotional state, First Digital Human can more accurately assess their abilities and qualities. This accurate assessment helps companies select talent more scientifically. First Digital Human can also act as a remote interviewer, enabling interviews across geographical boundaries. This remote interviewing method not only saves time and costs for both companies and interviewees but also increases the flexibility and convenience of the interview process. By analyzing interviewees' performance and feedback, First Digital Human can also provide personalized training suggestions and guidance. This personalized training helps improve interviewees' abilities and qualities, cultivating more outstanding talent for companies.
[0034] Optionally, determining a first emotional state based on the similarity between the interviewee's answer and the preset answer, and determining a second emotional state based on the interviewee's answer emotion, and determining the response state of the first digital human based on the first emotional state and the second emotional state includes:
[0035] Natural language processing technology is used to calculate the similarity between the answer content and the preset answer. A first emotional state is matched based on the similarity. Sentiment analysis technology is used to analyze the emotional words in the answer content to determine the interviewee's answer emotion. A second emotional state is matched based on the answer emotion.
[0036] Based on the first emotional state and the second emotional state, a corresponding target state is matched from a preset emotion table, and the target state is used as the response state of the first digital human.
[0037] By employing the aforementioned technical solutions, the similarity between the interviewee's answer and the preset answer can be calculated to objectively assess how close the interviewee's answer is to the standard answer. This text-based similarity calculation reduces the subjectivity of human judgment and improves the accuracy of the assessment. Sentiment analysis technology can identify emotional words in the answer content and judge the interviewee's emotional state accordingly. This technology makes the assessment of the interviewee's emotional state more detailed and accurate, helping to gain a more comprehensive understanding of the interviewee's psychological state. Matching the corresponding target state from the preset sentiment table based on the first emotional state (based on the similarity between the answer content and the preset answer) and the second emotional state (based on the emotion of the answer), this process ensures that the first digital human's response can be personalized according to the interviewee's specific performance and emotional state. This targeted response helps to establish a more positive and effective communication atmosphere. Since the first digital human's response state is dynamically adjusted according to the interviewee's real-time performance and emotional state, it can better adapt to changes in the interview process, improving the flexibility and adaptability of the response. The entire assessment and response process is automated, reducing the need for human intervention and improving the efficiency of the interview. At the same time, automated processing also reduces the possibility of human error and improves the reliability of the interview. Through personalized responses and targeted assessments, First Digital Human provides interviewees with a more attentive and professional interview experience. This experience helps enhance interviewee engagement and satisfaction, improving overall interview effectiveness. The process combines natural language processing and sentiment analysis technologies, showcasing the integrated application of different technological fields. This technological fusion not only enriches the means and methods of interview assessment but also provides valuable lessons and inspiration for technological innovation in other fields. Applying natural language processing and sentiment analysis technologies to interview scenarios is an innovative application attempt. This innovative application not only promotes the development and progress of related technologies but also brings new opportunities and challenges to the field of interviewing.
[0038] A second aspect of this application provides a rapid digital human development system, comprising a feature module, a construction module, a scene module, and a training module, wherein:
[0039] The feature module is configured to, when receiving a first person's request and a first scene's request, determine a first appearance feature and a personality feature based on the first person's request, and determine a second appearance feature and a preset function based on the first scene's request. The first appearance feature includes gender, age, face shape, and skin color; the second appearance feature includes facial features, hairstyle, and clothing; the personality feature includes humor, seriousness, and gentleness; and the preset function includes voice interaction, action performance, and emotional communication.
[0040] The construction module is configured to construct a first digital human based on the first appearance feature, the second appearance feature, the personality feature, and the preset functions;
[0041] The scenario module is configured to construct activity scenarios and test scenarios based on the requirements of the first scenario. The test scenarios include triggering conditions, interactive objects, and expected results.
[0042] A training module is configured to train the first digital human to obtain a second digital human based on the test scenario.
[0043] A third aspect of this application provides an electronic device including a processor, a memory, a user interface, and a network interface, wherein the memory is used to store instructions, the user interface and the network interface are both used to communicate with other devices, and the processor is used to execute the instructions stored in the memory to cause the electronic device to perform the method as described in any of the foregoing.
[0044] A fourth aspect of this application provides a computer-readable storage medium storing instructions that, when executed, perform the method described in any of the preceding descriptions.
[0045] In summary, one or more technical solutions provided in the embodiments of this application have at least the following technical effects or advantages:
[0046] 1. Capable of quickly generating digital humans that meet user requirements for characters and scenarios. By clearly defining primary appearance features, secondary appearance features, personality traits, and preset functions, the development process of digital humans is made more efficient; users can customize the gender, age, face shape, skin color, facial features, hairstyle, clothing, and personality traits of the digital human as needed, ensuring that the digital human accurately matches the requirements of specific scenarios and roles.
[0047] 2. Digital humans not only possess basic physical characteristics but also integrate preset functions such as voice interaction, motion expression, and emotional communication. These functions enable digital humans to interact with users more naturally and richly in practical applications. By constructing activity scenarios and test situations, digital humans are trained to exhibit corresponding behaviors and reactions in different scenarios, enhancing their adaptability and flexibility.
[0048] 3. Through the detailed setting of appearance and personality traits, as well as the rich preset functions, the digital human is made to be closer to the real person in appearance and behavior, thereby enhancing the user's sense of immersion and realism; the integration of functions such as voice interaction, action performance and emotional communication enables the digital human to have a deeper and more meaningful interaction with the user, enhancing the richness and fun of the user experience.
[0049] 4. By building and training automated processes for digital humans, the degree of human intervention is reduced, thereby improving development efficiency and reducing development costs. Once the digital human is developed and trained, its model and parameters can be reused in other similar projects or scenarios, further improving development efficiency and resource utilization. Attached Figure Description
[0050] Figure 1 This is a flowchart illustrating the rapid development method for digital humans disclosed in an embodiment of this application;
[0051] Figure 2 This is a schematic diagram of the modules of the rapid digital human development system disclosed in the embodiments of this application;
[0052] Figure 3 This is a schematic diagram of the structure of an electronic device disclosed in an embodiment of this application.
[0053] Explanation of reference numerals in the attached figures: 201, Feature module; 202, Construction module; 203, Scene module; 204, Training module; 301, Processor; 302, Communication bus; 303, User interface; 304, Network interface; 305, Memory. Detailed Implementation
[0054] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments.
[0055] In the description of the embodiments of this application, the words "for example" or "for instance" are used to indicate examples, illustrations, or explanations. Any embodiment or design that is described as "for example" or "for instance" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design options. Rather, the use of the words "for example" or "for instance" is intended to present the relevant concepts in a specific manner.
[0056] In the description of the embodiments of this application, the term "multiple" means two or more. For example, multiple systems means two or more systems, and multiple screen terminals means two or more screen terminals. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the indicated technical features. Thus, a feature defined with "first" or "second" may explicitly or implicitly include one or more of that feature. The terms "comprising," "including," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.
[0057] This embodiment discloses a rapid development method for digital humans, applied to a digital human development platform. Figure 1 This is a flowchart illustrating the rapid development method for digital humans disclosed in an embodiment of this application, as shown below. Figure 1 As shown, the method includes the following steps:
[0058] S110. When receiving the first person's requirements and the first scene requirements, determine the first appearance feature and personality feature based on the first person's requirements, and determine the second appearance feature and preset function based on the first scene requirements. The first appearance feature includes gender, age, face shape and skin color. The second appearance feature includes facial features, hairstyle and clothing. The personality feature includes humor, seriousness and gentleness. The preset function includes voice interaction, action performance and emotional communication.
[0059] The digital human development platform includes multiple models, which can be used to implement the following steps. The first step, character requirements, refers to the setting of the character's basic attributes and personality, typically derived from the storyline, game settings, or user needs. Based on these requirements, the character's basic appearance is first determined. This includes:
[0060] Gender: Whether the character is male or female.
[0061] Age: The age range of the character, such as child, youth, middle-aged or elderly.
[0062] Face shape: The character's facial contours, such as round face, square face, oval face, etc.
[0063] Skin color: The character's skin color, which can be any color within the range of natural skin tones, or it can include special colors (such as the skin color of non-human characters in science fiction or fantasy works).
[0064] Personality Traits: Based on the character's needs, further develop the character's personality traits, which will influence the character's behavior and language style, including:
[0065] Humor: Does the character possess a sense of humor? Can they express their views or handle situations in a lighthearted and witty manner?
[0066] Rigorous: Does the character pay attention to detail, act meticulously, and think logically?
[0067] Gentleness: Does the character show consideration and care for others?
[0068] The first scenario requirement refers to the character's environment, tasks, and interactions with other elements (such as other characters, objects, and events). Based on the scenario requirements, the character's appearance is further refined and adjusted to better integrate into the environment or represent the character's state. This includes:
[0069] Facial features, such as eye size, eyebrow shape, and lip thickness, can reflect a character's emotional state or racial characteristics.
[0070] Hairstyle: Hairstyles designed according to the historical context, cultural customs, or the character's personal preferences.
[0071] Clothing: A character's clothing style not only reflects their status and position, but is also influenced by the atmosphere of the scene and seasonal changes.
[0072] Pre-defined functions: Specific functions or abilities designed for a character to meet the needs of a given scenario. These functions can be:
[0073] Voice interaction: Characters can communicate with users or other characters via voice, increasing immersion and interactivity.
[0074] Action performance: Depending on the requirements of the scene, the character can perform specific actions, such as walking, running, jumping, fighting, etc.
[0075] Emotional communication: Expressing the character's emotional state through facial expressions, body language, and voice changes enhances the character's realism and appeal.
[0076] S120. Construct a first digital human based on the first appearance feature, the second appearance feature, the personality feature, and the preset function;
[0077] The basic appearance framework of the digital human is determined based on the primary physical characteristics (gender, age, face shape, and skin tone). This is the most intuitive and fundamental identifier for the digital human. Secondary physical characteristics (facial features, hairstyle, and clothing) are then incorporated to further refine and enrich the digital human's appearance. Adjustments to facial features can make it more consistent with specific races, cultural backgrounds, or personality traits; the choice of hairstyle and clothing can showcase the character's status, professional characteristics, or personal style. Personality traits are the expression of the digital human's inherent qualities, determining its behavioral responses and language expression in different situations. During the construction process, personality traits such as humor, seriousness, and gentleness need to be endowed to the digital human, giving it a unique personal charm. This is usually achieved through programming, allowing the digital human to react accordingly to its personality traits when interacting with users or other characters. Preset functions are the foundation for the digital human to realize its functions and roles. Preset functions such as voice interaction, action performance, and emotional communication are configured for the digital human according to scenario requirements. The implementation of these functions requires advanced technologies, such as speech recognition and synthesis technology, animation technology, and affective computing technology. Voice Interaction: Through voice recognition technology, the digital human can understand the user's voice commands or questions and generate corresponding answers or feedback through voice synthesis technology. Motion Performance: Utilizing animation technology, a rich motion library is designed for the digital human, including walking, running, jumping, fighting, and subtle changes in facial expressions to showcase its emotional state and personality traits. Emotional Communication: Through affective computing technology, the emotional state of the user or other characters is analyzed, and the digital human's expressions, tone, and behavior are adjusted accordingly to achieve more natural and realistic emotional communication. After completing the above steps, the first digital human needs comprehensive debugging and optimization. This includes checking whether the appearance features are consistent and harmonious, whether the personality traits are distinct and prominent, and whether the preset functions are stable and reliable. Through continuous debugging and optimization, it is ensured that the digital human perfectly presents the designer's intentions and desired effects.
[0078] Optionally, constructing the first digital human based on the first appearance feature, the second appearance feature, the personality feature, and the preset function includes:
[0079] Build an initial 3D model and create a skeletal system, which includes multiple joints and bones;
[0080] The skeletal system is bound to the initial 3D model so that the initial 3D model can undergo corresponding changes when the skeleton moves.
[0081] The initial 3D model is adjusted based on the first appearance feature, the second appearance feature, and the personality feature to obtain an initial digital human;
[0082] The preset functions are added to the initial digital human to construct the first digital human.
[0083] Building the initial 3D model is the first step in creating a digital human, involving using 3D modeling software (such as Maya, Blender, etc.) to sculpt a basic 3D model. This model is typically the general outline of the digital human, possibly containing basic body structure, but not yet detailed. The initial 3D model provides the basic framework for subsequent detailing and functional implementation. The skeletal system is central to the digital human's animation and motion performance. It consists of multiple joints and bones, simulating the skeletal structure of the real human body. In 3D modeling software, specialized skeletal rigging tools are typically used to create this system. Each joint and bone is precisely placed to ensure they can correctly simulate the range of motion of the human body. Rigging the skeletal system to the initial 3D model is a crucial step, ensuring that the initial 3D model changes accordingly when the bones move. This rigging is achieved through weighting, which determines the degree of association between each part of the model (such as skin, muscles) and the bones. The correctness of the weighting directly affects the realism and smoothness of the animation. After the skeletal system is bound to the initial 3D model, the model is meticulously adjusted based on primary appearance features (gender, age, face shape, and skin tone), secondary appearance features (facial features, hairstyle, and clothing), and personality traits. This includes refining facial features (such as the shape and position of the eyes, nose, and mouth), designing hairstyles, and adding clothing. Simultaneously, the model's expressions and postures are adjusted according to personality traits to better match the character's inner qualities. After the initial digital human model is adjusted, preset functions need to be added to achieve its specific purpose. These functions may include voice interaction, motion expression, and emotional communication. Depending on design requirements, these functions are implemented through programming and integration of relevant technologies (such as speech recognition and synthesis, animation engines, and affective computing). For example, a speech recognition module is configured for the digital human to understand user voice commands; an animation engine is used to drive the skeletal system to achieve complex movements and expressions; and affective computing technology is used to simulate the digital human's emotional responses. The final step is to comprehensively debug and optimize the completed first digital human. This includes checking whether the model's appearance meets expectations, whether the animation is smooth and natural, and whether the functions are stable and reliable. Through continuous testing and adjustments, we ensure that the digital human perfectly reflects the designer's intentions and desired effects.
[0084] By creating an initial 3D model and equipping it with a complex skeletal system, including multiple joints and bones, the digital human can simulate the structure and movement of the real human body. The skeletal system provides the foundation for the digital human's dynamic performance, ensuring natural and fluid movements and postures. Precisely binding the skeletal system to the initial 3D model ensures that bone movement drives corresponding changes in the model. This step is crucial for achieving the digital human's dynamic effects, enabling it to perform various complex movements and expressions, enhancing its realism and interactivity. The initial 3D model is meticulously adjusted based on primary appearance features (gender, age, face shape, and skin tone) and secondary appearance features (facial features, hairstyle, and clothing), giving the digital human a unique appearance. This process reflects a high degree of personalization and customization, meeting the needs of different scenarios and roles. Through programming and design, personality traits such as humor, rigor, and gentleness are integrated into the digital human, giving it a unique charm. This integration of personality traits not only enhances the digital human's expressiveness but also makes it more in line with user expectations and preferences. By incorporating pre-set functions such as voice interaction, motion expression, and emotional communication into the initial digital human, the digital human can interact with users or other characters in various forms. The addition of these functions not only enriches the application scenarios of the digital human but also enhances the user experience and immersion. By implementing functions such as voice interaction and motion expression, the digital human can interact with users more naturally. This enhanced interactivity not only strengthens user participation and immersion but also makes the digital human a more intelligent and engaging interactive medium. The entire construction process involves knowledge and technical means from multiple technical fields, including 3D modeling, skeletal rigging, animation design, speech recognition and synthesis, and affective computing. The comprehensive application of these technologies not only reflects the achievements of technological innovation but also demonstrates the potential and value of technology in practical applications. The constructed first digital human not only possesses high realism and dynamic expressiveness but also rich functionality and powerful interactivity. This gives the digital human broad application prospects and practical value in multiple fields such as games, animation, virtual reality, and education.
[0085] S130. Construct an activity scenario and test scenario based on the requirements of the first scenario. The test scenario includes triggering conditions, interactive objects, and expected results.
[0086] An activity scene is the specific environment in which the digital human will interact and perform. It is constructed based on the requirements of the primary scene, including the physical layout of the scene, lighting conditions, background elements, other characters or objects, etc. These elements together constitute a complete, context-sensitive environment, providing a stage for the digital human's behavior and performance.
[0087] When constructing an event scenario, the following aspects need to be considered:
[0088] Scenario realism: Ensure that the scenario matches the actual or expected situation so that the digital human can interact and behave naturally within it.
[0089] Diversity of scenarios: Construct multiple different scenarios to test the adaptability and flexibility of digital humans in various environments.
[0090] Scenario scalability: The scenario is designed with future changes in requirements in mind, making it easy to modify and extend.
[0091] A test scenario is a set of specific events or situations defined within an activity scenario to test whether the digital human's reactions and behaviors meet expectations. Each test scenario contains the following three key elements:
[0092] Triggering condition: This refers to the condition or event that causes the test scenario to occur. It defines the conditions that begin the test and provides a starting point for the digital human's actions. For example, a user uttering a specific voice command or the digital human detecting the presence of an object.
[0093] Interaction objects: Entities that interact with the digital human in the test scenario. These can be users, other digital humans, objects, or the virtual environment itself. These interaction objects define the targets that the digital human needs to interact with.
[0094] Expected Outcome: The anticipated outcome of the digital human's behavior or performance after the test scenario concludes. It is used to evaluate whether the digital human reacted or performed correctly as expected. For example, the digital human should provide a corresponding response or perform a corresponding action after receiving a user's voice command.
[0095] When constructing test scenarios, the following aspects should be considered:
[0096] Comprehensive coverage: Ensure that the test scenarios cover all situations and behaviors that the digital human may encounter in order to fully evaluate its performance and reliability.
[0097] Repeatability: When designing test scenarios, ensure that they can be executed repeatedly so that they can be tested multiple times at different times or under different conditions.
[0098] Measurability: Define clearly defined expected results so that the performance of the digital human can be objectively evaluated to ensure that it meets the requirements.
[0099] Optionally, constructing the activity scenario and test scenario based on the requirements of the first scenario includes:
[0100] When the activity scenario is an interview, the interviewee is constructed as the interactive object in the test scenario, the interviewee's self-introduction is constructed as the trigger condition, and the preset questions and preset answers are determined based on the self-introduction;
[0101] A first emotional state is determined based on the similarity between the interviewee's answer and the preset answer, and a second emotional state is determined based on the interviewee's answer emotion. The response state of the first digital human is determined based on the first emotional state and the second emotional state.
[0102] Based on the requirements of the first scenario, construct a simulated interview scenario. This scenario should include the physical environment required for the interview (such as the meeting room, table and chair layout, lighting, etc.) and virtual elements (such as the interviewee's digital representation, the display interface for interview questions, etc.). It is important to ensure that the atmosphere and details of the scenario closely resemble a real interview environment, in order to provide a realistic interactive platform for the digital human and the interviewee. In the interview scenario, the construction of test situations mainly revolves around the interviewee's behavior and the digital human's reactions. Specifically:
[0103] Interaction Objects: Interviewees are constructed as interaction objects in the testing scenario. They are the subjects with whom the digital human needs to communicate and evaluate.
[0104] Trigger condition: The interviewee's self-introduction is set as the trigger condition. This action marks the formal start of the interview and prompts the digital human to prepare subsequent questions and assessments based on the content of the self-introduction.
[0105] Pre-set questions and answers: Based on the interviewee's self-introduction, DigitalMan will pre-set a series of relevant questions and prepare corresponding pre-set answers. These questions aim to further understand the interviewee's abilities, experience, and attitude, while the pre-set answers serve as a reference standard for evaluating the quality of the interviewee's responses.
[0106] During the interview, the digital human dynamically adjusts its response based on the interviewee's answers. This process involves two key emotional state judgments:
[0107] The first emotional state is determined based on the similarity between the interviewee's answer and the pre-set answer. If the interviewee's answer is highly consistent with the pre-set answer, it may indicate that they have a deep understanding and preparation in this field, and the digital human may exhibit a positive emotional state (such as satisfaction or encouragement). Conversely, if the answer differs significantly from the pre-set answer, the digital human may exhibit a more cautious or skeptical emotional state.
[0108] The second emotional state is determined based on the interviewee's emotional expression during the interview. The interviewee's tone of voice, facial expressions, and other nonverbal cues can also convey their emotional state. Digital humans need to possess emotion recognition capabilities to accurately capture these emotional signals and adjust their responses accordingly. For example, if the interviewee exhibits nervousness or anxiety, the digital human might adopt a gentler, more reassuring response.
[0109] By simulating real interview scenarios and processes, the first digital human can communicate and interact with interviewees in real time. This interactivity not only enhances the realism of the test but also helps assess the digital human's performance in different situations. A first emotional state (e.g., accuracy, comprehensiveness) is determined based on the similarity between the interviewee's answers and preset answers, and a second emotional state is determined based on the interviewee's emotional state (e.g., nervousness, confidence). The combination of these two emotional states makes the assessment more comprehensive and in-depth, helping to more accurately understand the interviewee's psychological state and performance level. The first digital human's response state is determined based on the first and second emotional states. This dynamic response mechanism allows the digital human to flexibly adjust its response content and method according to the interviewee's actual situation, thus better simulating the interaction process in a real interview. Pre-set questions and answers can significantly save interview time and costs. Simultaneously, the digital human's real-time assessment capabilities make the assessment process more efficient and accurate. Based on the performance of different interviewees, the first digital human's response strategy and assessment criteria can be personalized. This personalized development helps to better meet the needs of different users, improving user satisfaction and experience quality.
[0110] Optionally, determining a first emotional state based on the similarity between the interviewee's answer and the preset answer, and determining a second emotional state based on the interviewee's answer emotion, and determining the response state of the first digital human based on the first emotional state and the second emotional state includes:
[0111] Natural language processing technology is used to calculate the similarity between the answer content and the preset answer. A first emotional state is matched based on the similarity. Sentiment analysis technology is used to analyze the emotional words in the answer content to determine the interviewee's answer emotion. A second emotional state is matched based on the answer emotion.
[0112] Based on the first emotional state and the second emotional state, a corresponding target state is matched from a preset emotion table, and the target state is used as the response state of the first digital human.
[0113] Natural Language Processing (NLP) techniques are used to calculate the similarity between an interviewee's response and a pre-set answer. NLP is a technique in computer science used to understand and process human language. In this process, NLP analyzes the textual features of the response and the pre-set answer, such as vocabulary, grammatical structure, and semantic meaning, and then calculates a similarity score. Similarity scores can be derived using various algorithms, such as cosine similarity, Jaccard similarity, or deep learning-based methods. Based on the calculated similarity score, a first emotional state is matched from a pre-set list of emotional states. These emotional states may be based on a range of similarity scores; for example, high similarity might represent an "accurate" or "professional" emotional state, while low similarity might represent a "vague" or "uncertain" emotional state. This matching process allows the digital human to assess the quality or accuracy of the interviewee's response based on how closely it resembles the pre-set answer. Simultaneously, sentiment analysis techniques are used to analyze the emotional vocabulary in the interviewee's response to determine its emotional tone. Sentiment analysis is an NLP technique used to identify the emotional tendency in text, such as positive, negative, or neutral. In interview scenarios, sentiment analysis helps digital humans understand the emotional state of interviewees when they answer questions, such as confidence, nervousness, or hesitation. By analyzing sentiment words and phrases in the answers, sentiment analysis technology can generate a sentiment score or label to describe the interviewee's emotional state. Based on the sentiment analysis-derived answer sentiment, a second sentiment state is matched from a pre-set list of sentiment states. This second sentiment state is determined based on the interviewee's actual emotion during the answer and differs from the first sentiment state (based on similarity). For example, if sentiment analysis indicates that the interviewee expressed confidence during the answer, then the second sentiment state might be "confident." Based on the first and second sentiment states, a corresponding target state is matched from a pre-set sentiment table and used as the digital human's response state. The pre-set sentiment table is a database or mapping table containing different sentiment states and their corresponding response strategies. By combining the first sentiment state (based on the accuracy of the answer content) and the second sentiment state (based on the answer sentiment), the digital human can choose the most appropriate response strategy to address the interviewee's answer. This strategy might include providing further clarification questions, giving positive feedback, adjusting the difficulty of the questions, etc., to ensure the smooth progress of the interview process.
[0114] Using NLP technology to calculate the similarity between the interviewee's response and the pre-set answer can accurately capture the accuracy and relevance of the interviewee's response, thereby inferring the first emotional state (e.g., accurate, vague, off-topic). Sentiment analysis technology can further analyze the emotional vocabulary in the response, identifying the interviewee's emotional tendency (e.g., positive, negative, neutral), providing strong support for determining the second emotional state. By combining the first emotional state (based on content accuracy) and the second emotional state (based on sentiment tendency), a more comprehensive understanding of the interviewee's response can be achieved, leading to more flexible and appropriate response strategies. This strategy not only considers the accuracy of the content but also takes into account the interviewee's emotional feelings, helping to create a more positive and harmonious interview atmosphere. The first digital avatar can dynamically adjust its response state based on the interviewee's actual response, making the interaction process more natural and smooth. This personalized response method helps improve the interviewee's user experience, making them feel more valued and respected. Through pre-set sentiment tables and matching mechanisms, the first digital avatar's response state can be quickly determined, avoiding the tedious and time-consuming manual evaluation. This not only improves evaluation efficiency but also reduces the influence of human factors on the evaluation results, making the evaluation results more objective and accurate. This process integrates NLP and sentiment analysis techniques, demonstrating the application potential of artificial intelligence in complex scenarios. This technological fusion not only drives the development and progress of related technologies but also provides valuable exploration and reference for intelligent applications in more similar scenarios in the future.
[0115] S140. Train the first digital human according to the test conditions to obtain the second digital human.
[0116] The training objectives need to be clearly defined. For example, in the context of an interview scenario, training objectives might include improving the digital human's ability to accurately understand interviewees' answers, its emotion recognition capabilities, and its ability to respond appropriately based on these capabilities. Additionally, it might include optimizing the digital human's reaction speed and enhancing its natural language processing fluency. Based on the training objectives, a specific training plan should be designed. The training plan should include the following aspects:
[0117] Dataset preparation: Based on the testing scenarios, collect and organize a large amount of interview dialogue data. This data should cover different interview scenarios, answers, and emotional states to simulate the diversity of real interviews.
[0118] Model tuning: Adjust the algorithmic model behind the first digital human based on the training objectives. This may include optimizing the natural language processing model to improve text understanding, adjusting the sentiment analysis model to more accurately identify emotional states, and modifying the response generation model to generate more appropriate responses.
[0119] Training strategy: Develop a training strategy, including setting parameters such as training period, learning rate, and batch size. It is also necessary to determine how to evaluate the training effect and whether model tuning and parameter adjustments are needed.
[0120] The first digital human is trained according to the training plan. The training process may include the following steps:
[0121] Data preprocessing: The collected interview dialogue data is preprocessed, including noise removal, word segmentation, and sentiment annotation.
[0122] Model training: The preprocessed data is input into the adjusted algorithm model for training. During training, the model continuously learns how to better understand and analyze the interviewee's answers, and how to generate appropriate responses.
[0123] Performance evaluation: Regularly evaluate the model's performance during training. This can be achieved by comparing the model with data from the test scenario, assessing its performance in areas such as answer accuracy, sentiment recognition, and response appropriateness.
[0124] Model tuning: The model is tuned based on the evaluation results. If the model is found to perform poorly in certain aspects, performance can be improved by adjusting model parameters, improving the algorithm, or increasing training data.
[0125] After training, the second digital human (i.e., the trained first digital human) needs to be validated and tested. This can be achieved by placing the second digital human in real or simulated interview scenarios to evaluate its performance under various conditions. The results of validation and testing will be used to further confirm the training effectiveness and provide feedback for subsequent optimization. Based on the results of validation and testing, the second digital human is iteratively optimized. This may include further adjusting model parameters, improving the algorithm, increasing training data, etc. Through continuous iterative optimization, the performance of the second digital human can be gradually improved, making it more adaptable to complex interview scenarios and diverse interview requirements.
[0126] Optionally, training the first digital human to obtain the second digital human based on the test scenario includes:
[0127] Acquire first limb movement data of a real human in the test scenario, and perform a first process on the first limb movement data to obtain second limb movement data. The first process includes noise removal, smoothing of movement curves, and adjustment of movement speed.
[0128] The joint angle and position information in the second limb movement data are mapped to the corresponding position of the first digital human to obtain the second digital human.
[0129] Acquire real-world human body movement data in specific testing scenarios (such as interviews). This is typically achieved using motion capture technology, which records and tracks the human body's movement trajectory, joint angles, and positions in real time. This data forms the basis for subsequent processing and mapping. The raw body movement data often contains noise, uneven movement curves, or inconsistent movement speeds, which directly affect the mapping effect on the digital human. Therefore, a series of processing steps are required—the first processing—to obtain cleaner, smoother, and more uniformly paced second body movement data.
[0130] Noise removal: Noise in the data is removed through filtering and other methods to ensure the accuracy of the motion curve.
[0131] Smooth motion curves: Smooth the motion curves to eliminate jitter and abrupt changes, making the motion look more natural.
[0132] Adjust movement speed: Adjust the overall speed of movement as needed to ensure that the digital human's movement speed matches the speed expected by humans.
[0133] The processed second-limb motion data contains the angle and position information of each joint, which will be mapped to the corresponding positions on the first digital human. This step is the core of the training process, involving accurately converting human motion data into motion instructions that the digital human can understand and execute.
[0134] Joint mapping: Ensure that each joint in the second limb motion data matches the corresponding joint in the first digital human model.
[0135] Motion synchronization: The processed motion data is synchronized to the digital human model in chronological order to achieve continuous playback of the motion.
[0136] Detailed adjustments: Make detailed adjustments to the digital human's movements as needed, such as facial expressions and strength, to enhance the realism and expressiveness of the movements.
[0137] After the above steps, the first digital human has learned to mimic the body movements of real humans in test scenarios. These movements are not only more natural and fluid, but also conform to human habits. Therefore, the first digital human at this stage can be regarded as a trained and optimized second digital human.
[0138] Acquired real-world human body movement data often contains noise introduced by factors such as device precision and environmental interference. The noise removal step in the first processing stage can significantly reduce the impact of this noise on the digital human's movements, making the movements cleaner and more accurate. Smoothing processing eliminates abrupt changes and jitters in the movement data, resulting in smoother and more natural movement curves. This helps improve the fluidity and visual appeal of the digital human's movements, avoiding user discomfort. Adjusting the movement speed according to actual needs allows the digital human's movements to better match the rhythm and atmosphere of a specific scenario. For example, in an interview scenario, appropriately slowing down the movement speed can project a more composed and professional image. Accurately mapping the angle and position information of each joint in the second body movement data to the corresponding position on the first digital human is key to achieving a high degree of motion fidelity in the digital human. This step ensures that the digital human's movements accurately reflect the characteristics of real human movements, including posture, force, and speed. By adjusting mapping parameters and algorithms, personalized customization of the digital human's movements can also be achieved. For example, based on different users' preferences and needs, the digital human's movement style and facial expressions can be adjusted to better meet user expectations. The optimized second digital human exhibits more natural and fluid movements, significantly enhancing the user experience. In interview scenarios, this improvement creates a more realistic and welcoming atmosphere. Through training and adjustments, the second digital human can better adapt to different testing situations and scenario requirements. Whether it's a simple self-introduction or a complex Q&A session, the digital human displays appropriate movements and expressions, interacting smoothly with the user. The optimized digital human can be applied in more fields, such as education, entertainment, and healthcare. In education, it can serve as a virtual lecturer; in entertainment, it can perform as a virtual idol; and in healthcare, it can serve as an auxiliary tool for rehabilitation training. These applications will greatly enrich people's lives and work.
[0139] Optionally, training the first digital human to obtain the second digital human based on the test scenario includes:
[0140] The first facial expression data of real humans under various emotions in the test scenario is obtained, and the first facial expression data is processed to obtain the second facial expression data. The second processing includes noise removal, classification of expressions according to emotion category, definition of a unified expression description language, expression coding standard and expression evaluation standard.
[0141] Identify the current emotional state of the first digital human, select a matching facial expression based on the current emotional state, and map the matching facial expression onto the first digital human to obtain the second digital human.
[0142] In testing scenarios, the first step is to collect facial expression data from real humans under different emotional states. This can be achieved in various ways, such as recording real human reactions to various emotional stimuli using a high-definition camera, or extracting data from existing facial expression databases. This data should cover a wide range of emotion types, such as happiness, sadness, anger, surprise, fear, disgust, and neutrality. Facial expression data may contain noise introduced by factors such as changes in lighting, facial occlusion, and device vibration. Image processing techniques, such as filtering and denoising algorithms, can effectively remove this noise, improving the purity and accuracy of the data. The expression data should be categorized according to emotion type to ensure that each emotion has a sufficient sample size for subsequent analysis and training. This helps to build a more refined emotion recognition model. To achieve cross-platform and cross-system exchange and sharing of facial expression data, a unified expression description language needs to be defined. This language should include a comprehensive description of expression features, such as eye shape, mouth corner curvature, and eyebrow position, to facilitate understanding and conversion between different systems. An expression coding standard should be developed to convert facial expression data into standardized digital or symbolic forms for easy computer processing and storage. This also facilitates the automated recognition and classification of facial expression data. Establish expression evaluation criteria to assess the authenticity and validity of facial expression data. This can be based on various methods, such as expert scoring, user feedback, or automated evaluation algorithms.
[0143] The current emotional state of the first digital human is identified using emotion recognition technology (such as machine learning-based emotion recognition models). This can be achieved by analyzing the digital human's speech, text, or behavioral data. For example, in a dialogue scenario, the emotional state can be determined based on features such as the digital human's tone of voice, word choice, and pauses. Based on the identified current emotional state, a matching facial expression is selected from processed second facial expression data. This requires establishing a mapping relationship between emotional states and facial expressions to ensure that the selected expression accurately reflects the digital human's emotional state. The selected facial expression is then mapped onto the first digital human's facial model to generate the second digital human. This involves the parsing, rendering, and animation techniques used in facial expression data. Fine-grained animation control ensures subtle changes and natural transitions in facial expressions, thereby enhancing the realism and expressiveness of the digital human.
[0144] Denoising the acquired facial expression data eliminates noise introduced by factors such as acquisition equipment and environmental interference, ensuring the purity of the facial expression data and laying a solid foundation for subsequent processing. Categorizing expressions according to emotion categories and defining unified expression description languages, expression coding standards, and expression evaluation standards helps to standardize and regulate facial expressions, improving the accuracy and consistency of digital human facial expressions. Advanced artificial intelligence technology identifies the current emotional state of the digital human, ensuring that the digital human can accurately perceive and reflect its own emotional changes. Selecting matching facial expressions based on the current emotional state and mapping them onto the digital human achieves precise matching and real-time interaction between emotion and expression. This process allows the digital human to more vividly display emotional changes, enhancing the realism and immersion of the user experience. Digital humans with rich and natural facial expressions can interact with users more vividly and naturally, increasing user engagement and satisfaction. This technology is particularly effective in fields such as virtual customer service, online education, and entertainment games. Through continuous learning and optimization, digital humans can gradually master more emotional expression methods and skills, improving their intelligence and adaptability. This intelligence is not only reflected in the accurate expression of facial expressions, but also in the keen perception and timely response to changes in the user's emotions. This training process involves technologies from multiple fields, including facial expression recognition, emotion computing, and digital human modeling, accumulating valuable experience for the development of digital human technology. Simultaneously, through continuous technological innovation and optimization, the application and expansion of digital human technology in more fields can be promoted. Digital humans with rich emotional expression capabilities can be applied in more areas, such as virtual idols, virtual anchors, and virtual customer service. These applications can not only bring users a more vivid and interesting experience, but also create more commercial and social value for enterprises.
[0145] Optionally, training the first digital human to obtain the second digital human based on the test scenario includes:
[0146] The speech rate of the first digital human is set according to the personality traits, the timbre of the first digital human is set according to the age and gender, and the tone of the first digital human is set according to the current emotional state and dialogue content.
[0147] The second digital human is obtained by training the first digital human to engage in dialogue according to a preset dialogue script.
[0148] The speaking speed of the digital human should be adjusted according to its personality traits. For example, an introverted digital human might speak more slowly, while an extroverted one might speak faster. This setting helps enhance the expressiveness of the digital human's personality. The timbre should be set according to the digital human's age and gender. Different ages and genders often have different vocal characteristics; for example, children's voices are usually clearer, while older people's voices may be deeper; male voices are generally more rugged, while female voices are softer. Appropriate timbre settings can make the digital human's voice more closely match its identity characteristics. The tone of voice should be adjusted according to the digital human's current emotional state and the content of the conversation. Tone is one of the important means of expressing emotions; by changing the pitch, speed, and strength of the tone, different emotional colors and attitudes can be conveyed. For example, when expressing happiness, the tone might rise; when expressing sadness, the tone might fall. To train the digital human's conversational abilities, a series of pre-set dialogue scripts need to be prepared. These scripts should cover a variety of scenarios and topics to comprehensively test the digital human's conversational and responsiveness. Simultaneously, the script should include rich emotional expressions and contextual variations to simulate the complexities of real-world dialogue. During training, the pre-configured first digital human is placed within a pre-set dialogue script, and its speech rate, timbre, and tone are tested for appropriateness through simulated dialogue. Simultaneously, adjustments and optimizations are made in real-time based on the digital human's performance in the dialogue to ensure accurate and fluent expression of the dialogue content and to demonstrate an expressive style consistent with its personality traits and emotional state. After training, the dialogue performance of the second digital human needs to be evaluated. Evaluation criteria may include the fluency, naturalness, accuracy, and richness and accuracy of emotional expression. The evaluation results determine whether the training has achieved its intended goals, and the training program is further adjusted and optimized accordingly.
[0149] The speech rate and timbre of the digital human are set according to its personality traits, age, and gender, making its voice expression more consistent with its established identity. This personalized setting helps enhance the realism and credibility of the digital human, making it easier for users to resonate and connect emotionally. The tone of voice is adjusted based on the digital human's current emotional state and the content of the dialogue, making its voice expression more vivid and natural. This fusion of emotion and tone can convey richer emotional information, enhancing the digital human's emotional expression capabilities. Training the digital human with pre-set dialogue scripts can simulate real-life dialogue scenarios, helping it become familiar with and master various dialogue skills and expressions. This training method helps improve the naturalness and fluency of the digital human in actual conversations, enabling it to interact with users more naturally. As training progresses, the digital human's dialogue capabilities will be continuously optimized and improved. It can more accurately understand the user's intentions and needs and provide more appropriate and useful responses. This continuous optimization will allow the digital human to play its due role and value in more scenarios. Digital humans with rich and natural expressive abilities can be applied to various scenarios, such as virtual customer service, online education, and entertainment games. In these scenarios, digital humans can play different roles, interacting and communicating with users, bringing them a more vivid and engaging experience. Through the multi-dimensional setup and training of the first digital human, a wealth of technical experience and data resources can be accumulated. These resources and experience will provide strong support for subsequent innovation and development of digital human technology. With the continuous innovation and development of digital human technology, its application areas will continue to expand. In the future, digital human technology is expected to be applied and promoted in more fields, bringing greater convenience and value to society.
[0150] This embodiment also discloses a rapid development system for digital humans. Figure 2 This is a schematic diagram of the modules of the rapid digital human development system disclosed in the embodiments of this application, such as... Figure 2 As shown, the system includes a feature module 201, a construction module 202, a scene module 203, and a training module 204, wherein:
[0151] The feature module 201 is configured to, when receiving a first person's request and a first scene's request, determine a first appearance feature and a personality feature based on the first person's request, and determine a second appearance feature and a preset function based on the first scene's request. The first appearance feature includes gender, age, face shape, and skin color. The second appearance feature includes facial features, hairstyle, and clothing. The personality feature includes humor, seriousness, and gentleness. The preset function includes voice interaction, action performance, and emotional communication.
[0152] Construction module 202 is configured to construct a first digital human based on the first appearance feature, the second appearance feature, the personality feature, and the preset function;
[0153] Scene module 203 is configured to construct activity scenes and test scenarios according to the requirements of the first scene, wherein the test scenarios include triggering conditions, interactive objects, and expected results;
[0154] Training module 204 is configured to train the first digital human to obtain a second digital human based on the test scenario.
[0155] Optionally, the training module 204 is configured to:
[0156] Acquire first limb movement data of a real human in the test scenario, and perform a first process on the first limb movement data to obtain second limb movement data. The first process includes noise removal, smoothing of movement curves, and adjustment of movement speed.
[0157] The joint angle and position information in the second limb movement data are mapped to the corresponding position of the first digital human to obtain the second digital human.
[0158] Optionally, the training module 204 is configured to:
[0159] The first facial expression data of real humans under various emotions in the test scenario is obtained, and the first facial expression data is processed to obtain the second facial expression data. The second processing includes noise removal, classification of expressions according to emotion category, definition of a unified expression description language, expression coding standard and expression evaluation standard.
[0160] Identify the current emotional state of the first digital human, select a matching facial expression based on the current emotional state, and map the matching facial expression onto the first digital human to obtain the second digital human.
[0161] Optionally, the training module 204 is configured to:
[0162] The speech rate of the first digital human is set according to the personality traits, the timbre of the first digital human is set according to the age and gender, and the tone of the first digital human is set according to the current emotional state and dialogue content.
[0163] The second digital human is obtained by training the first digital human to engage in dialogue according to a preset dialogue script.
[0164] Optionally, the building module 202 is configured to:
[0165] Build an initial 3D model and create a skeletal system, which includes multiple joints and bones;
[0166] The skeletal system is bound to the initial 3D model so that the initial 3D model can undergo corresponding changes when the skeleton moves.
[0167] The initial 3D model is adjusted based on the first appearance feature, the second appearance feature, and the personality feature to obtain an initial digital human;
[0168] The preset functions are added to the initial digital human to construct the first digital human.
[0169] Optionally, the scene module 203 is configured to:
[0170] When the activity scenario is an interview, the interviewee is constructed as the interactive object in the test scenario, the interviewee's self-introduction is constructed as the trigger condition, and the preset questions and preset answers are determined based on the self-introduction;
[0171] A first emotional state is determined based on the similarity between the interviewee's answer and the preset answer, and a second emotional state is determined based on the interviewee's answer emotion. The response state of the first digital human is determined based on the first emotional state and the second emotional state.
[0172] Optionally, the scene module 203 is configured to:
[0173] Natural language processing technology is used to calculate the similarity between the answer content and the preset answer. A first emotional state is matched based on the similarity. Sentiment analysis technology is used to analyze the emotional words in the answer content to determine the interviewee's answer emotion. A second emotional state is matched based on the answer emotion.
[0174] Based on the first emotional state and the second emotional state, a corresponding target state is matched from a preset emotion table, and the target state is used as the response state of the first digital human.
[0175] It should be noted that the above embodiments of the apparatus are only illustrated by the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.
[0176] This embodiment also discloses an electronic device, as shown in the reference. Figure 3 The electronic device may include: at least one processor 301, at least one communication bus 302, user interface 303, network interface 304, and at least one memory 305.
[0177] The communication bus 302 is used to enable communication between these components.
[0178] The user interface 303 may include a display screen and a camera. Optionally, the user interface 303 may also include a standard wired interface and a wireless interface.
[0179] The network interface 304 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface).
[0180] The processor 301 may include one or more processing cores. The processor 301 connects to various parts of the server using various interfaces and lines, and performs various server functions and processes data by running or executing instructions, programs, code sets, or instruction sets stored in memory 305, and by calling data stored in memory 305. Optionally, the processor 301 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). The processor 301 may integrate one or a combination of several of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), and modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the content required for display; and the modem handles wireless communication. It is understood that the modem may also not be integrated into the processor 301 and may be implemented as a separate chip.
[0181] The memory 305 may include random access memory (RAM) or read-only memory. Optionally, the memory 305 may include a non-transitory computer-readable storage medium. The memory 305 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 305 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-described method embodiments, etc.; the data storage area may store data involved in the above-described method embodiments, etc. Optionally, the memory 305 may also be at least one storage device located remotely from the aforementioned processor 301. Figure 3 As shown, the memory 305, which serves as a computer storage medium, may include an operating system, a network communication module, a user interface module, and an application program for a rapid development method for digital humans.
[0182] exist Figure 3 In the electronic device shown, the user interface 303 is mainly used to provide an input interface for the user and to obtain the user input data; while the processor 301 can be used to call the application program stored in the memory 305 for rapid development of digital humans. When executed by one or more processors 301, the electronic device performs one or more methods as described in the above embodiments.
[0183] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0184] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0185] In the several embodiments provided in this application, it should be understood that the disclosed apparatus can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the shown or discussed mutual couplings or direct couplings or communication connections may be through some service interfaces; indirect couplings or communication connections between apparatuses or units may be electrical or other forms.
[0186] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0187] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0188] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage device (CMD). Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory 305 and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned memory 305 includes various media capable of storing program code, such as a USB flash drive, external hard drive, magnetic disk, or optical disk.
[0189] The foregoing description is merely an exemplary embodiment of this disclosure and should not be construed as limiting the scope of this disclosure. Any equivalent changes and modifications made in accordance with the teachings of this disclosure shall still fall within the scope of this disclosure. Other embodiments of this disclosure will be readily apparent to those skilled in the art upon consideration of the disclosure in this specification. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not described in this disclosure. The specification and embodiments are to be considered exemplary only, and the scope and spirit of this disclosure are defined by the claims.
Claims
1. A rapid development method for digital humans, characterized in that, The rapid development method for digital humans, applied to a digital human development platform, includes: When the first character's requirements and the first scene requirements are received, the first appearance characteristics and personality characteristics are determined according to the first character's requirements, and the second appearance characteristics and preset functions are determined according to the first scene requirements. The first appearance characteristics include gender, age, face shape and skin color, the second appearance characteristics include facial features, hairstyle and clothing, the personality characteristics include humor, seriousness and gentleness, and the preset functions include voice interaction, action performance and emotional communication. A first digital human is constructed based on the first appearance feature, the second appearance feature, the personality feature, and the preset functions; Based on the requirements of the first scenario, an activity scenario and a test scenario are constructed, wherein the test scenario includes triggering conditions, interactive objects, and expected results; The first digital human is trained according to the test conditions to obtain the second digital human. The construction of activity scenarios and test scenarios based on the requirements of the first scenario includes: When the activity scenario is an interview, the interviewee is constructed as the interactive object in the test scenario, the interviewee's self-introduction is constructed as the trigger condition, and preset questions and preset answers are determined based on the self-introduction; A first emotional state is determined based on the similarity between the interviewee's answer and the preset answer; a second emotional state is determined based on the interviewee's answer emotion; and the response state of the first digital human is determined based on the first and second emotional states. The process of determining a first emotional state based on the similarity between the interviewee's answer and the preset answer, determining a second emotional state based on the interviewee's answer emotion, and determining the response state of the first digital human based on the first and second emotional states includes: Natural language processing technology is used to calculate the similarity between the answer content and the preset answer. A first emotional state is matched based on the similarity. Sentiment analysis technology is used to analyze the emotional words in the answer content to determine the interviewee's answer emotion. A second emotional state is matched based on the answer emotion. Based on the first emotional state and the second emotional state, a corresponding target state is matched from a preset emotion table, and the target state is used as the response state of the first digital human. The step of training the first digital human to obtain the second digital human based on the test scenario includes: The first facial expression data of real humans under various emotions in the test scenario is obtained, and the first facial expression data is processed to obtain the second facial expression data. The second processing includes noise removal, classification of expressions according to emotion category, definition of a unified expression description language, expression coding standard and expression evaluation standard. The current emotional state of the first digital human is identified, a matching facial expression is selected based on the current emotional state, and the matching facial expression is mapped onto the first digital human to obtain the second digital human. The step of training the first digital human to obtain the second digital human based on the test scenario includes: The speech rate of the first digital human is set according to the personality traits, the timbre of the first digital human is set according to the age and gender, and the tone of the first digital human is set according to the current emotional state and dialogue content. The second digital human is obtained by training the first digital human to engage in dialogue according to a preset dialogue script.
2. The rapid development method for digital humans according to claim 1, characterized in that, The step of training the first digital human to obtain the second digital human based on the test scenario includes: Acquire first limb movement data of a real human in the test scenario, and perform a first process on the first limb movement data to obtain second limb movement data. The first process includes noise removal, smoothing of movement curves, and adjustment of movement speed. The joint angle and position information in the second limb movement data are mapped to the corresponding position of the first digital human to obtain the second digital human.
3. The rapid development method for digital humans according to claim 1, characterized in that, The construction of the first digital human based on the first appearance feature, the second appearance feature, the personality feature, and the preset functions includes: Build an initial 3D model and create a skeletal system, which includes multiple joints and bones; The skeletal system is bound to the initial 3D model so that the initial 3D model can undergo corresponding changes when the skeleton moves. The initial 3D model is adjusted based on the first appearance feature, the second appearance feature, and the personality feature to obtain an initial digital human; The preset functions are added to the initial digital human to construct the first digital human.
4. A rapid development system for digital humans, characterized in that, It includes a feature module, a construction module, a scene module, and a training module, among which: The feature module is configured to, when receiving a first person's request and a first scene's request, determine a first appearance feature and a personality feature based on the first person's request, and determine a second appearance feature and a preset function based on the first scene's request. The first appearance feature includes gender, age, face shape, and skin color; the second appearance feature includes facial features, hairstyle, and clothing; the personality feature includes humor, seriousness, and gentleness; and the preset function includes voice interaction, action performance, and emotional communication. The construction module is configured to construct a first digital human based on the first appearance feature, the second appearance feature, the personality feature, and the preset functions; The scenario module is configured to construct activity scenarios and test scenarios based on the requirements of the first scenario. The test scenarios include triggering conditions, interactive objects, and expected results. The training module is configured to train the first digital human to obtain a second digital human based on the test conditions. The construction of activity scenarios and test scenarios based on the requirements of the first scenario includes: When the activity scenario is an interview, the interviewee is constructed as the interactive object in the test scenario, the interviewee's self-introduction is constructed as the trigger condition, and preset questions and preset answers are determined based on the self-introduction; A first emotional state is determined based on the similarity between the interviewee's answer and the preset answer; a second emotional state is determined based on the interviewee's answer emotion; and the response state of the first digital human is determined based on the first and second emotional states. The process of determining a first emotional state based on the similarity between the interviewee's answer and the preset answer, determining a second emotional state based on the interviewee's answer emotion, and determining the response state of the first digital human based on the first and second emotional states includes: Natural language processing technology is used to calculate the similarity between the answer content and the preset answer. A first emotional state is matched based on the similarity. Sentiment analysis technology is used to analyze the emotional words in the answer content to determine the interviewee's answer emotion. A second emotional state is matched based on the answer emotion. Based on the first emotional state and the second emotional state, a corresponding target state is matched from a preset emotion table, and the target state is used as the response state of the first digital human. The step of training the first digital human to obtain the second digital human based on the test scenario includes: The first facial expression data of real humans under various emotions in the test scenario is obtained, and the first facial expression data is processed to obtain the second facial expression data. The second processing includes noise removal, classification of expressions according to emotion category, definition of a unified expression description language, expression coding standard and expression evaluation standard. The current emotional state of the first digital human is identified, a matching facial expression is selected based on the current emotional state, and the matching facial expression is mapped onto the first digital human to obtain the second digital human. The step of training the first digital human to obtain the second digital human based on the test scenario includes: The speech rate of the first digital human is set according to the personality traits, the timbre of the first digital human is set according to the age and gender, and the tone of the first digital human is set according to the current emotional state and dialogue content. The second digital human is obtained by training the first digital human to engage in dialogue according to a preset dialogue script.
5. An electronic device, characterized in that, The device includes a processor, a memory, a user interface, and a network interface. The memory is used to store instructions. The user interface and the network interface are both used to communicate with other devices. The processor is used to execute the instructions stored in the memory to cause the electronic device to perform the method as described in any one of claims 1-3.
6. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions that, when executed, perform the method as described in any one of claims 1-3.
Citation Information
Patent Citations
Digital human adjustment method and device based on text emotion, equipment and storage medium
CN116304183A
Virtual image generation method and device and nonvolatile storage medium
CN117292089A
Digital interaction method and system based on artificial intelligence, and medium
CN117348736A
Virtual digital human-based interaction processing method, system, terminal, equipment and medium
CN117520498A
Processing method, device and equipment for intelligent simulation interview and medium
CN118093826A