A method for the rapid creation of virtual digital humans

By combining text models and script-based action models for training, the problems of unstable and unresponsive virtual digital human movements have been solved, achieving stable and flexible control under text logic and enabling the rapid creation of complex actions.

CN117151159BActive Publication Date: 2026-01-06CHINA UNICOM WO MUSIC & CULTURE CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311131420.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-04
Publication Date
2026-01-06
Estimated Expiration
2043-09-04

AI Technical Summary

Technical Problem

Existing technologies lack efficient text logic control methods when virtual digital humans perform complex actions and respond to events, resulting in unstable actions and insufficient adaptability.

Method used

By acquiring motion data of virtual digital humans, a control strategy is generated using a text model. The model parameters are then adjusted through overflow and differential feedback to train the text model and achieve stable control of the virtual digital humans. This is combined with script-based motion models for multi-scenario adaptive training.

Benefits of technology

It has improved the stability and adaptability of virtual digital humans under the control of text models, enabling the rapid creation of complex and continuous actions that conform to real-world situations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117151159B_ABST
    Figure CN117151159B_ABST
Patent Text Reader

Abstract

The application discloses a virtual digital person quick creation generation method, and relates to the technical field of computer content generation.The virtual digital person quick creation generation method provided by the application can control and drive each part of the virtual digital person, the control and driving is according to joint actions formed according to a text model, and reference is made to time action data, so that the actions of the digital person are under the guidance of a strategy, and the action characteristics are more in line with real conditions, the digital person has better stability and adaptability under the control of the text model, and under the trained architecture, the virtual digital person can quickly create complex continuous actions according to recombination of the text.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer content generation technology, and in particular to a method for quickly creating virtual digital humans. Background Technology

[0002] With the development of natural language processing tools such as ChatGPT, which are driven by large-scale model artificial intelligence technology, the problem of content generation in human-computer interaction has been solved. These AI language tools are trending towards personalization, which will drive digital interaction towards a virtual universe. Further integration of AI language tools with virtual digital humans will enable better simulation of virtual universe scenarios.

[0003] With the development of artificial intelligence technology, virtual digital humans are increasingly appearing in scenarios such as virtual customer service, short videos, and live streaming. Therefore, virtual digital humans need to evolve from simply performing facial expressions or actions to responding to text or logic and processing large volumes of content. Summary of the Invention

[0004] In view of this, in order to solve the problems existing in the prior art, the present invention provides a method for generating virtual digital humans quickly, which can combine simple controls on the expressions and movements of virtual digital humans to enable virtual digital humans to quickly create a series of actions based on text logic.

[0005] To solve the above-mentioned technical problems, the present invention adopts the following technical solution:

[0006] A method for quickly creating virtual digital humans, the method comprising:

[0007] The current state action data of a virtual digital human in an event scenario is obtained, and the current state action data is input into a text model to be trained to obtain a control strategy generated by the text model based on the current state action data.

[0008] Using the obtained control strategy, the virtual digital human is controlled to interact with the event scene, and the action operation data of the next state of the virtual digital human and the overflow feedback of the event scene are determined.

[0009] Based on the event action data and the action operation data of the current state, determine the digital human action strategy corresponding to the control strategy, and determine the spillover feedback between the control strategy and the corresponding digital human action strategy.

[0010] Based on the overflow and differential feedback, the parameters of the text model are adjusted. The action data of the next state corresponding to the virtual digital human is input into the adjusted text model, and the text model is trained until the set training termination condition is reached, thus obtaining the trained text model.

[0011] Optionally, the event scenario includes the scene view and interactive object view of the current activity area of ​​the virtual digital human;

[0012] The interactive object screen includes interactive objects in the current activity area of ​​the digital virtual human in the scene screen, and indicates the interactive space of the activity area;

[0013] Image features are extracted from the scene images and interactive object images of the event scene to obtain the scene features of the event scene;

[0014] Based on the scene characteristics of the event scenario, the current posture characteristics and current facial expression characteristics of the virtual digital human, the action operation data of the virtual digital human in the current state of the event scenario is generated.

[0015] Optionally, the event scenario in this method includes a second-person digital human interacting with the virtual digital human.

[0016] The aforementioned control strategy, which controls the interaction between the virtual digital human and the event scene, determines the action execution data of the next state of the virtual digital human and the overflow feedback from the event scene, includes:

[0017] Using the obtained control strategy, the virtual digital human is controlled to interact with the second-role digital human in the event scene, and the interaction process data and the overflow feedback obtained from the event scene are obtained.

[0018] Obtain the action execution data of the next state corresponding to the virtual digital human in the interaction process data.

[0019] Optionally, this method can execute the steps of the virtual digital human interacting with the second-person digital human in parallel using multiple independent training resources to obtain interaction process data.

[0020] Optionally, the method further includes: after obtaining the trained text model, saving the virtual digital human corresponding to the trained text model to an interactive model pool for saving historical versions of virtual digital humans, so as to serve as a new virtual digital human when retraining the text model.

[0021] Optionally, the method further includes: after obtaining the trained text model, acquiring the action data of the virtual digital human in a virtual environment, inputting the action data of the virtual digital human into the trained text model, and obtaining the control strategy generated by the text model based on the action data of the virtual digital human.

[0022] The obtained control strategy is used to control the virtual digital human to perform actions.

[0023] Optionally, the method includes determining the digital human action strategy corresponding to the control strategy based on the event action data and the action execution data of the current state, which includes:

[0024] The action data of the virtual digital human in the current state of the event scene is input into the trained script action model to obtain the digital human action strategy output by the script action model; the script action model is trained based on the selected event action data.

[0025] The differential feedback between the determined control strategy and the corresponding digital human motion strategy includes:

[0026] Based on the posture differences between the control strategy and the corresponding digital human motion strategy, spillover feedback is determined.

[0027] Optionally, the training process of the script action model includes:

[0028] Acquire selected event action data; the event action data includes the action execution data and action data of the selected corresponding virtual digital human;

[0029] The motion data of the virtual digital human is input into the script action model to be trained, and the prediction strategy output by the script action model is obtained to control the virtual digital human to perform actions.

[0030] The motion data of the virtual digital human based on the prediction strategy is compared with the motion data of the virtual digital human in the event motion data to determine the real-time posture data;

[0031] The parameters of the script action model are adjusted based on the determined real-time posture data, and the script action model with adjusted parameters is trained until the real-time posture data converges to the set expected value, thus obtaining the trained script action model.

[0032] In a secondary aspect, the present invention provides the following technical solution: a computer device, including a processor, the processor being configured to execute a computer program stored in a memory to implement the method for generating virtual digital humans quickly, as described above.

[0033] In a secondary aspect, the present invention provides the following technical solution: a computer-readable storage medium having a computer program stored thereon, wherein a processor is used to execute the computer program stored in the storage medium to implement the above-mentioned method for generating virtual digital humans quickly.

[0034] Compared with existing technologies, the virtual digital human creation method provided by this invention allows each part of the virtual digital human to be controlled and driven. The control and driving are based on the joint actions formed by the text model and refer to time action data, so that the actions of the digital human are guided by strategy and are more in line with the action characteristics of real-world situations. This makes the digital human more stable and adaptable under the control of the text model. Under the trained architecture, the virtual digital human can quickly create complex continuous actions according to the recombination of text.

[0035] To make the above-mentioned objects, features and advantages of this disclosure more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0036] Figure 1 This is a flowchart illustrating the method for quickly creating virtual digital humans according to an embodiment of the present invention.

[0037] Figure 2 This diagram illustrates the implementation architecture of the method for quickly creating virtual digital humans according to an embodiment of the present invention. Detailed Implementation

[0038] For clarity, the invention has been described with reference to specific embodiments; however, it should be understood that the invention is not limited to the described embodiments. Rather, the invention encompasses substitutions, modifications, and equivalents that may be included within the scope defined by any of the patent claims.

[0039] Figure 1 The flowchart shown is a method for quickly creating virtual digital humans according to an embodiment of this application. This method can be performed by... Figure 2 The process can be executed by server 20 or other electronic devices. This application scenario includes multiple terminals 10 and server 20. Terminals 10 and server 20 can connect and transmit data via wired or wireless connections. The virtual digital human controlled by terminal 10 undergoes training, while server 20 or other electronic devices collect data during the training process.

[0040] For example, the following describes the specific implementation process of the virtual digital human quick creation generation method of this application embodiment, using a server for generating text models and corresponding joint actions of virtual digital humans as the execution subject. Figure 1 As shown, the method includes the following steps:

[0041] Step S101: Obtain the current state action data of the virtual digital human in the event scene, input the current state action data into the text model to be trained, and obtain the control strategy generated by the text model based on the current action data.

[0042] The state of a virtual digital human in an event scenario changes over time, and the action data corresponding to the current moment is called the action data of the current state. The action data of a certain state may include, but is not limited to, the posture characteristics of the virtual digital human, the scene characteristics of the event scenario, etc.

[0043] The text model to be trained can be a neural network model, including an input layer, hidden layers, and an output layer. This neural network model can employ a policy function to execute a control policy based on the input action, and adjusting the parameters of the hidden layers can adjust the output control policy.

[0044] Step S102: Using the obtained control strategy, control the virtual digital human to interact with the event scene, and determine the action operation data of the next state of the virtual digital human and the overflow feedback of the event scene.

[0045] In some embodiments, a control strategy output by a text model can be used to control the virtual digital human to interact with a second-person digital human in the event scene, obtaining interaction process data and spillover feedback from the event scene. This interaction process data includes scene features of the event scene in which the virtual digital human is located, the virtual digital human's current posture features, and its current facial expression features. Based on this interaction process data, the action execution data for the virtual digital human's next state can be obtained.

[0046] Step S103: Based on the event action data and the current state action operation data, determine the digital human action strategy corresponding to the control strategy, and determine the differential feedback between the control strategy and the corresponding digital human action strategy.

[0047] The event-action data can be pre-collected data that controls the actions of real people in a scene. For example, a digital human action strategy can be generated based on the event-action data to guide the text model in generating control strategies.

[0048] Step S104: Based on the overflow feedback and differential feedback, adjust the parameters of the text model, input the action data of the next state corresponding to the virtual digital human into the adjusted text model, and continue to train the text model until the set training termination condition is reached, and obtain the trained text model.

[0049] Specifically, the model parameters of the text model are adjusted based on the augmentation feedback determined in step S102 and the differential feedback determined in step S103 in order to train the text model.

[0050] The process involves obtaining the trained text model, acquiring the action data of the virtual digital human in the virtual environment, inputting the action data of the virtual digital human into the trained text model, and obtaining the control strategy generated by the text model based on the action data of the virtual digital human. The obtained control strategy is then used to control the virtual digital human to perform actions.

[0051] The process of inputting the action data for the next state into the text model and continuing to train the text model can be performed according to steps S101 to S105 above. Executing steps S101 to S105 once can be considered as training the text model once. The training termination conditions may include, but are not limited to, reaching a set number of training iterations, or the convergence of incremental and differential feedback to a set expected value.

[0052] In one embodiment, training the text model a preset number of times can be considered as performing one round of training on the text model. After completing one round of training, the overall reward can be determined based on the feedback obtained from each training session in this round. The reward is used to evaluate the ability of the text model. If the change in the reward obtained from N consecutive rounds of training is within a set range or the reward reaches a set threshold, it indicates that the text model has reached its upper limit of ability, and the training process of the text model can be stopped. Otherwise, steps S101 to S105 are repeated to continue training the text model.

[0053] The text model obtained through the above training method can be used to control virtual digital humans in a virtual environment. Because this method, in training the text model, not only adapts the virtual digital human to the event scenario but also references event action data, and under the guidance of the digital human's action strategy, the text model controlling the virtual digital human can learn multiple control strategies. This results in a text model with better stability and adaptability, enabling it to output control strategies in the virtual environment that better meet the needs of the event scenario and achieve better results in controlling the virtual digital human.

[0054] It should be noted that the virtual digital human creation method provided by this invention allows for the control and driving of various parts of the virtual digital human. This control and driving is based on a combined action formed by a text model and references time-based action data, ensuring that the digital human's actions are strategically guided and more consistent with real-world movement characteristics. For example, when the text control module is set to "get up," the virtual digital human will interact within the room's event scene, performing actions such as rotating its body relative to its legs (sitting upright), flipping one hand from one side of its body to the other (lifting the blanket), bending its legs, rotating its body (to the bedside), and finally placing its legs on the ground. This series of coordinated actions is achieved through the text control module.

[0055] Therefore, the virtual digital human rapid creation generation method in this embodiment can enable the digital human to have better stability and adaptability under the control of the text model. Under the trained architecture, the virtual digital human can quickly create complex continuous actions according to the recombination of text.

[0056] Furthermore, to enable the virtual digital human to quickly create more complex content, in this embodiment, a second-role digital human is set up in the event scene, and the aforementioned virtual digital human becomes the first-role digital human. The first-role digital human and the second-role digital human are controlled by different text control modules and employ different control strategies. The first-role digital human and the second-role digital human control the virtual digital human to interact with the event scene, determining the action execution data of the virtual digital human's next state and the overflow feedback from the event scene. Specifically, the action execution data includes interaction process data, obtaining the action execution data of the virtual digital human's next state from the interaction process data. Overflow feedback is obtained from the posture similarity between the action execution data of the current state and the action execution data of the next state.

[0057] It is understandable that the text-based control model can combine the posture control variables of the first and second digital human characters based on the same control. Specifically, however, the interaction process data is obtained by executing the steps of the first and second digital human characters interacting with each other in parallel using multiple independent training resources. In some embodiments, the second digital human character is obtained by saving the first digital human character corresponding to the trained text model into an interaction model pool used to store historical versions of virtual digital humans.

[0058] Furthermore, to make the training of the virtual digital human more realistic, in this embodiment, the event scene includes the scene image of the virtual digital human's current activity area and the interactive object image. The scene image is a preset event scene, such as a room, shopping mall, street, tourist attraction, etc. The interactive object image includes the interactive objects in the scene image representing the current activity area of ​​the virtual digital human, and indicates the interactive space within the activity area. It can be understood that the interactive objects associated with specific scenes are diverse, but generally can be divided into relatively simple types such as non-touchable and touchable. Specifically, the scene image and interactive object image of the event scene can be obtained by extracting image features from the data image to obtain the scene features of the event scene. Based on the environmental features of the event scene, the current posture features and current facial expression features of the virtual digital human, motion operation data of the virtual digital human's current state in the event scene is generated. This motion operation data serves as the data source for training.

[0059] Furthermore, to make the training of the virtual digital human more realistic, this embodiment allows the trained text model to be logically combined to form a script action model. For example, multiple event scenarios and multiple virtual digital humans can be combined as needed. The script action model is then further trained. The action execution data of the virtual digital human in the current state of the event scenario is input into the trained action script model to obtain the digital human action strategy output by the action script model. The action script model is trained based on selected event action data and determines the differential feedback between the control strategy and the corresponding digital human action strategy.

[0060] Specifically, the current state and action data of the virtual digital human in the event scene are input into a trained script action model to obtain the digital human action strategy output by the script action model. The action data of the virtual digital human is input into a script action model to be trained to obtain a prediction strategy output by the script action model for controlling the virtual digital human's actions. The action data of the virtual digital human based on the prediction strategy is compared with the action data of the virtual digital human in the event action data to determine the real-time posture data. The parameters of the script action model are adjusted according to the determined real-time posture data, and the script action model with adjusted parameters is trained again until the real-time posture data converges to the set expected value, thus obtaining the trained script action model.

[0061] The above embodiments of this application can be implemented using a computer-readable storage medium that provides a method for quickly creating virtual digital humans. This computer-readable storage medium stores a computer program, which, when executed by a processor, performs the steps of the special effects display method in the above method embodiments. The storage medium can be either volatile or non-volatile. The computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0062] The above embodiments of this application can form a computer program product for a method of quickly creating virtual digital humans. The computer program product carries program code, and the instructions included in the program code can be used to execute the steps of the special effects display method in the above method embodiments.

[0063] The aforementioned computer program product can be implemented through hardware, software, or a combination thereof. In one optional embodiment, the computer program product is specifically embodied in a computer storage medium; in another optional embodiment, the computer program product is specifically embodied in a software product, such as a software development kit (SDK), etc.

[0064] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the division of units is only a logical functional division of a method for quickly creating virtual digital humans. In actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some communication interfaces. The indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0065] The embodiments provided in this application are merely specific implementations of this application, used to illustrate the technical solutions of this application, and are not intended to limit it. The protection scope of this application is not limited thereto. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments within the scope of the technology disclosed in this application, or make equivalent substitutions for some of the technical features; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be covered within the protection scope of this application.

Claims

1. A method for generating a virtual digital human quick creation, characterized in that, The method comprises: obtaining action running data of a current state of a virtual digital person in an event scene, inputting the action running data of the current state into a text model to be trained to obtain a control strategy generated by the text model according to the action running data of the current state; controlling the virtual digital person to interact with the event scene by using the obtained control strategy, determining action running data of a next state corresponding to the virtual digital person and overflow feedback fed back by the event scene; determining a digital person action strategy corresponding to the control strategy according to event action data and the action running data of the current state, and determining difference overflow feedback between the control strategy and the corresponding digital person action strategy; the event action data is data of actions of a real person in a scene collected in advance; and the determination of the digital person action strategy corresponding to the control strategy according to the event action data and the action running data of the current state comprises: inputting the action running data of the current state of the virtual digital person in the event scene into a trained script action model to obtain a digital person action strategy output by the script action model; the training process of the script action model comprises: inputting action running data of a virtual digital person into a script action model to be trained to obtain a prediction strategy output by the script action model for controlling the virtual digital person to perform actions; comparing action data of the virtual digital person based on the prediction strategy with action data of the virtual digital person in the event action data to determine real-time posture data; adjusting parameters of the script action model according to the determined real-time posture data, and continuing to train the script action model after the parameters are adjusted until the real-time posture data converges to a set expected value, thereby obtaining a trained script action model; adjusting parameters of the text model according to the overflow feedback and the difference overflow feedback, inputting action running data of a next state corresponding to the virtual digital person into the text model after the parameters are adjusted, and continuing to train the text model until a set training end condition is reached, thereby obtaining a trained text model.

2. The method of claim 1, wherein, the event scene comprises a scene picture of an activity area currently occupied by the virtual digital person and an interactive object picture; the interactive object picture contains an interactive object in the activity area currently occupied by the virtual digital person in the scene picture, and is marked with an interactable space of the activity area; image feature extraction is performed on the scene picture and the interactive object picture of the event scene to obtain scene features of the event scene; action running data of a current state of the virtual digital person in the event scene is generated according to the scene features of the event scene, current posture features and current expression features of the virtual digital person.

3. The method of claim 2, wherein, the event scene comprises a second role digital person that interacts with the virtual digital person; The obtained control strategy is used to control the virtual digital human to interact with the event scene, determine the action running data of the next state of the virtual digital human, and obtain the overflow feedback of the event scene, including: using the obtained control strategy to control the virtual digital human to interact with the second role digital human in the event scene, obtaining the interaction process data and the overflow feedback obtained from the event scene; The action running data of the next state of the virtual digital human in the interaction process data is obtained.

4. The method of claim 3, wherein, The step of the virtual digital human interacting with the second role digital human is executed in parallel by multiple independent training resources to obtain the interaction process data.

5. The method of claim 1, wherein, The method further comprises: after obtaining the trained text model, saving the virtual digital human corresponding to the trained text model to an interaction model pool for saving historical versions of virtual digital humans, for use as a new virtual digital human when training the text model again.

6. The method of claim 1, wherein, The method further comprises: after obtaining the trained text model, obtaining the action running data of the virtual digital human in the virtual environment, inputting the action running data of the virtual digital human into the trained text model, obtaining the control strategy generated by the text model according to the action running data of the virtual digital human, and using the obtained control strategy to control the virtual digital human to perform actions.

7. The method of claim 1, wherein, The script action model is obtained by training based on selected event action data; The difference overflow feedback between the control strategy and the corresponding digital human action strategy includes: determining the difference overflow feedback according to the posture difference between the control strategy and the corresponding digital human action strategy.

8. A computer apparatus, characterized in that, The processor is used to execute the computer program stored in the memory to realize the generation method of the virtual digital human quick creation according to any one of claims 1-7.

9. A computer readable storage medium having stored thereon a computer program, characterized in that, The processor is used to execute the computer program stored in the memory to realize the generation method of the virtual digital human quick creation according to any one of claims 1-7.

Citation Information

Patent Citations

  • Multi-modal interaction method, device and system based on virtual character, storage medium and terminal

    CN112162628A

  • Method and device for driving digital human and electronic equipment

    CN113689530A