Drawing method and device based on digital human, storage medium and program product

By introducing drawing modes and image drawing models in digital human interaction and generating and displaying drawing process videos, the problem of a single form of traditional digital human interaction is solved, and a richer and more diverse user interaction experience is achieved.

CN120017924APending Publication Date: 2025-05-16BEIJING 58 INFORMATION TTECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510174307.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

Traditional digital people have a single interaction form in live broadcasts, which is difficult to attract the audience's attention and cannot meet the audience's diverse interaction needs.

Method used

By introducing a drawing mode in the interaction between digital people and users, the image drawing model is used to analyze the drawing process of the user's video screen, generate target videos, and transform the digital people from the conversation form to the drawing form, showing the drawing process to the user.

Benefits of technology

It enriches the interaction forms between digital people and users, meets the diverse interaction needs of viewers, and improves the attractiveness and interactivity of live broadcasts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120017924A_ABST
    Figure CN120017924A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a drawing method and device based on a digital person, a storage medium and a program product, and the method comprises the steps: responding to a video connection request of a user, receiving and displaying a video image of the user, and controlling the target digital person to carry out the session with the user in a process that the target digital person carries out the live broadcast in a session form; in the session process with the user, determining a target object which needs to be drawn by the target digital person and a target image where the target object is located; in response to the image drawing event, inputting the target image into an image drawing model, and performing drawing process analysis on the target image to obtain a target video; and converting the target digital person from the session form into a drawing form, and displaying the target video to the user, so that the user can perceive the drawing process. Through the mode, a drawing mode is introduced in interaction between the digital human and the user, interaction with the user is carried out based on the video picture of the user, the interaction form is enriched, and diversified interaction requirements of audiences are met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a digital human-based drawing method, device, storage medium and program product. Background Art

[0002] In the field of live broadcasting, digital humans refer to virtual 3D humanoid models created through computer graphics, artificial intelligence and other technologies, which can appear as anchors or assistants in live broadcasts. These digital humans not only have realistic appearances, but can also interact with viewers through technologies such as voice recognition, natural language processing and emotional computing.

[0003] However, the current traditional digital humans mainly rely on voice or text to interact with users, resulting in a single form of interaction. During the live broadcast, this single interaction mode is difficult to attract the audience's attention and cannot meet the audience's diverse needs. Summary of the invention

[0004] Multiple aspects of the present application provide a drawing method, device, storage medium and program product based on a digital human, which are used to introduce a drawing mode in the interaction between a digital human and a user, and then interact with the user based on the user's video screen.

[0005] The embodiment of the present application provides a drawing method based on a digital human, comprising: in the process of a target digital human live broadcasting in a conversational form, responding to a user's video connection request, receiving and displaying the user's video screen, and controlling the target digital human to have a conversation with the user; in the process of having a conversation with the user, determining a target object that needs to be drawn by the target digital human and a target image where the target object is located; in response to an image drawing event, inputting the target image into an image drawing model, analyzing the drawing process of the target image to obtain a target video for describing the drawing process of the target image; converting the target digital human from the conversational form to the drawing form, and displaying the target video to the user so that the user can perceive the drawing process.

[0006] Further optionally, responding to a user's video connection request, receiving and displaying the user's video screen includes: responding to the user's video connection request, running a preset automation script, the automation script being used to establish an audio and video connection with the user according to the video connection request, receiving the user's video screen through the audio and video connection, and displaying the user's video screen.

[0007] Further optionally, during a conversation with the user, a target object that needs to be drawn by the target digital human and a target image where the target object is located are determined, including: when displaying the video screen of the user, taking a screenshot of the user's video screen at least once; using a human body detection model to detect whether each screenshot obtained by the screenshot meets the preset image requirements; stopping the screenshot when a target screenshot that meets the preset image requirements is detected, and taking the user as the target object and the target screenshot as the target image.

[0008] Further optionally, the image drawing model includes: a drawing process prediction network, an action prediction network and a video generation network; the target image is input into the image drawing model, and a drawing process analysis is performed on the target image to obtain a target video for describing the drawing process of the target image, including: inputting the target image into the drawing process prediction network, converting the target image into a line draft to obtain a target line draft, and predicting multiple drawing step information required to draw the target line draft; inputting the multiple drawing step information into the action prediction network, and predicting the action parameters of the target digital human performing the multiple drawing steps according to the multiple drawing step information; inputting the action parameters into the video generation network, and controlling the target digital human to perform the multiple drawing steps according to the action parameters to simulate the drawing process of the target image to obtain the target video.

[0009] Further optionally, the drawing process prediction network includes: a first extraction layer and a first prediction layer; predicting multiple drawing step information required to draw the target line draft, including: inputting the target line draft into the first extraction layer, extracting a first feature of the target line draft; inputting the first feature into the first prediction layer, and predicting multiple drawing step information required to draw the target line draft based on the first feature.

[0010] Further optionally, it also includes: providing a drawing style selection page to the user, and in response to a selection operation, determining a target drawing style selected by the user; inputting the first feature into the first prediction layer, and predicting multiple drawing step information required to draw the target line draft based on the first feature, including: determining a target prompt word corresponding to the target drawing style from multiple prompt words corresponding to multiple preset drawing styles; inputting the target prompt word and the first feature into the first prediction layer, and under the guidance of the target prompt word, predicting the multiple drawing step information that conforms to the target drawing style based on the first feature.

[0011] Further optionally, the action prediction network includes: a second extraction layer and a second prediction layer; based on the multiple drawing step information, predicting the action parameters of the target digital human performing the multiple drawing steps, including: inputting the multiple drawing step information into the second extraction layer, extracting the second features of the multiple drawing step information; inputting the second features of the multiple drawing step information into the second prediction layer, respectively predicting the action parameters of the target digital human performing the multiple drawing steps; the action parameters include: action start and end positions, action amplitude and / or action speed.

[0012] An embodiment of the present application also provides a terminal device, comprising: a memory and a processor; wherein the memory is used to: store one or more computer instructions; the processor is used to execute the one or more computer instructions to: execute the steps in the digital human-based drawing method.

[0013] The embodiment of the present application also provides a computer-readable storage medium, which, when the computer program is executed by a processor, enables the processor to implement the steps in the digital human-based drawing method.

[0014] The embodiment of the present application also provides a computer program product, including a computer program / instruction. When the computer program / instruction is executed by a processor, the processor is enabled to implement the steps in the digital human-based drawing method.

[0015] In this embodiment, when the target digital human is broadcasting live in a conversational form, it responds to the user's video connection request, receives and displays the user's video screen, and controls the target digital human to have a conversation with the user; during the conversation with the user, it determines the target object that the target digital human needs to draw and the target image where the target object is located; responds to the image drawing event, inputs the target image into the image drawing model, analyzes the drawing process of the target image, and obtains the target video; converts the target digital human from the conversational form to the drawing form, and displays the target video to the user so that the user can perceive the drawing process. In this way, the drawing mode is introduced in the interaction between the digital human and the user, and the digital human interacts with the user based on the user's video screen, enriching the interaction form and meeting the audience's diverse interaction needs. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0017] Figure 1 A flowchart of a digital human-based drawing method provided by an exemplary embodiment of the present application;

[0018] Figure 2 A schematic diagram of an electronic device provided by an exemplary embodiment of the present application. DETAILED DESCRIPTION

[0019] In order to make the purpose, technical solution and advantages of the present application clearer, the technical solution of the present application will be clearly and completely described below in combination with the specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present application.

[0020] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and provide corresponding operation portals for users to choose to authorize or refuse. In addition, the various models involved in this application (including but not limited to language models or large models) are in compliance with relevant laws and standards.

[0021] In the field of live broadcasting, digital humans refer to virtual 3D humanoid models created through computer graphics, artificial intelligence and other technologies, which can appear as anchors or assistants in live broadcasts. These digital humans not only have realistic appearances, but can also interact with viewers through technologies such as voice recognition, natural language processing and emotional computing. However, today's traditional digital humans mainly rely on voice or text to interact with users, resulting in a single form of interaction. During live broadcasts, when digital humans interact with users or during question-and-answer sessions, this single interaction mode is difficult to attract the audience's attention and cannot meet the diverse needs of the audience.

[0022] In response to the above technical problems, a technical solution is provided in an embodiment of the present application, which is used to provide users with a drawing mode based on the user's video screen, enrich the interaction form, and meet the audience's diverse interaction needs.

[0023] The technical solutions provided by various embodiments of the present application are described in detail below in conjunction with the accompanying drawings.

[0024] Figure 1 The digital human-based drawing method provided by the exemplary embodiment of the present application is as follows: Figure 1 As shown, the method may include the following steps:

[0025] Step 11: When the target digital human is broadcasting live in a conversational form, respond to the user's video connection request, receive and display the user's video screen, and control the target digital human to have a conversation with the user.

[0026] Step 12: during the conversation with the user, determine the target object that needs to be drawn by the target digital human and the target image where the target object is located.

[0027] Step 13: In response to the image drawing event, the target image is input into the image drawing model, and the drawing process of the target image is analyzed to obtain a target video for describing the drawing process of the target image.

[0028] Step 14: Convert the target digital human from a conversational form to a drawing form, and display the target video to the user so that the user can perceive the drawing process.

[0029] This embodiment may be applicable to a server or a terminal device, and may also be implemented by a combination of a server and a terminal device. The terminal device may be a mobile phone, a tablet computer, or a computer, etc., and this embodiment does not impose any limitation.

[0030] In this embodiment, the target digital person can broadcast live in a conversational form, where the conversational form refers to the form in which the target digital person is in a conversation with the user; from the user's perspective, the target digital person can have a conversation with himself as a conversational person. When the target digital person broadcasts live in a conversational form, he can conduct Q&A, conversation or other interactive forms with the user, which is not limited in this embodiment.

[0031] When the target digital human is broadcasting live in a conversational form, the user can send a video connection request through his or her own terminal device. Correspondingly, the user's video connection request can be responded to, the user's video screen can be received and displayed, and the target digital human can be controlled to have a conversation with the user.

[0032] Among them, the user's voice signal can be collected and reported to the server. The server matches the answer information adapted to the question information expressed by the voice signal, and feeds back the answer information to the digital human to drive the digital human to broadcast the answer information, and the conversation is repeated in this way.

[0033] During the conversation with the user, the user's video screen can be obtained, and the target object that needs to be drawn by the target digital human and the target image where the target object is located can be determined based on the video screen. The target object can be any object, such as an animal or clothing, etc. Preferably, the target object can be a user.

[0034] Afterwards, the target image can be input into the image drawing model in response to the image drawing event. The image drawing event can be the following events: ① The user triggers the drawing control. Correspondingly, the drawing control trigger event can be responded to, the digital human can be controlled to start the drawing mode, and the target image can be input into the image drawing model; ② During the conversation between the digital human and the user, information that meets the preset information requirements is monitored, the digital human is controlled to start the drawing mode, and the target image is input into the image drawing model. The preset information requirement can be that the frequency of specific text appearing in the user's conversation content is higher than a certain threshold, and the specific text can be text related to the drawing mode.

[0035] The image rendering model can analyze the rendering process of the target image to obtain a target video for describing the rendering process of the target image. The image rendering model has the ability to reconstruct / reproduce the rendering process of the image through pre-training, and can output a target video for describing the rendering process of the target image. In other words, the target video can include the rendering process of the target image, allowing users to perceive the rendering process.

[0036] After obtaining the target video, the target digital human can be transformed from the conversational form to the drawing form. Among them, the drawing form can be used to show the user that the target digital human is drawing; from the user's perspective, the target digital human is drawing.

[0037] The embodiment of the present application does not limit the specific implementation method of the target digital human's morphological transformation. Optionally, a preset control parameter can be obtained, and the target digital human can be controlled to paint based on the preset control parameter, that is, to transform from a conversational morphology to a drawing morphology. For example, the target digital human is talking to the user facing the screen, and based on the preset control parameter, the target digital human can be controlled to lower his head, pick up a paintbrush and start painting. Optionally, a transition video can be preset, which can be used to describe the process of the target digital human's morphological transformation. Based on this, the transition video can be shown to the user to allow the user to perceive the process of the target digital human's transformation from a conversational morphology to a drawing morphology.

[0038] After the target digital human completes the transformation, the target video can be shown to the user so that the user can perceive the drawing process.

[0039] In this embodiment, when the target digital human is broadcasting live in a conversational form, it responds to the user's video connection request, receives and displays the user's video screen, and controls the target digital human to have a conversation with the user; during the conversation with the user, it determines the target object that the target digital human needs to draw and the target image where the target object is located; responds to the image drawing event, inputs the target image into the image drawing model, analyzes the drawing process of the target image, and obtains the target video; converts the target digital human from the conversational form to the drawing form, and displays the target video to the user so that the user can perceive the drawing process. In this way, the drawing mode is introduced in the interaction between the digital human and the user, and the digital human interacts with the user based on the user's video screen, enriching the interaction form and meeting the audience's diverse interaction needs.

[0040] In some optional embodiments, step 11 of the aforementioned embodiment, “responding to the user's video connection request, receiving and displaying the user's video screen”, can be implemented based on the following methods:

[0041] In response to the user's video connection request, a preset automation script is run. The automation script can be used to establish an audio and video connection with the user according to the video connection request, receive the user's video screen through the audio and video connection, and display the user's video screen.

[0042] In this way, the user's video images can be received and displayed more efficiently.

[0043] Based on this, step 12 of the aforementioned embodiment, “determining the target object that needs to be drawn by the target digital human and the target image where the target object is located during the conversation with the user”, can be implemented based on the following steps:

[0044] Step 121: When displaying the video screen of the user, take at least one screenshot of the video screen of the user.

[0045] Step 122: Use the human body detection model to detect whether each screenshot obtained by the screenshot meets the preset image requirements. The image requirements can be set to any requirements according to actual design requirements, such as the requirements for image clarity, whether there is a user's head and neck in the screenshot, or whether there is background interference.

[0046] Optionally, the human body detection model can identify image parameters such as clarity or brightness of each screenshot, so as to determine whether the screenshot meets the preset image requirements. Optionally, the human body detection model can extract the main object, that is, the foreground object, from each screenshot and identify the type of the main object, so as to determine whether each screenshot meets the preset image requirements based on the type of the identified main object.

[0047] Step 123: Stop taking screenshots when a target screen screenshot that meets preset image requirements is detected, and use the user as a target object and the target screen screenshot as a target image.

[0048] Based on the above steps 121 to 123 , the target object that needs to be drawn by the target digital human and the target image where the target object is located can be determined more accurately based on the human body detection model.

[0049] The embodiment of the present application does not limit the specific model architecture of the image rendering model in the aforementioned embodiment. In some exemplary embodiments, the image rendering model may include: a rendering process prediction network, an action prediction network, and a video generation network. Based on this, step 13 in the aforementioned embodiment, "inputting the target image into the image rendering model, performing rendering process analysis on the target image to obtain a target video for describing the rendering process of the target image" can be implemented based on the following steps:

[0050] Step 131, input the target image into the drawing process prediction network, convert the target image into a line draft, obtain the target line draft, and predict multiple drawing step information required to draw the target line draft. The target image can be converted into a line draft using the edge detection algorithm in the drawing process prediction network. The drawing process prediction network can use the pre-learned drawing habits to split the target line draft and determine the order of each split line, thereby predicting multiple drawing step information required to draw the target line draft. Each drawing step information can correspond to the target line, which can be used to describe the corresponding target line and the drawing order of the target line.

[0051] Step 132: Input the multiple drawing step information into the action prediction network, and predict the action parameters of the target digital human for executing the multiple drawing steps according to the multiple drawing step information. The action prediction network can use the learned correspondence between the drawing step information and the action parameters to predict the action parameters of the target digital human for executing the multiple drawing steps according to the multiple drawing step information. For example, if a certain drawing step information is to draw a straight line from position A to position B, then the action parameters corresponding to the drawing step information can be used to drive the hand of the target digital human to move from position A to position B.

[0052] Step 133: Input the action parameters into the video generation network, and control the target digital human to perform multiple drawing steps according to the action parameters to simulate the drawing process of the target image and obtain the target video. The target digital human can be controlled to simulate the execution of multiple drawing steps according to the action parameters, and the simulation execution process can be recorded to obtain the target video, which can be used to describe the drawing process of the target image. Of course, the above method is only an example. In actual application, the simulation execution process may not be recorded, but the target digital human can be controlled to simulate the execution of multiple drawing steps to display the drawing process to the user in real time.

[0053] In this way, the image rendering model can be used to more accurately analyze the rendering process of the target image to obtain the target video.

[0054] The following will explain in detail the drawing process prediction network and action prediction network in the image drawing model:

[0055] 1. Drawing process prediction network:

[0056] The drawing process prediction network may include a first extraction layer and a first prediction layer. Based on this, when predicting multiple drawing step information required to draw the target line draft, the target line draft may be input into the first extraction layer to extract the first feature of the target line draft; the first feature may be input into the first prediction layer, and the multiple drawing step information required to draw the target line draft may be predicted according to the first feature.

[0057] Considering that there are usually different drawing methods, i.e., different drawing styles, for the same target image, the embodiment of the present application can also provide a drawing style selection page to the user, and respond to the user's selection operation to determine the target drawing style selected by the user. The specific implementation method of the drawing style can be set according to the actual needs of the user, such as children's painting style, Western painting style, sketch style, or ink painting style, etc.

[0058] Afterwards, a target prompt word corresponding to the target drawing style can be determined from a plurality of prompt words corresponding to a plurality of preset drawing styles. Different prompt words can be used to guide the first prediction layer to perform line splitting on the target line drawing according to different drawing habits. Afterwards, the target prompt word and the first feature can be input into the prediction layer, and under the guidance of the target prompt word, multiple drawing step information that conforms to the target drawing style is predicted according to the first feature.

[0059] In this way, based on the model prompt words, the drawing process prediction network can more accurately predict the multiple drawing step information required to draw the target line draft, and the predicted drawing step information is more in line with the user's desired painting style.

[0060] 2. Action Prediction Network:

[0061] The action prediction network may include: a second extraction layer and a second prediction layer. Based on this, when predicting the action parameters of the target digital human performing the multiple drawing steps according to the multiple drawing step information, the multiple drawing step information may be input into the second extraction layer to extract the second features of the multiple drawing step information; the second features of the multiple drawing step information may be input into the second prediction layer to predict the action parameters of the target digital human performing the multiple drawing steps. The action parameters may include: at least one of the action start and end positions, the action amplitude and the action speed.

[0062] In this way, the action prediction network can more accurately predict the action parameters of the target digital human performing multiple drawing steps based on the information of multiple drawing steps.

[0063] In some optional embodiments, each network in the image rendering model may also be trained based on the following method:

[0064] Plot the training process of the process prediction network:

[0065] Obtain sample images and sample drawing step information; input the sample images and the sample drawing step information into the drawing process prediction network to be trained, and under the supervision of the sample drawing step information, train the drawing process prediction network to be trained with the sample images with the goal of converging the first loss function of the drawing process prediction network to be trained to a first target range, so as to obtain a trained drawing process prediction model.

[0066] The first loss function is used to calculate the error between the target drawing step information predicted by the drawing process prediction network to be trained and the sample drawing step information.

[0067] Training process of action prediction network:

[0068] Get sample drawing step information and sample action parameters;

[0069] The sample drawing step information and the sample action parameters are input into the action prediction network to be trained. Under the supervision of the sample action parameters, the action prediction network to be trained is trained with the sample drawing step information with the goal of converging the second loss function of the action prediction network to be trained to a second target range, so as to obtain a trained action prediction network.

[0070] The second loss function is used to calculate the error between the target action parameters predicted by the action prediction network to be trained and the sample action parameters.

[0071] The training process of the video generation network:

[0072] Obtain sample action parameters and sample videos; input the sample action parameters and sample videos into the video generation network to be trained, and under the supervision of the sample videos, train the video generation network to be trained using the sample action parameters with the goal of converging the third loss function of the video generation network to be trained to a third target range, so as to obtain a trained video generation network.

[0073] Among them, the third loss function is used to calculate the error between the target video predicted by the video generation network to be trained and the sample video.

[0074] Through the above method, the image rendering model can be trained more accurately.

[0075] It should be noted that the execution subject of each step of the method provided in the above embodiment can be the same device, or the method can be executed by different devices. For example, the execution subject of steps 11 to 14 can be device A; for another example, the execution subject of steps 11 to 12 can be device A, and the execution subject of steps 13 to 14 can be device B; and so on.

[0076] In addition, in some of the processes described in the above embodiments and the accompanying drawings, multiple operations appearing in a specific order are included, but it should be clearly understood that these operations may not be executed in the order in which they appear in this document or may be executed in parallel, and the sequence numbers of the operations, such as 12, 13, etc., are only used to distinguish between different operations, and the sequence numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel.

[0077] It should be noted that the descriptions such as “first” and “second” in this article are used to distinguish different messages, devices, modules, etc., and do not represent the order of precedence, nor do they limit “first” and “second” to different types.

[0078] Figure 2 is a schematic diagram of the structure of an electronic device provided by an exemplary embodiment of the present application, and the electronic device is applicable to the digital human-based drawing method provided in the above-mentioned embodiment, such as Figure 2 As shown, the electronic device may include: a memory 201 , a processor 202 and a communication component 203 .

[0079] The memory 201 is used to store computer programs and can be configured to store various other data to support operations on the electronic device. Examples of such data include instructions for any application or method operating on the electronic device, contact data, phone book data, messages, pictures, videos, etc.

[0080] In some exemplary embodiments, the processor 202 is coupled to the memory 201 and is used to execute the computer program in the memory 201, so as to: respond to the user's video connection request, receive and display the user's video screen, and control the target digital human to have a conversation with the user during the target digital human broadcasts live in a conversation form; determine the target object that needs to be drawn by the target digital human and the target image where the target object is located during the conversation with the user; respond to the image drawing event, input the target image into the image drawing model, and analyze the drawing process of the target image to obtain a target video for describing the drawing process of the target image; convert the target digital human from the conversation form to the drawing form, and display the target video to the user so that the user can perceive the drawing process.

[0081] Further optionally, when the processor 202 responds to the user's video connection request and receives and displays the user's video screen, it is specifically used to: respond to the user's video connection request, run a preset automation script, the automation script is used to establish an audio and video connection with the user according to the video connection request, receive the user's video screen through the audio and video connection, and display the user's video screen.

[0082] Further optionally, when the processor 202 determines the target object that the target digital human needs to draw and the target image where the target object is located during the conversation with the user, it is specifically used to: take a screenshot of the user's video screen at least once when displaying the user's video screen; use a human body detection model to detect whether each screenshot obtained by the screenshot meets the preset image requirements; stop taking screenshots when a target screenshot that meets the preset image requirements is detected, and use the user as the target object and the target screenshot as the target image.

[0083] Further optionally, the image drawing model includes: a drawing process prediction network, an action prediction network and a video generation network; when the processor 202 inputs the target image into the image drawing model and performs a drawing process analysis on the target image to obtain a target video for describing the drawing process of the target image, it is specifically used to: input the target image into the drawing process prediction network, convert the target image into a line draft to obtain a target line draft, and predict multiple drawing step information required to draw the target line draft; input the multiple drawing step information into the action prediction network, and predict the action parameters of the target digital human performing the multiple drawing steps according to the multiple drawing step information; input the action parameters into the video generation network, and control the target digital human to perform the multiple drawing steps according to the action parameters to simulate the drawing process of the target image and obtain the target video.

[0084] Further optionally, the drawing process prediction network includes: a first extraction layer and a first prediction layer; when the processor 202 predicts multiple drawing step information required to draw the target line draft, it is specifically used to: input the target line draft into the first extraction layer to extract the first feature of the target line draft; input the first feature into the first prediction layer, and predict multiple drawing step information required to draw the target line draft based on the first feature.

[0085] Further optionally, the processor 202 is also used to: provide a drawing style selection page to the user, and in response to a selection operation, determine the target drawing style selected by the user; input the first feature into the first prediction layer, and predict multiple drawing step information required to draw the target line draft based on the first feature, including: determining a target prompt word corresponding to the target drawing style from multiple prompt words corresponding to multiple preset drawing styles; input the target prompt word and the first feature into the first prediction layer, and under the guidance of the target prompt word, predict the multiple drawing step information that conforms to the target drawing style based on the first feature.

[0086] Further optionally, the action prediction network includes: a second extraction layer and a second prediction layer; when the processor 202 predicts the action parameters of the target digital human executing each of the multiple drawing steps based on the multiple drawing step information, it is specifically used to: input the multiple drawing step information into the second extraction layer, extract the second features of each of the multiple drawing step information; input the second features of each of the multiple drawing step information into the second prediction layer, and respectively predict the action parameters of the target digital human executing each of the multiple drawing steps; the action parameters include: action start and end positions, action amplitude and / or action speed.

[0087] Further, if Figure 2As shown, the electronic device also includes: a display 204, a power component 205, an audio component 206 and other components. Figure 2 Only some components are shown schematically, which does not mean that the electronic device only includes Figure 2 Components shown.

[0088] The embodiment of the present application also provides a computer-readable storage medium, which, when the computer program is executed by a processor, enables the processor to implement the steps in the digital human-based drawing method.

[0089] The embodiment of the present application further provides a computer program product, including a computer program / instruction. When the computer program / instruction is executed by a processor, the steps in the digital human-based drawing method are executed.

[0090] In this embodiment, when the target digital human is broadcasting live in a conversational form, it responds to the user's video connection request, receives and displays the user's video screen, and controls the target digital human to have a conversation with the user; during the conversation with the user, it determines the target object that the target digital human needs to draw and the target image where the target object is located; responds to the image drawing event, inputs the target image into the image drawing model, analyzes the drawing process of the target image, and obtains the target video; converts the target digital human from the conversational form to the drawing form, and displays the target video to the user so that the user can perceive the drawing process. In this way, the drawing mode is introduced in the interaction between the digital human and the user, and the digital human interacts with the user based on the user's video screen, enriching the interaction form and meeting the audience's diverse interaction needs.

[0091] The above-mentioned memory can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read only memory (EEPROM), erasable programmable read only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk.

[0092] The above-mentioned communication component is configured to facilitate wired or wireless communication between the device where the communication component is located and other devices. The device where the communication component is located can access a wireless network based on a communication standard, such as WiFi, 2G, 3G, 4G / LTE, 5G and other mobile communication networks, or a combination thereof. In an exemplary embodiment, the communication component receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be based on Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra Wide Band (UWB) technology, Bluetooth (BT) technology and other technologies.

[0093] The above-mentioned display includes a screen, and the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from a user. The touch panel includes one or more touch sensors to sense touch, slide, and gestures on the touch panel. The touch sensor may not only sense the boundary of a touch or slide action, but also detect the duration and pressure associated with the touch or slide operation.

[0094] The power supply assembly provides power to various components of the device where the power supply assembly is located. The power supply assembly may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the device where the power supply assembly is located.

[0095] The above-mentioned audio component can be configured to output and / or input audio signals. For example, the audio component includes a microphone (Microphone, MIC), and when the device where the audio component is located is in an operating mode, such as a call mode, a recording mode, and a speech recognition mode, the microphone is configured to receive an external audio signal. The received audio signal can be further stored in a memory or sent via a communication component. In some embodiments, the audio component also includes a speaker for outputting an audio signal.

[0096] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-readable storage media (including but not limited to disk storage, compact disc read-only memory (Compact Disc Read-Only Memory, CD-ROM), optical storage, etc.) containing computer-usable program code.

[0097] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0098] These computer program instructions may also be stored in a computer readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0099] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0100] In a typical configuration, a computing device includes one or more processors (Central Processing Unit, CPU), input / output interface, network interface and memory.

[0101] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.

[0102] Computer readable media include permanent and non-permanent, removable and non-removable media that can be used to store information by any method or technology. Information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase-change random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0103] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0104] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included within the scope of the claims of the present application.

Claims

1. A drawing method based on digital human, characterized in that: include: During the process of the target digital human being broadcasting live in the form of conversation, responding to the user's video connection request, receiving and displaying the user's video screen, and controlling the target digital human to have a conversation with the user; During the conversation with the user, determining a target object that needs to be drawn by the target digital human and a target image where the target object is located; In response to an image drawing event, the target image is input into an image drawing model, and a drawing process analysis is performed on the target image to obtain a target video for describing the drawing process of the target image; The target digital human is transformed from the conversational form into a drawing form, and the target video is displayed to the user so that the user can perceive the drawing process.

2. The method according to claim 1, characterized in that Responding to a user's video connection request, receiving and displaying the user's video image, including: In response to the video connection request of the user, a preset automation script is run, wherein the automation script is used to establish an audio and video connection with the user according to the video connection request, receive the video screen of the user through the audio and video connection, and display the video screen of the user.

3. The method according to claim 2, characterized in that During the conversation with the user, determining the target object that needs to be drawn by the target digital human and the target image where the target object is located includes: When displaying the video screen of the user, taking at least one screenshot of the video screen of the user; Using the human body detection model to detect whether each screenshot obtained by the screenshot meets the preset image requirements; When a target screen screenshot that meets the preset image requirement is detected, the screenshot is stopped, and the user is taken as the target object, and the target screen screenshot is taken as the target image.

4. The method according to claim 1, characterized in that: The image rendering model includes: a rendering process prediction network, an action prediction network and a video generation network; Inputting the target image into an image rendering model, and performing rendering process analysis on the target image to obtain a target video for describing the rendering process of the target image, including: Inputting the target image into the drawing process prediction network, converting the target image into a line draft to obtain a target line draft, and predicting a plurality of drawing step information required to draw the target line draft; Inputting the plurality of drawing step information into the action prediction network, and predicting the action parameters of the target digital human executing the plurality of drawing steps according to the plurality of drawing step information; The action parameters are input into the video generation network, and the target digital human is controlled to execute the multiple drawing steps according to the action parameters to simulate the drawing process of the target image and obtain the target video.

5. The method according to claim 4, characterized in that The rendering process prediction network includes: a first extraction layer and a first prediction layer; Predicting multiple drawing step information required to draw the target line draft, including: Inputting the target line draft into the first extraction layer to extract a first feature of the target line draft; The first feature is input into the first prediction layer, and multiple drawing step information required for drawing the target line draft is predicted according to the first feature.

6. The method according to claim 5, characterized in that Also includes: Providing a drawing style selection page to the user, and in response to a selection operation, determining a target drawing style selected by the user; Inputting the first feature into the first prediction layer, and predicting multiple drawing step information required for drawing the target line draft according to the first feature, including: Determining a target prompt word corresponding to the target drawing style from a plurality of prompt words corresponding to a plurality of preset drawing styles; The target prompt word and the first feature are input into the first prediction layer, and under the guidance of the target prompt word, the plurality of drawing step information that conforms to the target drawing style is predicted according to the first feature.

7. The method according to claim 4, characterized in that The action prediction network includes: a second extraction layer and a second prediction layer; Predicting, based on the plurality of drawing step information, the action parameters of the target digital human executing the plurality of drawing steps, respectively, includes: Inputting the plurality of drawing step information into the second extraction layer, and extracting second features of the plurality of drawing step information; The second features of the plurality of drawing step information are input into the second prediction layer to respectively predict the action parameters of the target digital human executing the plurality of drawing steps; the action parameters include: action start and end positions, action amplitude and / or action speed.

8. A terminal device, characterized in that: include: A memory and a processor; wherein the memory is used to: store one or more computer instructions; The processor is configured to execute the one or more computer instructions to perform the steps in the method according to any one of claims 1 to 7.

9. A computer-readable storage medium, characterized in that: When the computer program is executed by a processor, the processor is enabled to implement the steps of the method according to any one of claims 1 to 7.

10. A computer program product, characterized in that The method comprises a computer program / instruction, which, when executed by a processor, enables the processor to implement the steps of the method according to any one of claims 1 to 7.