Digital human generation method, intelligent agent, device, equipment and storage medium
By segmenting the target object from the image, filtering the adaptive digital people and generating their dress textures, the problem of insufficient anthropomorphization of the target object in the existing technology is solved, personalized digital life generation is achieved, and user experience is improved.
Patent Information
- Application Number
- CN202411963903.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-28
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art is difficult to anthropomorphize the target object in the image into a personalized digital person, and the user experience is insufficient.
By segmenting the target object from the image to be processed, filtering the adaptive digital person, generating its dress texture and driving the digital person, combining big models and rendering techniques, the target object is achieved.
Improve user experience, provide more emotional value through anthropomorphic digital people, and enhance user interaction and emotional linkage with digital people.
Smart Images

Figure CN119991936A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, in particular to computer vision, deep learning, large models, augmented reality and other technical fields, and can be applied to scenarios such as digital humans. Background Art
[0002] With the rapid development of AIGC (Artificial Intelligence Generated Content) technology, digital humans, as an emerging human-computer interaction method, are gradually entering the public eye and receiving widespread attention. Digital humans can not only imitate human appearance and behavior, but also achieve natural dialogue and emotional communication with humans through deep learning and natural language processing technology. Digital human technology will play an important role in various applications. Summary of the invention
[0003] The present disclosure provides a method, an intelligent agent, an apparatus, a device and a storage medium for generating a digital human.
[0004] According to one aspect of the present disclosure, a method for generating a digital human is provided, comprising:
[0005] Segment the target object from the image to be processed to obtain a target sub-image;
[0006] Based on the target subgraph, the digital human to be optimized that is compatible with the target object is selected from the digital human set;
[0007] Generate the clothing texture of the digital human to be optimized based on the appearance features of the target object in the target sub-image;
[0008] Apply the clothing texture to the digital human to be optimized to obtain the target digital human;
[0009] Driving the target digital human.
[0010] According to another aspect of the present disclosure, there is provided a digital human generation device, comprising:
[0011] A segmentation module is used to segment the target object from the image to be processed to obtain a target sub-image;
[0012] A screening module, used to screen out a digital human to be optimized that is compatible with the target object from the digital human set based on the target subgraph;
[0013] A generation module, used for generating the clothing texture of the digital human to be optimized based on the appearance features of the target object in the target sub-image;
[0014] A determination module is used to apply the clothing texture to the digital human to be optimized to obtain a target digital human;
[0015] A driving module is used to drive the target digital human.
[0016] According to another aspect of the present disclosure, there is provided an intelligent agent, comprising:
[0017] An interactive interface for obtaining images to be processed;
[0018] The artificial intelligence module is used to segment the target object from the image to be processed to obtain a target sub-graph; based on the target sub-graph, a digital human to be optimized that is compatible with the target object is selected from a digital human set; based on the appearance features of the target object in the target sub-graph, a clothing texture of the digital human to be optimized is generated;
[0019] A rendering engine is used to apply the clothing texture to the digital human to be optimized to obtain a target digital human; and drive the target digital human.
[0020] According to another aspect of the present disclosure, there is provided an electronic device, comprising:
[0021] at least one processor; and
[0022] a memory communicatively connected to the at least one processor; wherein,
[0023] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute any method in the embodiments of the present disclosure.
[0024] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute any method according to the embodiments of the present disclosure.
[0025] According to another aspect of the present disclosure, a computer program product is provided, including a computer program, which implements any method according to the embodiments of the present disclosure when executed by a processor.
[0026] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure.
[0028] Figure 1 is a flowchart of a method for generating a digital human according to the first embodiment of the present disclosure;
[0029] Figure 2is a schematic diagram of a process for screening and adapting a digital human to be optimized according to the second embodiment of the present disclosure;
[0030] Figure 3 is a schematic diagram of a process of generating a clothing texture of a digital human to be optimized according to the third embodiment of the present disclosure;
[0031] Figure 4 is a schematic diagram of a process of generating a clothing texture of a digital human to be optimized according to the fourth embodiment of the present disclosure;
[0032] Figure 5 is a schematic diagram of the structure of an intelligent agent provided according to the fifth embodiment of the present disclosure;
[0033] Figure 6 is a schematic diagram of a framework of a method for generating a digital human according to a sixth embodiment of the present disclosure;
[0034] Figure 7 is a structural schematic diagram of a digital human generation device provided according to a seventh embodiment of the present disclosure;
[0035] Figure 8 It is a block diagram of an electronic device used to implement the digital human generation method of the embodiment of the present disclosure. DETAILED DESCRIPTION
[0036] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for the sake of clarity and conciseness, the description of well-known functions and structures is omitted in the following description.
[0037] The terms "first", "second", etc. in this disclosure are used to distinguish similar objects, and are not necessarily used to describe a specific order or precedence. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, including a series of steps or units. The method, system, product or device is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0038] People have a wide range of interests and hobbies, ranging from electronic products to cute pets. For example, cat lovers are found in all age groups, and they like to record the daily lives of their pet cats by taking photos or recording videos.
[0039] Although simply taking photos or recording videos can record the wonderful moments of cute pets, new technologies are still needed to enhance user experience and provide emotional value to users.
[0040] With the development of artificial intelligence technology and digital human technology, many applications can be improved on this basis to enhance user experience.
[0041] In view of this, the embodiments of the present disclosure provide a method for generating a digital human, which can generate a personalized digital human based on an image and drive it to improve the user experience. Figure 1 As shown, this is a flow chart of the method:
[0042] S101, segmenting a target object from an image to be processed to obtain a target sub-image.
[0043] The image to be processed may be an image uploaded by a user, such as a photographed image of a cute pet.
[0044] The target object in the image to be processed can be any one of animals, plants, and man-made objects. Animals are, for example, cute pets, and plants are, for example, jasmine, roses, tiger plants, etc. Man-made objects are objects made by humans through certain techniques and processes. They can be simple tools or complex high-tech products. Man-made objects in the embodiments of the present disclosure refer to objects with specific appearances, such as shields, hanging signs, dolls, etc.
[0045] During implementation, the image to be processed can be input into the target detection model to detect the position of the target object, thereby obtaining a 2D (Two Dimensions) position frame of the target object in the image to be processed. Then, the target object is cropped out from the image to be processed based on the 2D position frame to achieve the purpose of segmenting the target object from the image to be processed and obtaining a target sub-image.
[0046] S102, based on the target subgraph, a digital human to be optimized that is compatible with the target object is selected from the digital human set.
[0047] In the disclosed embodiment, the target object and the digital human to be optimized may belong to different species. For example, the target object is a pet cat, and the digital human to be optimized is an anthropomorphic virtual image. It is understandable that a digital human refers to an anthropomorphic virtual image generated by computer technology and having vision, voice, behavior, etc.
[0048] For example, if the target object is an image of a cute pet, at least some of its characteristics can be transferred to the digital human to be optimized, thus realizing the anthropomorphization of the cute pet.
[0049] In some embodiments, the digital human set may be a cartoon digital human set. Cartoon digital humans have common characteristics similar to cute pets that are well-liked by users, so digital humans selected from the cartoon digital human set can provide users with positive and active emotional value and improve user experience.
[0050] In some embodiments, a digital human that is aesthetically compatible with the target object can be selected from the digital human collection as the digital human to be optimized.
[0051] Aesthetic adaptation is a concept involving the application of aesthetic principles in different fields, which emphasizes the coordination and integration between aesthetics and practical applications. Aesthetic adaptation can be understood as applying aesthetic principles and aesthetic values to the target object and digital human field, so as to smoothly transfer some of the user's intuitive feelings from the target object to the digital human. It can be understood as anthropomorphizing the target object to provide users with more emotional value, and extending the user's love for the target object through scientific and technological means, so as to provide users with more emotional value and improve user experience.
[0052] S103, generating a clothing texture of the digital human to be optimized based on the appearance features of the target object in the target sub-image.
[0053] Among them, appearance features are the external manifestations of products or objects, which may include at least one or a combination of visual elements such as shape, pattern, color, etc. In computer vision, appearance features can refer to the visual attributes of an object in an image, such as at least one of color, texture, shape, etc.
[0054] The digital human to be optimized in the embodiment of the present disclosure may already have a stylized face, an anthropomorphic body, and a clothing style, but the clothing texture needs to be generated according to the target object to achieve the anthropomorphization of the target object.
[0055] S104, applying the clothing texture to the digital human to be optimized to obtain a target digital human.
[0056] S105, driving the target digital human.
[0057] In summary, in the disclosed embodiments, the target object is segmented from the image to be processed, and irrelevant or minor information can be removed, so that subsequent processing can focus on the target object. By screening the digital human to be optimized that is compatible with the target object, and generating the clothing texture of the digital human to be optimized based on the target object, the appearance features of the target object can be transferred to the digital human, and a personalized digital human can be generated based on the target object in the image, so that the target object can be anthropomorphized. By further driving the target digital human, the target digital human can provide users with more emotional value through anthropomorphic operations, further improving the user experience.
[0058] In addition, if the digital human that is aesthetically compatible is selected, it can meet the aesthetic requirements in terms of aesthetics, so as to smoothly transfer the important features of the target object to the target digital human, so that the target object and the target digital human can echo each other and produce linkage. Furthermore, the user's feelings about the target object can be aesthetically transferred to the target digital human, thereby providing users with more emotional value and further improving the user experience.
[0059] In the disclosed embodiment, the digital human to be optimized includes a two-dimensional character model. Among them, the two-dimensional culture has formed a unique community culture, and fans establish social connections and cultural identity by sharing their love for specific works or characters. The audience group of the two-dimensional image covers a wide range, and it is one of the consumption hotspots for contemporary young people to meet emotional values. Therefore, using the two-dimensional character model as the digital human to be optimized can meet people's aesthetic needs, especially when the target object is a pet, the two-dimensional character model and the pet have similar characteristics, such as cuteness and softness. The selected digital human to be optimized can associate the target object with the digital human in some sensory characteristics, thereby realizing the generation of a personalized target digital human that matches the target object in the image, providing users with more emotional value, thereby improving the user experience.
[0060] In some embodiments, usually, the target object and the digital human to be optimized are essentially different species. In this case, in order to improve the accuracy of screening the digital human to be optimized and improve the adaptability of the digital human to be optimized and the target object, in the embodiment of the present disclosure, based on the target sub-graph, the digital human to be optimized that is compatible with the target object is screened from the cartoon digital human set, which can be implemented as follows: Figure 2 The operations shown include:
[0061] S201, based on the species category corresponding to the target object in the target subgraph, sub-classify the target object to obtain a sub-classification category of the target object.
[0062] For example, if the target object is a pet cat, the corresponding species category is the cat category. If the target object is a dog, the corresponding species category is the dog category. Cats and dogs are just large categories, which include more detailed subcategories.
[0063] In the disclosed embodiments, a classification model corresponding to the species category of the target object may be used to sub-classify the target object. Taking cats as an example, training samples may be collected, and the training samples include images of cats and their corresponding sub-classification labels. By using the training samples to perform supervised training on the initial classification model of cats, a classification model applied to the species category of cats may be obtained. For dogs, by analogy, training samples may be collected, and the training samples include images of dogs and their corresponding sub-classification labels. By using the training samples to perform supervised training on the initial classification model of dogs, a classification model applied to the species category of dogs may be obtained.
[0064] Different subcategories contain objects with different appearances and habits. Subcategories can help you better understand the target object, so that you can choose the appropriate digital human to be optimized.
[0065] To this end, in the embodiment of the present disclosure, each sub-classification category is pre-associated with a candidate object set. The candidate object set includes multiple objects of the same sub-classification category as the target object. In addition, in the embodiment of the present disclosure, each candidate object can be aesthetically designed and associated with a corresponding 3D digital human model in advance. It can be understood that each candidate object has a corresponding 3D digital human model, and multiple candidate objects can correspond to the same 3D digital human model.
[0066] S202, in a candidate object set corresponding to a sub-classification category, select candidate objects having object features similar to the target object as similar objects; the candidate objects in the candidate object set and the target object belong to the same species.
[0067] In some embodiments, texture features of the target object and the candidate objects may be extracted and matched to obtain similar objects.
[0068] In other embodiments, in order to improve the screening accuracy of similar objects, in the embodiments of the present disclosure, suitable similar objects can be screened out based on the reasoning and understanding ability of the large model, which can be implemented as follows:
[0069] Step A1: extracting a first feature of a target object from a target subgraph based on a feature extraction model, and extracting second features of a plurality of candidate objects in a candidate object set.
[0070] "Big models" generally refer to models in the field of machine learning and artificial intelligence that have large parameter scales and usually require a lot of computing resources to train and run. Such models can include language models, image models, reinforcement learning models, etc. They are large in scale and can handle more complex tasks and data.
[0071] The feature extraction model in the embodiment of the present disclosure is a large model that can extract features from an image and has computer vision processing capabilities. A transformer-based encoder of a CLIP (Contrastive Language-Image Pre-training, multimodal pre-trained neural network) model can be used to extract the first feature of the target object and the second feature of the candidate object.
[0072] The advantage of the CLIP model lies in its strong versatility and zero-sample learning ability, which enables it to perform well in practical applications, especially in areas such as image retrieval. It can mine the association between target objects and candidate objects and improve the accuracy of retrieving similar objects.
[0073] Step A2: Determine the similarity between the second features and the first features of the plurality of candidate objects.
[0074] Step A3: Select candidate objects corresponding to the second feature with the highest similarity as similar objects.
[0075] The disclosed embodiment can accurately understand the differences and similarities between the target object and each candidate object through the feature extraction big model, and can more accurately retrieve similar objects with the help of the reasoning and understanding ability of the big model, so as to improve the effect of screening the digital human to be optimized, improve the compatibility of the finally generated digital human with the target object, and thus improve the user experience.
[0076] S203, obtaining 3D digital human models corresponding to similar objects in the digital human set as the digital human to be optimized.
[0077] Among them, in the embodiment of the present disclosure, in order to improve the effect of the initially screened digital human and improve the adaptability of the target object and the digital human to be optimized, the 3D digital human model is established based on the principle of aesthetic adaptation to similar objects.
[0078] For example, for objects of different subcategories, artists draw aesthetically adapted 3D digital human models based on the characteristics of the objects of different subcategories.
[0079] Alternatively, the appearance requirements of cartoon digital humans can be generated for objects of different subcategories. For example, prompt words can be generated based on appearance and personality (lively or cute), and the large model can fine-tune the standard digital human model based on the prompt words to generate an aesthetically suitable 3D digital human model. Even under each subcategorie, corresponding 3D digital human models can be provided for different postures and / or emotions of the same object. In this way, different postures and / or emotions can be finely distinguished to select aesthetically suitable 3D digital human models.
[0080] Therefore, in order to solve the problem that it is difficult to directly match features between the target object and the digital human due to different species, in the embodiment of the present disclosure, a set of candidate objects in the same subdivision category as the target object is cited as an intermediate medium, which can match features between the same species and improve the accuracy of screening the digital human to be optimized. In addition, by screening on the subdivision category, the accuracy of retrieving the appropriate digital human to be optimized can be improved.
[0081] After the digital human to be optimized is selected, in order to transfer the appearance features of the target object to the clothes of the digital human to be optimized to realize the anthropomorphism of the target object, in some embodiments, it can be implemented as follows: Figure 3 As shown:
[0082] S301, processing the target subgraph based on the graph-to-text model to obtain appearance description text of the target object.
[0083] The appearance description text is the text that describes the appearance features of the target object. Taking a pet cat as an example, the pattern and fur color of the pet cat can be accurately described with the help of the understanding and reasoning ability of the large-scale image-based text model. The appearance description text can then be input into the texture generation model, which helps the subsequent texture generation model to generate similar clothing textures.
[0084] In order to further improve the similarity between the generated clothing texture and the target object in the embodiment of the present disclosure, in S302, the appearance description text and the target sub-image may be input into the texture generation model to generate the clothing texture of the digital human to be optimized.
[0085] Therefore, the texture generation model can accurately understand the characteristics of the required clothing texture through the appearance expression of the image modality and the appearance expression of the text modality, improve the adaptability between the generated clothing texture and the target object, and further associate the target object with the target digital human, thereby improving the adaptability of the generated target digital human and the target object, thereby improving the user experience.
[0086] For example, the large model of graphene can be used as follows Figure 3 The VLM (Visual Language Model) 303 shown in the figure allows the VLM 303 to describe the pattern and overall color of the cat on the target sub-image to obtain the cat's appearance description text. Then, the appearance description text and the cat's target sub-image are input into the texture generation model 304 constructed based on the diffusion model to obtain the clothing texture.
[0087] Among them, the diffusion model can select LDM (Latent Diffusion, latent diffusion model). LDM is a generator based on a diffusion model. It reduces the computational complexity and improves the efficiency and quality of clothing texture generation by operating in latent space rather than pixel space.
[0088] In some other embodiments, the clothing texture of the digital human to be optimized is generated based on the appearance features of the target object in the target sub-image, which can also be implemented as follows: Figure 4 As shown:
[0089] S401, processing the target subgraph based on the graph-to-text model to obtain appearance description text of the target object.
[0090] This step is the same as S301 and will not be described again here.
[0091] S402, processing the appearance description text based on the Wenshengtu large model to generate a character image; the clothing of the character object in the character image is generated under the constraint of the appearance description text; the style of the character object is the same as the character style of the digital human to be optimized.
[0092] That is, the appearance description text can be converted into a flat character image based on the Wensheng graph model. Since the character object in the character image has the same style as the digital human to be optimized, it can better align the digital human to be optimized and better describe the requirements for the generated texture.
[0093] S403, inputting the appearance description text and the character image into a texture generation model to obtain the clothing texture of the digital human to be optimized.
[0094] Therefore, in the disclosed embodiment, first obtaining the appearance description text of the target object can obtain rich appearance details of the target object to help the texture generation model understand and generate clothing textures. In addition, by converting the appearance description text into a human object of the same style, the image input to the texture generation model can better align the digital human to be optimized, improve the generation effect and quality of clothing textures, and thus improve the quality of the generated target digital human.
[0095] For example, the large model of graphene can be used as follows Figure 4 The VLM404 shown in FIG. 4 describes the pattern and overall color of the cat on the target sub-image to obtain the cat's appearance description text. Then, the appearance description text is input into the Wenshengtu large model 405 to obtain a two-dimensional character image. After that, the appearance description text and the character image are input into the texture generation model 406 constructed based on the diffusion model to obtain the clothing texture of the two-dimensional character model.
[0096] In the disclosed embodiment, after the target digital human is constructed based on the clothing texture, the target digital human can be driven by preset parameters, which can be implemented as follows: determining the target digital human based on at least one of the following preset driving parameters: skeleton driving parameters, facial expression driving parameters, and text-to-speech driving parameters.
[0097] The target digital human can be determined by combining skeleton point drive, facial drive and TTS (Text-to-Speech) technology.
[0098] For example, after the target digital human is generated, it is not displayed to the user statically, but is animated through preset driving parameters to interact with the user in an anthropomorphic manner. For example, a greeting method can be selected, using body language and combined with TTS technology to drive the target digital human to greet the user. For another example, the target digital human can also be driven to perform a singing and dancing program.
[0099] In the disclosed embodiment, by presetting driving parameters, the target digital human can be driven from different aspects such as skeleton, face and voice, so as to perform anthropomorphic operations, so as to improve the user experience.
[0100] In addition, the embodiments of the present disclosure can also rely on the large model to interact with the user. It can be implemented as follows:
[0101] Step B1: Obtain interaction information.
[0102] During implementation, users can send interactive information through a single mode such as voice, text, video, or a combination of multiple modes.
[0103] Step B2: Process the interactive information based on the large language model to obtain a response text for the interactive information.
[0104] The large language model can understand the interactive information, and thus generate personalized response text with the help of the logical reasoning ability of the large language model.
[0105] For multimodal interactive information, the content of the user's interactive information can be first understood through a large multimodal model, so that the features of the media modality (such as voice and image) are mapped to the text feature space, and then processed by a large language model to obtain the response text.
[0106] Step B3: generating driving parameters of the target digital human based on the response text.
[0107] Based on TTS technology, the response text can be presented through the target digital human to achieve humanized intelligent interaction.
[0108] Step B4: driving the target digital human based on the driving parameters.
[0109] During implementation, the driving parameters may include not only the lip shape required for the target digital human to express the response text, but also appropriate facial expressions and / or movements to make the target digital human more humanized and improve the user experience.
[0110] In summary, the embodiments of the present disclosure support flexible and personalized interaction between users and target digital humans, so that the target digital humans are presented to users in the form of intelligent entities, thereby improving user experience.
[0111] Based on the same technical concept, the embodiment of the present disclosure also provides an intelligent agent, such as Figure 5 As shown, including:
[0112] Interaction interface 501, used to obtain the image to be processed;
[0113] The artificial intelligence module 502 is used to segment the target object from the image to be processed to obtain a target sub-image;
[0114] Based on the target sub-graph, a digital human to be optimized that matches the target object is selected from the digital human set; based on the appearance features of the target object in the target sub-graph, a clothing texture of the digital human to be optimized is generated;
[0115] The rendering engine 503 is used to apply the clothing texture to the digital human to be optimized to obtain the target digital human; and drive the target digital human.
[0116] In addition, in order to facilitate front-end display, the driving results of the target digital human can also be output to the front-end through an interactive interface.
[0117] Among them, the various operations performed by the intelligent agent have been explained in the aforementioned method embodiments and will not be repeated here.
[0118] In summary, a pet cat is taken as an example and anthropomorphized to illustrate the method for generating a digital human provided by the embodiment of the present disclosure. Figure 6 As shown, including the following:
[0119] In S601, the user may upload a pet cat image through a smart terminal.
[0120] In S602, the agent performs target detection on the pet cat image through the target detection model to obtain the 2D box of the pet cat, and cuts out the 2D box from the pet cat image to obtain the target sub-image.
[0121] In S603, the agent performs asset retrieval, which may specifically include:
[0122] S6031, the intelligent agent subclassifies the pet cat through the classification model to obtain the subclassification category of the pet cat, and detects similar cats that look similar to the pet cat from the cat set corresponding to the subclassification category.
[0123] S6032, the intelligent agent retrieves the character assets corresponding to the similar cat, and then obtains the two-dimensional character model corresponding to the similar cat.
[0124] In S604, the intelligent agent generates clothing textures of the two-dimensional character model, including:
[0125] S6041, the intelligent agent generates a pet cat's appearance description text through VLM, and generates a character image of the same two-dimensional style based on the appearance description text. The character clothing parameters in the two-dimensional character image are generated based on the appearance description text.
[0126] S6042, the agent inputs the appearance description text and the character image into the texture generation model,
[0127] Get the clothing texture of the two-dimensional character model.
[0128] In S605, the intelligent agent applies the generated clothing texture to the two-dimensional character model through the rendering engine, and drives the two-dimensional character model in combination with the background, music, script and action to facilitate interaction with the user.
[0129] The present disclosure provides an innovative AI (Artificial Intelligence) technology application. The application anthropomorphizes the user's pet cat into a two-dimensional character, and presents it to the user in the form of an intelligent entity powered by AI. This function is designed to deepen the emotional bond between users and bring users a better companionship experience.
[0130] In summary, the present invention integrates pattern recognition, large model reasoning and 3D rendering to create a new and attractive virtual human intelligent agent application. The intelligent agent application technology chain includes multiple core modules: location detection, asset retrieval, clothing texture generation, driving and engine rendering.
[0131] The digital human model and clothing texture map used in the disclosed embodiment conform to the general model storage format, so any traditional rendering engine can be used to achieve the driving and real-time interaction of "character body and expression". In practical applications, you can also use an engine embedded in an app (Application) or an engine in the cloud for driving, rendering, and interaction. In addition, it supports outputting the driving results in the form of video. For example, according to actual needs, record a video of a two-dimensional character dancing, greeting the New Year, saying hello, and other actions, and output it to the user as a result.
[0132] Based on the same technical concept, the present disclosure also proposes a digital human generation device 700. Figure 7 As shown, including:
[0133] The segmentation module 701 is used to segment the target object from the image to be processed to obtain a target sub-image;
[0134] A screening module 702 is used to screen out a digital human to be optimized that is compatible with the target object from the digital human set based on the target subgraph;
[0135] A generating module 703, used for generating the clothing texture of the digital human to be optimized based on the appearance features of the target object in the target sub-image;
[0136] A determination module 704 is used to apply the clothing texture to the digital human to be optimized to obtain a target digital human;
[0137] The driving module 705 is used to drive the target digital human.
[0138] In some embodiments, the screening module comprises:
[0139] A classification unit is used to sub-classify the target object based on the species category corresponding to the target object in the target sub-graph to obtain a sub-classification category of the target object;
[0140] A screening unit is used to screen candidate objects with similar object features to the target object from the candidate object set corresponding to the sub-classification category as similar objects; the candidate objects in the candidate object set and the target object belong to the same species;
[0141] The acquisition unit is used to acquire the 3D digital human model corresponding to the similar object in the digital human set as the digital human to be optimized.
[0142] In some embodiments, the screening unit is specifically used for:
[0143] Extracting a first feature of a target object from a target subgraph based on a feature extraction macromodel, and extracting second features of a plurality of candidate objects in a candidate object set;
[0144] Determining similarities between the second feature and the first feature of the plurality of candidate objects;
[0145] The candidate objects corresponding to the second feature with the highest similarity are selected as similar objects.
[0146] In some embodiments, the generating module comprises:
[0147] The image-to-text unit is used to process the target sub-graph based on the image-to-text large model to obtain the appearance description text of the target object;
[0148] The first texture generation unit is used to input the appearance description text and the target sub-image into the texture generation model to generate the clothing texture of the digital human to be optimized.
[0149] In some embodiments, the generating module comprises:
[0150] The image-to-text unit is used to process the target sub-graph based on the image-to-text large model to obtain the appearance description text of the target object;
[0151] The text image unit is used to process the appearance description text based on the text image large model to generate a character image; the clothing of the character object in the character image is generated under the constraint of the appearance description text; the style of the character object is the same as the character style of the digital human to be optimized;
[0152] The second texture generation unit is used to input the appearance description text and the character image into the texture generation model to obtain the clothing texture of the digital human to be optimized.
[0153] In some embodiments, the driving module includes:
[0154] The first driving unit is used to determine the target digital human based on at least one of the following preset driving parameters:
[0155] Skeleton driven parameters, facial expression driven parameters, and text-to-speech driven parameters.
[0156] In some embodiments, the driving module includes:
[0157] An acquisition unit, used for acquiring interaction information;
[0158] An inference unit, used to process the interactive information based on a large language model to obtain a response text for the interactive information;
[0159] A synthesis unit, used for generating driving parameters of a target digital human based on the response text;
[0160] The second driving unit is used to drive the target digital human based on the driving parameters.
[0161] In some embodiments, the digital human to be optimized includes a two-dimensional character model.
[0162] For the description of specific functions and examples of each module and submodule of the device in the embodiment of the present disclosure, reference can be made to the relevant description of the corresponding steps in the above method embodiment, which will not be repeated here.
[0163] In the technical solution disclosed herein, the acquisition, storage and application of user personal information involved are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0164] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium and a computer program product.
[0165] Figure 8A schematic block diagram of an example electronic device 800 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.
[0166] like Figure 8 As shown, the device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0167] A number of components in the device 800 are connected to the I / O interface 805, including: an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a disk, an optical disk, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the device 800 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0168] The computing unit 801 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 801 performs the various methods and processes described above, such as a digital human generation method. For example, in some embodiments, the digital human generation method may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 808. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 800 via ROM 802 and / or communication unit 809. When the computer program is loaded into RAM 803 and executed by the computing unit 801, one or more steps of the digital human generation method described above may be performed. Alternatively, in other embodiments, the computing unit 801 may be configured to perform the digital human generation method in any other appropriate manner (e.g., by means of firmware).
[0169] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0170] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.
[0171] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0172] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0173] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0174] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server combined with a blockchain.
[0175] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.
[0176] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the principles of the present disclosure shall be included in the protection scope of the present disclosure.
Claims
1. A method for generating a digital human, comprising: Segment the target object from the image to be processed to obtain a target sub-image; Based on the target subgraph, a digital human to be optimized that is compatible with the target object is selected from the digital human set; Generate the clothing texture of the digital human to be optimized based on the appearance features of the target object in the target sub-image; Applying the clothing texture to the digital human to be optimized to obtain a target digital human; Driving the target digital human.
2. The method according to claim 1, wherein: The step of selecting a digital human to be optimized that is compatible with the target object from a digital human set based on the target subgraph includes: Based on the species category corresponding to the target object in the target subgraph, subclassify the target object to obtain a subclassified category of the target object; In the candidate object set corresponding to the sub-classification category, candidate objects having object features similar to those of the target object are selected as similar objects; the candidate objects in the candidate object set and the target object belong to the same species; In the digital human set, the 3D digital human model corresponding to the similar object is obtained as the digital human to be optimized.
3. The method according to claim 2, wherein: The step of selecting, from the candidate object set corresponding to the sub-classification category, candidate objects with similar object features to the target object as similar objects includes: Extracting a first feature of the target object from the target subgraph based on a feature extraction macromodel, and extracting second features of a plurality of candidate objects in the candidate object set; Determining similarities between second features of the plurality of candidate objects and the first feature; The candidate object corresponding to the second feature with the highest similarity is selected as the similar object.
4. The method according to any one of claims 1 to 3, wherein: The step of generating the clothing texture of the digital human to be optimized based on the appearance features of the target object in the target sub-image comprises: Processing the target subgraph based on the graph-to-text model to obtain appearance description text of the target object; The appearance description text and the target sub-image are input into a texture generation model to generate the clothing texture of the digital human to be optimized.
5. The method according to any one of claims 1 to 3, wherein: The step of generating the clothing texture of the digital human to be optimized based on the appearance features of the target object in the target sub-image comprises: Processing the target subgraph based on the graph-to-text model to obtain appearance description text of the target object; The appearance description text is processed based on the Wenshengtu large model to generate a character image; the clothing of the character object in the character image is generated under the constraint of the appearance description text; the style of the character object is the same as the character style of the digital human to be optimized; The appearance description text and the character image are input into a texture generation model to obtain the clothing texture of the digital human to be optimized.
6. The method according to any one of claims 1 to 5, wherein: The driving of the target digital human comprises: The target digital human is determined based on at least one of the following preset driving parameters: Skeleton driven parameters, facial expression driven parameters, and text-to-speech driven parameters.
7. The method according to any one of claims 1 to 5, wherein: The driving of the target digital human comprises: Obtain interaction information; Processing the interaction information based on the large language model to obtain a response text for the interaction information; Generate driving parameters of the target digital human based on the response text; The target digital human is driven based on the driving parameters.
8. The method according to any one of claims 1 to 7, wherein: The digital human to be optimized includes a two-dimensional character model.
9. An intelligent agent, comprising: An interactive interface for obtaining images to be processed; An artificial intelligence module is used to segment the target object from the image to be processed to obtain a target sub-image; Based on the target subgraph, a digital human to be optimized that is compatible with the target object is selected from the digital human set; Generate the clothing texture of the digital human to be optimized based on the appearance features of the target object in the target sub-image; A rendering engine is used to apply the clothing texture to the digital human to be optimized to obtain a target digital human; and drive the target digital human.
10. A digital human generation device, comprising: A segmentation module is used to segment the target object from the image to be processed to obtain a target sub-image; A screening module, used for screening out a digital human to be optimized that is compatible with the target object from the digital human set based on the target subgraph; A generating module, used for generating the clothing texture of the digital human to be optimized based on the appearance features of the target object in the target sub-image; A determination module, used for applying the clothing texture to the digital human to be optimized to obtain a target digital human; A driving module is used to drive the target digital human.
11. The device according to claim 10, wherein: The screening module comprises: A classification unit, configured to sub-classify the target object based on the species category corresponding to the target object in the target sub-graph, to obtain a sub-classification category of the target object; a screening unit, configured to screen, from a candidate object set corresponding to the sub-classification category, candidate objects with similar object features to the target object as similar objects; the candidate objects in the candidate object set and the target object belong to the same species; The acquisition unit is used to acquire the 3D digital human model corresponding to the similar object in the digital human set as the digital human to be optimized.
12. The device according to claim 11, wherein The screening unit is specifically used for: Extracting a first feature of the target object from the target subgraph based on a feature extraction macromodel, and extracting second features of a plurality of candidate objects in the candidate object set; Determining similarities between second features of the plurality of candidate objects and the first feature; The candidate object corresponding to the second feature with the highest similarity is selected as the similar object.
13. The device according to any one of claims 10 to 12, wherein: The generating module comprises: A picture-to-text unit, used for processing the target sub-graph based on the picture-to-text large model to obtain an appearance description text of the target object; The first texture generation unit is used to input the appearance description text and the target sub-image into a texture generation model to generate the clothing texture of the digital human to be optimized.
14. The device according to any one of claims 10 to 12, wherein: The generating module comprises: A picture-to-text unit, used for processing the target sub-graph based on the picture-to-text large model to obtain an appearance description text of the target object; A Wensheng graph unit, used for processing the appearance description text based on the Wensheng graph large model to generate a character image; the clothing of the character object in the character image is generated under the constraint of the appearance description text; the style of the character object is the same as the character style of the digital human to be optimized; The second texture generation unit is used to input the appearance description text and the character image into a texture generation model to obtain the clothing texture of the digital human to be optimized.
15. The device according to any one of claims 10 to 14, wherein: The driving module comprises: The first driving unit is configured to determine the target digital human based on at least one of the following preset driving parameters: Skeleton driven parameters, facial expression driven parameters, and text-to-speech driven parameters.
16. The device according to any one of claims 10 to 14, wherein: The driving module comprises: An acquisition unit, used for acquiring interaction information; An inference unit, configured to process the interaction information based on a large language model to obtain a response text for the interaction information; A synthesis unit, used for generating driving parameters of the target digital human based on the response text; The second driving unit is used to drive the target digital human based on the driving parameters.
17. The device according to any one of claims 10 to 16, wherein: The digital human to be optimized includes a two-dimensional character model.
18. An electronic device, comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 8.
19. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-8.
20. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 8.
Citation Information
Cited By
Three-dimensional digital human detection method and device based on large model, electronic equipment, medium, program product and three-dimensional digital human
CN121169888A