Digital human generation method, platform, electronic device and storage medium
The method of fusing a target object model with a point cloud of head key features automates digital human generation, addressing manual inefficiencies and skill-dependent accuracy issues, resulting in cost-effective and high-quality digital human images.
Patent Information
- Application Number
- JP2024100035
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2023-10-16
- Filing Date
- 2024-06-20
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2044-06-20
AI Technical Summary
The traditional production of 3D ultra-realistic digital human images requires highly specialized input, involves multiple manual stages, is cost-prohibitive, and struggles to accurately replicate real-world features like eyes and lips, with the final effect heavily dependent on individual skill levels.
A method and platform for generating digital humans by obtaining a target object model and fusing it with a point cloud of head key features from a pre-defined feature library, utilizing pre-trained models and avatar creation operations to enhance accuracy and efficiency.
Automates the creation of digital humans, reducing costs and time, improving accuracy, and enhancing the realism of generated images, while lowering entry barriers for digital human creation.
Smart Images

Figure 0007735628000001 
Figure 0007735628000002 
Figure 0007735628000003
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to the fields of computer technology and artificial intelligence, particularly to technical fields such as augmented reality, virtual reality, computer vision, and deep learning, and can be applied to scenes such as meta-universes and virtual digital humans. Specifically, the present disclosure relates to a method, platform, electronic device, and storage medium for generating a digital human. [Background technology]
[0002] In the traditional production of digital human images, especially 3D ultra-realistic digital human images, highly specialized input is required in multiple stages, including original drawing, modeling, binding, and animation. The manual production cycle usually takes several months, the input costs are extremely high, and the effect can only be displayed in stages. Operations such as correction and optimization all require a large amount of additional input.
[0003] In addition, with conventional technologies, it is difficult to restore the ideal features of a real-world person, such as the shape of the eyes and lips, in an artificially created digital human image, and the final effect is largely limited by the individual skill level at each stage of digital human creation. Summary of the Invention [Problem to be solved by the invention]
[0004] The present disclosure provides a method, platform, electronic device, and storage medium for generating a digital human. [Means for solving the problem]
[0005] According to one aspect of the present disclosure, there is provided a method for generating a digital human, comprising: Obtaining a corresponding target object model based on an image of the digital human to be generated; obtaining a point cloud of corresponding head key features from a pre-defined feature library based on the head key features in the image; and fusing the point cloud of the head key features with the target object model to obtain a digital human image.
[0006] According to another aspect of the present disclosure, there is provided a digital human generation platform, comprising: a model acquisition module for acquiring a corresponding target object model based on an image of the digital human to be generated; a feature acquisition module for acquiring a point cloud of corresponding head key features from a pre-defined feature library based on the head key features in the image; a fusion module for fusing the point cloud of the head key features with the target object model to obtain a digital human image.
[0007] According to another aspect of the present disclosure, there is provided an electronic device, comprising: at least one processor; a memory communicatively coupled to the at least one processor; The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enable the at least one processor to perform the method of the aspect and any possible implementation.
[0008] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium having stored thereon computer instructions that cause the computer to perform the method of the aspect and any possible implementation thereof.
[0009] According to another aspect of the present disclosure, a computer program product is provided, comprising a computer program which, when executed by a processor, implements the method of the aspect and any possible implementation manner.
[0010] The technology disclosed herein can automatically realize the creation of digital humans and obtain digital human images, overcoming the problems of the prior art, which require multiple manual steps to create digital humans, high costs, and poor accuracy, and can effectively save the cost and time required to create digital human images, as well as effectively improve the accuracy of the generated digital human images and the generation efficiency of digital human images.
[0011] It should be understood that the contents described herein are not intended to identify key or important features of the embodiments of the present disclosure, nor should they be used to limit the scope of the present disclosure. Other features of the present disclosure can be readily understood through the following specification. [Brief explanation of the drawings]
[0012] The drawings are for a better understanding of the present application and are not intended to limit the present application. [Figure 1] FIG. 1 is a schematic diagram according to a first embodiment of the present disclosure. [Figure 2] FIG. 10 is a schematic diagram according to a second embodiment of the present disclosure. [Figure 3] 1 is a schematic diagram of an interface display of attribute information of head key features provided by the present disclosure; [Figure 4] FIG. 10 is a schematic diagram according to a third embodiment of the present disclosure. [Figure 5] FIG. 10 is a schematic diagram according to a fourth embodiment of the present disclosure. [Figure 6] FIG. 1 is a block diagram of an electronic device for implementing a method according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0013] Hereinafter, exemplary embodiments of the present application will be described with reference to the drawings. For ease of understanding, various details of the embodiments of the present application are included and should be considered as merely examples. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present application. Similarly, for the sake of clarity, the following description will omit descriptions of well-known functions and structures.
[0014] Obviously, the described embodiments are some of the embodiments of the present application, but not all of the embodiments, and all other embodiments that a person skilled in the art can obtain without creative effort according to the embodiments of the present application all belong to the scope of protection of the present application.
[0015] Note that terminal devices according to the embodiments of the present disclosure may include, but are not limited to, smart devices such as mobile phones, personal digital assistants (PDAs), wireless handheld devices, and tablet computers. Display devices may include, but are not limited to, devices with display capabilities such as personal computers and televisions.
[0016] Furthermore, the term "and / or" in this specification describes only the relation between related objects and indicates that three types of relations can exist, for example, A and / or B can indicate three cases: only A exists, A and B exist simultaneously, or only B exists. It should be understood that the symbol " / " generally indicates that the related objects before and after it are in an "or" relation.
[0017] The creation of digital humans in the prior art is basically done manually, which not only results in high production costs, but also in low accuracy and poor performance of the created digital humans.
[0018] 1 is a schematic diagram according to a first embodiment of the present disclosure. As shown in FIG. 1, this embodiment provides a method for generating a digital human, which specifically can include the following steps: S101, obtaining a corresponding target object model according to an image of the digital human to be generated; S102, based on the head key features in the image, obtain a point cloud of corresponding head key features from a pre-defined feature library; S103, the point cloud of head key features is fused with the target object model to obtain a digital human image.
[0019] The execution body of the digital human generation method of this embodiment may be a digital human generation platform, which may be a software integrated platform that can automatically realize the generation of a digital human, or may configure a corresponding electronic device based on the platform to realize the generation of a digital human.
[0020] In this embodiment, a corresponding digital human image can be automatically generated based on the image of the digital human to be generated, and the image of the digital human to be generated is a two-dimensional (2D) image, or the digital human image can be a 3D digital human image.
[0021] In the generated digital human image, the features of the digital human are mainly concentrated in the head and facial region of the head, so in this embodiment, the image of the digital human to be generated may be a frontal image including a person's head, for example, only a head image, or an upper body image including the person's head features.
[0022] The target object model in this embodiment can be a prototype of a pre-created digital human model. To improve the efficiency of digital human generation, this embodiment does not need to create a digital human model prototype, and directly generates a corresponding digital human image based on the acquired target object model.
[0023] Since the features of a digital human are mainly concentrated in the head, in this embodiment, for each key head feature in the image, the point cloud of the corresponding key head feature is obtained from the feature library, and then merged into the target object model, so as to finally obtain a digital human image.
[0024] In this embodiment, the head key features may be limited to specific areas according to actual needs, including at least one feature such as eyebrows, eyes, nose, mouth, and face shape.
[0025] The digital human generation method of this embodiment uses the above-mentioned technical solution to automatically realize the creation of digital humans and obtain digital human images, overcoming the problems of the prior art, which require multiple manual steps to create digital humans, high costs, and poor accuracy, and can effectively save the cost and time required to create digital human images. It can also effectively improve the accuracy of the generated digital human images and the generation efficiency of digital human images, allowing the generated digital human images to have ultra-realistic expression effects. At the same time, it can effectively lower the entry requirements for digital human creation and promote the progress of the digital industry.
[0026] FIG. 2 is a schematic diagram of a second embodiment of the present disclosure. As shown in FIG. 2, the digital human generation method of this embodiment is based on the technical solution of the embodiment shown in FIG. 1 above, and further describes the technical solution of this disclosure in more detail. As shown in FIG. 2, the digital human generation method of this embodiment can specifically include the following steps: S201, extracting attribute features of the digital human based on an image of the digital human to be generated; Since the main features of a digital human are all concentrated in the head, the image of the digital human to be generated in this embodiment may be a single frontal image including the head, which can clearly represent each key head feature area on the front of the person's head.
[0027] For example, in this embodiment, a pre-trained attribute feature extraction model can be used to extract attribute features of a digital human. The attribute feature extraction model may be a pre-trained neural network model. When used, an image of a digital human is input to the attribute feature extraction model, and the attribute feature extraction model can extract attribute features of the person in the image based on the input image. For example, the attribute features in this embodiment can include middle-aged woman, middle-aged man, girl, boy, elderly man, elderly woman, etc.
[0028] Alternatively, in this embodiment, other algorithms can be used to analyze the image of the digital human to be generated to obtain the attribute features of the digital human, and are not limited thereto.
[0029] S202, based on the attribute features of the digital human, obtain a corresponding target object model from a pre-defined model library; Because people with different attribute features have different facial region features, in order to generate digital human images more accurately, in this embodiment, a model library can be preset, and the model library can include multiple object models, such as at least one of a middle-aged female model, a middle-aged male model, a female model, a male model, an elderly male model, and an elderly female model. For each object model, corresponding attribute features can be set during pre-setting. Therefore, based on the attribute features of the digital human, a corresponding target object model can be obtained from the preset model library. Compared with the prior art, this can directly eliminate the processes of original drawing and modeling, and can directly use the preset matching target object model, thereby effectively improving the production efficiency of digital human images.
[0030] Optionally, in one embodiment of the present disclosure, if attribute features of a digital human cannot be extracted based on an image of the digital human to be generated, a preset standard model can be used as the target object model. The standard model can be one model in a preset model library, such as a middle-aged female model. Alternatively, a separate individual standard model can be preset and used when attribute features cannot be extracted, ensuring that a reasonable target object model can be obtained in any case.
[0031] S203: Based on the plurality of head key features contained in the image, detect whether there is an unfused head key feature; if there is, execute step S204; if not, execute step S210; Specifically, since an image can contain multiple head key features, in order to facilitate one-by-one fusion into the target object model, one head key feature can be fused and then labeled each time to avoid repeated operations.
[0032] S204, obtain the target attribute information of one unfused head key feature in the image, and then execute step S205; In this embodiment, the head key features may be eyebrows, eyes, nose, mouth, ears, or face shape, etc. Corresponding target attribute information can be used to define the head key features, for example, eye attribute information can include apricot eyes, red phoenix eyes, hawk eyes, elongated eyes, narrow eyes, etc.
[0033] In this embodiment, the attribute extraction model for each head key feature is pre-trained to extract the target attribute information of the head key feature, or the target attribute information of the head key feature in the image selected by the user can be obtained.
[0034] S205, based on the target attribute information of the head key feature, obtain the point cloud of the corresponding head key feature from the pre-defined feature library, and execute step S206; Before step S205, a step of collecting point clouds of multiple key head features of each person among the multiple people and attribute information of each key head feature and storing them in a feature library may be included.
[0035] In this embodiment, the collected multiple people include people with different attribute information, such as a middle-aged woman, a middle-aged man, a girl, a boy, an elderly man, or an elderly woman. The attribute information of the same key head feature among the multiple people should also be as different as possible. To provide effective feature data support for digital human generation, the feature library should contain a wide variety of key head features with various types of attribute information. For example, for eyebrows, the eyebrow point clouds and corresponding eyebrow shape information of people with as many different eyebrow shapes as possible should be collected. For eyes, the eye point clouds and corresponding eye shape information of people with as many different eyebrow shapes as possible should be collected. For mouths, the mouth point clouds and corresponding mouth shape information of people with as many different mouth shapes as possible should be collected. Based on the above-mentioned established feature library, when step S205 is specifically performed, the digital human generation platform can automatically accurately and efficiently match the target attribute information of the key head feature with the point cloud of the key head feature corresponding to the target attribute information based on all the target attribute information of the key head feature in the feature library.
[0036] For example, Figure 3 is a schematic diagram of an interface displaying attribute information of key head features provided by the present disclosure. As shown in Figure 3, in a digital human generation platform, the key head feature is the nose as an example. The various nose shapes shown in Figure 3 are nose shapes supported and used in the feature library of the digital human generation platform. When a user clicks on a nose shape, the interface shown in Figure 3 can be displayed, allowing the user to select a nose shape that best matches the nose shape in the image of the digital human to be generated. The user can also reset the currently selected nose shape to 0 and reselect a nose shape.
[0037] S206: Perform registration between the point cloud of the head key features and the target object model, and then execute step S207; The registration in this embodiment can be adjusted by at least one of the manipulation methods of displacement, rotation, scaling, etc. For example, specifically, by adjusting the point cloud of the head key feature, the coordinate system of the point cloud of the head key feature can be made to match with that of the target object model, and the size of the point cloud of the head key feature can be made to match with the area size of the head key feature of the target object model, so that the point cloud of the head key feature can be accurately and efficiently matched with the target object model.
[0038] S207, transfer the point cloud of the head key feature to the corresponding region of the head key feature in the target object model, and then execute step S208; S208: Detect whether the similarity between the head key feature in the processed target object model and the head key feature in the image is equal to or greater than a preset similarity threshold; if so, return to step S203; if not, execute step S209; For example, the head key feature in the front view of the fused target object model and the head key feature in the image can be extracted, and then a pre-trained similarity calculation model can be used to calculate the similarity between them. Alternatively, a pre-trained feature representation model can be used to extract the feature representation of the head key feature region in the front view of the fused target object model and the feature representation of the head key feature in the image, and then the similarity between the two feature representations can be calculated as the similarity between them.
[0039] Based on this step, in actual applications, there may be cases where an ideal effect is achieved after combining key head features, and in this case, it is understood that the avatar creation operation is not performed for the key head features.
[0040] S209, based on the user's trigger, perform an avatar creation operation on the point cloud of the head key features fused with the target object model, and return to step S208; The avatar creation operation of this embodiment is performed on the point cloud of the head key features merged with the target object model based on the user's trigger, with the pre-established implicit constraint surface and the pre-set dynamic curve as constraints. Specifically, the avatar creation operation can include the following steps: (1) obtaining motion information of a first controller set on an implicit constraint surface triggered by a user; Before adjusting the first controller, the implicit constraint surface is a surface that matches the topological structure of the surface of the target object model, and the identifiers and positions of the included points are exactly the same. Multiple first controllers are set on the implicit constraint surface, but the implicit constraint surface is invisible to the user. A first motion mapping relationship exists between each first controller and the points on the implicit constraint surface, and can drive the movement of the points on the implicit constraint surface when the first controller operates.
[0041] In the embodiment of the present disclosure, since the main features of the digital human image are concentrated in the head, the fusion and avatar creation operations are also mainly performed on the key features of the head of the target object model. In this case, the implicit constraint surface and the topology structure of the target object model only need to be limited to the head of the digital human. In this embodiment, the motion information of the first controller can include the direction and displacement of the first controller movement.
[0042] (2) obtaining motion information of the triggered point on the implicit constraint surface based on a first motion mapping relationship between the first controller on the implicit constraint surface and the point on the implicit constraint surface established in advance and motion information of the first controller; Alternatively, in this embodiment, more than 85 usable first controllers can be designed for the implicit constraint surface based on facial features, muscle flow, and skeletal root positions. A target object model can have more than 20,000 points. Therefore, one first controller can trigger the movement of multiple points on the target object model. When a user triggers a first controller, the user can trigger the movement of only one first controller, or multiple first controllers simultaneously. For example, the control interface can first prompt the user to select the first controller to be triggered. After the selection is made, the user can simultaneously trigger the movement of all selected first controllers. In this embodiment, the point movement information can also include the movement direction and displacement of the point.
[0043] Specifically, the motion information of each point is determined by a weighted addition method based on the motion information of each first controller among the multiple first controllers that triggers the point and the weight of each first controller, and the final motion information of the point is then determined, and the motion position of the point is also determined.
[0044] The main purpose of the implicit constraint surface in this embodiment is to resolve unreasonable situations caused by large shape changes, such as changes in the skull vertex, which are transmitted to the forehead, sides, back of the head, and neck through this constraint, thereby ensuring the physics and aesthetics of the target object model to a certain extent and effectively improving the aesthetics of the generated digital human.
[0045] (3) determining, based on the motion information of the triggered point on the implicit constraint surface, motion information of a second controller that is located at the same position as the triggered point on the target object model; In this embodiment, in the target object model, some second controllers are deployed at points of the target object model. In the control process, after a point on the implicit constraint surface is triggered to move, the second controllers deployed at points with the same identifiers in the target object model will perform synchronous actions according to the movement of the point on the implicit constraint surface.
[0046] (4) controlling the motion information of the point on the target object model based on a second motion mapping relationship between the second controller and the point on the target object model that is previously established for the target object model and the motion information of the second controller; In this embodiment, more than 550 second controllers are set on the target object model. Similarly, each second controller also establishes a binding relationship with multiple points on the target object model, for example, establishing the above-mentioned second motion mapping relationship. Compared with the number of first controllers, the number of second controllers is greater, which can realize more precise control of points on the target object model.
[0047] Based on step (3), it can be seen that after the motion of the triggered point on the implicit constraint surface, the second controller deployed on the point with the same identifier on the target object model will perform a synchronized motion according to the motion of the point, and further, based on the second motion mapping relationship between the second controller set on the target object model and the point on the target object model, can obtain the motion information of all points bound by the second controller on the target object model, and further, based on the motion information of the points on the target object model, control the motion of the corresponding points. In this process, the motion of the first controller on the implicit constraint surface is transmitted to the motion of the points on the target object model to achieve accurate adjustment to the target object model.
[0048] (5) Using a preset dynamic curve as a constraint, position adjustment is performed on the point cloud of the head key features fused to the target object model.
[0049] For example, when specifically implementing step (5), the following steps may be included: (a) detecting whether the second controller is on a preset dynamic curve, and setting at least two second controllers for each preset dynamic curve; The preset dynamic curves in this embodiment may be some natural curves of the target object model's face, such as the jawline, jawline, or nasolabial fold. In the avatar creation optimization process, when adjusting areas such as the jawline, jawline, or nasolabial fold, it is difficult to achieve a continuous, smooth, and natural topology effect using a single controller. Therefore, this embodiment introduces dynamic curves, which can be made more continuous, smooth, and natural by controlling the movement of at least two second controllers on the dynamic curves.
[0050] (b) in response to the second controller being on the preset dynamic curve, obtaining operation information of another second controller on the dynamic curve based on the operation information of the second controller; (c) Based on the motion information of the second controller and the motion information of other second controllers in the dynamic curve, a position adjustment is performed on the point cloud of the head key features fused into the target object model.
[0051] Furthermore, the method may further include a step of, in response to the second controller not being on the predetermined dynamic curve, performing position adjustment on the point cloud of the head key features fused to the target object model based on the motion information of the second controller.
[0052] Based on the above, we can understand that the essence of dynamic curve constraint is that multiple consecutive, same-level secondary controllers are connected in series using a single spatial curve. When the endpoints are fixed, the position and rotation changes of each secondary controller are all propagated to its surroundings, with the closer the secondary controller, the greater the weight. Conversely, the farther the endpoint, the smaller the weight of the secondary controller. The relationship between weight and distance is also preset.
[0053] In this embodiment, when performing an avatar creation operation on the point cloud of the head key features integrated into the target object model based on the motion information of the second controller and the motion information of other second controllers on the dynamic curve, smoothing processing is further performed on the points around the adjusted points. In this embodiment, the feature curve on which the points are located is adjusted as a constraint.
[0054] For example, the surface of the target object model itself is a grid structure containing edges between multiple points and neighboring points. When adjusting a point, the feature curve on which the point is located can be used as the topological center. To ensure a uniform and smooth topology, the neighboring topology lines also need to be updated. This process primarily uses dynamic weights to transition from the feature curve to the outline, while simultaneously using tangents, normals, and UVs for auxiliary smoothness detection. During the transition from the feature curve to the outline, the adjustment amplitude is attenuated outward. That is, to achieve a smoother adjustment, in actual applications, the change in one feature curve can be attenuated outward by a preset number of feature curves. The preset number can be set according to actual needs, for example, 6, 8, 10, or other values. The adjustment amplitude of the feature curve on which the point is located is the largest. As the distance between the feature curve on which the point is located increases, the adjustment amplitude gradually decreases for each of the neighboring feature curves in the preset number of feature curves until the final adjustment step of the adjustment is completed.
[0055] The dynamic curve constraints of this embodiment ensure smoothness throughout all secondary controller manipulations on the curve. Furthermore, secondary controller control can be propagated to surfaces and points, ensuring the uniformity of the target object model's topology. In real-time manipulation, users can complete specialized and wide-ranging manipulations, ensuring consistent modeling levels and topology.
[0056] In this embodiment, the points on the implicit constraint surface establish binding logic with the target object model. This logic is simple but very effective. For example, a binding relationship is established between a first controller on the implicit constraint surface and a point on the implicit constraint surface, and a binding relationship is also established between a second controller on the target object model and a point on the target object model. The movement of a point on the implicit constraint surface can be transmitted to a second controller located at that point on the target object model. Furthermore, a binding relationship also exists between a second controller on the target object model and a point on the implicit constraint surface. The position change of the point on the implicit constraint surface is transmitted to the second controller located at the corresponding point under the influence of weights. At the same time, the tangent and normal changes of the point are also transmitted to the rotation of the corresponding second controller via weights, which further causes the movement of other points on the target object model bound by the second controller.
[0057] Based on the above, it can be understood that in this embodiment, the 85 or more first controllers established on the implicit constraint surface can be mapped to the 550 or more second controllers in the target object model through binding relationships after updating the point cloud, and further updated to the 20,000 or more point cloud in the target object model, thereby realizing finer control and making the generated digital human more realistic and accurate.
[0058] In this embodiment, when creating an avatar, implicit constraint surfaces and dynamic curve constraints are used to ensure that the topology effect of the target object model is always constrained during operation, ensuring the speed of product use and the versatility and standardity of product output, effectively ensuring the accuracy and standardity of the generated digital human, and effectively improving the efficiency of digital human generation.
[0059] S210: determining whether the attached template of the target object model is suitable for the generated digital human image; if so, obtaining the digital human image; if not, executing step S211; S211, adjust the attached template in the digital human image so that the added template is suitable for the digital human image.
[0060] In this embodiment, step S211 is performed in response to the attached template of the target object model not being suitable for the generated digital human image, and if so, a perfect digital human image is obtained.
[0061] The attached template in this embodiment includes at least one of eyeballs, teeth, AO, and lacrimal glands. AO refers to the eyeshell, and is abbreviated as AO in the digital human object model. For example, if the eyeball template is too large, the digital human may not be able to close his or her eyes when closing them, requiring a shrinking adjustment of the eyeball template. If the eyeball template is too small, a gap may exist between the eyeball template and the digital human image when the digital human opens his or her eyes, requiring a scaling adjustment of the eyeball template. Similarly, if the tooth template is too large, the digital human may not be able to close his or her mouth when closing them, requiring a scaling adjustment of the tooth template. If the tooth template is too small, a gap may be too large between the mouth and the tooth template when the digital human opens his or her mouth, requiring a scaling adjustment of the tooth template. Similarly, the adjustment principles for the AO template and the lacrimal gland template are the same, and will not be described in detail here. Through this process, the final digital human image can be more accurate, more representative, and more complete.
[0062] By step S211, a perfect digital human image can be obtained, which can be matched with the features in the image through the above-mentioned feature fusion and avatar creation operations, and the topology structure of the generated digital human image can be effectively standardized through the avatar creation adjustment of the implicit constraint surface and dynamic curve constraint, thereby ensuring the accuracy of the generated digital human image.
[0063] Note that the above step S203 takes as an example the case where multiple key head features contained in the image are all fused. In actual applications, the feature library may not contain key head features corresponding to the target attribute information. In this case, an avatar creation operation can be performed on the point cloud of key head features in the target object model based on a user's trigger. The avatar creation operation method is described in the related description above and will not be described in detail here. After the avatar creation, the avatar creation is stopped until the similarity between the key head features in the target object model and the key head features in the image is equal to or greater than a predetermined similarity threshold.
[0064] The digital human generation method of this embodiment significantly reduces image production costs compared with the prior art: the traditional digital human image production process requires a large amount of manpower and can take several months, but by using this technical solution, the characteristics of the scanned body can be quickly and reasonably combined, and then modified step by step using an avatar creation tool, significantly shortening the image generation cycle to within a few days.
[0065] Furthermore, the digital human generation method of this embodiment can effectively improve the quality of the digital human image. Based on a preset target object model, the digital human image is generated through fusion and avatar creation operations, which significantly improves the efficiency of digital human generation while still maintaining the quality of the generated digital human. Furthermore, the generated digital human image is checked for compatibility with the attached template, and if it is not compatible, adjustments are made to further ensure the quality of the generated digital human.
[0066] 4 is a schematic diagram according to a third embodiment of the present disclosure. As shown in FIG. 4, this embodiment provides a digital human generation platform 400, which includes a model acquisition module 401, a feature acquisition module 402, and a fusion module 403. The model acquisition module 401 is used to acquire a corresponding target object model based on an image of the digital human to be generated; The feature acquisition module 402 is used to acquire a point cloud of corresponding head key features from a pre-defined feature library based on the head key features in the image; The fusion module 403 is used to fuse the point cloud of the head key features with the target object model to obtain a digital human image.
[0067] The digital human generation platform 400 of this embodiment is the same as the implementation of the above-mentioned related method embodiment by using the above-mentioned modules to realize the principles and technical effects of digital human generation. For details, please refer to the description of the above-mentioned related method embodiment and will not be described in detail here.
[0068] Figure 5 is a schematic diagram of a fourth embodiment of the present disclosure. The digital human generation platform 500 of this embodiment will be described in more detail based on the embodiment shown in Figure 4. First, as shown in Figure 5, the digital human generation platform 500 of this embodiment includes the same modules with the same functions as the digital human generation platform shown in Figure 4: a model acquisition module 501, a feature acquisition module 502, and a fusion module 503.
[0069] As shown in FIG. 5 , the digital human generation platform 500 of this embodiment further includes a first detection module 504 and an avatar creation module 505; a first detection module 504 for detecting whether a similarity between a head key feature in the target object model after fusion and a head key feature in the image is equal to or greater than a predetermined similarity threshold; The avatar creation module 505 is used to perform an avatar creation operation on the point cloud of the head key features fused to the target object model based on a user trigger in response to the similarity being less than a preset similarity threshold.
[0070] Further optionally, in one embodiment of the present disclosure, the avatar creation module 505: It is used to perform avatar creation operations on the point cloud of the head key features fused with the target object model based on a user trigger, using pre-established implicit constraint surfaces and pre-set dynamic curves as constraints.
[0071] Further optionally, in one embodiment of the present disclosure, the avatar creation module 505: Obtain operation information of a first controller set on the implicit constraint surface triggered by the user, the implicit constraint surface being a surface that completely matches the surface topological structure of the target object model before being controlled, and the implicit constraint surface being invisible to the user; Obtaining motion information of a triggered point on the implicit constraint surface based on a first motion mapping relationship between a first controller on the implicit constraint surface and a point on the implicit constraint surface that has been established in advance and motion information of the first controller; determining motion information of a second controller that is co-located with the point on the target object model based on motion information of the triggered point on the implicit constraint surface; controlling the motion information of the points on the target object model based on a second motion mapping relationship between a second controller and the points on the target object model that is established in advance and motion information of the second controller; A pre-defined dynamic curve is used as a constraint to perform position adjustment on the point cloud of the head key features fused to the target object model.
[0072] Further optionally, in one embodiment of the present disclosure, the avatar creation module 505: Detect whether the second controller is on a preset dynamic curve, and set at least two second controllers for each preset dynamic curve; In response to the second controller being on a preset dynamic curve, acquiring operation information of another second controller on the dynamic curve based on operation information of the second controller; It is used to perform position adjustment on the point cloud of the head key features fused to the target object model based on the motion information of the second controller and the motion information of other second controllers in the dynamic curve.
[0073] Further optionally, in one embodiment of the present disclosure, the avatar creation module 505 further comprises: In response to the second controller not being on a preset dynamic curve, a position adjustment is made to the point cloud of the head key features fused to the target object model based on the motion information of the second controller.
[0074] Optionally, as shown in FIG. 5 , in an embodiment of the present disclosure, the digital human generation platform 500 further includes a second detection module 506 and an adjustment module 507; a second detection module 506 for detecting whether the attached template of the target object model is suitable for the digital human image; The adjustment module 507 is used to adjust the attached template in the digital human image in response to the attached template of the target object model being appropriate for the digital human image.
[0075] Further optionally, in one embodiment of the present disclosure, the model acquisition module 501: Extracting attribute features of the digital human based on the image of the digital human to be generated; Based on the attribute features of the digital human, a corresponding target object model is obtained from a pre-defined model library, where the model library includes a plurality of object models.
[0076] Further optionally, in one embodiment of the present disclosure, the model acquisition module 501: Based on the image of the digital human to be generated, attribute features of the digital human are not extracted, but are used to set a pre-defined standard model as a target object model.
[0077] Further optionally, in one embodiment of the present disclosure, the feature acquisition module 502: obtaining target attribute information of head key features in the image; It is used to obtain a point cloud of the corresponding head key feature from the feature library based on the target attribute information of the head key feature.
[0078] Further optionally, in one embodiment of the present disclosure, the avatar creation module 505 further comprises: If the feature library does not contain a head key feature corresponding to the target attribute information, it is used to perform an avatar creation operation on the point cloud of the head key feature in the target object model based on a user trigger.
[0079] Optionally, as shown in FIG. 5 , in one embodiment of the present disclosure, the digital human generation platform 500 further includes a collection module 508: The collection module 508 is used to collect point clouds of multiple head key features of each person in the multiple people and attribute information of each head key feature, and store them in the feature library.
[0080] Further optionally, in one embodiment of the present disclosure, the fusion module 503: registering the head key feature point cloud with the target object model; It is used to transfer the point cloud of the head key feature to the region of the corresponding head key feature in the target object model. The digital human generation platform 500 of this embodiment is the same as the implementation of the above-mentioned related method embodiment by using the above-mentioned modules to realize the principles and technical effects of digital human generation. For details, please refer to the description of the above-mentioned related method embodiment and will not be described in detail here.
[0081] According to embodiments of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium, and a computer program product.
[0082] As shown in Figure 6, a block diagram of an electronic device 600 for implementing an embodiment of the present disclosure is shown. The electronic device is intended to represent various types of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. The electronic device may also represent various types of mobile devices, such as personal digital assistants, mobile phones, smartphones, wearable devices, and other similar computing devices. The components, their connections and relationships, and their functions shown herein are merely examples and are not intended to limit the description herein and / or the practice of the present disclosure as claimed.
[0083] 6, the device 600 includes a computing unit 601, which can perform various appropriate operations and processes based on a computer program stored in a read-only memory (ROM) 602 or loaded from a storage unit 608 into a random access memory (RAM) 603. The RAM 603 can also store various programs and data required for the device 600 to operate. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0084] Multiple components in device 600 are connected to I / O interface 605, including input units 606 such as a keyboard, mouse, etc., output units 607 such as various types of displays, speakers, etc., storage units 608 such as a disk, optical disk, etc., and communication units 609 such as a network card, modem, wireless communication transceiver, etc. The communication unit 609 enables device 600 to exchange information / data with other devices via computer networks such as the Internet and / or various telecommunication networks.
[0085] The computing unit 601 is a general-purpose and / or special-purpose processing component equipped with various processing and computational capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, computing units that execute various machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 performs the various methods and processes described above, such as the method for constructing a terrain map. For example, in some embodiments, the method for constructing a terrain map may be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 608. In some embodiments, some or all of the computer program may be loaded and / or installed into the device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the computing unit 601, it may perform one or more steps of the methods described above. Alternatively, in other embodiments, the computing unit 601 may be configured to perform the method for constructing a terrain map via any other suitable manner (e.g., via firmware).
[0086] Various implementations of the systems and techniques described herein may be realized in digital electronic circuitry systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-a-chip systems (SOCs), chip programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include being implemented in one or more computer programs that can be executed and / or interpreted by a programmable system that includes at least one programmable processor, which may be a special-purpose or general-purpose programmable processor, and that can receive data and instructions from, and transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0087] Program codes for implementing the methods of the present disclosure can be written using any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus such that, when executed by the processor or controller, the functions / acts specified in the flowcharts and / or block diagrams are performed. The program code can be executed entirely on the machine, partially on the machine, partially on the machine as a separate software package and partially on a remote machine, or entirely on a remote machine or server.
[0088] In the context of this disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use with or in connection with an instruction execution system, device, or apparatus. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium includes, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0089] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to a user, and a keyboard and pointing device (e.g., a mouse or trackball) through which a user can provide input to the computer. Other types of devices can also be used to provide interaction with a user; for example, feedback provided to the user can be any form of sensing feedback (e.g., visual feedback, auditory feedback, or haptic feedback) and can receive input from the user in any form (including acoustic input, voice input, and tactile input).
[0090] The systems and techniques described herein can be implemented in a computing system including a back-end component (e.g., a data server), a computing system including a middleware component (e.g., an application server), a computing system including a front-end component (e.g., a user computer having a graphical user interface or a web browser through which a user interacts with an implementation of the systems and techniques described herein), or any combination of such back-end, middleware, and front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), the Internet, and a blockchain network.
[0091] The computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The relationship between the client and the server is created by computer programs running on corresponding computers and having a client-server relationship to each other. The server may be a cloud server, a server in a distributed system, or a server that combines blockchains.
[0092] It should be understood that steps can be rearranged, added, or deleted using the various types of flows shown above. For example, the steps described in the present disclosure may be performed in parallel, sequentially, or in a different order, but this specification is not limited thereto as long as the technical solution disclosed in the present disclosure can achieve the desired results.
[0093] The above specific implementation methods do not constitute limitations on the scope of protection of the present disclosure. Those skilled in the art may make various modifications, combinations, subcombinations, and substitutions based on design requirements and other factors. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present disclosure shall fall within the scope of protection of the present disclosure.
Claims
1. 1. A method for generating a digital human, comprising: Obtaining a corresponding target object model based on an image of the digital human to be generated; obtaining a point cloud of corresponding head key features from a pre-defined feature library based on the head key features in the image; Fusing the point cloud of the head key features with the target object model to obtain a digital human image; After fusing the point cloud of the head key features into the target object model, and before acquiring a digital human image, the method includes: detecting whether a similarity between a head key feature in the target object model after fusion and a head key feature in the image is equal to or greater than a predetermined similarity threshold; and performing a create avatar operation on the point cloud of the head key features fused to the target object model based on a user trigger in response to the similarity being less than a predetermined similarity threshold. How to generate a digital human.
2. performing an avatar creation operation on the point cloud of the head key features fused to the target object model based on a user trigger, and performing an avatar creation operation on the point cloud of the head key features fused with the target object model based on a user trigger, using a pre-established implicit constraint surface and a pre-set dynamic curve as constraints. The method for generating a digital human according to claim 1 .
3. performing an avatar creation operation on the point cloud of the head key features fused to the target object model based on a user trigger, using a pre-established implicit constraint surface and a pre-set dynamic curve as constraints; Obtaining operation information of a first controller set on the implicit constraint surface triggered by the user, the implicit constraint surface being a surface that completely matches the surface topological structure of the target object model before being controlled, and the implicit constraint surface being invisible to the user; obtaining motion information of a triggered point on the implicit constraint surface based on a first motion mapping relationship between a first controller on the implicit constraint surface and a point on the implicit constraint surface that has been established in advance, and motion information of the first controller; determining motion information of a second controller that is co-located with the point on the target object model based on motion information of the triggered point on the implicit constraint surface; controlling motion information of points on the target object model based on a second motion mapping relationship between a second controller and points on the target object model that is previously established for the target object model and motion information of the second controller; and performing position adjustment on the point cloud of the head key features merged with the target object model using a predetermined dynamic curve as a constraint. The method for generating a digital human according to claim 2 .
4. The step of performing position adjustment on the point cloud of the head key features merged with the target object model using a predetermined dynamic curve as a constraint includes: detecting whether the second controller is on a preset dynamic curve, and setting at least two second controllers for each preset dynamic curve; In response to the second controller being on a preset dynamic curve, acquiring operation information of another second controller on the dynamic curve based on operation information of the second controller; and performing a position adjustment on the point cloud of the head key features fused to the target object model based on the motion information of the second controller and motion information of other second controllers in the dynamic curve. The method for generating a digital human according to claim 3 .
5. The step of performing position adjustment on the point cloud of the head key features merged with the target object model using a predetermined dynamic curve as a constraint includes: and, in response to the second controller not being on a predetermined dynamic curve, performing a position adjustment on the point cloud of the head key features fused to the target object model based on motion information of the second controller. The method for generating a digital human according to claim 4 .
6. After performing an avatar creation operation on the point cloud of the head key features fused to the target object model based on a user trigger, the method further comprises: detecting whether an attached template of the target object model is suitable for the digital human image; and adjusting the attachment template in the digital human image in response to the attachment template of the target object model being inappropriate for the digital human image. A method for generating a digital human according to any one of claims 1 to 5.
7. The step of obtaining a corresponding target object model based on an image of the digital human to be generated includes: extracting attribute features of the digital human based on an image of the digital human to be generated; and obtaining the corresponding target object model from a preset model library based on the attribute features of the digital human, the model library including a plurality of object models. A method for generating a digital human according to any one of claims 1 to 5.
8. The step of obtaining a corresponding target object model based on an image of the digital human to be generated includes: and selecting a predetermined standard model as a target object model when an attribute feature of the digital human is not extracted based on the image of the digital human to be generated. The method for generating a digital human according to claim 7.
9. The step of obtaining a point cloud of corresponding head key features from a predetermined feature library based on the head key features in the image includes: obtaining target attribute information of head key features in the image; and acquiring a point cloud of corresponding head key features from the feature library based on target attribute information of the head key features. A method for generating a digital human according to any one of claims 1 to 5.
10. A method for generating a digital human, comprising: Obtaining a corresponding target object model based on an image of the digital human to be generated; obtaining a point cloud of corresponding head key features from a pre-defined feature library based on the head key features in the image; Fusing the point cloud of the head key features with the target object model to obtain a digital human image; The step of obtaining a point cloud of corresponding head key features from a predetermined feature library based on the head key features in the image includes: obtaining target attribute information of head key features in the image; and acquiring a point cloud of the corresponding head key feature from the feature library based on target attribute information of the head key feature; The method comprises: and if the feature library does not include a head key feature corresponding to the target attribute information, performing an avatar creation operation on a point cloud of the head key feature in the target object model based on a user's trigger. How to generate a digital human.
11. Before the step of obtaining a point cloud of corresponding head key features from a preset feature library based on the head key features in the image, the method includes: further comprising collecting a point cloud of a plurality of head key features and attribute information of each head key feature of each person among the plurality of people, and storing the point cloud in the feature library; A method for generating a digital human according to any one of claims 1 to 5.
12. The step of fusing the head key feature point cloud to the target object model includes: registering the head key feature point cloud with the target object model; and transferring the point cloud of the head key feature to a region of a corresponding head key feature in the target object model. A method for generating a digital human according to any one of claims 1 to 5.
13. A digital human generation platform, a model acquisition module for acquiring a corresponding target object model based on an image of the digital human to be generated; a feature acquisition module for acquiring a point cloud of corresponding head key features from a pre-defined feature library based on the head key features in the image; a fusion module for fusing the point cloud of the head key features with the target object model to obtain a digital human image; The generation platform comprises: a first detection module for detecting whether a similarity between a head key feature in the target object model after fusion and a head key feature in the image is equal to or greater than a predetermined similarity threshold; an avatar creation module for performing an avatar creation operation on the point cloud of the head key features fused to the target object model based on a user trigger in response to the similarity being less than a predetermined similarity threshold. A platform for generating digital humans.
14. The avatar creation module: and performing an avatar creation operation on the point cloud of the head key features fused to the target object model based on a user trigger, using a pre-established implicit constraint surface and a pre-set dynamic curve as constraints. The digital human generation platform of claim 13.
15. The avatar creation module: Obtain operation information of a first controller set on the implicit constraint surface triggered by the user, the implicit constraint surface being a surface that completely matches the surface topological structure of the target object model before being controlled, and the implicit constraint surface is invisible to the user; obtaining motion information of a triggered point on the implicit constraint surface based on a first motion mapping relationship between a first controller on the implicit constraint surface and a point on the implicit constraint surface that has been established in advance and motion information of the first controller; determining motion information of a second controller that is co-located with the point on the target object model based on motion information of the triggered point on the implicit constraint surface; controlling the motion information of the point on the target object model based on a second motion mapping relationship between a second controller set in the target object model in advance and the point on the target object model, and motion information of the second controller; a predetermined dynamic curve is used as a constraint to perform position adjustment on the point cloud of the head key features fused to the target object model; The digital human generation platform of claim 14.
16. The avatar creation module: Detect whether the second controller is on a preset dynamic curve, and set at least two second controllers for each preset dynamic curve; in response to the second controller being on a preset dynamic curve, acquiring operation information of another second controller on the dynamic curve based on operation information of the second controller; and performing position adjustment on the point cloud of the head key features fused to the target object model based on the motion information of the second controller and the motion information of another second controller in the dynamic curve. The digital human generation platform of claim 15.
17. The avatar creation module further includes: and in response to the second controller not being on a preset dynamic curve, performing position adjustment on the point cloud of the head key features fused to the target object model based on motion information of the second controller. The digital human generation platform of claim 16.
18. The generation platform comprises: a second detection module for detecting whether an attached template of the target object model is suitable for the digital human image; an adjustment module for adjusting the attached template in the digital human image in response to the attached template of the target object model being appropriate for the digital human image; A digital human generation platform according to any one of claims 13 to 17.
19. The model acquisition module: Extracting attribute features of the digital human based on the image of the digital human to be generated; and acquiring the corresponding target object model from a pre-defined model library based on the attribute features of the digital human, the model library including a plurality of object models. A digital human generation platform according to any one of claims 13 to 17.
20. The model acquisition module: When the attribute features of the digital human are not extracted based on the image of the digital human to be generated, a preset standard model is used as the target object model. The digital human generation platform of claim 19.
21. The feature acquisition module: obtaining target attribute information of head key features in the image; used to obtain a point cloud of corresponding head key features from the feature library based on target attribute information of the head key features; A digital human generation platform according to any one of claims 13 to 17.
22. The avatar creation module further includes: If the feature library does not include a head key feature corresponding to the target attribute information, a point cloud of the head key feature in the target object model is used to perform an avatar creation operation based on a user's trigger. The digital human generation platform of claim 21.
23. The generation platform comprises: a collection module for collecting point clouds of a plurality of head key features of each person among the plurality of people and attribute information of each head key feature, and storing the point clouds in the feature library; A digital human generation platform according to any one of claims 13 to 17.
24. The fusion module comprises: registering the head key feature point cloud with the target object model; used to transfer the point cloud of the head key features to the region of the corresponding head key features in the target object model. A digital human generation platform according to any one of claims 13 to 17.
25. An electronic device, at least one processor; a memory communicatively coupled to the at least one processor; The memory stores instructions executable by the at least one processor, the instructions, when executed by the at least one processor, causing the at least one processor to perform the method of any one of claims 1 to 5. electronic equipment.
26. A non-transitory computer-readable storage medium having computer instructions stored thereon, comprising: The computer instructions cause a computer to perform the method of any one of claims 1 to 5. A non-transitory computer-readable storage medium.
27. A computer program comprising: The computer program, when executed by a processor, implements the method according to any one of claims 1 to 5. Computer program.
Citation Information
Patent Citations
Image creation system, image creation application server, image creation method and program
JP2014006873A