3D digital human generation method and apparatus, and computing device cluster

By redrawing the target area of ​​the 3D digital human image and combining it with the style information set by the user, the problem of strong CG feel in the existing technology is solved, generating a more realistic 3D digital human and improving the user experience.

CN121170089APending Publication Date: 2025-12-19HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411073770.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-06-19
Filing Date
2024-08-06
Publication Date
2025-12-19

AI Technical Summary

Technical Problem

Existing 3D digital human generation technologies have strong computer graphics but poor user experience.

Method used

By acquiring multiple frames of digital human images corresponding to the 3D digital human, determining each target region, extracting control signals and user-set style information, redrawing each target region, and finally merging the redrawn region with the original 3D digital human.

Benefits of technology

It reduces the CG feel of 3D digital humans, making the generated 3D digital humans more realistic and improving the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121170089A_ABST
    Figure CN121170089A_ABST
Patent Text Reader

Abstract

The invention provides a 3D digital person generation method and device and a computing device cluster, and relates to the technical field of computers.The method comprises the steps that firstly, a first 3D digital person (corresponding to multiple frames of digital person images) is obtained; determining each target area (including at least one of a skin area, a dressing area or a background area) of the first 3D digital person in the image; then extracting control signals (including depth information and edge information) of each target area; and redrawing each target area according to the control signal and style information of each target area set by the user, and fusing each redrawn target area with the first 3D digital human. According to the technical scheme provided by the invention, the CG feeling of the 3D digital human can be reduced, and the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer, in particular to a 3D digital person generation method and device and a computing device cluster. BACKGROUND

[0002] 3D digital person is a new virtual interactive technology, which aims to realize the interaction between natural person and digital person by driving the 3D digital person model through large language model combined with 3D pose estimation. Among them, the generation of 3D digital person is an important part of the technology.

[0003] However, the 3D digital person generated by the current technology has strong computer graphics (CG) and poor user experience. SUMMARY

[0004] Therefore, the present application provides a 3D digital person generation method, device and computing device cluster, which can reduce the CG of 3D digital person and improve user experience.

[0005] In order to achieve the above purpose, in a first aspect, the present application provides a 3D digital person generation method, comprising:

[0006] Obtaining a first 3D digital person, the first 3D digital person corresponding to a plurality of digital person images;

[0007] For any one of the plurality of digital person images, determining the target regions of the first 3D digital person in the image, the target regions including at least one of skin region, clothing region or background region;

[0008] Extracting the control signals of the target regions, the control signals including depth information and edge information;

[0009] According to the control signals and the style information of the target regions set by the user, redrawing the target regions;

[0010] Fusing the redrawn target regions with the first 3D digital person.

[0011] The 3D digital person generation method provided by the present application can reduce the CG of the target regions and make the redrawn target regions more real for any one of the plurality of digital person images corresponding to the first 3D digital person. Then, the target regions are redrawn combined with the style information of the target regions set by the user. Finally, the redrawn target regions are fused with the first 3D digital person to generate a new 3D digital person, so that the finally generated 3D digital person is more real and the user experience is improved.

[0012] In a possible implementation of the first aspect, the style information includes at least one of the following:

[0013] a face style, a hairstyle, a skin texture style, a dressing style, or a background style.

[0014] In a possible implementation of the first aspect, redrawing the target region according to the control signal and the style information of the target region set by the user includes:

[0015] adding noise to the image to obtain a first image;

[0016] adding the control signal and the style information to the first image to redraw the target region, to obtain a second image;

[0017] removing the noise in the second image to obtain the redrawn target region.

[0018] In a possible implementation of the first aspect, the background region is redrawn after the skin region and the dressing region are redrawn.

[0019] In a possible implementation of the first aspect, the control signal further includes normal information and / or optical flow information.

[0020] In a possible implementation of the first aspect, obtaining the first 3D digital human includes:

[0021] generating the first 3D digital human in response to a first operation of inputting character data on a modeling interface.

[0022] In a possible implementation of the first aspect, the character data includes multi-angle character pictures.

[0023] In a possible implementation of the first aspect, the character data further includes at least one of the following data: height, weight, gender, and age.

[0024] In a second aspect, an embodiment of the present application provides a 3D digital human generation device, including an obtaining module, a determining module, an extracting module, a redrawing module, and a fusing module.

[0025] The obtaining module is configured to obtain a first 3D digital human, the first 3D digital human corresponding to a plurality of digital human images.

[0026] The determining module is configured to determine, for any one of the digital human images, each target region of the first 3D digital human in the image, the target region including at least one of a skin region, a dressing region, or a background region.

[0027] The extracting module is configured to extract a control signal of each target region, the control signal including depth information and edge information.

[0028] The redrawing module is configured to redraw each target region according to the control signal and the style information of each target region set by the user;

[0029] The fusion module is configured to fuse the redrawn target regions with the first 3D digital person.

[0030] In a possible implementation of the second aspect, the style information includes at least one of the following:

[0031] a face style, a hairstyle, a skin texture style, a dressing style, or a background style.

[0032] In a possible implementation of the second aspect, the redrawing module is specifically configured to:

[0033] add noise to the image to obtain a first image;

[0034] add the control signal and the style information to the first image to redraw the target region, and obtain a second image;

[0035] remove the noise in the second image to obtain the redrawn target region.

[0036] In a possible implementation of the second aspect, the background region is redrawn after the skin region and the dressing region are redrawn.

[0037] In a possible implementation of the second aspect, the control signal further includes normal information and / or optical flow information.

[0038] In a possible implementation of the second aspect, the obtaining module is specifically configured to:

[0039] generate the first 3D digital person in response to a first operation of inputting the character data on the modeling interface.

[0040] In a possible implementation of the second aspect, the character data includes multi-angle character pictures.

[0041] In a possible implementation of the second aspect, the character data further includes at least one of the following data: height, weight, gender, and age.

[0042] In a third aspect, an embodiment of the present application provides a computing device cluster, including: at least one computing device, each computing device including a processor and a memory;

[0043] The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the method of the first aspect or any implementation of the first aspect.

[0044] Fourthly, embodiments of this application provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described in the first aspect or any embodiment of the first aspect.

[0045] Fifthly, embodiments of this application provide a computer program product that, when run on an electronic device, causes the electronic device to perform the method described in the first aspect or any embodiment of the first aspect.

[0046] Sixthly, embodiments of this application provide a chip system including a processor coupled to a memory. The processor executes a computer program stored in the memory to implement the method described in the first aspect or any embodiment thereof. The chip system may be a single chip or a chip module composed of multiple chips.

[0047] It is understood that the beneficial effects of the second to sixth aspects mentioned above can be found in the relevant descriptions in the first aspect mentioned above, and will not be repeated here. Attached Figure Description

[0048] Figure 1 A flowchart illustrating a 3D digital human generation method provided in an embodiment of this application;

[0049] Figures 2 to 5 A schematic diagram of the modeling interface provided in an embodiment of this application;

[0050] Figure 6 This is a schematic diagram of the structure of the 3D digital human generation device provided in the embodiments of this application;

[0051] Figure 7 This is a schematic diagram of the structure of a computing device cluster provided in an embodiment of this application;

[0052] Figure 8 This is a schematic diagram of another computing device cluster provided in an embodiment of this application. Detailed Implementation

[0053] The embodiments of this application are described below with reference to the accompanying drawings. The terminology used in the implementation section of this application is only for explaining specific embodiments and is not intended to limit the application. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.

[0054] 3D digital humans are an emerging virtual interaction technology that aims to enable interaction between natural and digital humans by combining large language models with 3D pose technology to drive 3D digital human models. The generation of 3D digital humans is a crucial part of this technology.

[0055] One approach to generating 3D digital humans is to first extract a person's image from video footage, then create a 3D model based on that image to obtain a 3D digital human model associated with the person in the image. After generating the 3D digital human model, the model can be driven by user-input commands, including actions and voice commands, to generate a driving video. This driving video is then rendered to obtain the final 3D digital human video.

[0056] However, this method of directly modeling 3D digital humans from 2D samples is limited by the current technology of generating 3D models from 2D. The resulting 3D digital humans have a strong CG feel (unrealistic reflections, textures, and materials unique to 3D modeling), resulting in a poor user experience.

[0057] To this end, this application provides a 3D digital human generation method that can redraw the original 3D digital human based on the style information set by the user, and then merge the redrawing result with the original 3D digital human to obtain a new 3D digital human. This can reduce the CG feel of the final generated 3D digital human, make the final generated 3D digital human more realistic, and improve the user experience.

[0058] The 3D digital human generation method provided in this application embodiment can be applied to electronic devices, including but not limited to personal computers (PCs), smartphones, netbooks, tablets, PDAs, smart screens, etc. For ease of explanation, this application embodiment will use a mobile phone as an example for illustrative purposes.

[0059] Figure 1 This is a flowchart illustrating a 3D digital human generation method provided in an embodiment of this application, as shown below. Figure 1 As shown, the 3D digital human generation method provided in this application embodiment may include the following steps:

[0060] S110. Obtain the first 3D digital human, which corresponds to multiple frames of digital human images.

[0061] In some embodiments, the electronic device may generate a first 3D digital human in response to a first operation of inputting human data on a modeling interface.

[0062] The data on people can include multi-angle photos of people. These photos can be of real people, cartoon / anime characters, or artistic representations. The first step can be for the user to upload multi-angle photos of people. For example, such as... Figure 2As shown in (a) and (b), users can upload front-facing and side-view images of a person in the modeling interface. The electronic device can perform 3D modeling based on the multi-angle images uploaded by the user, generating a first 3D digital human (i.e., the original 3D digital human). The first 3D digital human can correspond to multiple frames of digital human images, and each frame of digital human images can display different poses of the first 3D digital human.

[0063] The naming conventions for user operations in this application are merely examples and should not be construed as limiting the scope of this application. In some embodiments, the same operations may use other names. Similarly, the naming conventions for various functions, interfaces, and interface elements in this application are also merely examples and are not intended to limit this application. In other embodiments, other names may also be used.

[0064] In some embodiments, the person data may also be a video including person images; correspondingly, the first operation may be a user uploading a video. The electronic device can first extract person images from the video, and then perform 3D modeling based on the extracted person images to generate a first 3D digital human. The embodiments of this application will subsequently use the example of person data including multi-angle person images for illustrative purposes.

[0065] In some embodiments, the person data may also include the person's height, weight, gender, age, etc. For example, such as... Figure 3 As shown, users can set the character's height, weight, gender, age, etc. in the modeling interface. Users can set the character's height, weight, gender, and age in the modeling interface first, or they can upload multi-angle images of the character beforehand.

[0066] Electronic devices can create a first 3D digital human by combining multi-angle images uploaded by the user with information such as the user-defined height, weight, gender, and age. The height, weight, gender, and age of the first 3D digital human are matched with the user-defined data.

[0067] The modeling interface shown in this embodiment is merely an example and is not intended to limit this application. In some embodiments, the modeling interface displayed by the electronic device may include more or fewer interface elements than illustrated to achieve more or fewer functions; the position of each interface element may be adjusted as needed; each function may also be implemented using other interface elements, or may also be implemented in other user interfaces, and this embodiment does not particularly limit these aspects.

[0068] In some embodiments, users can also directly upload a 3D digital human generated in other ways as the first 3D digital human in the modeling interface.

[0069] S120. For any frame of digital human image, determine each target region of the first 3D digital human in the image.

[0070] The target region may include one or more of the following: a binary image (mask) region of the 3D digital human skin (i.e., the skin region), a clothing mask region (i.e., the clothing region), and a background mask region (i.e., the background region). The skin region may include the head, face, and skin of the first 3D digital human.

[0071] Specifically, for each frame of the first 3D digital human, the target regions of the first 3D digital human in the image can be extracted by semantic segmentation or depth recognition.

[0072] S130. Extract the control signals of each target region in the image.

[0073] The control signals for the target region can be used to determine the position, depth, and motion direction of the redrawn target region. The control signals may include depth information and edge information. In some embodiments, the control signals may also include normal information and optical flow information.

[0074] Electronic devices can extract control signals for each target region using image-controlled editing generation algorithms, such as Prompt-to-prompt and InstructPix2Pix. For example, an electronic device can extract control signals for the skin of a first 3D digital human from the skin region.

[0075] S140. Obtain style information for each target area set by the user.

[0076] Style information can include at least one of the following: facial style, hairstyle style, skin texture style, clothing style, or background style. Correspondingly, for example... Figure 4 As shown in (a) and (b), the modeling interface may include one or more of the following: face shape options, hairstyle options, skin texture options, clothing options, and background options. The target action may be an action in which the user selects at least one of the following: face shape options, hairstyle options, skin texture options, clothing options, and background options.

[0077] The face shape option allows users to select various face shapes. The hairstyle option allows users to select various hairstyles. The skin texture option allows users to select various skin textures. The clothing option allows users to select various styles and materials of clothing. In some embodiments, users can further select the size of clothing through the clothing option. The background option allows users to select various backgrounds.

[0078] In some embodiments, the face shape option, hairstyle option, skin texture option, clothing option, and background option may each include a preset option. For any of the above options, if the user does not select any of the options, the electronic device can use the style information corresponding to the preset option of that option as the style information set by the user. For example, if the preset option for the skin texture option is skin texture 1, and the user does not select a skin texture, the electronic device can use skin texture 1 as the skin texture set by the user.

[0079] In some embodiments, for any of the above options, if the user does not select any of the options, the electronic device may also use the style information corresponding to the option with the highest matching degree with the multi-angle image uploaded by the user as the style information set by the user. For example, if the user does not select skin texture, and the skin texture option with the highest matching degree with the multi-angle image uploaded by the user is skin texture 2, then the electronic device may use skin texture 2 as the skin texture set by the user.

[0080] In some embodiments, such as Figure 5 As shown in (a) and (b), users can also upload style information such as face shape, hairstyle, skin texture, clothing, and background. The electronic device can use the uploaded style information as the style information set by the user. For example, if a user uploads a face image in the face shape option, the electronic device can use the face shape in the uploaded face image as the face shape set by the user.

[0081] S150. Based on the control signals of each target area and the style information of each target area set by the user, redraw each target area.

[0082] An electronic device can redraw a first 3D digital human based on various style information set by the user, and then merge the redrawing result with the first 3D digital human to generate a new 3D digital human. The redrawing result may include at least one of the following: face redrawing result, hairstyle redrawing result, skin texture redrawing result, clothing redrawing result, and background redrawing result.

[0083] An electronic device can redraw the first 3D digital human based on an image redrawing model and according to style information set by the user. Alternatively, the electronic device can input the first 3D digital human and the user-set style information into a video redrawing model to redraw the first 3D digital human. The embodiments of this application will subsequently use the example of an electronic device redrawing the first 3D digital human using a video redrawing model for illustrative purposes.

[0084] Video redrawing models can include generative models such as diffusion models, autoregressive models, and generative adversarial networks (GANs). This application will subsequently use a diffusion model as an example for illustrative explanation.

[0085] In some embodiments, for any frame image corresponding to the first 3D digital human, the electronic device can add control signals for each target region and style information of each target region set by the user to the frame image, and redraw each target region of the frame image to obtain the redrawing result of each target region of the frame image.

[0086] In other embodiments, the electronic device may also add noise to each frame of the image before redrawing each target area and remove the noise after redrawing to achieve a better redrawing effect.

[0087] Specifically, for each frame of the first 3D digital human, the electronic device can first add noise (such as Gaussian noise, white noise, etc.) to the frame to obtain the first image.

[0088] Next, for any target region of the frame image, the electronic device can add the control signal of the target region and the style information of the target region set by the user to the first image corresponding to the frame image, redraw the target region, and then remove the noise added before redrawing to obtain the second image, which includes the redrawing result of the target region.

[0089] In some embodiments, the electronic device may add control signals and style information of each target region included in the frame image to the first image corresponding to the frame image, so as to redraw each target region simultaneously (that is, a single full image redraw), and then remove the noise added before redrawing after redrawing to obtain the second image.

[0090] For example, image A includes a skin area, a clothing area, or a background area. An electronic device can add control signals for the skin area and encoded user-defined skin texture, hairstyle, and face shape to a first image corresponding to image A; control signals for the clothing area and encoded user-defined clothing; and control signals for the background area and encoded user-defined background. Simultaneously, the skin area, clothing area, or background area is redrawn, and then the noise added before redrawing is removed to obtain a second image.

[0091] In some embodiments, the electronic device may also redraw a portion of the target area each time to reduce the probability of problems such as clothing deformation, pattern loss, and character deformation after redrawing.

[0092] For example, image A includes a skin area, a clothing area, or a background area. An electronic device can first add control signals for the skin area and encoded user-defined skin texture, hairstyle, and face shape to the first image corresponding to image A, as well as control signals for the clothing area and encoded user-defined clothing. The skin and clothing areas are then redrawn. Next, noise added before redrawing is removed to obtain image 1. Then, noise is added to image 1, along with control signals for the background area and encoded user-defined background. The background area is redrawn, and the noise added before redrawing is removed to obtain the second image.

[0093] In some embodiments, the electronic device may also redraw a target area at a time to further reduce the probability of problems such as clothing deformation, pattern loss, and character deformation after redrawing.

[0094] For example, image A includes a skin area, a clothing area, or a background area. The electronic device can first add control signals for the skin area and encoded user-defined skin texture, hairstyle, and face shape to the first image corresponding to image A, then redraw the skin area. After redrawing, the noise added before redrawing is removed, resulting in image 2. Next, noise is added to image 2, along with control signals for the clothing area and encoded user-defined clothing. The clothing area is then redrawn, and the noise added before redrawing is removed, resulting in image 3. Then, noise is added to image 3, along with control signals for the background area and encoded user-defined background. The background area is then redrawn, and the noise added before redrawing is removed, resulting in the second image.

[0095] For example, an electronic device can first add control signals for the skin region and encoded user-defined skin texture, hairstyle, and face shape to the first image corresponding to image A, redraw the skin region, and remove the noise added before redrawing to obtain image 2. Next, control signals for the clothing region and encoded user-defined clothing are added to the first image, and the clothing region is redrawn. After redrawing, the noise added before redrawing is removed to obtain image 4. Then, control signals for the background region and encoded user-defined background are added to the first image, and the background region is redrawn. After redrawing, the noise added before redrawing is removed to obtain image 5. The second image can include images 2, 4, and 5.

[0096] Compared to the first 3D digital human generated by directly modeling from 2D samples, the redrawn skin area features more realistic facial features, hairstyle, and skin texture; the redrawn clothing area features more realistic clothing material; and the redrawn background area also features more realistic background.

[0097] S160, merge the redrawn target areas with the first 3D digital human.

[0098] After redrawing the 3D digital human, the electronic device can merge the redrawn target area with the first 3D digital human to generate a new 3D digital human. That is, the redrawn target area is merged with the digital human image corresponding to the first 3D digital human (e.g., the image in S130).

[0099] For example, image A is any frame image corresponding to the first 3D digital human. Redrawing the skin region of image A yields image 2, redrawing the clothing region of image 2 yields image 3, and redrawing the background region of image 3 yields the second image. Electronic devices can use personalized image algorithms, such as Faceswap or ROOP, to replace the skin region in image A with the skin region in the second image, the clothing region in image A with the clothing region in the second image, and the background region in image A with the background region in the second image, resulting in the fused image A. The frames corresponding to the new 3D digital human are thus the fused images obtained from the frames corresponding to the first 3D digital human.

[0100] For example, image 2 is obtained by redrawing the skin area of ​​image A, image 4 is obtained by redrawing the clothing area of ​​image A, and image 5 is obtained by redrawing the background area of ​​image A. The electronic device can replace the skin area of ​​image A with the skin area of ​​image 2, replace the clothing area of ​​image A with the clothing area of ​​image 4, and replace the background area of ​​image A with the background area of ​​image 5 to obtain the merged image A.

[0101] In some embodiments, for any frame image corresponding to the first 3D digital human, the electronic device may first fuse the redrawing results of each target region of the first 3D digital human in the frame image with the frame image. Then, the redrawing results of the background region in the frame image are fused again with the previous fusion results to obtain the fused image corresponding to the frame image.

[0102] For example, image 2 is obtained by redrawing the skin area of ​​image A, image 4 is obtained by redrawing the clothing area of ​​image A, and image 5 is obtained by redrawing the background area of ​​image A. An electronic device can then replace the skin area in image A with the skin area from image 2, and replace the clothing area in image A with the clothing area from image 4, to obtain image 6. Finally, the background area in image 6 is replaced with the background area from image 5 to obtain the merged image A.

[0103] Replacing the original skin area in an image with the redrawn skin area makes the hairstyle, face shape, and skin texture of the new 3D digital human more realistic; replacing the original clothing area in an image with the redrawn clothing area makes the clothing material of the new 3D digital human more realistic; replacing the original background area in an image with the redrawn background area makes the background corresponding to the new 3D digital human more realistic.

[0104] The 3D digital human generation method provided in this application first determines each target region of the first 3D digital human in any frame of a multi-frame digital human image corresponding to the first 3D digital human. Then, it extracts the control signals of each target region and, combined with the style information of each target region set by the user, redraws each target region. This reduces the CG-like appearance of each target region, making the redrawn target regions more realistic. Finally, it merges the redrawn target regions with the first 3D digital human to generate a new 3D digital human, thereby making the final generated 3D digital human more realistic and improving the user experience.

[0105] Those skilled in the art will understand that the above embodiments are exemplary and not intended to limit this application. Where possible, one or more of the above steps can be selectively combined to obtain one or more other embodiments. For example, in some embodiments, step S140 can be performed before step S120 or step S130. Those skilled in the art can arbitrarily select and combine the above steps as needed, and all combinations that do not depart from the essence of this application fall within the protection scope of this application.

[0106] Based on the same concept, as an implementation of the above method, this application provides a 3D digital human generation device. This device embodiment corresponds to the aforementioned method embodiment. For ease of reading, this device embodiment will not repeat the details of the aforementioned method embodiment one by one, but it should be clear that the device in this embodiment can implement all the contents of the aforementioned method embodiment.

[0107] Figure 6 This is a schematic diagram of the structure of the 3D digital human generation device provided in the embodiments of this application, as shown below. Figure 6 As shown, the 3D digital human generation device provided in this embodiment may include an acquisition module, a determination module, an extraction module, a redrawing module, and a fusion module.

[0108] The acquisition module is used to: acquire the first 3D digital human, which corresponds to multiple frames of digital human images;

[0109] The determination module is used to: for any frame of digital human image, determine each target region of the first 3D digital human in the image, the target region including at least one of skin region, clothing region or background region;

[0110] The extraction module is used to extract control signals for each target region, including depth information and edge information.

[0111] The redraw module is used to redraw each target area based on the control signals and the style information of each target area set by the user.

[0112] The fusion module is used to fuse the redrawn target areas with the first 3D digital human.

[0113] As an optional implementation, the style information includes at least one of the following:

[0114] Facial style, hairstyle style, skin texture style, clothing style, or background style.

[0115] As an optional implementation, the redraw module is specifically used for:

[0116] Noise is added to the image to obtain the first image;

[0117] Add control signals and style information to the first image, redraw the target area, and obtain the second image;

[0118] Noise is removed from the second image to obtain the redrawn target area.

[0119] As an alternative implementation, the background area is redrawn after the skin area and clothing area have been redrawn.

[0120] As an optional implementation, the control signal may also include normal information and / or optical flow information.

[0121] As an optional implementation, the acquisition module is specifically used for:

[0122] In response to the first operation of inputting character data on the modeling interface, the first 3D digital human is generated.

[0123] As an optional implementation, the people data includes multi-angle images of people.

[0124] As an optional implementation, the person data also includes at least one of the following: height, weight, gender, and age.

[0125] The acquisition module, determination module, extraction module, redrawing module, and fusion module can all be implemented in software or hardware. For example, the implementation of the acquisition module will be described below. Similarly, the implementation methods of the determination module, extraction module, redrawing module, and fusion module can refer to the implementation method of the acquisition module.

[0126] As an example of a software functional unit, a module can include code running on a computing instance. A computing instance can include at least one of a physical host (computing device), a virtual machine, or a container. Furthermore, the aforementioned computing instance can be one or more. For example, a module can include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code can be distributed within the same region or in different regions. Further, the multiple hosts / virtual machines / containers used to run the code can be distributed within the same availability zone (AZ) or in different AZs, each AZ comprising one or more geographically proximate data centers. Typically, a region can include multiple AZs.

[0127] Similarly, multiple hosts / virtual machines / containers used to run this code can be distributed within the same Virtual Private Cloud (VPC) or across multiple VPCs. Typically, a VPC is set up within a region. Communication between two VPCs within the same region, as well as between VPCs in different regions, requires a communication gateway to be set up within each VPC to enable interconnection between VPCs.

[0128] As an example of a hardware functional unit, an acquisition module may include at least one computing device, such as a server. Alternatively, an acquisition module may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The aforementioned PLD may be implemented using a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), generic array logic (GAL), or any combination thereof.

[0129] The acquisition module includes multiple computing devices that can be distributed within the same region or in different regions. Similarly, the acquisition module can be distributed within the same Availability Zone (AZ) or in different AZs. Likewise, the acquisition module can be distributed within the same Virtual Private Cloud (VPC) or multiple VPCs. These multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.

[0130] It should be noted that, in other embodiments, the acquisition module, determination module, extraction module, redrawing module, and fusion module can all be used to execute any step in the above method embodiments. The steps implemented by the acquisition module, determination module, extraction module, redrawing module, and fusion module can be specified as needed. By implementing different steps in the above method embodiments through the acquisition module, determination module, extraction module, redrawing module, and fusion module, all functions of the 3D digital human generation device can be realized.

[0131] Based on the same concept, embodiments of this application also provide a computing device cluster. This computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.

[0132] like Figure 7 As shown, the computing device cluster includes at least one computing device 100. The memory 106 of one or more computing devices 100 in the computing device cluster may store the same instructions for executing the above-described method embodiments.

[0133] In some possible implementations, the memory 106 of one or more computing devices 100 in the computing device cluster may also store partial instructions for executing the above-described method embodiments. In other words, a combination of one or more computing devices 100 can jointly execute the instructions for executing the above-described method embodiments.

[0134] It should be noted that the memory 106 in different computing devices 100 within the computing device cluster can store different instructions, each used to execute a portion of the functions of the 3D digital human generation device. That is, the instructions stored in the memory 106 of different computing devices 100 can implement the functions of one or more modules among the acquisition module, determination module, extraction module, redrawing module, and fusion module.

[0135] In some possible implementations, one or more computing devices in a computing device cluster can be connected via a network. This network can be a wide area network (WAN) or a local area network (LAN), etc. Figure 8 One possible implementation is shown. For example... Figure 8 As shown, the two computing devices 100A and 100B are connected via a network. Specifically, they are connected to the network through communication interfaces in each computing device. In this possible implementation, the memory 106 in computing device 100A stores instructions for executing the functions of the acquisition module. Meanwhile, the memory 106 in computing device 100B stores instructions for executing the functions of the determination module, extraction module, redrawing module, and fusion module.

[0136] It should be understood that Figure 8 The functions of the computing device 100A shown can also be performed by multiple computing devices 100. Similarly, the functions of the computing device 100B can also be performed by multiple computing devices 100.

[0137] This application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the methods described in the above-described method embodiments.

[0138] This application also provides a computer program product that, when run on an electronic device, causes the electronic device to implement the method described in the above-described method embodiments.

[0139] This application also provides a chip system including a processor coupled to a memory. The processor executes a computer program stored in the memory to implement the method described in the above-described method embodiments. The chip system may be a single chip or a chip module composed of multiple chips.

[0140] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted through the computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, or magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)).

[0141] Those skilled in the art will understand that implementing all or part of the processes in the above embodiments can be accomplished by a computer program instructing related hardware. This program can be stored in a computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The aforementioned storage medium can include various media capable of storing program code, such as ROM or random access memory (RAM), magnetic disks, or optical disks.

[0142] The naming or numbering of steps in this application does not mean that the steps in the method flow must be executed in the time / logical order indicated by the naming or numbering. The execution order of the named or numbered process steps can be changed according to the technical purpose to be achieved, as long as the same or similar technical effect can be achieved.

[0143] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0144] In the embodiments provided in this application, it should be understood that the disclosed apparatus / devices and methods can be implemented in other ways. For example, the apparatus / device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0145] It should be understood that in the description of this application and the appended claims, the terms "comprising," "including," "having," and any variations thereof are intended to cover a non-exclusive inclusion and mean "including but not limited to," unless otherwise specifically emphasized. For example, a process, method, system, product, or apparatus that includes a series of steps or modules is not necessarily limited to those steps or modules that are explicitly listed, but may include other steps or modules that are not explicitly listed or that are inherent to such process, method, product, or apparatus.

[0146] Furthermore, in the description of this application, unless otherwise stated, "multiple" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.

[0147] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."

[0148] Furthermore, in the description of this application and the appended claims, the terms "first," "second," etc., are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence, nor should they be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in a sequence other than that illustrated or described herein; features defined as "first" or "second" may explicitly or implicitly include at least one of those features.

[0149] In the embodiments described in this application specification, the words "exemplarily" or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as "exemplarily" or "for example" in the embodiments of this application specification should not be construed as being more preferred or advantageous than other embodiments or design solutions. Specifically, the use of the words "exemplarily" or "for example" is intended to present the relevant concepts in a specific manner.

[0150] References to "one embodiment" or "some embodiments" in this specification mean that one or more embodiments of this specification include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in still other embodiments" appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized.

[0151] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

Claims

1. A method for generating 3D digital humans, characterized in that, include: Acquire a first 3D digital human, which corresponds to multiple frames of digital human images; For any one of the multiple frames of digital human images, determine each target region of the first 3D digital human in the image, the target region including at least one of the skin region, clothing region or background region; The control signals of each target region are extracted respectively, and the control signals include depth information and edge information; Based on the control signals and the style information of each target area set by the user, redraw each target area; The redrawn target regions are then fused with the image.

2. The method according to claim 1, characterized in that, The style information includes at least one of the following: Facial style, hairstyle style, skin texture style, clothing style, or background style.

3. The method according to claim 1 or 2, characterized in that, The step of redrawing the target area based on the control signal and the style information of the target area set by the user includes: Noise is added to the image to obtain a first image; The control signal and style information are added to the first image, and the target area is redrawn to obtain the second image; Noise is removed from the second image to obtain the redrawn target region.

4. The method according to any one of claims 1-3, characterized in that, The background area is redrawn after the skin area and the clothing area have been redrawn.

5. The method according to any one of claims 1-4, characterized in that, The control signal also includes normal information and / or optical flow information.

6. The method according to any one of claims 1-5, characterized in that, The process of obtaining the first 3D digital human includes: In response to the first operation of inputting character data on the modeling interface, the first 3D digital human is generated.

7. The method according to claim 6, characterized in that, The data on individuals includes photos of individuals from multiple angles.

8. The method according to claim 7, characterized in that, The personal data also includes at least one of the following: height, weight, gender, and age.

9. A 3D digital human generation device, characterized in that, include: The modules include: Acquisition Module, Determination Module, Extraction Module, Redraw Module, and Fusion Module. The acquisition module is used to: acquire a first 3D digital human, wherein the first 3D digital human corresponds to multiple frames of digital human images; The determining module is used to: for any one frame of the multi-frame digital human image, determine each target region of the first 3D digital human in the image, wherein the target region includes at least one of the skin region, clothing region or background region; The extraction module is used to: extract control signals for each target region, wherein the control signals include depth information and edge information; The redrawing module is used to: redraw each target region according to the control signal and the style information of each target region set by the user; The fusion module is used to fuse the redrawn target regions with the image.

10. A computing device cluster, characterized in that, It includes at least one computing device, each computing device including a processor and memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device to cause the cluster of computing devices to perform the method as described in any one of claims 1-8.

11. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-8.

12. A computer program product, characterized in that, When the computer program product is run on an electronic device, it causes the electronic device to perform the method as described in any one of claims 1-8.

13. A chip system, characterized in that, The chip system includes a processor coupled to a memory, the processor executing a computer program stored in the memory to implement the method as described in any one of claims 1-8.