Method, apparatus, storage medium and tutoring machine for generating three-dimensional virtual characters

By extracting the key point information and clothing parameters in the user's whole body image and adjusting it correspondingly with the three-dimensional virtual character model, the problem of generating three-dimensional virtual characters in the prior art requires multiple face photos and is limited to face similarity, and the effect of generating three-dimensional virtual characters is achieved.

CN114519773BActive Publication Date: 2025-06-27BEIJING ZHIYUAN HANGCHENG SOFTWARE TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202210101429.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-27
Publication Date
2025-06-27
Estimated Expiration
2042-01-27

AI Technical Summary

Technical Problem

The prior art requires multiple two-dimensional face photos to generate three-dimensional virtual characters, and can only make the face similar to the user, and cannot effectively generate three-dimensional virtual characters with the same whole body.

Method used

By obtaining the user's full body image, the image feature extraction model is used to extract key point information and clothing parameters, and correspond to the three-dimensional virtual character model, adjust the model position to generate candidate three-dimensional virtual characters, and finally generate the target three-dimensional virtual characters based on the clothing parameters.

Benefits of technology

A two-dimensional full-body image is used to generate three-dimensional virtual characters that are highly similar to the user's image, including faces and body parts, and their clothes are also similar to the images entered by the user.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114519773B_ABST
    Figure CN114519773B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a method, apparatus, storage medium, and tutoring machine for generating a three-dimensional virtual character, so that a three-dimensional virtual character can be generated by using a two-dimensional full-body image, and the generated three-dimensional virtual character includes a face part and a body part. The method includes: obtaining a full-body image input by a user; obtaining key point information corresponding to key points in the full-body image and clothing parameters corresponding to the user in the full-body image through an image feature extraction model, where the key point information includes the serial number of the key point and the three-dimensional coordinates of the key point; corresponding the key points of the full-body image to preset key points with the same serial number in the three-dimensional virtual character model according to the serial number of the key point, and adjusting the positions of the preset key points in the three-dimensional virtual character model according to the three-dimensional coordinates of the key point to obtain a candidate three-dimensional virtual character; generating a target three-dimensional virtual character according to the clothing parameters corresponding to the user in the full-body image and the candidate three-dimensional virtual character.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer image processing, and in particular, to a method, apparatus, storage medium, and tutoring machine for generating a three-dimensional virtual character. Background Art

[0002] With the development of technology, people have higher and higher requirements for various image resources. In order to create an immersive experience for users, more and more image resources use three-dimensional display technology. Therefore, the three-dimensional display technology has achieved unprecedented development and application. When users use applications such as games, teaching, and singing through electronic products, they hope that a three-dimensional virtual character highly similar to their own image can be displayed on the application interface to obtain an immersive usage experience.

[0003] In related technologies, it is usually necessary to take multiple two-dimensional face photos from multiple angles to generate a three-dimensional virtual face, and then fuse the three-dimensional virtual face with a pre-set three-dimensional virtual character model in the system to obtain a full-body three-dimensional virtual character. Finally, only the face of the obtained three-dimensional virtual character is similar to the user. Summary of the Invention

[0004] The purpose of the present disclosure is to provide a method, apparatus, storage medium, and tutoring machine for generating a three-dimensional virtual character, so as to solve the technical problem that generating a three-dimensional virtual character in related technologies requires multiple two-dimensional face photos and only the face is similar to the user.

[0005] To achieve the above purpose, a first aspect of the present disclosure provides a method for generating a three-dimensional virtual character, the method comprising:

[0006] Obtain a full-body image input by a user;

[0007] Obtain key point information corresponding to key points in the full-body image and clothing parameters corresponding to the user in the full-body image through an image feature extraction model, where the key point information includes the serial number of the key point and the three-dimensional coordinates of the key point;

[0008] Correspond the key points of the full-body image with the pre-set key points having the same serial number in the three-dimensional virtual character model according to the serial number of the key point, and adjust the positions of the pre-set key points in the three-dimensional virtual character model according to the three-dimensional coordinates of the key point to obtain a candidate three-dimensional virtual character;

[0009] Generate a target three-dimensional virtual character according to the clothing parameters corresponding to the user in the full-body image and the candidate three-dimensional virtual character.

[0010] Optionally, the image feature extraction model includes an image preprocessing network and a feature extraction network. The process of obtaining the key point information corresponding to the key points in the full-body image and the clothing parameters corresponding to the user in the full-body image through the image feature extraction model includes:

[0011] Input the full-body image into the image feature extraction model, and perform image preprocessing on the full-body image through the image preprocessing network in the image feature extraction model based on at least one preset image processing strategy to obtain a target image;

[0012] Perform feature extraction on the target image through the feature extraction network in the image feature model to obtain the key point information corresponding to the key points in the full-body image and the clothing parameters corresponding to the user in the full-body image.

[0013] Optionally, the process of performing image preprocessing on the full-body image through the image preprocessing network in the image feature extraction model based on at least one preset image processing strategy to obtain a target image includes:

[0014] Perform image preprocessing on the full-body image through the image preprocessing network in the image feature extraction model based on at least one of the following preset image processing strategies to obtain multiple target sub-images: RGB channel separation processing, HSI channel separation processing, YCrCb channel separation processing, grayscale processing, histogram equalization processing, and image sharpening processing;

[0015] Overlay the multiple target sub-images to obtain the target image.

[0016] Optionally, the training process of the feature extraction network includes:

[0017] Obtain a sample image labeled with sample key point information and sample clothing parameters;

[0018] Input the sample image into the image feature extraction network to obtain the predicted key point information and predicted clothing parameters of the sample image;

[0019] Calculate a loss function based on the sample key point information and the predicted key point information in the sample image, and the sample clothing parameters and the predicted clothing parameters, and adjust the parameters of the feature extraction network according to the calculation result of the loss function.

[0020] Optionally, the process of adjusting the position of the preset key points in the three-dimensional virtual character model according to the three-dimensional coordinates of the key points includes:

[0021] Execute a global adjustment strategy and a local adjustment strategy according to the three-dimensional coordinates of the key points;

[0022] Among them, the global adjustment strategy is used to adjust the positions of the preset key points corresponding to the key points in the three-dimensional virtual character model according to the three-dimensional coordinates of each key point;

[0023] The local adjustment strategy is used to adjust the positions of the preset key points corresponding to the target key points in the three-dimensional virtual character model according to the three-dimensional coordinates of the target key points in the key points after the global adjustment strategy is executed, where the target key point is a key point that can be used to uniquely determine the user's face contour.

[0024] Optionally, obtaining the key point information corresponding to the key points in the full-body image through the image feature extraction model includes:

[0025] Obtaining all key points, the serial numbers of all key points, the three-dimensional coordinates of all key points, and the first normal vector of the target key point in the full-body image through the image feature extraction model;

[0026] Adjusting the positions of the preset key points corresponding to the target key points in the three-dimensional virtual character model according to the three-dimensional coordinates of the target key points in the key points includes:

[0027] Determining the second normal vector of the target key point according to the three-dimensional coordinates of the target key point and the three-dimensional coordinates of a preset number of key points closest to the target key point among all key points;

[0028] Taking the average of the first normal vector and the second normal vector of the target key point to obtain the target normal vector of the target key point, and adjusting the position of the preset key point corresponding to the target key point in the three-dimensional virtual character model according to the target normal vector.

[0029] Optionally, obtaining the clothing parameters corresponding to the user in the full-body image through the image feature extraction model includes:

[0030] Obtaining the clothing style parameters and clothing color parameters corresponding to the user in the full-body image through the image feature extraction model;

[0031] Generating a target three-dimensional virtual character according to the clothing parameters corresponding to the user in the full-body image and the candidate three-dimensional virtual character includes:

[0032] Generating corresponding clothing for the candidate three-dimensional virtual character according to the clothing style parameters and clothing color parameters corresponding to the user in the full-body image to obtain a target three-dimensional virtual character.

[0033] Optionally, obtaining the full-body image input by the user includes:

[0034] In response to a user's three-dimensional virtual character creation operation on a tutoring machine, obtain a full-body image input by the user;

[0035] After generating the target three-dimensional virtual character, it further includes:

[0036] Display the target three-dimensional virtual character and target learning content on the tutoring machine;

[0037] In response to collecting a learning action of the user for the target learning content, drive the target three-dimensional virtual character to perform corresponding expressions and / or actions based on the learning action.

[0038] The second aspect of the present disclosure further provides a device for generating a three-dimensional virtual character, and the device includes:

[0039] An acquisition module, configured to acquire a full-body image input by a user;

[0040] An extraction module, configured to obtain key point information corresponding to key points in the full-body image and clothing parameters corresponding to the user in the full-body image through an image feature extraction model, where the key point information includes the serial number of the key point and the three-dimensional coordinates of the key point;

[0041] An adjustment module, configured to correspond the key points of the full-body image with preset key points having the same serial number in the three-dimensional virtual character model one by one according to the serial number of the key point, and adjust the positions of the preset key points in the three-dimensional virtual character model according to the three-dimensional coordinates of the key point to obtain a candidate three-dimensional virtual character;

[0042] A generation module, configured to generate a target three-dimensional virtual character according to the clothing parameters corresponding to the user in the full-body image and the candidate three-dimensional virtual character.

[0043] The third aspect of the present disclosure further provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the method described in any one of the first aspects above are implemented.

[0044] The fourth aspect of the present disclosure further provides a tutoring machine, including:

[0045] A memory, on which a computer program is stored;

[0046] A processor, configured to execute the computer program in the memory to implement the steps of the method described in any one of the first aspects above.

[0047] Through the above technical solutions, at least the following technical effects can be achieved:

[0048] A full-body image input by a user is obtained, and key point information corresponding to key points in the full-body image and clothing parameters corresponding to the user in the full-body image are obtained through an image feature extraction model. The key point information includes the key point serial number and the three-dimensional coordinates of the key point. Then, according to the key point serial number, the key points of the full-body image are matched one by one with the preset key points with the same serial number in the three-dimensional virtual character model, and the positions of the preset key points in the three-dimensional virtual character model are adjusted according to the three-dimensional coordinates of the key points to obtain a candidate three-dimensional virtual character. Finally, according to the clothing parameters corresponding to the user in the full-body image and the candidate three-dimensional virtual character, a target three-dimensional virtual character is generated. Through this method, a three-dimensional virtual character can be generated using a two-dimensional full-body image, and the three-dimensional virtual character generated based on the full-body image includes not only a face part but also a body part. In addition, the user's clothing parameters are obtained, thereby establishing a three-dimensional virtual character that is highly similar to the user's own image.

[0049] Other features and advantages of the present disclosure will be described in detail in the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] The accompanying drawings are used to provide a further understanding of the present disclosure and constitute a part of the specification. Together with the following specific embodiments, they are used to explain the present disclosure but do not constitute a limitation of the present disclosure. In the accompanying drawings:

[0051] Figure 1 It is a flowchart of a method for generating a three-dimensional virtual character provided by an embodiment of the present disclosure;

[0052] Figure 2 is a schematic diagram of a three-dimensional virtual character model provided by an embodiment of the present disclosure;

[0053] Figure 3 is a block diagram of a device for generating a three-dimensional virtual character provided by an embodiment of the present disclosure;

[0054] Figure 4 It is a block diagram of a tutoring machine provided in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0055] The specific implementation of the present disclosure is described in detail below in conjunction with the accompanying drawings. It should be understood that the specific implementation described herein is only used to illustrate and explain the present disclosure, and is not used to limit the present disclosure.

[0056] It should be understood that the various steps described in the method embodiments of the present disclosure may be executed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard. The term "including" and its variations used herein are open-ended, that is, "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.

[0057] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units. In addition, the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that unless clearly specified otherwise in the context, it should be understood as "one or more".

[0058] Currently, in order to enable users to establish a three-dimensional virtual character highly similar to their own image and display it on the application interface when using applications such as games, teaching, and singing through electronic products, so as to obtain an immersive usage experience. However, in the related art, it is usually necessary to take multiple two-dimensional face photos from multiple angles to generate a three-dimensional virtual face, and then fuse the three-dimensional virtual face with a pre-set three-dimensional virtual character model in the system to obtain a full-body three-dimensional virtual character. Finally, the obtained three-dimensional virtual character is only similar to the user in terms of the face.

[0059] In view of this, the present disclosure provides a method, device, storage medium and tutoring machine for generating a three-dimensional virtual character to solve the above problems.

[0060] Before giving a detailed embodiment description of the technical solution of the present disclosure, the application scenarios of the technical solution of the present disclosure will be described first.

[0061] A tutoring machine refers to an electronic teaching product that stores rich learning materials and learning methods and can assist children's learning. It usually includes functions such as teaching courses, homework tutoring, and learning games. The method for generating a three-dimensional virtual character provided by the present disclosure can generate a three-dimensional virtual character highly similar to the user's own image by obtaining a full-body photo of the user and display it on the application interface of the tutoring machine, and can also control the three-dimensional virtual character to perform actions, such as raising hands, speaking, walking, etc. in a learning game, so that the user obtains an immersive usage experience.

[0062] The following gives a detailed embodiment description of the technical solution of the present disclosure.

[0063] Referring to Figure 1 , an embodiment of the present disclosure provides a method for generating a three-dimensional virtual character, and the method includes:

[0064] S101. Obtain a full-body image input by a user.

[0065] S102. Obtain key-point information corresponding to key points in the full-body image and clothing parameters corresponding to the user in the full-body image through an image feature extraction model, where the key-point information includes the serial numbers of the key points and the three-dimensional coordinates of the key points.

[0066] S103. Correspond the key points of the full-body image with preset key points having the same serial numbers in the three-dimensional virtual character model according to the serial numbers of the key points, and adjust the positions of the preset key points in the three-dimensional virtual character model according to the three-dimensional coordinates of the key points to obtain a candidate three-dimensional virtual character.

[0067] S104. Generate a target three-dimensional virtual character according to the clothing parameters corresponding to the user in the full-body image and the candidate three-dimensional virtual character.

[0068] By using the above method, key-point information corresponding to key points in the full-body image of the user and clothing parameters corresponding to the user in the full-body image are obtained through an image feature extraction model, and a target three-dimensional virtual character is generated. Through this method, a three-dimensional virtual character can be generated by using a single two-dimensional full-body image, and the three-dimensional virtual character generated based on the full-body image includes not only the face part but also the body part, and the clothing of the three-dimensional virtual character is also similar to the full-body image input by the user, so as to establish a three-dimensional virtual character highly similar to the user's own image.

[0069] In order to enable those skilled in the art to better understand the method for generating a three-dimensional virtual character provided by the present disclosure, the above steps will be described in detail with examples below.

[0070] In a possible manner, the image feature extraction model includes an image preprocessing network and a feature extraction network. Obtaining key-point information corresponding to key points in the full-body image and clothing parameters corresponding to the user in the full-body image through the image feature extraction model includes the following steps: First, input the full-body image into the image feature extraction model, and perform image preprocessing on the full-body image through the image preprocessing network in the image feature extraction model based on at least one preset image processing strategy to obtain a target image. Then, perform feature extraction on the target image through the feature extraction network in the image feature model to obtain key-point information corresponding to key points in the full-body image and clothing parameters corresponding to the user in the full-body image.

[0071] Among them, the image preprocessing of the full-body image by the image preprocessing network in the image feature extraction model based on at least one preset image processing strategy to obtain the target image includes: the image preprocessing network in the image feature extraction model performs image preprocessing on the full-body image based on at least one preset image processing strategy to obtain a plurality of target sub-images, and then superimposes the plurality of target sub-images to obtain the target image. Among them, the preset image processing strategies include at least one of RGB channel separation processing, HSI channel separation processing, YCrCb channel separation processing, grayscale processing, histogram equalization processing, and image sharpening processing.

[0072] Exemplarily, the image input by the user is usually a color image with RGB three channels. Through RGB channel separation processing, the R (red) channel sub-image, G (green) channel sub-image, and B (blue) channel sub-image can be obtained respectively. Correspondingly, through HSI channel separation processing, the H (hue) channel sub-image, S (saturation) channel sub-image, and I (intensity) channel sub-image can be obtained respectively. Through YCrCb channel separation processing, the Y (luminance) channel sub-image, Cr (the difference between the red part of the RGB image and the luminance value) channel sub-image, and Cb (the difference between the blue part of the RGB image and the luminance value) channel sub-image can be obtained respectively. Also, through grayscale processing, a grayscale sub-image is obtained, through histogram equalization processing, an equalized sub-image with enhanced contrast effect is obtained, and through image sharpening processing, a sharpened sub-image with enhanced image edge contrast effect is obtained. For details, reference can be made to the related art, and the present disclosure will not elaborate herein.

[0073] Further, after performing image preprocessing on the full-body image, a plurality of target sub-images on different channels are obtained, and then the plurality of target sub-images are superimposed in the channel dimension to obtain the final target image.

[0074] It should be noted that one preset image processing strategy can be used to perform image preprocessing on the full-body image, or multiple preset image processing strategies can be used to perform image preprocessing on the full-body image. Preferably, multiple preset image processing strategies are used. Performing image preprocessing on the full-body image through multiple preset image processing strategies can obtain more image information dimensions, which is beneficial for the subsequent feature extraction network to extract more effective features, and thus significantly increases the accuracy of the key points output by the image feature extraction model.

[0075] After obtaining the target image, the target image can be input into the feature extraction network in the image feature model for feature extraction. Furthermore, after obtaining the key point information corresponding to the key points in the full-body image, the position of the preset key points in the three-dimensional virtual character model can be adjusted through the key point information to obtain a candidate three-dimensional virtual character.

[0076] Among possible ways, the position of a preset key point in a three-dimensional virtual character model can be adjusted according to the three-dimensional coordinates of a key point in the following manner: a global adjustment strategy and a local adjustment strategy are executed based on the three-dimensional coordinates of the key point, where the global adjustment strategy is used to adjust the position of the preset key point corresponding to the key point in the three-dimensional virtual character model according to the three-dimensional coordinates of each key point, and the local adjustment strategy is used to adjust the position of the preset key point corresponding to the target key point in the three-dimensional virtual character model according to the three-dimensional coordinates of the target key point in the key points after the global adjustment strategy is executed, and the target key point is a key point that can be used to uniquely determine the user's face contour.

[0077] Optionally, first, all key points, the serial numbers of all key points, the three-dimensional coordinates of all key points, and the first normal vector of the target key point are obtained through an image feature extraction model. Then, the position of the preset key point corresponding to the target key point in the three-dimensional virtual character model is adjusted according to the three-dimensional coordinates of the target key point in the key points in the following manner: First, according to the three-dimensional coordinates of the target key point and the three-dimensional coordinates of the preset number of key points closest to the target key point among all key points, the second normal vector of the target key point is determined; the mean value of the first normal vector and the second normal vector of the target key point is obtained to get the target normal vector of the target key point, and the position of the preset key point corresponding to the target key point in the three-dimensional virtual character model is adjusted according to the target normal vector.

[0078] Exemplarily, a three-dimensional coordinate system can be pre-constructed to determine the three-dimensional coordinate information of all key points. Among them, the origin of the three-dimensional coordinate system can be the top of the head, the face, the waist, the feet, etc., and the present disclosure does not make specific limitations on this. Refer to Figure 2 , the three-dimensional virtual character model has predefined key points and target key points. Each key point has corresponding serial number information and three-dimensional coordinate information. The target key point also includes normal vector information. Among them, the key points cover the whole body of the three-dimensional virtual character model, and the general contour of the three-dimensional virtual character model can be restored through the key points. The target key point is a key point that can uniquely determine the user's face contour, such as parts of the nose, eyes, mouth, eyebrows, cheekbones, etc. The differences in the height and size of these parts have a greater impact on the similarity between the three-dimensional virtual character and the real person. Taking the eye part as an example, target key points can be defined at the positions of the inner canthus, outer canthus, eyeball, upper eyelid, and lower eyelid. The width of the eye can be determined by the target key points of the inner canthus and the outer canthus, the height of the eye can be determined by the target key points of the upper eyelid and the lower eyelid, and the size and position of the eye can be determined by adding the target key point of the eyeball.

[0079] It should be noted that since the face can determine the similarity between the generated 3D virtual character and the real person more than the body part, the number of key points defined for the face is more than that for the body part. In addition, the number of key points depends on different requirements for the similarity between the 3D virtual character and the real person. Generally speaking, the more the number of key points, the higher the similarity between the 3D virtual character and the real person. In addition, in addition to setting target key points on the face, some target key points can also be appropriately set on the body part to improve the similarity between the body part of the 3D virtual character and the body part of the real person. Figure 2 The key points and target key points in Figure 2 are only exemplary descriptions of the embodiments of the present disclosure. The number and position of the key points and target key points can be adjusted according to requirements, and the present disclosure does not make specific limitations thereon.

[0080] All key points in the full-body image are obtained through the image feature extraction model. Each key point has corresponding serial number information and three-dimensional coordinate information. The target key point also includes a first normal vector. The key points of the full-body image are corresponded to the preset key points of the 3D virtual character model one by one according to the serial number information, and then the positions of the preset key points corresponding to the 3D virtual character model are adjusted based on the three-dimensional coordinates of each key point in the full-body image. Thus, a 3D virtual character model that is roughly similar to the real person in the full-body image can be obtained. Further, in order to make the 3D virtual character model closer to the real person in the full-body image, the target key points that determine the face contour can be adjusted.

[0081] Among them, when the image feature extraction model outputs the target key point, it can be represented by Euler angles, and then the first normal vector of the target key point can be obtained through the Euler angles. However, since the first normal vector of the target key point is the normal vector information recognized by the image feature extraction model, there is a certain error from the target key point of the actual real person image. In order to improve the similarity between the 3D virtual character and the real person, the normal vector of the target key point can be corrected, and then the position of the target key point can be adjusted to make the 3D virtual character closer to the real person.

[0082] Exemplarily, first, according to the three-dimensional coordinates of the target key point and the three-dimensional coordinates of a preset number of key points closest to the target key point among all key points, for example, it can be three key points closest to the target key point, the second normal vector of the target key point is determined by using the method of finding the normal vector of a plane from three points. The method of finding the normal vector of a plane from three points can refer to related technologies, which will not be elaborated in this disclosure. Further, the average value of the first normal vector and the second normal vector of the target key point is calculated to obtain the target normal vector of the target key point, and the position of the preset key point corresponding to the target key point in the three-dimensional virtual character model is adjusted according to the target normal vector. By taking the average value of the second normal vector of the target key point estimated from the surrounding key points and the first normal vector of the target key point recognized by the image feature extraction model, the target normal vector of the target key point is obtained, thereby correcting the normal vector of the target key point, and then adjusting the position of the target key point, making the three-dimensional virtual character closer to a real person.

[0083] It should be noted that the number of key points that the image feature extraction model can recognize is the same as the number of preset key points defined in the three-dimensional virtual character model. However, due to factors such as partial occlusion of the person in the input full-body image or the shooting angle, the image feature extraction model may not be able to recognize all key points. In this case, the three-dimensional coordinates of the surrounding key points and the coordinates of the corresponding preset key points in the three-dimensional virtual character model can be obtained, and by setting the weights of the surrounding key points and the preset key points, the three-dimensional coordinates of the unrecognized key points are calculated, and then the three-dimensional virtual character model is adjusted to obtain a three-dimensional virtual character. In this way, even in the case of missing some key points, a three-dimensional virtual character can be generated, and the missing part of the key points and the surrounding part have a natural transition.

[0084] After obtaining a candidate three-dimensional virtual character similar to the real person in the full-body image, the clothing parameters of the full-body image can be further obtained to generate the clothing of the candidate three-dimensional virtual character, and a final target three-dimensional virtual character highly similar to the real person is obtained.

[0085] In a possible way, the clothing style parameters and clothing color parameters corresponding to the user in the full-body image are obtained through the image feature extraction model. Then, according to the clothing style parameters and clothing color parameters corresponding to the user in the full-body image, the corresponding clothing is generated for the candidate three-dimensional virtual character to obtain the target three-dimensional virtual character.

[0086] Exemplarily, the clothing style parameters can be divided into upper clothing style parameters and lower clothing style parameters, and the clothing color parameters can be divided into upper clothing color parameters and lower clothing color parameters. For example, if the user's clothing in the full-body image is a red short-sleeved top and blue pants, the corresponding output parameters can be expressed as [(255, 0, 0), 1, (0, 0, 255), 2]. Among them, (255, 0, 0) represents red, 1 represents a short-sleeved top, (0, 0, 255) represents blue, and 2 represents pants. The color parameters refer to the parameter design of the RGB image, and the style parameters are mainly user-defined. The present disclosure does not make specific limitations on this. When the clothing style of the user in the input full-body image is a dress, if the dress style parameter is set to 3, the output upper clothing style parameter and lower clothing style parameter can both be output as 3, so as to indicate that the user's clothing style is a dress.

[0087] After obtaining the clothing style parameters and clothing color parameters corresponding to the user in the full-body image, generate corresponding clothing for the candidate three-dimensional virtual character, where the generated clothing can be pre-stored clothing. For example, a clothing library is established in advance to store clothing of different styles and colors. According to the clothing style parameters and clothing color parameters corresponding to the user in the full-body image, select the clothing with corresponding parameters from the clothing library and fuse it with the candidate three-dimensional virtual character to obtain a target three-dimensional virtual character highly similar to the real person.

[0088] In addition, in addition to generating clothing similar to that worn by the user in the full-body image, the clothing in the clothing library can also be displayed to the user, and the user can change the clothing of the three-dimensional virtual character according to needs, so as to meet the user's need to change clothing and increase the fun.

[0089] The training process of the feature extraction network in the image feature extraction model in the present disclosure will be described below.

[0090] In a possible way, the feature extraction network can be trained in the following way: First, obtain a sample image marked with sample key point information and sample clothing parameters, then input the sample image into the image feature extraction network to obtain the predicted key point information and predicted clothing parameters of the sample image, and finally calculate the loss function according to the sample key point information and predicted key point information in the sample image, as well as the sample clothing parameters and predicted clothing parameters, and adjust the parameters of the feature extraction network according to the calculation result of the loss function.

[0091] Exemplarily, first, the sample key point information and sample clothing parameters of the sample image are labeled. Among them, the sample key point information of the sample image refers to the category, serial number, and three-dimensional coordinates of each key point of the person in the sample image, and the sample clothing parameters refer to the clothing style parameters and clothing color parameters of the person wearing in the sample image. The sample image is input into the image feature extraction network to obtain the predicted key point information and predicted clothing parameters of the sample image, and then the loss function is calculated. Among them, the loss function represents the degree of difference between the sample key point information and the predicted key point information, and the degree of difference between the sample clothing parameters and the predicted clothing parameters. According to the calculation result of the loss function, the parameters of the feature extraction network are adjusted to retrain the model, so that the predicted key point information output by the feature extraction network is getting closer and closer to the sample key point information, and the predicted clothing parameters are getting closer and closer to the sample clothing parameters, thereby completing the training of the feature extraction network.

[0092] It should be noted that the feature extraction network used in the embodiments of the present disclosure is a neural network, which can be, for example, a convolutional neural network, a residual network, etc. The present disclosure does not make specific limitations thereto. In addition, the image obtained after preprocessing the sample image by inputting it into the image preprocessing network can be input into the feature extraction network for training, which is beneficial for the feature extraction network to extract more image features and perform model training more effectively.

[0093] By training the feature extraction network through the above method, the feature extraction network can identify the key point information and clothing parameters from the user's full-body image, which helps to generate a three-dimensional virtual character similar to the user in the subsequent steps.

[0094] Next, an example of the method for generating a three-dimensional virtual character provided by the present disclosure in combination with the application scenario of a tutoring machine will be described.

[0095] In a possible way, the user's input full-body image is obtained through the following method: in response to the user's three-dimensional virtual character creation operation in the tutoring machine, the user's input full-body image is obtained. Then, after generating the target three-dimensional virtual character, the target three-dimensional virtual character and the target learning content are displayed in the tutoring machine, and in response to the learning action of the user collected for the target learning content, the target three-dimensional virtual character is driven to perform corresponding expressions and / or actions based on the learning action.

[0096] Exemplarily, the user can obtain a full-body image of the user through the camera built in the tutoring machine, or obtain a full-body image of the user from other devices or the historical photo library. After obtaining the full-body image of the user, a target 3D virtual character similar to the user's full-body image is generated and displayed on the display screen of the tutoring machine. The user can modify the target 3D virtual character, such as replacing the clothing, modifying the hairstyle, etc., and the present disclosure does not make specific limitations thereto. Further, the user can enter different function modules by clicking on the display screen of the tutoring machine. For example, when entering the course learning module and the user reads the text aloud, the tutoring machine collects the user's voice and can drive the target 3D virtual character to make a mouth-opening reading action, and so on. In addition, for different function modules, the target 3D virtual character can be driven to make corresponding expressions or actions, and the present disclosure will not elaborate herein.

[0097] Through the above method, the generated target 3D virtual character is applied to the tutoring machine. By generating a 3D virtual character highly similar to the user's image, the user has an immersive feeling in different learning scenarios, increasing the fun in the learning process. And, in response to collecting the learning actions of the user for the target learning content, the 3D virtual character can be driven to perform corresponding actions based on the learning actions, and real-time interaction with the user can be carried out during the learning process, which is beneficial for the user to focus on the learning scenario.

[0098] Figure 3 is a block diagram of a device for generating a 3D virtual character shown according to an exemplary embodiment. As Figure 3 shown, the device 400 includes:

[0099] An acquisition module 401, configured to acquire a full-body image input by a user;

[0100] An extraction module 402, configured to obtain key point information corresponding to key points in the full-body image and clothing parameters corresponding to the user in the full-body image through an image feature extraction model, where the key point information includes the serial number of the key point and the three-dimensional coordinates of the key point;

[0101] An adjustment module 403, configured to correspond the key points of the full-body image with preset key points having the same serial number in the 3D virtual character model according to the serial number of the key point, and adjust the positions of the preset key points in the 3D virtual character model according to the three-dimensional coordinates of the key point to obtain a candidate 3D virtual character;

[0102] A generation module 404, configured to generate a target 3D virtual character according to the clothing parameters corresponding to the user in the full-body image and the candidate 3D virtual character.

[0103] Using the above device, key point information corresponding to key points in the user's full-body image and clothing parameters corresponding to the user in the full-body image are obtained through an image feature extraction model, and a target three-dimensional virtual character is generated. With this device, a three-dimensional virtual character can be generated using a single two-dimensional full-body image. Moreover, the three-dimensional virtual character generated based on the full-body image includes not only the face part but also the body part, and the clothing of the three-dimensional virtual character is also similar to the full-body image input by the user, thus establishing a three-dimensional virtual character highly similar to the user's own image.

[0104] Optionally, the image feature extraction model includes an image preprocessing network and a feature extraction network, and the extraction module 402 includes:

[0105] A preprocessing module, configured to input the full-body image into the image feature extraction model, and perform image preprocessing on the full-body image based on at least one preset image processing strategy through the image preprocessing network in the image feature extraction model to obtain a target image;

[0106] An extraction sub-module, configured to perform feature extraction on the target image through the feature extraction network in the image feature model to obtain key point information corresponding to key points in the full-body image and clothing parameters corresponding to the user in the full-body image.

[0107] Optionally, the preprocessing module is configured to:

[0108] Perform image preprocessing on the full-body image based on at least one of the following preset image processing strategies through the image preprocessing network in the image feature extraction model to obtain multiple target sub-images: RGB channel separation processing, HSI channel separation processing, YCrCb channel separation processing, grayscale processing, histogram equalization processing, and image sharpening processing;

[0109] Overlay the multiple target sub-images to obtain the target image.

[0110] Optionally, the device 400 further includes a training module, and the training module is configured to:

[0111] Obtain a sample image labeled with sample key point information and sample clothing parameters;

[0112] Input the sample image into the image feature extraction network to obtain predicted key point information and predicted clothing parameters of the sample image;

[0113] Calculate a loss function based on the sample key point information and the predicted key point information in the sample image, and the sample clothing parameters and the predicted clothing parameters, and adjust the parameters of the feature extraction network according to the calculation result of the loss function.

[0114] Optionally, the adjustment module 403 is configured to:

[0115] Execute a global adjustment strategy and a local adjustment strategy according to the three-dimensional coordinates of the key points;

[0116] Wherein, the global adjustment strategy is used to adjust the positions of the preset key points corresponding to the key points in the three-dimensional virtual character model according to the three-dimensional coordinates of each key point;

[0117] The local adjustment strategy is used to adjust the positions of the preset key points corresponding to the target key points in the three-dimensional virtual character model according to the three-dimensional coordinates of the target key points in the key points after the global adjustment strategy is executed, and the target key points are the key points that can be used to uniquely determine the user's face contour.

[0118] Optionally, the extraction module 402 is configured to:

[0119] Obtain all key points, the serial numbers of all key points, the three-dimensional coordinates of all key points, and the first normal vector of the target key point in the whole body image through an image feature extraction model.

[0120] The adjustment module 403 is configured to:

[0121] Determine the second normal vector of the target key point according to the three-dimensional coordinates of the target key point and the three-dimensional coordinates of a preset number of key points closest to the target key point among all key points;

[0122] Calculate the mean value of the first normal vector and the second normal vector of the target key point to obtain the target normal vector of the target key point, and adjust the position of the preset key point corresponding to the target key point in the three-dimensional virtual character model according to the target normal vector.

[0123] Optionally, the extraction module 402 is configured to:

[0124] Obtain the clothing style parameters and clothing color parameters corresponding to the user in the whole body image through an image feature extraction model.

[0125] The generation module 404 is configured to:

[0126] Generate corresponding clothing for the candidate three-dimensional virtual character according to the clothing style parameters and clothing color parameters corresponding to the user in the whole body image to obtain a target three-dimensional virtual character.

[0127] Optionally, the acquisition module 401 is configured to:

[0128] In response to a user's three-dimensional virtual character creation operation on a tutoring machine, obtain a full-body image input by the user.

[0129] After generating the target three-dimensional virtual character, the device 400 is further configured to:

[0130] Display the target three-dimensional virtual character and target learning content on the tutoring machine;

[0131] In response to learning actions of the user for the target learning content being collected, drive the target three-dimensional virtual character to perform corresponding expressions and / or actions based on the learning actions.

[0132] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.

[0133] The embodiments of the present disclosure further provide a non-transitory computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the method for generating a three-dimensional virtual character provided in the above embodiments are implemented.

[0134] The embodiments of the present disclosure further provide a tutoring machine, including:

[0135] A memory, on which a computer program is stored;

[0136] A processor, configured to execute the computer program in the memory to implement the steps of the method for generating a three-dimensional virtual character provided in the above embodiments.

[0137] Figure 4 is a block diagram of a tutoring machine 500 shown according to an exemplary embodiment. As Figure 4 shown, the tutoring machine 500 may include: a processor 501, a memory 502. The tutoring machine 500 may further include one or more of a multimedia component 503, an input / output (I / O) interface 504, and a communication component 505.

[0138] Among them, the processor 501 is used to control the overall operation of the tutoring machine 500 to complete all or part of the steps in the above method for generating a three-dimensional virtual character. The memory 502 is used to store various types of data to support the operation of the tutoring machine 500. These data may include, for example, instructions for any application or method operating on the tutoring machine 500, as well as application-related data, such as contact data, received and sent messages, pictures, audio, video, and so on. The memory 502 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disc. The multimedia component 503 may include a screen and an audio component. Among them, the screen may be, for example, a touch screen, and the audio component is used to output and / or input audio signals. For example, the audio component may include a microphone, and the microphone is used to receive external audio signals. The received audio signals may be further stored in the memory 502 or sent through the communication component 505. The audio component further includes at least one speaker for outputting audio signals. The I / O interface 504 provides an interface between the processor 501 and other interface modules. The above other interface modules may be a keyboard, a mouse, buttons, etc. These buttons may be virtual buttons or physical buttons. The communication component 505 is used for wired or wireless communication between the tutoring machine 500 and other devices. Wireless communication, such as Wi-Fi, Bluetooth, Near Field Communication (NFC), 2G, 3G, 4G, NB-IOT, eMTC, or other 5G, etc., or a combination of one or more of them, is not limited herein. Therefore, the corresponding communication component 505 may include: a Wi-Fi module, a Bluetooth module, an NFC module, and so on.

[0139] In an exemplary embodiment, the tutoring machine 500 may be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components, and is used to execute the method for generating a three-dimensional virtual character described above.

[0140] In another exemplary embodiment, a computer-readable storage medium including program instructions is further provided. When the program instructions are executed by a processor, the steps of the method for generating a three-dimensional virtual character described above are implemented. For example, the computer-readable storage medium may be the memory 502 including the program instructions described above, and the program instructions may be executed by the processor 501 of the tutoring machine 500 to complete the method for generating a three-dimensional virtual character described above.

[0141] The preferred embodiments of the present disclosure have been described in detail above with reference to the accompanying drawings. However, the present disclosure is not limited to the specific details in the above embodiments. Within the scope of the technical concept of the present disclosure, various simple modifications can be made to the technical solutions of the present disclosure, and these simple modifications all fall within the protection scope of the present disclosure.

[0142] In addition, it should be noted that, in the above specific embodiments, the various specific technical features described can be combined in any suitable manner without conflict. To avoid unnecessary repetition, the present disclosure will not separately describe various possible combination methods.

[0143] Furthermore, any combination can be made between different embodiments of the present disclosure as long as it does not violate the idea of the present disclosure, and it should also be regarded as the content disclosed by the present disclosure.

Claims

1. A method for generating a three-dimensional virtual character, characterized in that, The method includes: obtaining a full-body image input by a user; obtaining key point information corresponding to key points in the full-body image and clothing parameters corresponding to the user in the full-body image through an image feature extraction model, where the key point information includes the serial number of the key point and the three-dimensional coordinates of the key point; corresponding the key points of the full-body image to preset key points with the same serial numbers in a three-dimensional virtual character model according to the serial numbers of the key points, and adjusting the positions of the preset key points in the three-dimensional virtual character model according to the three-dimensional coordinates of the key points to obtain a candidate three-dimensional virtual character; generating a target three-dimensional virtual character according to the clothing parameters corresponding to the user in the full-body image and the candidate three-dimensional virtual character; The adjusting the positions of the preset key points in the three-dimensional virtual character model according to the three-dimensional coordinates of the key points includes: executing a global adjustment strategy and a local adjustment strategy according to the three-dimensional coordinates of the key points; wherein, the global adjustment strategy is used to adjust the positions of the preset key points corresponding to the key points in the three-dimensional virtual character model according to the three-dimensional coordinates of each key point; the local adjustment strategy is used to adjust the positions of the preset key points corresponding to the target key points in the three-dimensional virtual character model according to the three-dimensional coordinates of the target key points in the key points after executing the global adjustment strategy, and the target key point is a key point that can be used to uniquely determine the facial contour of the user; The obtaining key point information corresponding to key points in the full-body image through an image feature extraction model includes: obtaining all key points, the serial numbers of all key points, the three-dimensional coordinates of all key points and the first normal vector of the target key point in the full-body image through an image feature extraction model; The adjusting the positions of the preset key points corresponding to the target key points in the three-dimensional virtual character model according to the three-dimensional coordinates of the target key points in the key points includes: determining the second normal vector of the target key point according to the three-dimensional coordinates of the target key point and the three-dimensional coordinates of a preset number of key points closest to the target key point among all key points; taking the average of the first normal vector and the second normal vector of the target key point to obtain the target normal vector of the target key point, and adjusting the position of the preset key point corresponding to the target key point in the three-dimensional virtual character model according to the target normal vector.

2. The method according to claim 1, wherein The image feature extraction model includes an image preprocessing network and a feature extraction network. The obtaining key point information corresponding to key points in the full-body image and clothing parameters corresponding to the user in the full-body image through the image feature extraction model includes: inputting the full-body image into the image feature extraction model, and performing image preprocessing on the full-body image through the image preprocessing network in the image feature extraction model based on at least one preset image processing strategy to obtain a target image; Feature extraction is performed on the target image through the feature extraction network in the image feature model to obtain key point information corresponding to key points in the full-body image and clothing parameters corresponding to the user in the full-body image.

3. The method according to claim 2, wherein The image preprocessing of the full-body image through the image preprocessing network in the image feature extraction model based on at least one preset image processing strategy to obtain a target image includes: Performing image preprocessing on the full-body image through the image preprocessing network in the image feature extraction model based on at least one of the following preset image processing strategies to obtain multiple target sub-images: RGB channel separation processing, HSI channel separation processing, YCrCb channel separation processing, grayscale processing, histogram equalization processing, and image sharpening processing; Superimposing the multiple target sub-images to obtain the target image.

4. The method according to claim 2, wherein The training process of the feature extraction network includes: Obtaining a sample image labeled with sample key point information and sample clothing parameters; Inputting the sample image into the image feature extraction network to obtain predicted key point information and predicted clothing parameters of the sample image; Calculating a loss function based on the sample key point information and the predicted key point information in the sample image, and the sample clothing parameters and the predicted clothing parameters, and adjusting the parameters of the feature extraction network according to the calculation result of the loss function.

5. The method according to any one of claims 1-4, characterized in that, The obtaining of the clothing parameters corresponding to the user in the full-body image through the image feature extraction model includes: Obtaining clothing style parameters and clothing color parameters corresponding to the user in the full-body image through the image feature extraction model; The generating of a target three-dimensional virtual character according to the clothing parameters corresponding to the user in the full-body image and the candidate three-dimensional virtual character includes: Generating corresponding clothing for the candidate three-dimensional virtual character according to the clothing style parameters and clothing color parameters corresponding to the user in the full-body image to obtain a target three-dimensional virtual character.

6. The method according to any one of claims 1-4, characterized in that The obtaining of the full-body image input by the user includes: Responding to a three-dimensional virtual character creation operation by the user in the tutoring machine to obtain the full-body image input by the user; After generating the target three-dimensional virtual character, it further includes: Displaying the target three-dimensional virtual character and target learning content in the tutoring machine; Responding to a learning action of the user for the target learning content, driving the target three-dimensional virtual character to perform corresponding expressions and / or actions based on the learning action.

7. An apparatus for generating a three-dimensional virtual character, characterized in that The device includes: An acquisition module for acquiring a full-body image input by the user; An extraction module for obtaining key point information corresponding to key points in the full-body image and clothing parameters corresponding to the user in the full-body image through an image feature extraction model, where the key point information includes the serial number of the key point and the three-dimensional coordinates of the key point; An adjustment module, configured to correspond the key points of the full-body image with the preset key points having the same serial numbers in the three-dimensional virtual character model according to the serial numbers of the key points, and adjust the positions of the preset key points in the three-dimensional virtual character model according to the three-dimensional coordinates of the key points, so as to obtain a candidate three-dimensional virtual character; A generation module, configured to generate a target three-dimensional virtual character according to the clothing parameters corresponding to the user in the full-body image and the candidate three-dimensional virtual character; The adjustment module is configured to: Execute a global adjustment strategy and a local adjustment strategy according to the three-dimensional coordinates of the key points; Wherein, the global adjustment strategy is configured to adjust the positions of the preset key points corresponding to the key points in the three-dimensional virtual character model according to the three-dimensional coordinates of each key point; The local adjustment strategy is configured to, after executing the global adjustment strategy, adjust the positions of the preset key points corresponding to the target key points in the three-dimensional virtual character model according to the three-dimensional coordinates of the target key points in the key points, and the target key points are the key points that can be used to uniquely determine the facial contour of the user; The extraction module is configured to: Obtain all key points, the serial numbers of all key points, the three-dimensional coordinates of all key points and the first normal vector of the target key point in the full-body image through an image feature extraction model; The adjustment module is configured to: Determine the second normal vector of the target key point according to the three-dimensional coordinates of the target key point and the three-dimensional coordinates of the preset number of key points closest to the target key point among all key points; Calculate the average value of the first normal vector and the second normal vector of the target key point to obtain the target normal vector of the target key point, and adjust the position of the preset key point corresponding to the target key point in the three-dimensional virtual character model according to the target normal vector.

8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the steps of the method according to any one of claims 1-6.

9. A tutoring machine, characterized in that, Including: A memory, on which a computer program is stored; A processor, configured to execute the computer program in the memory to implement the steps of the method according to any one of claims 1-6.

Citation Information

Patent Citations

  • 3D virtual image generation method and device

    CN108305312A

  • Image key point extraction method and device, readable storage medium and electronic equipment

    CN109697446A

  • Three-dimensional face reconstruction network training and virtual face image generation method and device

    CN111354079A

  • Virtual clothing try-on method and device, terminal equipment and storage medium

    CN111508079A