Image generation method and apparatus, device and medium
By integrating eye area images and expression control information during image generation, the problem of difficult control of the eye sight direction of the character's eyeball in the prior art is solved, and the vividness and authenticity of the character in the generated video are improved.
Patent Information
- Application Number
- PCT/CN2024/141416
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-28
- Filing Date
- 2024-12-23
- Publication Date
- 2025-07-03
AI Technical Summary
Existing image driving technology is difficult to effectively control the eye-catching direction of the character in the source image, resulting in the generated video characters lacking vividness.
By determining the facial rendering image based on the expression coefficient of the preset facial expression and the facial morphology information of the source image, the facial rendering image is determined, and combined with the target eye model and facial eye texture, expression control information is generated to control the eye direction and texture of the character in the target image, and to improve the authenticity and vividness of the character.
Accurate control of the eye direction and texture of the character in the target image is achieved, improving the vividness and realism of the character in the generated video.
Smart Images

Figure CN2024141416_03072025_PF_FP_ABST
Abstract
Description
Image generation method, device, equipment and medium
[0001] This application claims priority to Chinese Patent Application No. 202311840691.1 filed on December 28, 2023, and the contents of the above-mentioned Chinese patent application disclosure are hereby incorporated by reference in their entirety as a part of this application. Technical Field
[0002] The present disclosure relates to an image generation method, apparatus, device and medium. Background Art
[0003] Image driving, also known as motion transfer, uses a driving video to drive a source image to generate a video whose appearance is consistent with the source image, but the main motion is consistent with the driving video.
[0004] In the related art, after the source image is driven by image driving technology, it is difficult to drive the eye sight direction of the character included in the source image, resulting in the lack of vividness of the character included in the generated video. Summary of the Invention
[0005] According to one aspect of the present disclosure, there is provided an image generation method, comprising:
[0006] determining a facial rendering image based on an expression coefficient of a preset facial expression and facial morphology information included in the source image;
[0007] Determining expression control information based on the facial rendering image, the target eyeball model corresponding to the preset facial expression, and the facial eyeball texture of the source image;
[0008] A target image is generated based on the expression control information, the source image, the three-dimensional facial information of the source image, and the three-dimensional facial information corresponding to the preset facial expression.
[0009] According to another aspect of the present disclosure, there is provided an image generating apparatus, comprising:
[0010] a determination module, configured to determine a facial rendering image based on an expression coefficient of a preset facial expression and facial morphology information included in a source image, and determine expression control information based on the facial rendering image, a target eyeball model corresponding to the preset facial expression, and a facial eyeball texture of the source image;
[0011] A generation module is used to generate a target image based on the expression control information, the source image, the three-dimensional facial information of the source image, and the three-dimensional facial information corresponding to the preset facial expression.
[0012] According to another aspect of the present disclosure, there is provided an electronic device, comprising:
[0013] processor; and,
[0014] Memory for storing programs;
[0015] The program includes instructions, which, when executed by the processor, cause the processor to perform the method according to the exemplary embodiment of the present disclosure.
[0016] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium is provided, wherein the non-transitory computer-readable storage medium stores computer instructions for causing the computer to execute the method according to the exemplary embodiments of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Further details, features and advantages of the present disclosure are disclosed in the following description of exemplary embodiments in conjunction with the accompanying drawings, in which:
[0018] FIG1 is a schematic flow chart showing an image generating method according to an exemplary embodiment of the present disclosure;
[0019] FIG2 shows a schematic diagram of the architecture of an image driving model of an exemplary embodiment of the present disclosure;
[0020] FIG3 is a schematic diagram showing eye movement directions according to an exemplary embodiment of the present disclosure;
[0021] FIG4 shows a schematic diagram of a flow chart of visually determining expression control information using the eye region as an example in an exemplary embodiment of the present disclosure;
[0022] FIG5 shows a schematic diagram of a process for obtaining expression sample control information according to an exemplary embodiment of the present disclosure;
[0023] FIG6 shows a schematic block diagram of functional modules of an image generating apparatus according to an exemplary embodiment of the present disclosure;
[0024] FIG7 shows a schematic block diagram of a chip according to an exemplary embodiment of the present disclosure; and
[0025] FIG8 shows a block diagram of an exemplary electronic device that can be used to implement the embodiments of the present disclosure. DETAILED DESCRIPTION
[0026] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0027] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.
[0028] The term "including" and its variations used in this document are open inclusions, that is, "including but not limited to". The term "based on" means "based at least in part on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one other embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the description below. It should be noted that the concepts of "first", "second", etc. mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0029] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".
[0030] In related technologies, after driving a source image containing a source character using image driving technology such as single-image driving technology, although the target character contained in the generated target video is basically the same as the target character in appearance, it is still difficult to control the facial eye direction (i.e., the line of sight direction) of the source image, resulting in the character in the generated target image lacking vividness.
[0031] The inventors found that when training a single-image driven model, the area of the eye region is relatively small, and the single-image driven model is difficult to learn the eye movement of a smaller area. Therefore, in the inference stage, the single-image driven model is difficult to drive the eye direction of the face in a single image, resulting in the target character contained in the generated target image lacking vividness.
[0032] In response to the above problems, an exemplary embodiment of the present disclosure provides an image generation method, which can integrate an eye area image into a facial rendering image, so that expression control information can guide various source images that are authorized for use. On the one hand, it ensures that the facial expression of the target character in the final generated target image presents a preset facial expression, and on the other hand, it controls the eye direction and eye texture of the eye area of the target character included in the target image, thereby improving the authenticity and vividness of the target character in the generated target image.
[0033] The source image of the exemplary embodiment of the present disclosure may refer to a character image required for use in the image generation process, and the character included in the source image may be defined as a source character. The source image may be a character image collected by the client, or a character image obtained by the client and authorized for use. After the client obtains the character image, the image generation method of the exemplary embodiment of the present disclosure may be executed by an electronic device.
[0034] When the source image is guided by the expression control information, the generated target image is guided by the expression control information, and the eye direction of the target character's eye area included in the target image matches the preset facial expression. The difference is that the target character included in the target image is associated with the appearance of the source character included in the source image, and the eye texture of the target character included in the target image matches the eye texture of the source character included in the source image.
[0035] The source image can be used to provide morphological constraints for the target character displayed in the target image. These morphological constraints can be texture constraints or basic shape constraints for the target character. These constraints ensure that the target character in the target image and the source character in the source image are visually correlated. This appearance correlation can mean that the appearance of the source character in the source image and the target character in the target image are identical, or that the appearance differences between the source character in the source image and the target character in the target image are within a controllable range, but is not limited to these.
[0036] When the electronic device is a user terminal installed on a client, after acquiring the character image, the client can use the character image as the source image and execute the image generation method on the user terminal. When the user terminal is a terminal with a display function, after the user terminal generates the target image using the image generation method, the target image can be displayed on the user terminal.
[0037] When the electronic device is a cloud server, the cloud server can be connected to a user terminal installed with a client via a network. In this case, the client uploads the character image obtained by the user terminal to the cloud server via a wired or wireless network. The cloud server uses the character image as a source image and executes the image generation method to obtain the target image. Finally, the cloud server returns the target image to the user terminal via the network for display on the user terminal.
[0038] For example, the user terminal of the exemplary embodiments of the present disclosure may be a mobile phone, a tablet computer, a wearable device, an in-vehicle device, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), and a wearable device based on augmented reality (AR) and / or virtual reality (VR) technology, etc.
[0039] For example, when the user terminal is a wearable device, the wearable device can also be a general term for wearable devices developed by applying wearable technology to intelligently design everyday wearables, such as glasses, gloves, watches, clothing, and shoes. A wearable device is a portable device that is worn directly on the body or integrated into the user's clothing or accessories.
[0040] These wearable devices are more than just hardware; they achieve powerful functionality through software support, data exchange, and cloud-based interaction. Broadly speaking, these smart wearables include devices that are comprehensive, large, and can function completely or partially independently of a smartphone, such as smartwatches and smart glasses. These devices also focus on a specific application and require integration with other devices, such as smartphones, such as various smart wristbands and smart jewelry for vital sign monitoring.
[0041] The network of the exemplary embodiments of the present disclosure may include one or more networks, and any suitable network is contemplated. By way of example and not limitation, one or more portions of a network may include an ad hoc network, an intranet, an extranet, a virtual private network (VPN), a local area network (LAN), a wireless local area network (WLAN), a wide area network (WAN), a wireless wide area network (WWAN), a metropolitan area network (MAN), a portion of the Internet, a portion of a public switched telephone network (PSTN), a cellular telephone network, or a combination of two or more of these.
[0042] FIG1 is a flow chart showing an image generation method according to an exemplary embodiment of the present disclosure. As shown in FIG1 , the image generation method according to an exemplary embodiment of the present disclosure may include:
[0043] Step 101: Determine a facial rendering image based on the facial expression coefficient of a preset facial expression and the facial morphology information included in the source image. It should be understood that the source image of the exemplary embodiment of the present disclosure can be a photo of a character, or a frame of a character image included in a video containing changes in the character's expression. Regardless of whether it is a photo of a character or a frame of a character image, the source character displayed as the source image can be a real character, such as a real person or various real animals, or various virtual characters, such as virtual characters or virtual animals designed through various licensed software or self-developed software. These virtual animals can be virtual animals that exist in the real world, or they can be conceived virtual animals that do not exist in the real world.
[0044] For example, in the exemplary embodiments of the present disclosure, the facial morphology information included in the source image can represent facial mapping information included in the source image. Driven by the facial expression coefficient of a preset facial expression, the source image can be controlled so that the face of the resulting facial rendering image has both the facial texture of the character in the source image and the preset facial expression. In this case, the character morphology included in the facial rendering image is the same as the source character morphology, but the character expression included in the facial rendering image matches the preset facial expression.
[0045] For example, the facial morphology information included in the source image of the exemplary embodiment of the present disclosure can be determined by the facial texture features of the source image and the facial 3D information of the source image. The facial texture features of the source image can be extracted from the source image using various feature extraction algorithms such as image segmentation algorithms. The source image can also be detected using various target detection algorithms to obtain a facial 3D mesh model as the facial 3D information of the source image. For example, the target detection algorithm can be used to obtain multiple facial position coordinates, and the facial 3D information, i.e., the facial 3D mesh model, can be generated based on the multiple facial position coordinates.
[0046] For example, the facial texture features of the source image of the exemplary embodiment of the present disclosure can represent the facial texture of the source character, such as detailed features such as the character's facial color and facial lines. The facial three-dimensional mesh model of the source image can represent the source character's facial three-dimensional mesh data. Using this facial three-dimensional mesh data, a basic facial model of the character in the source image can be constructed, which can represent the specific shape and surface topology of the face. For example, the basic facial model of the character can include a two-dimensional mesh composed of a series of polygonal (such as triangle) face patches. By connecting different polygonal fragments (taking triangles as an example, each triangular face patch consists of three vertices and three edges), a complex basic facial model of the character can be obtained.
[0047] The exemplary embodiments of the present disclosure combine the three-dimensional facial information of a source image with the facial texture features of the source image through texture mapping, thereby obtaining facial topography information included in the source image. For example, during texture mapping, the facial texture features can be used to map each fragment included in the character's facial basic model. The mapping method can be, for example, UV mapping, but is not limited to this.
[0048] When the facial expression coefficients of the preset facial expressions in the exemplary embodiment of the present disclosure include at least the mouth expression in an open mouth state, the character's mouth state in the facial rendering image appears open mouthed. Of course, the facial expression coefficients of the preset facial expressions may also include expression coefficients of other facial parts. Thus, there is a corresponding relationship between the facial expression coefficients of the preset facial expressions in the exemplary embodiment of the present disclosure and the facial rendering image.
[0049] For example, the expression coefficients of the aforementioned preset facial expressions may include not only the expression coefficients of the mouth in an open mouth state, but also the expression coefficients of the eyes, jaw, eyebrow, cheek, nose, and tongue. When a facial rendering image is generated by combining the 52-dimensional facial expression coefficients with the facial morphology information included in the source image, the generated facial rendering image may present the facial expressions of the eyes, jaw, eyebrows, cheeks, nose, and tongue, and the mouth expressions of the facial expressions may include the mouth expression in an open mouth state.
[0050] Step 102: Determine expression control information based on the facial rendering image, the target eyeball model corresponding to the preset facial expression, and the facial eyeball texture of the source image.
[0051] In practical applications, an eye region image can be generated using a target eye model corresponding to a preset facial expression and the eye texture of the source image. This eye region image is then fused with the rendered facial image to produce a fused facial image that can serve as expression control information. The eye orientation of the character's eye region in this fused facial image is determined by the preset facial expression, while the eye texture of the character's eye region in this fused facial image is determined by the eye texture of the source image.
[0052] When the expression coefficient of the preset facial expression of the exemplary embodiment of the present disclosure includes an eye expression coefficient, the target eyeball model corresponding to the predicted facial expression of the exemplary embodiment of the present disclosure may refer to an eyeball model that presents the eye posture indicated by the eye expression coefficient, and the texture of the target eyeball model is determined by the facial eyeball texture of the source image.
[0053] The above-mentioned process of determining the facial fusion image actually includes two aspects. On the one hand, the target eyeball model corresponding to the preset facial expression is fused into the facial rendering image in the form of a target eyeball image to ensure that the eyeball direction of the character's eyeball area included in the obtained facial fusion image matches the facial expression. On the other hand, it ensures that the character's eyeball texture included in the facial fusion image is consistent with the facial eyeball texture included in the source image.
[0054] Step 103: Generate a target image based on the expression control information, the source image, the 3D facial information of the source image, and the 3D facial information corresponding to the facial rendering image. Here, the exemplary embodiment of the present disclosure can generate the target image using an image-driven model.
[0055] FIG2 shows a schematic diagram of the architecture of an image-driven model of an exemplary embodiment of the present disclosure. As shown in FIG2 , the image-driven model 200 of the exemplary embodiment of the present disclosure may include: a motion estimation network 201 and an image generation network 202. In this case, based on expression control information, a source image, three-dimensional facial information of the source image, and three-dimensional facial information corresponding to a preset facial expression, a target image is generated, including: inputting the three-dimensional facial posture information of the source image, the three-dimensional facial posture information corresponding to the preset facial expression, and the source image into the motion estimation network 201 to obtain expression change estimation information, and inputting the expression change estimation information, the expression control information, and the source image into the image generation network 202 to obtain the target image.
[0056] The exemplary embodiments of the present disclosure can obtain the three-dimensional facial information corresponding to the preset facial expression and the three-dimensional facial information of the source image through a three-dimensional facial reconstruction method.
[0057] In some embodiments, the method for obtaining the facial three-dimensional information corresponding to the preset facial expression may include: obtaining a facial rendering image corresponding to the preset facial expression based on the facial morphology information included in the source image and the facial expression coefficient of the preset facial expression, and then performing facial reconstruction on the facial rendering image to obtain the facial three-dimensional information corresponding to the preset facial expression.
[0058] In other embodiments, the method for obtaining three-dimensional facial information corresponding to a preset facial expression may include: obtaining three-dimensional facial information of a source image from facial topography information included in a source image, and then driving the three-dimensional facial information of the source image using a facial expression coefficient corresponding to the preset facial expression, thereby obtaining three-dimensional facial information corresponding to the preset facial expression, i.e., three-dimensional facial information corresponding to the preset facial expression. It can be seen that both the three-dimensional facial information of the source image and the three-dimensional facial information corresponding to the preset facial expression can reflect the facial appearance features included in the source image, but there are individual differences in expression characteristics.
[0059] The three-dimensional facial information of the source image and the three-dimensional facial information corresponding to the preset facial expression can both be three-dimensional facial mesh models. The three-dimensional facial mesh model included in the source image can describe the state of the facial features included in the source image in three-dimensional space, while the three-dimensional facial mesh model corresponding to the preset facial expression can describe the state of the preset facial features in three-dimensional space. Facial features herein include not only static facial features but also dynamic facial features. Static facial features can be understood as facial features that barely change or undergo minor changes in a short period of time, such as facial appearance features, while dynamic features can be understood as facial features that undergo significant changes in a short period of time, such as facial expression features.
[0060] When the facial three-dimensional information of the source image is driven by the facial expression coefficient of the preset facial expression, the facial expression coefficient of the preset facial expression is actually used to drive the movement of each vertex included in the facial three-dimensional mesh model of the source image, so that the facial expression features of the facial three-dimensional mesh model corresponding to the preset facial expression finally obtained are close to or the same as the preset facial expression.
[0061] After obtaining the 3D facial information corresponding to the preset facial expression and the 3D facial information of the source image, the 3D facial pose information of the source image, the 3D facial pose information corresponding to the preset facial expression, and the source image can be spliced together and then input into a motion estimation network to obtain expression change estimation information. This expression change estimation information can represent the expression change information required to change from the facial expression in the original image to the preset facial expression.
[0062] Exemplarily, as shown in FIG2 , the image generation network 202 of the exemplary embodiment of the present disclosure may include an encoder 2021 and a decoder 2022. Expression control information and a source image may be input into the encoder 2021, and the encoder may be used to encode the expression control information and the source image to obtain facial coding features of the source image. The facial coding features combine features of both the source image and the expression control information. Therefore, the decoder 2022 may be used to decode the facial coding features of the source image to obtain a target image.
[0063] As shown in Figure 2, in order to further improve the correlation between the facial expression of the target image and the preset facial expression, the facial coding features and expression change estimation information of the source image can be input into the decoder 2022 at the same time, so that when the decoder 2022 decodes the facial coding features of the source image, it refers to the expression change estimation information to ensure that the facial expression of the generated target image is closer to the preset facial expression.
[0064] The above-mentioned expression control information is determined by the facial rendering image, the target eyeball model corresponding to the preset facial expression, and the facial eyeball texture of the source image, so that the expression control information not only has the function of controlling the overall facial expression, but also has the function of controlling the eyeball direction and eyeball texture. Therefore, when the exemplary embodiment of the present disclosure inputs the expression control information and the source image into the encoder 2021, the encoder 2021 can encode the source image based on the expression control information to obtain the encoded image features, so that the facial expression, eyeball direction and eyeball texture contained in the encoded image features all match the information carried by the expression control information. Finally, the decoder 222 can ensure that the target character included in the obtained target image not only matches the preset facial expression, but also the eyeball direction of the eyeball area matches the eyeball expression of the preset facial expression, and the eyeball texture of the eyeball area matches the eyeball texture of the source character included in the source image (for example, the same).
[0065] It can be seen that under the control of the expression control information, the exemplary embodiment method of the present disclosure can not only control the facial expression of the character in the target image to present the preset facial expression, but also control the eye direction and eye texture of the eye area of the target character included in the target image, thereby improving the authenticity and vividness of the target character in the generated target image.
[0066] As a possible implementation method, the exemplary embodiment of the present disclosure can determine eye rotation control parameters based on the expression coefficient of a preset facial expression, control the posture of the basic eye model based on the eye rotation control parameters, and obtain a target eye model.
[0067] In practical applications, eye movement description coefficients can be obtained from the expression coefficients of a preset facial expression. Eye rotation control parameters can be determined based on the eye movement description coefficients and the eye movement angle threshold. The basic eye model can be a three-dimensional eye mesh model or a conventional three-dimensional geometric model. The pose of the basic eye model can be a default standard pose, such as looking straight ahead.
[0068] For example, when the eye rotation control parameters include the vertical motion control angle of the basic eye model and the horizontal motion control angle of the basic eye model, the eye motion angle threshold may include the vertical motion angle threshold and the vertical motion angle threshold. In this case, the description coefficient of the eye movement in the vertical direction and the description coefficient of the eye movement in the horizontal direction can be obtained from the expression coefficient of the 52-dimensional preset facial expression. Then, the description coefficient of the eye movement in the vertical direction and the eye movement in the vertical direction are combined to obtain the vertical motion control angle of the eye. The description coefficient of the eye movement in the horizontal direction and the eye movement angle threshold are combined to obtain the horizontal motion control angle of the eye.
[0069] Figure 3 shows a schematic diagram of eye movement directions in accordance with an exemplary embodiment of the present disclosure. As shown in Figure 3, an eye movement coordinate system can be established in accordance with an exemplary embodiment of the present disclosure, with the pupil center of the eye as the origin O, the direction of eye movement toward the upper eyelid as the positive direction of the vertical direction H, and the direction of eye movement toward the inner canthus as the positive direction of the horizontal direction P.
[0070] When the eyeball rotates while moving from the lower eyelid to the upper eyelid, it can be considered to be rotating in the positive vertical direction, defining the eye as looking up. Conversely, the eyeball rotates in the negative vertical direction, defining the eye as looking down. When the eyeball rotates while moving from the outer corner of the eye to the inner corner of the eye, it can be considered to be rotating in the positive horizontal direction, defining the eye as looking inward. Conversely, the eyeball rotates in the negative horizontal direction, defining the eye as looking outward. The following describes how to determine the eye rotation control parameters, using the right eye as an example.
[0071] When the right eyeball is in the horizontal direction, the expression coefficient includes the right eyelid movement coefficient eyeLookIn when the right eyeball looks inward. right and the right eyelid movement coefficient eyeLookOut when the right eyeball looks outward right In order to determine the rotation angle of the right eyeball during horizontal movement, the right eyelid motion coefficient eyeLookIn is obtained. right and the right eyelid movement coefficient eyeLookOut right , then according to Eye Yaw right =30°*eyeLookIn right +(-30°)*eyeLookOut right Determine the horizontal rotation control angle of the right eyeball Eye Yaw right .
[0072] When the right eyeball is in the vertical direction, the expression coefficient includes the right eyelid movement coefficient eyeLookUp when the right eyeball looks up right and the right eyelid movement coefficient eyeLookDown when the right eyeball looks down right In order to determine the rotation angle of the right eyeball in the vertical movement process, the right eyelid movement coefficient eyeLookUp is obtained when the right eyeball looks up. right and the right eyelid movement coefficient eyeLookDown right , then according to Eye Pitch right =30°*eyeLookUp right +(-30°)*eyeLookDown rightDetermine the vertical rotation control angle of the right eyeball Eye Pitch right .
[0073] As a possible implementation method, when determining expression control information, the exemplary embodiment of the present disclosure can fuse the target eyeball model with the facial rendering image in the form of an image to obtain a facial fusion image, and then migrate the facial eyeball texture included in the source image to the facial eyeball contained in the facial fusion image to obtain expression control information.
[0074] The aforementioned process of fusing the target eye model with the rendered facial image essentially involves mapping the target eye model to the eye region of the face included in the rendered facial image. Because the target eye model corresponds to a preset facial expression, and the character's facial expression included in the rendered facial image matches the preset facial expression, fusing the target eye model with the rendered facial image ensures that the resulting fused facial image's eye posture closely matches the facial expression, eliminating the issue of mismatch between eye posture and facial expression caused by significant posture differences.
[0075] In practical applications, when obtaining a facial fusion image, the exemplary embodiment of the present disclosure can project the target eyeball model onto the facial eyeball area included in the facial rendering image based on the camera parameters corresponding to the source image and the three-dimensional position information of the target eyeball model to obtain a facial fusion image.
[0076] When the three-dimensional position information of the basic eyeball model is known, the three-dimensional position information of the target eyeball model determined by the above-mentioned related methods is also known. When the three-dimensional position of the basic eyeball model is unknown, the three-dimensional position information of the target eyeball model in the exemplary embodiment of the present disclosure can be obtained by the following method:
[0077] Based on the facial three-dimensional information and the basic eyeball model corresponding to the source image, an initial facial three-dimensional model containing the basic eyeball model is determined; based on the initial facial three-dimensional information containing the basic eyeball model and the expression coefficient of the preset facial expression, a target facial three-dimensional model containing the target eyeball model is obtained; based on the target facial three-dimensional model containing the target eyeball model, the three-dimensional position information of the target eyeball model is determined.
[0078] Exemplarily, the facial three-dimensional information corresponding to the above-mentioned source image can be regarded as the facial three-dimensional model corresponding to the source image, and the basic eyeball model can be placed in the preset eyeball area of the facial three-dimensional model corresponding to the source image to obtain an initial facial three-dimensional model containing the basic eyeball model, and then the expression of the initial facial three-dimensional model is regulated by the expression coefficient of the preset facial expression until the expression of the initial facial three-dimensional model matches the preset facial expression.
[0079] During the process of manipulating the expression of the initial 3D facial model using the expression coefficients of the preset facial expression, the posture of the basic eye model also changes accordingly with the expression of the initial 3D facial model. After manipulating the expression of the initial 3D facial model using the expression coefficients of the preset facial expression, the manipulated initial 3D facial model can be defined as the target 3D facial model. In this case, the basic eye model has essentially been transformed into the target eye model, so that the target 3D facial model includes the target eye model.
[0080] It can be seen that the exemplary embodiment of the present disclosure can perform pupil key point detection on the target facial three-dimensional model, thereby obtaining the target three-dimensional position information of the target eyeball model, and then, based on the camera projection principle, use the target three-dimensional position information of the target eyeball model and the camera parameters corresponding to the source image to obtain a facial fusion image.
[0081] In the exemplary embodiment of the present disclosure, based on initial 3D facial information containing a basic eyeball model and expression coefficients of preset facial expressions, a target 3D position model containing a target eyeball model is obtained. On the one hand, the facial expression is controlled by using the preset facial expression coefficients to transform the initial 3D facial model into a target 3D facial model. On the other hand, eye rotation control parameters can be obtained by referring to the relevant content described above. These eye rotation control parameters are used to control the model posture of the basic eyeball, thereby obtaining a target eyeball model located on the target 3D facial model. In this case, the target 3D facial model is detected to obtain the 3D position information of the target eyeball model.
[0082] For example, considering that the positions of facial attributes change after the three-dimensional facial information corresponding to the source image is adjusted by the preset facial expression, the exemplary embodiment of the present disclosure can first obtain the pupil key point mapping relationship between the facial fusion image and the source image when obtaining expression control information, and then use the pupil key point mapping relationship to migrate the facial eyeball texture to the facial eyeball contained in the facial fusion image to obtain expression control information.
[0083] For example, the pupil key point position information model pupil landmarkers of the facial fusion image and the pupil key point position information source pupil landmarks of the source image are obtained, and then based on the pupil key point position information with the same attributes included in the facial fusion image and the source image, the pupil key point mapping relationship AffineTransform Matrix between the facial fusion image and the source image is determined, which can be expressed as AffineTransform Matrix = getAffineTransform(source pupil landmarks, model pupil landmarkers).
[0084] Considering that both the facial fusion image and the source image are two-dimensional images, the pupil keypoint mapping relationship can include translation parameters in the x- and y-directions on the image coordinate system. When the pupil keypoint mapping relationship is used to transfer the facial eyeball texture to the facial eyeball contained in the facial fusion image, the texture coordinates of the facial eyeball texture are essentially converted through the pupil keypoint mapping relationship, thereby transferring the facial eyeball texture to the facial eyeball contained in the facial fusion image.
[0085] The expression control information of the exemplary embodiment of the present disclosure is essentially a facial fusion image, which has facial eyeballs as the content of the expression control information, and is used to control the generation of facial eyeballs of the target character included in the target image during the image generation process.
[0086] To facilitate understanding of the process of determining the expression control information of the exemplary embodiment of the present disclosure, the following describes the process of generating the expression control information by taking the eye region image of the facial fusion image as an example.
[0087] FIG4 illustrates a schematic diagram of a visual process for determining expression control information using an eye region image as an example, according to an exemplary embodiment of the present disclosure. As shown in FIG4 , the exemplary embodiment of the present disclosure can obtain an initial eye region model 403 containing a basic eye model based on an eye region model 401 and a basic eye model 402 corresponding to a source image. Subsequently, a target three-dimensional eye model 404 can be determined based on expression coefficients of a preset facial expression and the initial eye region model 403. The expression coefficients of the preset facial expression include descriptive coefficients for eye movement. Therefore, based on the descriptive coefficients for eye movement and an eye movement angle threshold, an eye rotation control parameter is determined. The eye rotation control parameter is then used to control the posture of the basic eye model contained in the initial eye region model 403, thereby obtaining a target three-dimensional eye model 404 containing a target eye model 405.
[0088] After obtaining the target eye 3D model 404 including the target eye model 405, target detection can be performed on the target eye 3D model 404 to obtain the 3D position information of the target eye model. The target eye model is then projected onto the eye frame region of the rendered facial image using a camera projection model, thereby merging the target eye model with the eye region included in the rendered facial image in the form of an image, thereby obtaining a fused facial image. Finally, using the pupil key point mapping relationship between the fused facial image and the source image, the facial eye texture of the eye region image 406 corresponding to the source image is projected onto the eye included in the fused facial image, thereby obtaining the eye region image 407 shown in FIG. 4 .
[0089] It can be seen that if the expression control information is determined directly based on the facial rendering image, the determined expression control information lacks explicit eye movement control information. Therefore, the exemplary embodiment of the present disclosure controls the posture of the basic eyeball model by the expression coefficient of the preset facial expression on the basis of the facial rendering image, thereby obtaining the target eyeball model corresponding to the expression coefficient of the preset facial expression, and then uses the target eyeball model and the facial eyeball texture of the source image to render the eyeball area image in the facial rendering image, so that the expression control information finally obtained contains explicit eyeball control information (i.e., the eyeball area image), thereby ensuring that the facial expression of the target character included in the generated target image matches the preset facial expression, and the eyeball posture occupying a very small area matches the eyeball posture included in the preset facial expression, thereby improving the authenticity and vividness of the target character.
[0090] As a possible implementation method, the training set of the image-driven model of the exemplary embodiment of the present disclosure in the training stage includes source image samples, facial three-dimensional posture information of the source image samples, facial three-dimensional posture information of the driving image samples, and expression sample control information.
[0091] FIG5 shows a schematic diagram of a process for obtaining expression sample control information according to an exemplary embodiment of the present disclosure. As shown in FIG5 , the method for obtaining expression sample control information according to an exemplary embodiment of the present disclosure may include:
[0092] Step 501: Obtain a character facial rendering sample based on the facial morphology information of the source image sample and the facial expression information of the driving image sample. The facial morphology information of the source image sample can be referred to in the previous description and will not be repeated here. It should be understood that the driving image sample of the exemplary embodiment of the present disclosure can be a video containing an arbitrary character (the arbitrary character is defined as a reference character) or a picture containing a reference character.
[0093] In some examples, the source image sample of the exemplary embodiments of the present disclosure may be any frame in a video, and the driving image sample may be any frame in the video except the source image sample.
[0094] The facial expression information of the driving image sample may be facial expression coefficients of a reference character included in the driving image sample. When the facial topography information included in the source image is determined by facial texture features and three-dimensional facial information of the source image, the facial expression coefficients of the reference character may be used to drive the three-dimensional facial information of the source image, so that the facial expression included in the obtained facial rendering sample is correlated with the facial expression of the reference character included in the driving image.
[0095] When the driving image sample is a video containing a reference character, the character's facial expression coefficients for each frame of the video can be obtained. Then, based on the facial morphology information of the source image sample and the facial expression coefficients for each frame of the video image, a corresponding facial rendering sample is determined for each frame of the video image. This facial rendering sample essentially adds a facial rendering image sample, containing the same character appearance as the source character included in the source image sample, but with facial expressions related to the facial expressions of the reference character included in the driving image.
[0096] Step 502: Determine expression sample control information based on the eye area image and character face rendering sample included in the driving image sample.
[0097] The driving image sample of the exemplary embodiment of the present disclosure can be a video containing a reference character or a single picture containing a reference character. Pupil key point detection can be performed on the driving image sample to obtain the left and right pupil key points of the reference character.
[0098] For each eye's pupil key points, there can actually be multiple pupil key points, and the multiple pupil key points are distributed around the pupil. Therefore, the eyeball area image of the eye can be obtained through the multiple pupil key points of each eye, and then the eyeball texture of the eye can be combined with the character's facial rendering sample to obtain the expression sample control information.
[0099] For example, pupil key point detection algorithm LandmarkDetect can be used to detect pupil key points of the driving image sample driver to obtain multiple pupil landmarks {left, right} of the reference character, which can be divided into left and right eyes and include multiple left eye pupil landmarks of the reference character. left and multiple right eye pupil landmarks right , the process can be expressed as: pupil landmarks{left, right}=LandmarkDetect(driver).
[0100] In order to obtain the left and right pupil masks of the reference character, pupil mask {left, right}, considering multiple left pupil key points pupil landmarks left Multiple pupil landmarks of the right eye distributed circumferentially around the left eyeball right Distributed around the right eyeball, multiple pupil landmarks of the left eye can be left and multiple right eye pupil landmarksright As two point sets, use cv2.fillConvexPoly to solve the minimum convex hull surrounding each set of points. The pupil mask of the left eye can be determined based on the minimum convex hull corresponding to the left pupil key point. left , based on the minimum convex hull corresponding to the right pupil key points, determine the right eye mask pupil landmarks right , you can mask the left eye pupil mask left and right eye mask pupil landmarks right Defined as left and right eye masks pupil mask {left, right}.
[0101] Use pupil mask{left, right}=cv2.fillConvexPoly(pupil landmarks{left, right}) to get the left and right eye masks pupil mask{left, right}. Finally, you can use pupil area{left, right}=driver*pupil mask{left, right} to get the left and right eye images pupil area{left, right}, that is, the left eye image pupil area left and pupil area of the right eyeball image right When the left and right eyeball images pupil area {left, right} are obtained, the expression sample control information Control img can be obtained based on Control img = Render + pupil area2 {left, right}, where Render represents the character's facial rendering sample.
[0102] One or more technical solutions provided in exemplary embodiments of the present disclosure can determine a facial rendering image based on facial expression coefficients of a preset facial expression and facial topography information included in a source image, such that the facial features of the rendered image combine the character's facial texture in the source image with the preset facial expression. In this case, expression control information is determined based on the rendered image, the target eye model corresponding to the preset facial expression, and the facial eye texture of the source image. This essentially involves fusing an eye region image into the rendered image. This process involves two aspects: first, fusing the target eye model corresponding to the preset facial expression into the rendered image in image form, ensuring that the eye orientation in the rendered image matches the facial expression in the rendered image; second, ensuring that the eye texture in the rendered image is consistent with the eye texture in the source image. Based on this, exemplary embodiments of the present disclosure, under the control of expression control information, can control the facial expression of the character in the target image to reflect the preset facial expression, and can also control the eye orientation and texture of the target character's eye region in the target image, thereby enhancing the authenticity and vividness of the target character in the generated target image.
[0103] Taking single-image driving as an example, during the single-image driving training stage, the eyeball area image can be directly obtained from the driving image. The eyeball area image contains not only eyeball shape information, but also eyeball texture information. By combining these two types of information, the expression sample control information is guaranteed to provide more accurate line of sight information for the single-image driving model, thereby ensuring that the trained single-image driving model can flexibly drive the line of sight direction of the character included in the source image. During the single-image driving test or reasoning stage, the basic eyeball model, preset facial expressions, and facial eyeball texture included in the source image can be used to fuse the eyeball area image on the facial rendering image, so that the expression control information obtained can provide accurate line of sight information for the single-image driving model, ensuring that the single-image driving model can flexibly drive the line of sight direction of the character included in the source image.
[0104] It can be seen that the exemplary embodiments of the present disclosure use different strategies to obtain eye area images in the single-image driven network training and testing (or inference) stages, ensuring that the single-image driven model can flexibly drive the character's line of sight direction included in the source image, thereby improving the flexibility and expressiveness of the single-image drive.
[0105] The above mainly introduces the solution provided by the embodiment of the present disclosure from the perspective of an electronic device. It is understandable that, in order to realize the above functions, the electronic device includes a hardware structure and / or software module corresponding to the execution of each function. Those skilled in the art should easily realize that, in combination with the units and algorithm steps of each example described in the embodiments disclosed herein, the present disclosure can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in a hardware or computer software driven hardware manner depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present disclosure.
[0106] The embodiments of the present disclosure can divide the functional units of the electronic device according to the above method examples. For example, each functional module can be divided according to each function, or two or more functions can be integrated into one processing module. The above integrated modules can be implemented in the form of hardware or in the form of software functional modules. It should be noted that the division of modules in the embodiments of the present disclosure is schematic and is only a logical function division. In actual implementation, there may be other division methods.
[0107] In the case of dividing each functional module according to each function, the exemplary embodiment of the present disclosure provides an image generation device, which can be an electronic device or a chip used in an electronic device. Figure 6 shows a schematic block diagram of the functional modules of the image generation device according to the exemplary embodiment of the present disclosure. As shown in Figure 6, the image generation device 600 includes:
[0108] Determination module 601, configured to determine a facial rendering image based on an expression coefficient of a preset facial expression and facial morphology information included in a source image, and determine expression control information based on the facial rendering image, a target eye model corresponding to the preset facial expression, and facial eye texture of the source image;
[0109] The generating module 602 is configured to generate a target image based on the expression control information, the source image, the three-dimensional facial information of the source image, and the three-dimensional facial information corresponding to the preset facial expression.
[0110] In one possible implementation, the determination module 601 is further used to determine eye rotation control parameters based on the expression coefficient of the preset facial expression, and control the posture of the basic eye model based on the eye rotation control parameters to obtain the target eye model.
[0111] In one possible implementation, the determination module 601 is used to obtain the description coefficient of the eye movement from the expression coefficient of the preset facial expression, and determine the eye rotation control parameter based on the description coefficient of the eye movement and the eye movement angle threshold.
[0112] In one possible implementation, the determination module 601 is used to fuse the target eyeball model and the facial rendering image to obtain a facial fusion image, and migrate the facial eyeball texture included in the source image to the facial eyeball included in the facial fusion image to obtain expression control information.
[0113] In one possible implementation, the determination module 601 is used to project the target eye model onto the facial eye area included in the facial rendering image based on the camera parameters corresponding to the source image and the three-dimensional position information of the target eye model to obtain a facial fusion image.
[0114] In one possible implementation, the determining module 601 is further configured to determine an initial three-dimensional facial model including a basic eye model based on the three-dimensional facial information corresponding to the source image and the basic eye model;
[0115] Based on the initial three-dimensional facial model including the basic eyeball model and the expression coefficient of the preset facial expression, a target three-dimensional facial model including the target eyeball model is obtained;
[0116] Based on the target facial three-dimensional model including the target eyeball model, the three-dimensional position information of the target eyeball model is determined.
[0117] In one possible implementation, the determination module 601 is used to migrate the facial eyeball texture to the facial eyeball contained in the facial fusion image based on the pupil key point mapping relationship between the facial fusion image and the source image, using the pupil key point mapping relationship to obtain the expression control information.
[0118] In one possible implementation, the target image is generated by an image-driven model, which includes a motion estimation network and an image generation network. The generation module 602 is used to input the facial three-dimensional posture information of the source image, the facial three-dimensional posture information corresponding to the preset facial expression, and the source image into the motion estimation network to obtain expression change estimation information, and input the expression change estimation information, the expression control information, and the source image into the image generation network to obtain the target image.
[0119] In one possible implementation, the training set of the image-driven model in the training phase includes source image samples, facial three-dimensional posture information of the source image samples, facial three-dimensional posture information of the driving image samples, and expression sample control information. The device also includes a training module 603, which is used to obtain character facial rendering samples based on the facial morphology information of the source image samples and the facial expression information of the driving image samples, and determine the expression sample control information based on the eye area image included in the driving image samples and the character facial rendering samples.
[0120] Figure 7 shows a schematic block diagram of a chip according to an exemplary embodiment of the present disclosure. As shown in Figure 7, the chip 700 includes one or more (including two) processors 701 and a communication interface 702. The communication interface 702 can support the electronic device to perform the data transmission and reception steps in the above method, and the processor 701 can support the electronic device to perform the data processing steps in the above method.
[0121] Optionally, as shown in FIG7 , the chip 700 further includes a memory 703 , which may include a read-only memory and a random access memory, and provides operation instructions and data to the processor. A portion of the memory may also include a non-volatile random access memory (NVRAM).
[0122] In some embodiments, as shown in FIG7 , the processor 701 performs corresponding operations by calling operation instructions stored in the memory (the operation instructions may be stored in the operating system). The processor 701 controls the processing operations of any one of the terminal devices, and the processor may also be referred to as a central processing unit (CPU). The memory 703 may include a read-only memory and a random access memory, and provides instructions and data to the processor 701. A portion of the memory 703 may also include NVRAM. For example, in an application, the memory, the communication interface, and the memory are coupled together through a bus system, wherein the bus system may include, in addition to the data bus, a power bus, a control bus, and a status signal bus, etc. However, for the sake of clarity, various buses are labeled as bus system 704 in FIG7 .
[0123] The methods disclosed in the above embodiments of the present disclosure can be applied to or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits in the processor or by software instructions. The above processor may be a general-purpose processor, a digital signal processor (DSP), an ASIC, a field-programmable gate array (FPGA), or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component. The methods, steps, and logic block diagrams disclosed in the embodiments of the present disclosure can be implemented or executed. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the methods disclosed in conjunction with the embodiments of the present disclosure can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium well-known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the memory, and the processor reads the information in the memory and, in conjunction with its hardware, completes the steps of the above method.
[0124] The exemplary embodiments of the present disclosure further provide an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor. The memory stores a computer program executable by the at least one processor, the computer program being configured to cause the electronic device to perform a method according to an exemplary embodiment of the present disclosure when executed by the at least one processor.
[0125] Exemplary embodiments of the present disclosure further provide a non-transitory computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor of a computer, is used to cause the computer to perform a method according to an embodiment of the present disclosure.
[0126] Exemplary embodiments of the present disclosure further provide a computer program product, including a computer program, wherein when the computer program is executed by a processor of a computer, it is used to cause the computer to perform the method according to the embodiment of the present disclosure.
[0127] With reference to Figure 8, a block diagram of an electronic device 800 that can serve as a server or client of the present disclosure will now be described, which is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer equipment, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.
[0128] As shown in Figure 8, electronic device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. In RAM 803, various programs and data required for the operation of electronic device 800 can also be stored. Computing unit 801, ROM 802, and RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to bus 804.
[0129] Multiple components within electronic device 800 are connected to I / O interface 805, including an input unit 806, an output unit 807, a storage unit 808, and a communication unit 809. Input unit 806 can be any type of device capable of inputting information into electronic device 800. Input unit 806 can receive input numeric or character information and generate key input signals related to user settings and / or function control of the electronic device. Output unit 807 can be any type of device capable of presenting information and may include, but is not limited to, a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. Storage unit 808 may include, but is not limited to, a magnetic disk or an optical disk. Communication unit 809 allows electronic device 800 to exchange information / data with other devices via computer networks such as the Internet and / or various telecommunication networks and may include, but is not limited to, a modem, a network card, an infrared communication device, a wireless communication transceiver and / or a chipset, such as a Bluetooth™ device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.
[0130] The computing unit 801 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 801 performs the various methods and processes described above. For example, in some embodiments, the method of the exemplary embodiments of the present disclosure may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 808. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 800 via the ROM 802 and / or the communication unit 809. In some embodiments, the computing unit 801 may be configured to perform the method of the exemplary embodiments of the present disclosure in any other appropriate manner (e.g., by means of firmware).
[0131] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0132] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the operation requested will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operations of the disclosed technical solution based on the prompt message.
[0133] As an optional but non-limiting implementation, in response to receiving a user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. Furthermore, the pop-up window may also contain a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0134] It is to be understood that the above-mentioned notification and user authorization acquisition process is only illustrative and does not constitute a limitation on the implementation of the present disclosure. Other methods that meet relevant laws and regulations may also be applied to the implementation of the present disclosure. The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that the functions / operations specified in the flowchart and / or block diagram are implemented when the program code is executed by the processor or controller. The program code can be executed entirely on the machine, partially on the machine, partially on the machine as a stand-alone software package and partially on a remote machine, or entirely on a remote machine or server.
[0135] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0136] As used in this disclosure, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, apparatus, and / or device (e.g., a magnetic disk, an optical disk, a memory, a programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including machine-readable media that receive machine instructions as machine-readable signals. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0137] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0138] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0139] Computer systems may include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The client and server relationship arises through computer programs running on the respective computers and having a client-server relationship to each other.
[0140] In the above embodiments, they can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, they can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present disclosure are performed in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, a terminal, a user device, or other programmable device. The computer program or instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer program or instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired or wireless means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium, such as a floppy disk, a hard disk, or a tape; it can also be an optical medium, such as a digital video disc (DVD); it can also be a semiconductor medium, such as a solid state drive (SSD).
[0141] Although the present disclosure has been described with reference to specific features and embodiments thereof, it will be apparent that various modifications and combinations may be made thereto without departing from the spirit and scope of the present disclosure. Accordingly, the present disclosure and the accompanying drawings are merely illustrative of the present disclosure as defined by the appended claims and are deemed to cover any and all modifications, variations, combinations or equivalents within the scope of the present disclosure. Obviously, those skilled in the art may make various modifications and variations to the present disclosure without departing from the spirit and scope of the present disclosure. Thus, the present disclosure is intended to include such modifications and variations if they fall within the scope of the claims of the present disclosure and their equivalents.
Claims
1. An image generation method, comprising: Determining a facial rendering image based on an expression coefficient of a preset facial expression and facial morphology information included in a source image; Determining expression control information based on the facial rendering image, a target eyeball model corresponding to the preset facial expression, and facial eyeball texture of the source image; Generating a target image based on the expression control information, the source image, three-dimensional facial information of the source image, and three-dimensional facial information corresponding to the preset facial expression.
2. The method according to claim 1, further comprising: Determining an eyeball rotation control parameter based on the expression coefficient of the preset facial expression; Controlling the posture of a basic eyeball model based on the eyeball rotation control parameter to obtain the target eyeball model.
3. The method according to claim 2, wherein, The determining the eyeball rotation control parameter based on the expression coefficient of the preset facial expression includes: Obtaining a description coefficient of eyeball movement from the expression coefficient of the preset facial expression; Determining the eyeball rotation control parameter based on the description coefficient of the eyeball movement and a movement angle threshold of the eyeball.
4. The method according to any one of claims 1 to 3, wherein, The determining the expression control information based on the facial rendering image, the target eyeball model corresponding to the preset facial expression, and the facial eyeball texture of the source image includes: Fusing the target eyeball model and the facial rendering image to obtain a facial fusion image; Transferring the facial eyeball texture included in the source image to the facial eyeball included in the facial fusion image to obtain the expression control information.
5. The method according to claim 4, wherein The fusing the target eyeball model and the facial rendering image to obtain a facial fusion image includes: Projecting the target eyeball model onto a facial eyeball area included in the facial rendering image based on camera parameters corresponding to the source image and three-dimensional position information of the target eyeball model to obtain the facial fusion image.
6. The method according to claim 5, further comprising: Determining an initial three-dimensional facial model containing a basic eyeball model based on three-dimensional facial information corresponding to the source image and the basic eyeball model; Obtaining a target three-dimensional facial model containing the target eyeball model based on the initial three-dimensional facial model containing the basic eyeball model and the expression coefficient of the preset facial expression; Determining three-dimensional position information of the target eyeball model based on the target three-dimensional facial model containing the target eyeball model.
7. The method according to any one of claims 4-6, wherein The transferring the facial eyeball texture included in the source image to the facial eyeball included in the facial fusion image to obtain the expression control information includes: Obtaining a pupil key point mapping relationship between the facial fusion image and the source image; Transferring the facial eyeball texture to the facial eyeball included in the facial fusion image by using the pupil key point mapping relationship to obtain the expression control information.
8. The method according to any one of claims 1-7, wherein, The target image is generated by an image driving model, and the generating the target image based on the expression control information, the source image, three-dimensional facial information of the source image, and three-dimensional facial information corresponding to the preset facial expression includes: Input the three-dimensional facial pose information of the source image, the three-dimensional facial pose information corresponding to the preset facial expression, and the source image into the motion estimation network to obtain expression change estimation information; Input the expression change estimation information, the expression control information, and the source image into the image generation network to obtain the target image.
9. The method according to claim 8, wherein The training set of the image driving model in the training phase includes source image samples, the three-dimensional facial pose information of the source image samples, the three-dimensional facial pose information of the driving image samples, and expression sample control information. The method for obtaining the expression sample control information includes: Based on the facial morphology information of the source image sample and the facial expression information of the driving image sample, obtain a character facial rendering sample; Based on the eyeball region image included in the driving image sample and the character facial rendering sample, determine the expression sample control information.
10. An image generation device, comprising: A determination module configured to determine a facial rendering image based on the expression coefficient of a preset facial expression and the facial morphology information included in the source image, and determine expression control information based on the facial rendering image, the target eyeball model corresponding to the preset facial expression, and the facial eyeball texture of the source image; A generation module configured to generate a target image based on the expression control information, the source image, the three-dimensional facial information of the source image, and the three-dimensional facial information corresponding to the preset facial expression.
11. An electronic device, comprising: A processor; And, A memory storing a program; Wherein, the program includes instructions that, when executed by the processor, cause the processor to execute the method according to any one of claims 1-9.
12. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause a computer to execute the method according to any one of claims 1-9.
Citation Information
Patent Citations
Method and device for generating image
CN111599002A
Image processing method and device, video processing method and device, equipment and storage medium
CN111882627A
Explicit eye model for avatar
US11010951B1
Face Image Generation With Pose And Expression Control
US20210097730A1