Image generation method and apparatus, and device and medium
By introducing oral content images during the image generation process and using expression control information, the problem of blurred oral content of the target character is solved, and the clarity and authenticity of oral state in the target image is improved.
Patent Information
- Application Number
- PCT/CN2024/141418
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-28
- Filing Date
- 2024-12-23
- Publication Date
- 2025-07-03
AI Technical Summary
When the existing image driving technology generates the target image, the oral content of the target character is blurred, resulting in poor video quality, especially when the source character is driven from the shut-up state to the mouth opening state, the oral content such as teeth is not clear enough.
By introducing oral content images during the image generation process, using expression control information to guide the target image, ensuring the improvement of oral content clarity of the target character in the mouth state, and combining facial rendering images and oral content images to generate the target image.
The authenticity and clarity of the oral state of the target character in the target image is improved, making the generated video more realistic.
Smart Images

Figure CN2024141418_03072025_PF_FP_ABST
Abstract
Description
Image generation method, device, equipment and medium
[0001] This application claims priority to Chinese patent application No. 202311841126.7 filed on December 28, 2023, and the contents of the above-mentioned Chinese patent application disclosure are hereby incorporated by reference in their entirety as a part of this application. Technical Field
[0002] The present invention relates to an image generation method, device, equipment and medium. Background Art
[0003] Image driving, also known as motion transfer, uses a driving video to drive a source image to generate a video whose appearance is consistent with the source image, but the main motion is consistent with the driving video. Summary of the Invention
[0004] According to one aspect of the present disclosure, there is provided an image generation method, comprising:
[0005] determining a facial rendering image based on a facial expression coefficient of a preset facial expression and facial morphology information included in the source image;
[0006] Determining expression control information based on the facial rendering image and the oral content image, wherein the preset facial expression at least includes a mouth expression in an open mouth state;
[0007] A target image is generated based on the expression control information, the source image, the three-dimensional facial information of the source image, and the three-dimensional facial information corresponding to the preset facial expression.
[0008] According to another aspect of the present disclosure, there is provided an image generating apparatus, comprising:
[0009] a determination module, configured to determine a facial rendering image based on a facial expression coefficient of a preset facial expression and facial morphology information included in a source image, and determine expression control information based on the facial rendering image and an oral content image, wherein the preset facial expression includes at least a mouth expression in an open mouth state;
[0010] A generation module is used to generate a target image based on the expression control information, the source image, the three-dimensional facial information of the source image, and the three-dimensional facial information corresponding to the preset facial expression.
[0011] According to another aspect of the present disclosure, there is provided an electronic device, comprising:
[0012] processor; and,
[0013] Memory for storing programs;
[0014] The program includes instructions, which, when executed by the processor, cause the processor to perform the method according to the exemplary embodiment of the present disclosure.
[0015] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium is provided, wherein the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to enable the computer to execute the method according to the exemplary embodiment of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Further details, features and advantages of the present disclosure are disclosed in the following description of exemplary embodiments in conjunction with the accompanying drawings, in which:
[0017] FIG1 is a schematic flow chart showing an image generating method according to an exemplary embodiment of the present disclosure;
[0018] FIG2 is a schematic diagram showing the generation principle of a facial fusion image, taking a mouth fusion image as an example, according to an exemplary embodiment of the present disclosure;
[0019] FIG3 shows a schematic diagram of the architecture of an image driving model of an exemplary embodiment of the present disclosure;
[0020] FIG4 shows a schematic diagram of a training set acquisition process of a parameter prediction model in a training phase according to an exemplary embodiment of the present disclosure;
[0021] FIG5 shows a schematic diagram of a geometric template of a mouth sample according to an exemplary embodiment of the present disclosure;
[0022] FIG6 shows a schematic diagram of a process for obtaining expression sample control information according to an exemplary embodiment of the present disclosure;
[0023] FIG7 shows a schematic block diagram of functional modules of an image generating apparatus according to an exemplary embodiment of the present disclosure;
[0024] FIG8 shows a schematic block diagram of a chip according to an exemplary embodiment of the present disclosure;
[0025] FIG9 shows a block diagram of an exemplary electronic device that can be used to implement the embodiments of the present disclosure. DETAILED DESCRIPTION
[0026] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0027] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.
[0028] The term "including" and its variations used in this document are open inclusions, that is, "including but not limited to". The term "based on" means "based at least in part on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one other embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the description below. It should be noted that the concepts of "first", "second", etc. mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0029] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".
[0030] In some embodiments, after using image-driven technology to drive a source image containing a source character, the generated target video may contain a target character that is essentially identical in appearance to the target character, but its oral cavity may be blurred. For example, if the source character in the source image has a closed mouth, the generated target image may not have a clear oral cavity if the source image is driven using single-image-driven technology. In this case, the oral cavity of the target image does not match the actual oral cavity, resulting in poor quality of the generated video.
[0031] The inventors discovered that when training a single-image driven model, the training data used came from various characters, and the oral contents of different characters varied greatly. As a result, when the trained single-image driven model generated a target image, the single-image driven model might average the oral contents of different characters, resulting in the generated target image including the target character's teeth in an open mouth state, which is blurred, and is not conducive to improving the authenticity of the target character in the target image.
[0032] In response to the above problems, an exemplary embodiment of the present disclosure provides an image generation method, which can add an oral content image to the expression control information, and guide various source images that are authorized for use through the expression control information, to ensure that the target character included in the final generated target image not only presents the preset facial expression, but also can improve the clarity of the oral content of the target character included in the target image, thereby improving the authenticity of the oral state of the target character included in the target image.
[0033] The source image of the exemplary embodiment of the present disclosure may refer to a character image required for use in the image generation process, and the character included in the source image may be defined as a source character. The source image may be a character image collected by the client, or a character image obtained by the client and authorized for use. After the client obtains the character image, the image generation method of the exemplary embodiment of the present disclosure may be executed by an electronic device.
[0034] The source image includes at least a facial image of a character, which may be an image with the mouth open or closed, without limitation. When the target character included in the target image, guided by the expression control information in the source image, is associated with the appearance of the source character in the source image, the target image is guided by the expression control information so that the target character's expression in the target image matches the preset facial expression while displaying a clearer representation of the oral cavity.
[0035] The source image can be used to provide appearance constraints for the target character displayed in the target image. For example, the facial texture of the character in the source image can be used to constrain the facial texture of the target character in the target image, so that the appearance of the source character in the source image is associated with the appearance of the target character in the target image. The appearance association here can mean that the appearance of the source character included in the source image is the same as the appearance of the target character included in the target image, or that the difference in appearance between the source character included in the source image and the target character included in the target image is within a controllable range, but is not limited to this.
[0036] When the electronic device is a user terminal installed on a client, after obtaining the character image, the client can use the character image as the source image and execute the image generation method on the user terminal. When the user terminal is a terminal with a display function, after the user terminal generates the target image using the image generation method, the target image can be displayed on the user terminal.
[0037] When the electronic device is a cloud server, the cloud server can be connected to a user terminal installed with a client via a network. In this case, the client uploads the character image obtained by the user terminal to the cloud server via a wired or wireless network. The cloud server uses the character image as a source image and executes the image generation method to obtain the target image. Finally, the cloud server returns the target image to the user terminal via the network for display on the user terminal.
[0038] For example, the user terminal of the exemplary embodiments of the present disclosure may be a mobile phone, a tablet computer, a wearable device, an in-vehicle device, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), and a wearable device based on augmented reality (AR) and / or virtual reality (VR) technology, etc.
[0039] For example, when the user terminal is a wearable device, the wearable device can also be a general term for wearable devices developed by applying wearable technology to intelligently design everyday wearables, such as glasses, gloves, watches, clothing, and shoes. A wearable device is a portable device that is worn directly on the body or integrated into the user's clothing or accessories.
[0040] These wearable devices are more than just hardware; they achieve powerful functionality through software support, data exchange, and cloud-based interaction. Broadly speaking, these smart wearables include devices that are comprehensive, large, and can function completely or partially independently of a smartphone, such as smartwatches and smart glasses. These devices also focus on a specific application and require integration with other devices, such as smartphones, such as various smart wristbands and smart jewelry for vital sign monitoring.
[0041] The network of the exemplary embodiments of the present disclosure may include one or more networks, and any suitable network is contemplated. By way of example and not limitation, one or more portions of a network may include an ad hoc network, an intranet, an extranet, a virtual private network (VPN), a local area network (LAN), a wireless local area network (WLAN), a wide area network (WAN), a wireless wide area network (WWAN), a metropolitan area network (MAN), a portion of the Internet, a portion of a public switched telephone network (PSTN), a cellular telephone network, or a combination of two or more of these.
[0042] FIG1 is a flow chart showing an image generation method according to an exemplary embodiment of the present disclosure. As shown in FIG1 , the image generation method according to an exemplary embodiment of the present disclosure may include:
[0043] Step 101: Determine a facial rendering image based on the facial expression coefficient of a preset facial expression and the facial morphology information included in the source image. It should be understood that the source image of the exemplary embodiment of the present disclosure can be a photo of a character, or a frame of a character image included in a video containing changes in the character's expression. Regardless of whether it is a photo of a character or a frame of a character image, the source character displayed as the source image can be a real character, such as a real person or various real animals, or various virtual characters, such as virtual characters or virtual animals designed through various licensed software or self-developed software. These virtual animals can be virtual animals that exist in the real world, or they can be conceived virtual animals that do not exist in the real world.
[0044] For example, in exemplary embodiments of the present disclosure, the facial morphology information included in the source image may represent facial mapping information included in the source image. Driven by the facial expression coefficient, the source image may be controlled so that the resulting facial rendering image has both the facial texture of the character in the source image and the characteristics of a preset facial expression. In this case, the character morphology included in the facial rendering image is the same as the source character, but the character expression included in the facial rendering image matches the preset facial expression.
[0045] For example, in exemplary embodiments of the present disclosure, facial topography information included in a source image can be determined by facial texture features and three-dimensional facial information of the source image. Various feature extraction algorithms, such as image segmentation algorithms, can be used to extract the facial texture features from the source image. Alternatively, various object detection algorithms can be used to detect the source image and obtain a three-dimensional facial mesh model as the three-dimensional facial information of the source image. For example, an object detection algorithm can be used to obtain multiple facial position coordinates, and based on these multiple facial position coordinates, three-dimensional facial information, namely, a three-dimensional facial mesh model, can be generated.
[0046] For example, the facial texture features of the source image of the exemplary embodiment of the present disclosure can represent the facial texture of the source character, such as detailed features such as the character's facial color and facial lines. The facial three-dimensional mesh model of the source image can represent the source character's facial three-dimensional mesh data. Using this facial three-dimensional mesh data, a basic facial model of the character in the source image can be constructed, which can represent the specific shape and surface topology of the face. For example, the basic facial model of the character can include a two-dimensional mesh composed of a series of polygonal (such as triangle) face patches. By connecting different polygonal fragments (taking triangles as an example, each triangular face patch consists of three vertices and three edges), a complex basic facial model of the character can be obtained.
[0047] The exemplary embodiments of the present disclosure combine the three-dimensional facial information of a source image with the facial texture features of the source image through texture mapping, thereby obtaining facial topography information included in the source image. For example, during texture mapping, the facial texture features can be used to map each fragment included in the character's facial basic model. The mapping method can be, for example, UV mapping, but is not limited to this.
[0048] When the facial expression coefficients of the preset facial expressions in the exemplary embodiment of the present disclosure include at least the mouth expression in an open mouth state, the character's mouth state in the facial rendering image appears open mouthed. Of course, the facial expression coefficients of the preset facial expressions may also include expression coefficients of other facial parts. Thus, there is a corresponding relationship between the facial expression coefficients of the preset facial expressions in the exemplary embodiment of the present disclosure and the facial rendering image.
[0049] For example, the facial expression coefficients of the aforementioned preset facial expressions may include not only the mouth expression coefficients in the open mouth state, but also the eye expression coefficients, jaw expression coefficients, eyebrow expression coefficients, cheek expression coefficients, nose expression coefficients, tongue expression coefficients, etc. When a facial rendering image is generated by combining the 52-dimensional facial expression coefficients with the facial morphology information included in the source image, the generated facial rendering image can present facial expressions of the eyes, jaw, eyebrows, cheeks, nose, and tongue, and the mouth expressions in the facial expressions can include mouth expressions in the open mouth state.
[0050] Step 102: Determine expression control information based on the facial rendering image and the oral content image. This can be accomplished by concatenating the facial rendering image and the oral content image. The resulting expression control information is essentially a facial fusion image. The oral content image can be a preset oral content image, retrieved from an electronic device or a storage device, or generated using a specific algorithm. For example, if Render1 represents the facial rendering image and mouth area1 represents the oral content image, the expression control information Control img1 = Render1 + mouth area1.
[0051] In practical applications, exemplary embodiments of the present disclosure can determine reconstruction reference information for an oral content image based on a facial expression coefficient of a preset facial expression. The oral content image can then be determined based on the reconstruction reference information and the dimensional information of the oral content image. This oral content image can include information such as teeth and even gums within the oral cavity, which will not be detailed here.
[0052] For example, the exemplary embodiments of the present disclosure can pre-store dimensional information of multiple preset facial expressions and multiple oral content images in a storage system. The electronic device can obtain the required reconstruction reference information of the oral content image and the dimensional information of the oral content image according to actual needs, thereby constructing a personalized oral content image and providing the required personalized guidance information for the oral content generation of the target character of the target image.
[0053] For example: An exemplary embodiment of the present disclosure can establish a mapping relationship between facial expression coefficients and reconstruction reference information of oral content images. After obtaining the facial expression coefficients of a preset facial expression, the reconstruction reference information of the oral content image is obtained through the mapping relationship, and then the reconstruction reference information of the oral content image and the dimensional information of the oral content image are used to determine the oral content image.
[0054] For another example, a parameter prediction model can be used to obtain reconstruction reference information for an oral content image based on the facial expression coefficients of a preset facial expression. In this case, the parameter prediction model can serve as a mapping bridge, establishing a corresponding relationship between the facial expression coefficients of the preset facial expression and the oral content image.
[0055] It can be seen that when generating the reconstruction reference information of the oral content image, the exemplary embodiment of the present disclosure can refer to the facial expression coefficient of the preset facial expression to establish a correspondence between the preset facial expression and the oral content image. At the same time, there is a correspondence between the facial expression coefficient of the preset facial expression and the facial rendering image. Therefore, the facial rendering image and the oral content image of the exemplary embodiment of the present disclosure both correspond to the same preset facial expression.
[0056] To reduce subsequent computational burden, the dimensional information of the oral content image can be configured to include the dimensional parameters of the reduced image in multiple dimensions, while the reconstruction reference information for the oral content image can include the coupling parameters of the reduced image in multiple dimensions. In this case, the oral content image represents a reduced-dimensional image of the mouth expression image corresponding to a preset facial expression, and this reduced-dimensional image is determined by the dimensional parameters and coupling parameters of the reduced image in multiple dimensions.
[0057] Exemplarily, when the dimensional information of the oral content image includes principal components of multiple dimensions, the reconstruction reference information of the oral content image may include principal component analysis (PCA) coefficients of multiple dimensions. The PCA coefficients and principal components correspond one to one. A dimensionality reduction image of the mouth expression image can be constructed using the PCA coefficients and principal components of multiple dimensions to obtain the oral content image.
[0058] For example, the n-dimensional image data of the oral content image can be decentralized, and then the covariance matrix of the n-dimensional image data of the oral content image can be obtained. The eigenvalues and eigenvectors of the covariance matrix of the n-dimensional image data are calculated, and the eigenvalues are sorted in descending order. The eigenvectors corresponding to the k features with the largest eigenvalues are obtained as the principal components of the oral content image in the k dimensions, and the top k eigenvalues are used as the PCA coefficients of the oral content image in the k dimensions.
[0059] The reconstructed reference information for the oral content image predicted by the parameter prediction model can be equivalent to the top k eigenvalues, while the dimensional information of the oral content image can be equivalent to the eigenvectors corresponding to the k features with the largest eigenvalues. Both n and k are positive integers, with n > k. K can be set as needed. For example, when k = 32, the image data dimension of the oral content image is determined to be 32-dimensional using the top k eigenvalues and the eigenvectors corresponding to the k features with the largest eigenvalues, thereby reducing the data processing pressure for subsequent target image generation.
[0060] FIG2 is a schematic diagram illustrating the generation principle of a facial fusion image, using a mouth fusion image as an example, in an exemplary embodiment of the present disclosure. As shown in FIG2 , when generating a mouth fusion image, the exemplary embodiment of the present disclosure can refer to the relevant description above and obtain a mouth rendering image 202 based on a mouth shape image 201 of a source image and the facial expression coefficient of an original preset facial expression. As can be seen from the mouth rendering image 202, it does not contain oral content. Therefore, the oral content image corresponding to the preset facial expression can be fused with the mouth rendering image 202 to obtain a mouth fusion image 203. This mouth fusion image 203 can be used as mouth expression control information and participate in the generation process of the target image.
[0061] Step 103: Generate a target image based on the expression control information, the source image, the 3D facial information of the source image, and the 3D facial information corresponding to the preset facial expression. Here, the exemplary embodiment of the present disclosure can generate the target image using an image-driven model.
[0062] FIG3 shows a schematic diagram of the architecture of an image-driven model of an exemplary embodiment of the present disclosure. As shown in FIG3 , the image-driven model 300 of the exemplary embodiment of the present disclosure may include: a motion estimation network 301 and an image generation network 302. In this case, based on expression control information, a source image, three-dimensional facial information of the source image, and three-dimensional facial information corresponding to a preset facial expression, a target image is generated, including: inputting the three-dimensional facial posture information of the source image, the three-dimensional facial posture information corresponding to the preset facial expression, and the source image into the motion estimation network 301 to obtain expression change estimation information, and inputting the expression change estimation information, the expression control information, and the source image into the image generation network 302 to obtain the target image.
[0063] The exemplary embodiments of the present disclosure can obtain the three-dimensional facial information corresponding to the preset facial expression and the three-dimensional facial information of the source image through a three-dimensional facial reconstruction method.
[0064] In some embodiments, the method for obtaining the facial three-dimensional information corresponding to the preset facial expression may include: obtaining a facial rendering image corresponding to the preset facial expression based on the facial morphology information included in the source image and the facial expression coefficient of the preset facial expression, and then performing facial reconstruction on the facial rendering image to obtain the facial three-dimensional information corresponding to the preset facial expression.
[0065] In other embodiments, the method for obtaining three-dimensional facial information corresponding to a preset facial expression may include: obtaining three-dimensional facial information of a source image from facial topography information included in a source image, and then driving the three-dimensional facial information of the source image using a facial expression coefficient corresponding to the preset facial expression, thereby obtaining three-dimensional facial information corresponding to the preset facial expression, i.e., three-dimensional facial information corresponding to the preset facial expression. It can be seen that both the three-dimensional facial information of the source image and the three-dimensional facial information corresponding to the preset facial expression can reflect the facial appearance features included in the source image, but there are individual differences in expression characteristics.
[0066] The three-dimensional facial information of the source image and the three-dimensional facial information corresponding to the preset facial expression can both be three-dimensional facial mesh models. The three-dimensional facial mesh model included in the source image can describe the state of the facial features included in the source image in three-dimensional space, while the three-dimensional facial mesh model corresponding to the preset facial expression can describe the state of the preset facial features in three-dimensional space. Facial features herein include not only static facial features but also dynamic facial features. Static facial features can be understood as facial features that barely change or undergo minor changes in a short period of time, such as facial appearance features, while dynamic features can be understood as facial features that undergo significant changes in a short period of time, such as facial expression features.
[0067] When the facial three-dimensional information of the source image is driven by the facial expression coefficient of the preset facial expression, the facial expression coefficient of the preset facial expression is actually used to drive the movement of each vertex included in the facial three-dimensional mesh model of the source image, so that the facial expression features of the facial three-dimensional mesh model corresponding to the preset facial expression finally obtained are close to or the same as the preset facial expression.
[0068] After obtaining the 3D facial information corresponding to the preset facial expression and the 3D facial information of the source image, the 3D facial pose information of the source image, the 3D facial pose information corresponding to the preset facial expression, and the source image can be spliced together and then input into a motion estimation network to obtain expression change estimation information. This expression change estimation information can represent the expression change information required to change from the facial expression in the original image to the preset facial expression.
[0069] Exemplarily, as shown in FIG3 , the image generation network 302 of the exemplary embodiment of the present disclosure may include an encoder 3021 and a decoder 3022, which may obtain expression control information based on facial texture features included in the source image, facial three-dimensional posture information corresponding to a preset facial expression, and an oral content image, and then input the expression control information and the source image into the encoder 3021, and use the encoder 3021 to encode the expression control information and the source image to obtain facial coding features of the source image, which integrate features of both the source image and the expression control information. Therefore, the decoder 3022 may be used to decode the facial coding features of the source image to obtain a target image.
[0070] As shown in Figure 3, in order to further improve the correlation between the facial expression of the target image and the preset facial expression, the facial coding features and expression change estimation information of the source image can be input into the decoder 3022 at the same time, so that when the decoder 3022 decodes the facial coding features of the source image, the expression change estimation information is referred to, so that the facial expression of the generated target image is closer to the preset facial expression.
[0071] When the preset facial expression includes at least the mouth expression in the open mouth state, the expression control information determined based on the facial rendering image and the oral content image has the dual control functions of facial expression guidance and oral content generation. Therefore, the exemplary embodiment of the present disclosure can, under the control of the expression control information, not only control the facial expression of the target character included in the target image to present the preset facial expression, but also improve the clarity of the oral content of the target character included in the target image, thereby improving the authenticity of the oral state of the target character included in the target image.
[0072] As a possible implementation method, when the exemplary embodiment of the present disclosure uses a parameter prediction model to obtain reconstruction reference information of oral content images, the parameter prediction model can be trained in advance to ensure that the trained parameter prediction model not only has a high reconstruction reference information prediction accuracy, but also has strong robustness.
[0073] When the measurement accuracy and robustness of the parameter prediction model are high, the trained parameter prediction model can accurately predict the reconstructed reference information of the oral content image corresponding to the facial expression coefficients of different preset facial expressions. Therefore, the facial expression coefficients of the preset facial expressions can be selected according to actual needs to generate oral content images with different expressions.
[0074] When the electronic device executing the image generation method of the exemplary embodiment of the present disclosure is a user terminal, the client can invoke the user terminal's image acquisition function to obtain a source image, or alternatively, obtain a permitted source image from a resource server via a network. In response to a selection of a preset facial expression, the client can determine a facial expression coefficient for the preset facial expression. Facial morphology information included in the source image can then be extracted from the source image. A rendered facial image can be determined based on the facial morphology information included in the source image and the facial expression coefficient for the preset facial expression. Simultaneously, the facial expression coefficient for the preset facial expression can be input into a parameter prediction model to predict reference information for reconstructing the oral content image.
[0075] As for the dimensional information of the oral content image, it can be the default dimensional information of the oral content image, or the client can obtain the required dimensional information of the oral content image in response to the selection operation of the dimensional information, and finally generate the target image based on the reconstruction reference information of the oral content image and the dimensional information of the oral content image.
[0076] When the electronic device that executes the image generation method of the exemplary embodiment of the present disclosure is a cloud server, the client can send an image generation request to the cloud server in response to the user's image generation instruction. The image generation request can carry description information and other additional information of the source image, such as preset facial expressions and dimensional information of the oral content image, etc. The cloud server can parse the image generation request and generate the target image based on the preset facial expressions, dimensional information of the oral content image and the source image in accordance with the relevant description above.
[0077] In one example, the client can invoke the image acquisition function of the user terminal to obtain the source image. When the client sends an image generation request to the cloud server, the client encapsulates the source image in the image generation request and sends it to the cloud server. In another example, the user image generation instruction can include the network address of the source image, and the client can encapsulate the link to the source image in the image generation request. In yet another example, the user image generation instruction can include required information about the source image, such as the source character attributes contained in the source image, such as gender, skin color, etc. The client can encapsulate the required source image information in the image generation request and send it to the cloud server.
[0078] The exemplary embodiment of the present disclosure may employ a supervised training method to train a parameter prediction model. During the training phase, the training set may include: facial expression sample coefficients and reconstructed reference samples corresponding to the facial expression sample coefficients. The reconstructed reference samples corresponding to the facial expression sample coefficients may serve as supervisory information to supervise the training of the parameter prediction model. The facial expression sample coefficients may include coefficients of facial expression samples in an open-mouth state. In this case, the trained parameter prediction model can predict the reconstructed reference samples corresponding to the facial expression coefficients in the open-mouth state, and using the reconstructed reference information and dimensional information of the oral content image, an oral content image in the open-mouth state can be constructed.
[0079] The following describes an example with reference to the flowchart of the training set acquisition process of the parameter prediction model in the training phase of the exemplary embodiment of the present disclosure shown in FIG4 .
[0080] As shown in FIG4 , the method for obtaining a training set in the training phase of the parameter prediction model of the exemplary embodiment of the present disclosure may include:
[0081] Step 401: Transfer the mouth sample texture corresponding to the facial expression coefficient sample to a geometric template of the mouth sample to obtain a first mouth image. Here, the mouth sample image can be a mouth sample image in an open mouth state, and the geometric template of the mouth sample can be a geometric template of a mouth sample in an open mouth state, to ensure that the trained parameter prediction model can predict reconstruction reference information for an oral content image suitable for an open mouth state.
[0082] In practical applications, exemplary embodiments of the present disclosure can use the image acquisition functions of user terminals such as VR, AR, and mobile phones to capture facial images with various mouth expressions, and upload these facial images to a cloud server. The cloud server can then extract facial expression coefficient samples for different mouth expressions and obtain facial image samples using a basic single-image driven model. Based on this, key point detection can be performed on the facial sample images to obtain mouth sample images.
[0083] In order to ensure the clear texture of the mouth samples and the effect of subsequent model training, the quality of the acquired facial image samples can be enhanced. For example, various machine models or resolution enhancement algorithms can be used to achieve super-resolution processing of the facial image samples, thereby obtaining clear facial image samples, ensuring clear mouth texture, and improving the parameter prediction accuracy of the subsequent parameter prediction model.
[0084] Exemplarily, the process of migrating the above-mentioned mouth sample texture to the geometric template of the mouth sample is essentially a process of mapping the geometric template of the mouth sample using the mouth sample texture. When the geometric template of the mouth sample includes multiple sub-regions, the first mouth image is determined by the geometric template of the mouth sample and the texture image located in each sub-region. Here, the texture image of each sub-region is determined by the sub-mouth texture corresponding to the sub-region in the mouth sample image. For example, the sub-mouth texture corresponding to each sub-region in the mouth sample texture can be obtained first, and then the sub-mouth texture corresponding to each sub-region in the mouth sample texture can be affine transformed to obtain the first mouth image.
[0085] For example, when performing region segmentation on the contour area of a mouth sample, the exemplary embodiments of the present disclosure can use multiple contour key points of multiple mouth samples as vertices to divide the contour area of the mouth sample into multiple sub-regions, each of which can be defined by at least three contour key points. For example, when using a triangulation algorithm to segment the contour area of a mouth sample, the resulting geometric template of the mouth sample can be as shown in Figure 5. In this case, each sub-region is a triangular sub-region defined by three contour key points.
[0086] Step 402: Determine the reconstruction reference samples corresponding to the facial expression sample coefficients based on the dimensionality reduction strategy of the first mouth image. Here, the specific forms of the extracted reduced-dimensionality image reconstruction reference information and reduced-dimensionality image dimension information vary depending on the dimensionality reduction strategy of the first mouth image.
[0087] Taking PCA dimensionality reduction as an example, the reconstructed reference samples corresponding to the facial expression sample coefficients can represent the eigenvalues of the image data covariance matrix of the first mouth image in k dimensions, that is, the PCA coefficients of the first mouth image. Correspondingly, the eigenvectors of the image data covariance matrix of the first mouth image in k dimensions can also be obtained and saved in advance as the dimensional information of the oral content image for generating the oral content image in the inference stage.
[0088] Considering that there is a one-to-one correspondence between the facial expression sample coefficients, the eigenvalues of the image data covariance matrix in k dimensions, and the eigenvectors of the image data covariance matrix in k dimensions, when the eigenvectors of the image data covariance matrix in k dimensions are saved in advance as the dimensional information of the oral content image, the facial expression sample coefficients can be saved at the same time, and a mapping relationship between the eigenvectors of the image data covariance matrix in k dimensions and the facial expression sample coefficients can be established.
[0089] In this case, when a preset facial expression is selected during the image generation process, the dimensional information of the oral content image can also be indirectly determined, thereby ensuring that the reconstructed oral content image meets the requirements.
[0090] For example, the n-dimensional image data of the first mouth image can be decentralized according to actual needs, and then the covariance matrix of the n-dimensional image data of the first mouth image can be calculated. The eigenvectors and eigenvalues of the covariance matrix of the n-dimensional image data are then obtained, and the eigenvectors corresponding to the k features with the largest eigenvalues are selected as principal components. The k largest eigenvalues are used as PCA coefficients of the first mouth image, where n and k both represent positive integers, and n>k.
[0091] When the reduced-dimensionality image reconstruction reference information includes the eigenvalues of the image data covariance matrix of the first mouth image in k dimensions, and the reduced-dimensionality image dimension information includes the eigenvectors of the image data covariance matrix of the first mouth image in k dimensions, the eigenvectors of the image data covariance matrix of the first mouth image in k dimensions can be combined based on the eigenvalues of the image data covariance matrix of the first mouth image in k dimensions to restore the first oral content image sample.
[0092] Compared to the first mouth image, the first oral content image sample requires a smaller amount of data. Therefore, using the PCA coefficients of the first mouth image as supervisory information to train the parameter prediction model can, on the one hand, reduce the computational complexity of the parameter prediction model, allowing the parameter prediction model to converge quickly during the training phase, and, on the other hand, reduce the complexity requirements of the parameter prediction model architecture, thereby making the parameter prediction model lightweight. For example, a model composed of a multi-layer perceptron can meet the data processing and training accuracy requirements of the parameter prediction model.
[0093] Finally, when training the parameter prediction model, by obtaining a large number of facial expression coefficient samples of different mouth expressions and generating corresponding facial image samples, the amount of data contained in the training set of the parameter prediction model in the training stage can be increased, and the prediction accuracy and robustness of the trained parameter prediction model can be improved, so that the parameter prediction model can be used to generate corresponding reconstruction reference information for different mouth expression coefficients.
[0094] In an exemplary embodiment of the present disclosure, the contour area of the mouth sample can be determined based on the multiple contour key points of the mouth sample, and then the contour area of the mouth sample can be divided into regions to determine the geometric template of the mouth sample. For example, any character image can be selected and the mouth contour key points of the character image can be detected. If the character image is a character video, the mouth key points can be detected for each frame of the character image included in the character video to obtain the position information of multiple candidate contour key points of the mouth sample, and the attributes of each candidate contour key point can be determined. The position information of the candidate contour key points with the same attributes of different frames of the character image can then be averaged to obtain multiple contour key points of the mouth sample.
[0095] To ensure the precision of mapping the geometric template of the mouth sample and enhance the authenticity of the resulting first mouth image, keypoint interpolation can be performed based on the multiple keypoints of the mouth sample's contour, increasing the number of keypoints. In this case, by dividing the mouth sample's contour area into multiple subregions using the multiple keypoints as vertices, the resulting number of subregions can be maximized, thereby achieving the goal of refined texture mapping and improving the clarity and authenticity of the resulting first mouth image.
[0096] Each mouth sample image can be divided into regions by referring to the region division method of the geometric template of the mouth sample to obtain multiple sub-mouth textures of the mouth sample image. The layout method of the sub-mouth texture in the mouth sample image is the same as the layout method of the sub-region in the geometric template of the mouth sample.
[0097] For example, using the same region partitioning strategy for both the mouth sample image and the mouth sample's contour region, the resulting sub-mouth textures are laid out in the same way within the mouth sample image as they are within the mouth sample's geometric template. Based on this, an affine transformation relationship can be established between the mouth sample image and the mouth sample's geometric template using their shared contour keypoints.
[0098] After determining the affine transformation relationship between the mouth sample image and the geometric template of the mouth sample, the sub-mouth texture corresponding to each sub-region in the mouth sample texture can be obtained first, and then the sub-mouth texture corresponding to each sub-region in the mouth sample texture can be affine transformed to obtain the first mouth image.
[0099] As a possible implementation method, the training set of the image-driven model of the exemplary embodiment of the present disclosure in the training stage includes source image samples, facial three-dimensional posture information of the source image samples, facial three-dimensional posture information of the driving image samples, and expression sample control information.
[0100] FIG6 shows a schematic diagram of a process for obtaining expression sample control information according to an exemplary embodiment of the present disclosure. As shown in FIG6 , the method for obtaining expression sample control information according to an exemplary embodiment of the present disclosure may include:
[0101] Step 601: Determine a second oral content image sample based on the mouth texture sample information and the geometric template of the mouth sample included in the driving image sample. The geometric template of the mouth sample can be generated by referring to the above description and will not be described in detail here.
[0102] In practical applications, when determining the second oral content image sample, the exemplary embodiment of the present disclosure can migrate the mouth texture sample of the driving image sample to the geometric template of the mouth sample to obtain the second mouth image, and then determine the second oral content image sample based on the second mouth image.
[0103] Exemplarily, when the geometric template of the mouth sample includes multiple sub-regions, the second mouth image includes the geometric template of the mouth sample and a texture image located in each sub-region, and the texture image of each sub-region is determined by the sub-mouth texture corresponding to the sub-region in the driving image sample.
[0104] Exemplarily, the driving image sample of the exemplary embodiment of the present disclosure can be a video containing an arbitrary character (the arbitrary character is defined as a reference character). For each frame of video image contained in the video, a mouth key detection can be performed. If a mouth contour key point is detected in a certain frame of video image, it means that the frame of video image includes the reference character. When the mouth key point detection of all frames of video image of the video is completed, the mouth contour key points included in each frame of video image can be obtained. At this time, the mouth contour key points included in each frame of video image can be attributed, and then the mouth contour key points with the same attributes included in different frames of video image can be aligned.
[0105] By acquiring the mouth contour key points of the same attributes for each video frame, the reference character's mouth texture region within each video frame can be determined. Therefore, referring to the method described above for mapping the geometric template of the mouth sample using the mouth sample texture, the reference character's mouth texture within each video frame is used to map the geometric template of the mouth sample, thereby obtaining the second oral content image sample corresponding to each video frame. For simplicity, the process of obtaining the second oral content image sample corresponding to each video frame will not be detailed here.
[0106] Step 602: Obtain a facial rendering sample based on the facial morphology information of the source image sample and the facial expression information of the driving image sample. The facial morphology information of the source image can be referred to in the previous description and will not be repeated here. It should be understood that in the exemplary embodiments of the present disclosure, the source image sample can be any frame in a video, and the driving image sample can be any frame in the video other than the source image sample.
[0107] The facial expression information of the driving image sample may be facial expression coefficients of a reference character included in the driving image sample. When the facial topography information included in the source image is determined by facial texture features and three-dimensional facial information of the source image, the facial expression coefficients of the reference character may be used to drive the three-dimensional facial information of the source image, so that the facial expression included in the obtained facial rendering sample is correlated with the facial expression of the reference character included in the driving image.
[0108] When the driving image sample is a video containing a reference character, the character's facial expression coefficients for each frame of the video can be obtained. Then, based on the facial morphology information of the source image sample and the facial expression coefficients for each frame of the video image, a corresponding facial rendering sample is determined for each frame of the video image. This facial rendering sample essentially adds a facial rendering image sample, containing the same character appearance as the source character included in the source image sample, but with facial expressions related to the facial expressions of the reference character included in the driving image.
[0109] Step 603: Obtain expression sample control information based on the facial rendering sample and the second oral content image sample. This may be achieved by fusing the character's facial rendering sample and the second oral content image sample. The expression sample control information obtained is essentially a fused facial image formed by the character's facial rendering sample and the second oral content image sample. For example, if Render2 represents the facial rendering sample and mouth area2 represents the second oral content image sample, the expression sample control information Control img2 = Render2 + mouth area2.
[0110] To improve the accuracy of target image generation during the inference phase, the geometric template of the mouth sample used in both the acquisition of expression sample control information and the training of the parameter prediction model during the inference phase is the same. This geometric template setting ensures that the expression sample control information input to the image-driven model during the training phase and the expression control information input during the inference phase are aligned, thereby improving the accuracy of the target character included in the target image.
[0111] To reduce the weight of the image-driven model and improve its convergence speed, the second oral content image sample in the exemplary embodiment of the present disclosure can represent a dimensionality-reduced image of the second mouth image. In other words, the second mouth image can be subjected to dimensionality reduction processing to reduce the data volume of the second oral content image sample.
[0112] In practical applications, based on the dimensionality reduction strategy of the second mouth image, the dimensionality reduction image reconstruction reference information and dimensionality reduction image information of the second mouth image can be determined, and then based on the dimensionality reduction image reconstruction reference information and dimensionality reduction image information of the second mouth image, the second oral content image sample can be obtained.
[0113] The dimensionality reduction strategy for the second mouth image in the exemplary embodiment of the present disclosure can be set according to actual needs. Taking PCA dimensionality reduction as an example, the n-dimensional image data of the second mouth image can be decentralized. The covariance matrix of the n-dimensional image data of the second mouth image can then be calculated. The eigenvectors and eigenvalues of the covariance matrix of the n-dimensional image data are then obtained. The eigenvectors corresponding to the k features with the largest eigenvalues are selected as the k-dimensional principal components of the second mouth image. The k largest eigenvalues are then used as the PCA coefficients of the second mouth image, where n and k both represent positive integers, and n>k.
[0114] On this basis, the k-dimensional PCA coefficients of the second mouth image can be used to combine the k-dimensional principal components of the second mouth image to obtain the second oral content image sample. This not only reduces the data volume of the second oral content image sample, but also removes noise data, reducing its impact on training accuracy. Furthermore, reducing the data volume also reduces the model complexity requirements for the image-driven model, accelerating the convergence of the image-driven model and completing model training as quickly as possible.
[0115] One or more technical solutions provided in the exemplary embodiments of the present disclosure determine a facial rendering image based on the facial expression coefficient of a preset facial expression and the facial morphology information included in the source image, so that the facial features of the facial rendering image have both the facial texture of the character in the source image and the preset facial expression. When the preset facial expression includes at least the mouth expression in the open mouth state, the expression control information can be determined based on the facial rendering image and the oral content image, so that the expression control information has the dual control functions of facial expression guidance and oral content generation. Based on this, the exemplary embodiments of the present disclosure can, under the control of the expression control information, not only control the facial expression of the character in the target image to present the preset facial expression, but also improve the clarity of the oral content of the character in the target image, thereby improving the authenticity of the oral state of the character in the target image.
[0116] In some embodiments, the fuzzy problem of oral content generation in the single-image driven problem is actually a control problem for oral content generation. To address this problem, the exemplary embodiments of the present disclosure adopt the same concept in the training and testing stages of the image-driven model, that is, in the training stage, by adding a second oral content image sample to the facial rendering sample, and in the inference stage, by adding the oral content image to the expression control information, it is ensured that the control information input to the image-driven model contains the oral content image, so as to use the control information (expression sample control information in the training stage, expression control information in the inference stage) to drive the source image sample or source image to generate oral content, such as a target image with clear oral texture.
[0117] Furthermore, during the training phase of the image-driven model, PCA analysis is performed on the second oral content image sample corresponding to the driving image sample to obtain k-dimensional PCA coefficients and k-dimensional principal components. These k-dimensional PCA coefficients and k-dimensional principal components are then used to reconstruct the second oral content image sample, thereby achieving dimensionality reduction of the second oral content image sample. During the testing phase of the image-driven model, image enhancement can be performed based on the first mouth image corresponding to the facial expression coefficient sample to obtain a high-resolution first mouth image. PCA analysis is then performed on the first mouth image to obtain PCA coefficients corresponding to the facial expression sample coefficients. These coefficients are used as supervisory information for the parameter prediction model, allowing the trained parameter prediction model to predict PCA coefficients based on the facial expression coefficients. Thus, during the testing phase of the image-driven model, the trained parameter prediction model can be used to predict the PCA coefficients corresponding to the preset facial expressions, thereby reconstructing the oral content image. This is then incorporated into the rendered facial image as expression control information, improving the clarity of the image-driven model and the stability of the oral content.
[0118] The above mainly introduces the solution provided by the embodiment of the present disclosure from the perspective of an electronic device. It is understandable that, in order to realize the above functions, the electronic device includes a hardware structure and / or software module corresponding to the execution of each function. Those skilled in the art should easily realize that, in combination with the units and algorithm steps of each example described in the embodiments disclosed herein, the present disclosure can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in a hardware or computer software driven hardware manner depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present disclosure.
[0119] The embodiments of the present disclosure can divide the functional units of the electronic device according to the above method examples. For example, each functional module can be divided according to each function, or two or more functions can be integrated into one processing module. The above integrated modules can be implemented in the form of hardware or in the form of software functional modules. It should be noted that the division of modules in the embodiments of the present disclosure is schematic and is only a logical function division. In actual implementation, there may be other division methods.
[0120] In the case of dividing each functional module according to each function, the exemplary embodiments of the present disclosure provide an image generation device, which can be an electronic device or a chip used in an electronic device. Figure 7 shows a schematic block diagram of the functional modules of the image generation device according to the exemplary embodiments of the present disclosure. As shown in Figure 7, the image generation device 700 includes:
[0121] A determination module 701 is configured to determine a facial rendering image based on a facial expression coefficient of a preset facial expression and facial morphology information included in a source image, and to determine expression control information based on the facial rendering image and an oral content image, wherein the preset facial expression includes at least a mouth expression in an open mouth state;
[0122] The generating module 702 is configured to generate a target image based on the expression control information, the source image, the three-dimensional facial information of the source image, and the three-dimensional facial information corresponding to the preset facial expression.
[0123] In a possible implementation, the determination module 701 is further used to determine reconstruction reference information of the oral content image based on the facial expression coefficient of the preset facial expression, and to determine the oral content image based on the reconstruction reference information of the oral content image and the dimensional information of the oral content image.
[0124] In a possible implementation, the determination module 701 is configured to obtain a preset mouth expression coefficient from a facial expression coefficient of a preset facial expression, and determine reconstruction reference information of the oral content image based on the preset mouth expression coefficient.
[0125] In a possible implementation, the oral content image represents a dimensionality-reduced image of the mouth expression image corresponding to the preset mouth expression coefficient, and the dimensionality-reduced image is determined by dimensionality parameters and coupling parameters of the dimensionality-reduced image in multiple dimensions;
[0126] The dimensional information of the oral content image includes dimensional parameters of the reduced-dimensional image in multiple dimensions, and the reconstruction reference information of the oral content image includes coupling parameters of the reduced-dimensional image in multiple dimensions.
[0127] In one possible implementation, the reconstruction reference information of the oral content image is obtained by a parameter prediction model based on the facial expression coefficient of the preset facial expression, and the training set of the parameter prediction model in the training stage includes: facial expression sample coefficients and reconstructed reference samples corresponding to the facial expression sample coefficients, and the facial expression sample coefficients include mouth expression coefficients in an open mouth state.
[0128] In one possible implementation, the device also includes an acquisition module 703, which is used to migrate the mouth sample texture corresponding to the facial expression coefficient sample to the geometric template of the mouth sample to obtain a first mouth image, and determine the reconstructed reference sample corresponding to the facial expression sample coefficient based on the dimensionality reduction strategy of the first mouth image.
[0129] In one possible implementation, the geometric template of the mouth sample includes multiple sub-regions, the first mouth image includes the geometric template of the mouth sample, and a texture image located in each of the sub-regions, and the texture image of each sub-region is determined by the sub-mouth texture corresponding to the sub-region in the mouth sample image.
[0130] In one possible implementation, the target image is generated by an image-driven model, which includes a motion estimation network and an image generation network. The generation module 702 is used to input the facial three-dimensional posture information of the source image, the facial three-dimensional posture information corresponding to the preset facial expression, and the source image into the motion estimation network to obtain expression change estimation information, and input the expression change estimation information, the expression control information, and the source image into the image generation network to obtain the target image.
[0131] In one possible implementation, the training set of the image-driven model in the training phase includes source image samples, facial three-dimensional posture information of the source image samples, facial three-dimensional posture information of the driving image samples, and expression sample control information. The device also includes an acquisition module 703, which is used to determine the second oral content image sample based on the mouth texture sample information and the geometric template of the mouth sample included in the driving image sample, obtain the character facial rendering sample based on the facial morphology information of the source image sample and the facial expression information of the driving image sample, and obtain the expression sample control information based on the character facial rendering sample and the second oral content image sample.
[0132] In one possible implementation, the acquisition module 703 is used to migrate the mouth texture sample of the driving image sample to the geometric template of the mouth sample to obtain the second mouth image, and determine the second oral content image sample based on the second mouth image, where the second oral content image sample represents a reduced-dimensional image of the second mouth image.
[0133] In one possible implementation, the geometric template of the mouth sample includes multiple sub-regions, the second mouth image includes the geometric template of the mouth sample, and a texture image located in each of the sub-regions, and the texture image of each sub-region is determined by the sub-mouth texture corresponding to the sub-region in the driving image sample.
[0134] In one possible implementation, the acquisition module 703 is used to extract the reduced-dimensionality image reconstruction reference information and reduced-dimensionality image dimension information of the second mouth image, and obtain a second oral content image sample based on the reduced-dimensionality image reconstruction reference information and reduced-dimensionality image dimension information of the second mouth image.
[0135] Figure 8 shows a schematic block diagram of a chip according to an exemplary embodiment of the present disclosure. As shown in Figure 8, the chip 800 includes one or more (including two) processors 801 and a communication interface 802. The communication interface 802 can support the electronic device to perform the data transmission and reception steps in the above method, and the processor 801 can support the electronic device to perform the data processing steps in the above method.
[0136] Optionally, as shown in FIG8 , the chip 800 further includes a memory 803 , which may include a read-only memory and a random access memory, and provides operating instructions and data to the processor. A portion of the memory may also include a non-volatile random access memory (NVRAM).
[0137] In some embodiments, as shown in FIG8 , the processor 801 performs corresponding operations by calling operation instructions stored in the memory (the operation instructions may be stored in the operating system). The processor 801 controls the processing operations of any one of the terminal devices, and the processor may also be referred to as a central processing unit (CPU). The memory 803 may include a read-only memory and a random access memory, and provides instructions and data to the processor 801. A portion of the memory 803 may also include NVRAM. For example, in an application, the memory, the communication interface, and the memory are coupled together through a bus system, wherein the bus system may include, in addition to the data bus, a power bus, a control bus, and a status signal bus, etc. However, for the sake of clarity, various buses are labeled as bus system 804 in FIG8 .
[0138] The methods disclosed in the above embodiments of the present disclosure can be applied to or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits in the processor or by software instructions. The above processor may be a general-purpose processor, a digital signal processor (DSP), an ASIC, a field-programmable gate array (FPGA), or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component. The methods, steps, and logic block diagrams disclosed in the embodiments of the present disclosure can be implemented or executed. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the methods disclosed in conjunction with the embodiments of the present disclosure can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium well-known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the memory, and the processor reads the information in the memory and, in conjunction with its hardware, completes the steps of the above method.
[0139] The exemplary embodiments of the present disclosure further provide an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor. The memory stores a computer program executable by the at least one processor, the computer program being configured to cause the electronic device to perform a method according to an exemplary embodiment of the present disclosure when executed by the at least one processor.
[0140] Exemplary embodiments of the present disclosure further provide a non-transitory computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor of a computer, is used to cause the computer to perform a method according to an embodiment of the present disclosure.
[0141] Exemplary embodiments of the present disclosure further provide a computer program product, including a computer program, wherein when the computer program is executed by a processor of a computer, it is used to cause the computer to perform the method according to the embodiment of the present disclosure.
[0142] With reference to Figure 9, a block diagram of an electronic device 900 that can serve as a server or client of the present disclosure will now be described, which is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.
[0143] As shown in Figure 9, electronic device 900 includes a computing unit 901, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 902 or a computer program loaded from a storage unit 908 into a random access memory (RAM) 903. In RAM 903, various programs and data required for the operation of electronic device 900 can also be stored. Computing unit 901, ROM 902 and RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to bus 904.
[0144] As shown in Figure 9, multiple components within electronic device 900 are connected to an I / O interface 905, including an input unit 906, an output unit 907, a storage unit 908, and a communication unit 909. Input unit 906 can be any type of device capable of inputting information into electronic device 900. Input unit 906 can receive input digital or character information and generate key input signals related to user settings and / or function control of the electronic device. Output unit 907 can be any type of device capable of presenting information and may include, but is not limited to, a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. Storage unit 908 may include, but is not limited to, a magnetic disk or an optical disk. Communication unit 909 allows electronic device 900 to exchange information / data with other devices via computer networks such as the Internet and / or various telecommunication networks, and may include, but is not limited to, a modem, a network card, an infrared communication device, a wireless communication transceiver, and / or a chipset, such as a Bluetooth™ device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.
[0145] As shown in Figure 9, the computing unit 901 can be various general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 901 performs the various methods and processes described above. For example, in some embodiments, the method of the exemplary embodiments of the present disclosure may be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as a storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 900 via the ROM 902 and / or the communication unit 909. In some embodiments, the computing unit 901 can be configured to perform the method of the exemplary embodiments of the present disclosure in any other appropriate manner (e.g., by means of firmware).
[0146] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0147] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0148] As used in this disclosure, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, apparatus, and / or device (e.g., a magnetic disk, an optical disk, a memory, a programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including machine-readable media that receive machine instructions as machine-readable signals. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0149] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0150] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0151] Computer systems may include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The client and server relationship arises through computer programs running on the respective computers and having a client-server relationship to each other.
[0152] In the above embodiments, they can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, they can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present disclosure are performed in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, a terminal, a user device, or other programmable device. The computer program or instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer program or instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired or wireless means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium, such as a floppy disk, a hard disk, or a tape; it can also be an optical medium, such as a digital video disc (DVD); it can also be a semiconductor medium, such as a solid state drive (SSD).
[0153] Although the present disclosure has been described with reference to specific features and embodiments thereof, it will be apparent that various modifications and combinations may be made thereto without departing from the spirit and scope of the present disclosure. Accordingly, this specification and the drawings are merely illustrative of the present disclosure as defined by the appended claims and are deemed to cover any and all modifications, variations, combinations or equivalents within the scope of the present disclosure. Obviously, those skilled in the art may make various modifications and variations to the present disclosure without departing from the spirit and scope of the present disclosure. Thus, the present disclosure is intended to include such modifications and variations if they fall within the scope of the claims of the present disclosure and their equivalents.
Claims
1. An image generation method, comprising: Determining a facial rendering image based on a facial expression coefficient of a preset facial expression and facial morphology information included in a source image; Determining expression control information based on the facial rendering image and an oral content image, where the preset facial expression at least includes a mouth expression in an open - mouth state; Generating a target image based on the expression control information, the source image, the facial three - dimensional information of the source image, and the facial three - dimensional information corresponding to the preset facial expression.
2. The method according to claim 1, further comprising: Determining reconstruction reference information of the oral content image based on the facial expression coefficient of the preset facial expression; Determining the oral content image based on the reconstruction reference information of the oral content image and the dimension information of the oral content image.
3. The method according to claim 2, wherein The oral content image represents a dimensionality - reduced image of a mouth expression image corresponding to the preset facial expression, and the dimensionality - reduced image is determined by dimensionality parameters and coupling parameters in multiple dimensions; The dimension information of the oral content image includes the dimensionality parameters of the dimensionality - reduced image in multiple dimensions, and the reconstruction reference information of the oral content image includes the coupling parameters of the dimensionality - reduced image in multiple dimensions.
4. The method according to claim 2, wherein, The reconstruction reference information of the oral content image is obtained by a parameter prediction model based on the facial expression coefficient of the preset facial expression. The training set of the parameter prediction model in the training stage includes: facial expression sample coefficients and reconstruction reference samples corresponding to the facial expression sample coefficients, and the facial expression sample coefficients include mouth expression coefficients in an open - mouth state.
5. The method according to claim 4, wherein, The method for obtaining the training set of the parameter prediction model in the training stage includes: Transferring the mouth sample texture corresponding to the facial expression coefficient sample to the geometric template of the mouth sample to obtain a first mouth image; Determining the reconstruction reference sample corresponding to the facial expression sample coefficient based on the dimensionality - reduction strategy of the first mouth image.
6. The method according to claim 5, wherein, The geometric template of the mouth sample includes multiple sub - regions, and the first mouth image is determined by the geometric template of the mouth sample and texture images located in each sub - region. Each texture image in the sub - region is determined by the sub - mouth texture corresponding to the sub - region in the mouth sample texture.
7. According to the method according to any one of claims 1 to 6, wherein, The target image is generated by an image - driving model, and the image - driving model includes a motion estimation network and an image generation network. The generating the target image based on the expression control information, the source image, the facial three - dimensional information of the source image, and the facial three - dimensional information corresponding to the preset facial expression includes: Inputting the facial three - dimensional pose information of the source image, the facial three - dimensional pose information corresponding to the preset facial expression, and the source image into the motion estimation network to obtain expression change estimation information; Inputting the expression change estimation information, the expression control information, and the source image into the image generation network to obtain the target image.
8. The method according to claim 7, wherein, The training set of the image - driving model in the training stage includes source image samples, facial three - dimensional pose information of source image samples, facial three - dimensional pose information of driving image samples, and expression sample control information. The method for obtaining the expression sample control information includes: Determine a second oral content image sample based on the mouth texture sample information and the geometric template of the mouth sample included in the driving image sample; Obtain a character facial rendering sample based on the facial morphology information of the source image sample and the facial expression information of the driving image sample; Obtain expression sample control information based on the character facial rendering sample and the second oral content image sample.
9. The method according to claim 8, wherein, The determining the second oral content image sample based on the mouth texture sample information and the geometric template of the mouth sample included in the driving image sample includes: Transfer the mouth texture sample of the driving image sample to the geometric template of the mouth sample to obtain a second mouth image; Based on the second mouth image, determine the second oral content image sample, and the second oral content image sample represents a dimensionality-reduced image of the second mouth image.
10. The method according to claim 8 or 9, wherein, The geometric template of the mouth sample includes a plurality of sub-regions, and the second mouth image is determined by the geometric template of the mouth sample and the texture images located in each of the sub-regions, and the texture image of each sub-region is determined by the corresponding sub-mouth texture of the sub-region in the driving image sample.
11. The method according to claim 8, wherein, The determining the second oral content image sample based on the second mouth image includes: Extract the dimensionality-reduced image reconstruction reference information and the dimensionality-reduced image dimension information of the second mouth image; Obtain a second oral content image sample based on the dimensionality-reduced image reconstruction reference information and the dimensionality-reduced image dimension information of the second mouth image.
12. An image generation device, comprising: A determination module configured to determine a facial rendering image based on the facial expression coefficient of a preset facial expression and the facial morphology information included in the source image, and determine expression control information based on the facial rendering image and the oral content image, where the preset facial expression includes at least a mouth expression in an open-mouth state; A generation module configured to generate a target image based on the expression control information, the source image, the three-dimensional facial information of the source image, and the three-dimensional facial information corresponding to the preset facial expression.
13. An electronic device, comprising: A processor; And, A memory storing a program; Wherein, the program includes instructions that, when executed by the processor, cause the processor to execute the method according to any one of claims 1 to 11.
14. A non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Method and device for generating image
CN111599002A
Face image replaying method and system, electronic equipment and storage medium
CN116310146A
Method and apparatus for three-dimensional reconstruction of a human head for rendering a human image
WO2023085624A1
Cited By
Method and device for generating mouth shape video of digital human
CN120640101A