Image generation method and device, equipment and medium
By determining the expression control information of the facial rendered image and oral content image, combining the facial three-dimensional information of the source image and the three-dimensional information of the preset facial expression, the target image is generated, and the problem of blurred oral content in the prior art is solved, and the authenticity and clarity of the oral state of the character in the target image is improved.
Patent Information
- Application Number
- CN202311841126.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-28
- Publication Date
- 2025-07-08
AI Technical Summary
The oral content of characters in videos generated by existing image-driven technologies is blurred, resulting in poor video quality and inability to match the real oral state.
By determining the expression control information of the facial rendered image and oral content image, combining the facial three-dimensional information of the source image and the three-dimensional information of the preset facial expression, the target image is generated to ensure that the facial expression of the target character matches the preset expression and improve the clarity of the oral content.
Improves the authenticity and clarity of the oral state of the character in the target image, ensuring a clear match for the oral content of the character in the generated video.
Smart Images

Figure CN120279157A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular, to an image generation method, apparatus, device, and medium. Background Art
[0002] Image driving, also known as Motion Transfer, uses a driving video to drive a source image to generate a video. The appearance of this video is the same as that of the source image, but the main body movement is the same as that of the driving video.
[0003] In the related art, after driving a source image using image driving technology, the oral content of the character in the generated video is blurred, and it does not match the oral state in the real situation, resulting in poor quality of the generated video. Summary of the Invention
[0004] According to one aspect of the present disclosure, there is provided an image generation method, including:
[0005] Determining a facial rendering image based on a facial expression coefficient of a preset facial expression and facial morphology information included in a source image;
[0006] Determining expression control information based on the facial rendering image and an oral content image, where the preset facial expression includes at least a mouth expression in an open - mouth state;
[0007] Generating a target image based on the expression control information, the source image, three - dimensional facial information of the source image, and three - dimensional facial information corresponding to the preset facial expression.
[0008] According to another aspect of the present disclosure, there is provided an image generation apparatus, including:
[0009] A determination module, configured to determine a facial rendering image based on a facial expression coefficient of a preset facial expression and facial morphology information included in a source image, and determine expression control information based on the facial rendering image and an oral content image, where the preset facial expression includes at least a mouth expression in an open - mouth state;
[0010] A generation module, configured to generate a target image based on the expression control information, the source image, three - dimensional facial information of the source image, and three - dimensional facial information corresponding to the preset facial expression.
[0011] According to another aspect of the present disclosure, there is provided an electronic device, including:
[0012] A processor; and,
[0013] A memory storing a program;
[0014] Among them, the program includes instructions, which when executed by the processor cause the processor to execute the method according to the exemplary embodiments of the present disclosure.
[0015] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium, characterized in that the non-transitory computer-readable storage medium stores computer instructions for causing the computer to execute the method according to the exemplary embodiments of the present disclosure.
[0016] One or more technical solutions provided in the exemplary embodiments of the present disclosure determine a facial rendering image based on a facial expression coefficient of a preset facial expression and facial morphology information included in a source image, so that the facial features of the facial rendering image have both the character facial texture of the source image and the preset facial expression. When the preset facial expression includes at least a mouth expression in an open-mouth state, expression control information can be determined based on the facial rendering image and an oral content image, so that the expression control information has a dual control function of facial expression guidance and oral content generation. Based on this, the exemplary embodiments of the present disclosure can, under the control of the expression control information, not only control the facial expression of the character in the target image to present the preset facial expression, but also improve the clarity of the oral content of the character in the target image, thereby enhancing the authenticity of the oral state of the character in the target image. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In the following description of the exemplary embodiments with reference to the drawings, more details, features, and advantages of the present disclosure are disclosed. In the drawings:
[0018] Figure 1 A flowchart showing the image generation method according to the exemplary embodiments of the present disclosure is shown;
[0019] Figure 2 A schematic diagram showing the generation principle of a facial fusion image taking a mouth fusion image as an example according to the exemplary embodiments of the present disclosure is shown;
[0020] Figure 3 A schematic diagram showing the architecture of an image-driven model according to the exemplary embodiments of the present disclosure is shown;
[0021] Figure 4 A flowchart showing the process of obtaining a training set in the training stage of a parameter prediction model according to the exemplary embodiments of the present disclosure is shown;
[0022] Figure 5 A schematic diagram showing the geometric template of a mouth sample according to the exemplary embodiments of the present disclosure is shown;
[0023] Figure 6 A flowchart showing the process of obtaining expression sample control information according to the exemplary embodiments of the present disclosure is shown;
[0024] Figure 7Shows a schematic block diagram of the functional modules of an image generation device according to an exemplary embodiment of the present disclosure;
[0025] Figure 8 Shows a schematic block diagram of a chip according to an exemplary embodiment of the present disclosure;
[0026] Figure 9 Shows a structural block diagram of an exemplary electronic device that can be used to implement the embodiments of the present disclosure. Detailed implementation manners
[0027] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.
[0028] It should be understood that the various steps recorded in the method embodiments of the present disclosure can be executed in different orders and / or executed in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.
[0029] As used herein, the term "including" and its variants are open-ended inclusions, that is, "including but not limited to". The term "based on" is "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description. It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order of the functions performed by these devices, modules or units or their interdependence.
[0030] It should be noted that the modifications of "one" and "plural" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that, unless clearly indicated otherwise in the context, it should be understood as "one or more".
[0031] In the related art, after driving a source image containing a source character using image driving technology, although the target character included in the generated target video is basically the same in appearance as the target character, the oral content thereof may be blurred. For example, when the source character in the source image is a character with a closed mouth, if the single-image driving technology is used to drive the source image, the oral content of the target character in the generated target image is not clear enough. In this case, the oral state of the target image does not match the oral state in the real situation, resulting in poor quality of the generated video.
[0032] The inventors found that when training a single-image driving model, the training data used comes from various characters, and the oral contents of different characters vary greatly. As a result, when the trained single-image driving model generates a target image, the single-image driving model may average the oral contents of different characters, resulting in blurred oral contents such as teeth in the target character included in the generated target image when the target character is in an open-mouth state, which is not conducive to improving the authenticity of the target character in the target image.
[0033] To address the above problems, an exemplary embodiment of the present disclosure provides an image generation method, which can add an oral content image in the expression control information, guide various source images that are permitted to be used through authorization by the expression control information, and ensure that the target character included in the finally generated target image not only presents a preset facial expression in terms of facial expression, but also can improve the clarity of the oral content of the target character included in the target image, thereby enhancing the authenticity of the oral state of the target character included in the target image.
[0034] The source image in the exemplary embodiment of the present disclosure may refer to a character image required in the image generation process. The character included in the source image can be defined as the source character. The source image can be a character image collected by a client or a character image obtained by the client and permitted to be used through authorization. After the client obtains the character image, the image generation method in the exemplary embodiment of the present disclosure can be executed by an electronic device.
[0035] The above source image at least includes a facial image of the character. The facial image can be a facial image in an open-mouth state or a facial image in a closed-mouth state, which is not limited herein. When the target image guided by the expression control information of the source image includes a target character that is associated with the appearance of the source character included in the source image, the difference is that the target image is guided by the expression control information, so that the facial expression of the target character included in the target image matches the preset facial expression while showing relatively clear oral content.
[0036] The above source image can be used to provide a morphological constraint for the target character shown in the target image. For example, the facial texture of the target character in the target image can be constrained by the facial texture of the character in the source image, so that the appearance of the source character in the source image is associated with the appearance of the target character in the target image. The appearance association of the characters here can mean that the appearance of the source character included in the source image is the same as the appearance of the target character included in the target image, or the difference in the appearance of the source character included in the source image and the target character included in the target image is within a controllable range, but is not limited thereto.
[0037] When the electronic device is a user terminal installed by the client, after the client obtains the character image, the character image can be used as the source image, and the image generation method is executed on the user terminal. When the user terminal is a terminal with a display function, after the target image is generated by the image generation method on the user terminal, the target image can be displayed through the user terminal.
[0038] When the electronic device is a cloud server, the cloud server can be communicatively connected to the user terminal installed with the client through a network. In this case, the client uploads the obtained character image to the cloud server through the user terminal via a wired network or a wireless network. The cloud server uses the character image as the source image and executes the image generation method to obtain the target image. Finally, the cloud server returns the target image to the user terminal through the network and displays it through the user terminal.
[0039] Exemplarily, the user terminal of the exemplary embodiment of the present disclosure can be a mobile phone, a tablet computer, a wearable device, a vehicle-mounted device, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), and a wearable device based on augmented reality (AR) and / or virtual reality (VR) technologies, etc.
[0040] Exemplarily, when the user terminal is a wearable device, the wearable device can also be a general term for devices that use wearable technology to intelligently design daily wear and develop wearable devices, such as glasses, gloves, watches, clothing, and shoes. A wearable device is a portable device that is directly worn on the body or integrated into the user's clothes or accessories.
[0041] The above wearable device is not just a hardware device, but also realizes powerful functions through software support, data interaction, and cloud interaction. Broadly speaking, wearable intelligent devices include those with complete functions and large sizes that can achieve complete or partial functions without relying on a smart phone, such as smart watches or smart glasses, as well as those that only focus on a certain type of application function and need to cooperate with other devices such as smart phones, such as various smart bracelets and smart jewelry for physical sign monitoring.
[0042] The network of the exemplary embodiments of the present disclosure may include one or more networks, and any suitable network may be considered. By way of example and not limitation, one or more parts of the network may include an ad hoc network, an intranet, an extranet, a virtual private network (VPN), a local area network (LAN), a wireless local area network (WLAN), a wide area network (WAN), a wireless wide area network (WWAN), a metropolitan area network (MAN), a part of the Internet, a part of the public switched telephone network (PSTN), a cellular telephone network, or a combination of two or more of these.
[0043] Figure 1 A flowchart showing the image generation method of the exemplary embodiments of the present disclosure is shown. As Figure 1 shown, the image generation method of the exemplary embodiments of the present disclosure may include:
[0044] Step 101: Determine a facial rendering image based on the facial expression coefficient of a preset facial expression and the facial morphology information included in the source image. It should be understood that the source image of the exemplary embodiments of the present disclosure may be a photo of a certain character, or a frame of a character image included in a video with character expression changes. Whether it is a character photo or a frame of a character image, the source character shown as the source image may be a real character, such as a real person or various real animals, or various virtual characters, such as virtual people or virtual animals designed through various licensed software or self-developed software. These virtual animals may be virtual animals existing in the real world or virtual animals conceived but not existing in the real world.
[0045] Exemplarily, the facial morphology information included in the source image of the exemplary embodiment of the present disclosure may represent the facial mapping information included in the source image. Driven by the facial expression coefficient, the source image may be at least controlled so that the face of the obtained facial rendering image has both the facial texture of the character in the source image and the preset facial expression. At this time, the character morphology included in the facial rendering image is the same as the source character, but the character expression included in it matches the preset facial expression.
[0046] For example: the facial morphology information included in the source image of the exemplary embodiment of the present disclosure can be determined by the facial texture features of the source image and the facial 3D information of the source image. The facial texture features of the source image can be extracted from the source image by various feature extraction algorithms such as image segmentation algorithms, and the source image can also be detected by various target detection algorithms to obtain a facial 3D mesh model as the facial 3D information of the source image. For example: the target detection algorithm can be used to obtain multiple facial position coordinates, and the facial 3D information, i.e., the facial 3D mesh model, can be generated based on the multiple facial position coordinates.
[0047] Exemplarily, the facial texture features of the source image of the exemplary embodiment of the present disclosure can represent the facial texture of the source character, such as the facial color of the character, the facial lines of the character, and other detailed features. The facial three-dimensional mesh model of the source image can represent the facial three-dimensional mesh data of the source character. Through the facial three-dimensional mesh data, a basic facial model of the character of the source image can be constructed, which can represent the specific shape and surface topology of the face. For example: the basic facial model of the character can include a two-dimensional mesh composed of a series of polygonal (such as triangle) face patches. By connecting different multi-deformable fragments (taking triangles as an example, each triangle face patch consists of three vertices and three edges), a complex basic facial model of the character can be obtained.
[0048] The exemplary embodiment of the present disclosure combines the facial 3D information of the source image and the facial texture features of the source image by means of texture mapping, thereby obtaining the facial morphology information included in the source image. Exemplarily, when performing texture mapping, the facial texture features can be used to map each fragment included in the basic model of the character's face, and the mapping method can be UV mapping, etc., but is not limited thereto.
[0049] When the facial expression coefficient of the preset facial expression of the exemplary embodiment of the present disclosure includes at least the mouth expression in the open mouth state, the mouth state of the character in the facial rendering image is in the open mouth state. Of course, the facial expression coefficient of the preset facial expression can also include expression coefficients of other parts of the face. It can be seen that there is a corresponding relationship between the facial expression coefficient of the preset facial expression of the exemplary embodiment of the present disclosure and the facial rendering image.
[0050] For example, the facial expression coefficients of the above-mentioned preset facial expressions may include not only the mouth expression coefficients in the open-mouth state, but also eye expression coefficients, jaw expression coefficients, eyebrow expression coefficients, cheek expression coefficients, nose expression coefficients, tongue expression coefficients, etc. When generating a facial rendering image by combining the 52-dimensional facial expression coefficients with the facial morphology information included in the source image, the generated facial rendering image can present the expression movements of parts such as the eyes, jaw, eyebrows, cheeks, nose, and tongue, and the mouth expression of the expression movement may include the mouth expression in the open-mouth state.
[0051] Step 102: Determine expression control information based on the facial rendering image and the oral content image. Here, the facial rendering image and the oral content image can be spliced, and the obtained expression control information is essentially a facial fusion image. The oral content image here can be a preset oral content image, which can be a preset oral content image obtained by the electronic device from the storage device, or an oral content image generated by a certain algorithm. For example: When Render1 represents the facial rendering image and mouth area1 represents the oral content image, the expression control information Control img1 = Render1 + moutharea1.
[0052] In practical applications, the exemplary embodiments of the present disclosure can determine the reconstruction reference information of the oral content image based on the facial expression coefficients of the preset facial expressions, and then determine the oral content image based on the reconstruction reference information of the oral content image and the dimension information of the oral content image. The oral content image may include information such as teeth and even gums in the oral cavity, which will not be elaborated here.
[0053] Exemplarily, the exemplary embodiments of the present disclosure can pre-store the dimension information of multiple preset facial expressions and multiple oral content images in the storage system. The electronic device can obtain the reconstruction reference information and the dimension information of the required oral content image according to actual needs, so as to construct the oral content image personalized and provide the required personalized guidance information for the generation of the oral content of the target role in the target image.
[0054] For example: The exemplary embodiments of the present disclosure can establish a mapping relationship between the facial expression coefficients and the reconstruction reference information of the oral content image. After obtaining the facial expression coefficients of the preset facial expression, the reconstruction reference information of the oral content image can be obtained through this mapping relationship, and then the oral content image can be determined by using the reconstruction reference information of the oral content image and the dimension information of the oral content image.
[0055] For another example, a parameter prediction model can be used to obtain reconstruction reference information of an oral cavity content image based on the facial expression coefficients of a preset facial expression. At this time, the parameter prediction model can serve as a mapping bridge, such that there is a corresponding relationship between the facial expression coefficients of the preset facial expression and the oral cavity content image.
[0056] It can be seen that when generating the reconstruction reference information of the oral cavity content image, the exemplary embodiments of the present disclosure can refer to the facial expression coefficients of the preset facial expression to establish a corresponding relationship between the preset facial expression and the oral cavity content image. At the same time, there is a corresponding relationship between the facial expression coefficients of the preset facial expression and the facial rendering image. Therefore, the facial rendering image and the oral cavity content image in the exemplary embodiments of the present disclosure both correspond to the same preset facial expression.
[0057] To reduce the subsequent calculation pressure, it can be set that the dimension information of the oral cavity content image includes the dimension parameters of the dimensionality-reduced image in multiple dimensions, and the reconstruction reference information of the oral cavity content image can include the coupling parameters of the dimensionality-reduced image in multiple dimensions. In this case, the oral cavity content image represents the dimensionality-reduced image of the mouth expression image corresponding to the preset facial expression, and the dimensionality-reduced image is determined by the dimension parameters and the coupling parameters of the dimensionality-reduced image in multiple dimensions.
[0058] Exemplarily, when the dimension information of the oral cavity content image includes the principal components in multiple dimensions, the reconstruction reference information of the oral cavity content image can include the principal component analysis (PCA) coefficients in multiple dimensions. The PCA coefficients and the principal components correspond one by one. The dimensionality-reduced image of the mouth expression image can be constructed through the PCA coefficients and the principal components in multiple dimensions, so as to obtain the oral cavity content image.
[0059] For example: the n-dimensional image data of the oral cavity content image can be centralized, then the covariance matrix of the n-dimensional image data of the oral cavity content image is obtained, the eigenvalues and eigenvectors of the covariance matrix of the n-dimensional image data are calculated, sorted according to the order of the eigenvalues from large to small, and the eigenvectors corresponding to the k largest eigenvalues are obtained as the principal components of the oral cavity content image in k dimensions, and the k largest eigenvalues sorted at the front are used as the PCA coefficients of the oral cavity content image in k dimensions.
[0060] The reconstruction reference information of the oral cavity content image predicted by the above parameter prediction model can be equivalent to the k largest eigenvalues sorted at the front, and the dimension information of the oral cavity content image can be equivalent to the eigenvectors corresponding to the k largest eigenvalues. Both n and k represent positive integers, and n>k. The value of k can be set according to actual needs. For example: when k = 32, the image data dimension of the oral cavity content image determined by using the k largest eigenvalues sorted at the front and the eigenvectors corresponding to the k largest eigenvalues is 32 dimensions, thereby reducing the data processing pressure for subsequent target image generation.
[0061] Figure 2 It shows a schematic diagram of the generation principle of a facial fusion image taking the mouth fusion image as an example in an exemplary embodiment of the present disclosure. As Figure 2 shown, in the exemplary embodiment of the present disclosure, when generating the mouth fusion image, reference can be made to the relevant description above. Based on the mouth morphology image 201 of the source image and the facial expression coefficient of the original preset facial expression, the mouth rendering image 202 can be obtained. It can be seen from the mouth rendering image 202 that it does not contain oral content. Therefore, the oral content image corresponding to the preset facial expression can be fused with the mouth rendering image 202 to obtain the mouth fusion image 203. This mouth fusion image 203 can be used as mouth expression control information and participate in the generation process of the target image.
[0062] Step 103: Generate a target image based on the expression control information, the source image, the three-dimensional facial information of the source image, and the three-dimensional facial information corresponding to the preset facial expression. Here, in the exemplary embodiment of the present disclosure, an image driving model can be used to generate the target image.
[0063] Figure 3 It shows a schematic diagram of the architecture of the image driving model in an exemplary embodiment of the present disclosure. As Figure 3 shown, the image driving model 300 in the exemplary embodiment of the present disclosure may include: a motion estimation network 301 and an image generation network 302. In this case, generating a target image based on the expression control information, the source image, the three-dimensional facial information of the source image, and the three-dimensional facial information corresponding to the preset facial expression includes: inputting the three-dimensional facial pose information of the source image, the three-dimensional facial pose information corresponding to the preset facial expression, and the source image into the motion estimation network 301 to obtain expression change estimation information, and inputting the expression change estimation information, the expression control information, and the source image into the image generation network 302 to obtain the target image.
[0064] In the exemplary embodiment of the present disclosure, the three-dimensional facial information corresponding to the preset facial expression and the three-dimensional facial information of the source image can be obtained through three-dimensional facial reconstruction.
[0065] In some embodiments, the method for obtaining the three-dimensional facial information corresponding to the preset facial expression may include: based on the facial morphology information included in the source image and the facial expression coefficient of the preset facial expression, obtaining the facial rendering image corresponding to the preset facial expression, and then performing facial reconstruction on the facial rendering image to obtain the three-dimensional facial information corresponding to the preset facial expression.
[0066] In some other embodiments, the method for obtaining the three-dimensional facial information corresponding to the preset facial expression may include: obtaining the three-dimensional facial information of the source image from the facial topography information included in the source image, and then driving the three-dimensional facial information of the source image by the facial expression coefficient of the preset facial expression, so as to obtain the three-dimensional facial information corresponding to the preset facial expression, that is, the three-dimensional facial information corresponding to the preset facial expression. It can be seen that both the three-dimensional facial information of the source image and the three-dimensional facial information corresponding to the preset facial expression can reflect the facial appearance features included in the source image, but there are individual differences in the expression features.
[0067] Both the three-dimensional facial information of the source image and the three-dimensional facial information corresponding to the preset facial expression described above may be three-dimensional facial mesh models. Among them, the three-dimensional facial mesh model included in the source image can describe the state of the facial features included in the source image in the three-dimensional space, while the three-dimensional facial mesh model corresponding to the preset facial expression can describe the state of the preset facial features in the three-dimensional space. The facial features here not only include facial static features, but also may include facial dynamic features. Facial static features can be understood as the facial characteristics that hardly change or change slightly in a short period of time, such as facial appearance features, while dynamic features can be understood as the facial features that change greatly in a short period of time, such as facial expression features.
[0068] When driving the three-dimensional facial information of the source image by the facial expression coefficient of the preset facial expression, it is actually to drive each vertex included in the three-dimensional facial mesh model of the source image to move by the facial expression coefficient of the preset facial expression, so that the facial expression features of the finally obtained three-dimensional facial mesh model corresponding to the preset facial expression are close to or the same as the preset facial expression.
[0069] After obtaining the three-dimensional facial information corresponding to the preset facial expression and the three-dimensional facial information of the source image, the three-dimensional facial pose information of the source image, the three-dimensional facial pose information corresponding to the preset facial expression, and the source image can be spliced and then input into the motion estimation network to obtain the expression change estimation information. The expression change estimation information can represent the expression change information required to change from the facial expression of the original image to the preset facial expression.
[0070] Exemplarily, such as Figure 3As shown, the image generation network 302 of the exemplary embodiment of the present disclosure may include an encoder 3021 and a decoder 3022. It can obtain expression control information based on the facial texture features included in the source image, the facial three-dimensional pose information corresponding to the preset facial expression, and the oral content image. Then, the expression control information and the source image are input into the encoder 3021, and the encoder 3021 encodes the expression control information and the source image to obtain the facial encoding features of the source image. The facial encoding features integrate the features of both the source image and the expression control information. Therefore, the decoder 3022 can decode the facial encoding features of the source image to obtain the target image.
[0071] As Figure 3 shown, in order to further improve the correlation between the facial expression of the target image and the preset facial expression, the facial encoding features of the source image and the expression change estimation information can be input into the decoder 3022 simultaneously. During the process of the decoder 3022 decoding the facial encoding features of the source image, it refers to the expression change estimation information, so that the facial expression of the generated target image is closer to the preset facial expression.
[0072] When the preset facial expression includes at least the mouth expression in the open - mouth state, the expression control information determined based on the facial rendering image and the oral content image has a dual control function of facial expression guidance and oral content generation. Therefore, in the exemplary embodiment of the present disclosure, under the control of the expression control information, it can not only control the facial expression of the target character included in the target image to present the preset facial expression, but also improve the clarity of the oral content of the target character included in the target image, and further enhance the authenticity of the oral state of the target character included in the target image.
[0073] As a possible implementation manner, in the exemplary embodiment of the present disclosure, when obtaining the reconstruction reference information of the oral content image using the parameter prediction model, the parameter prediction model can be pre - trained to ensure that the trained parameter prediction model not only has high prediction accuracy of the reconstruction reference information, but also has strong robustness.
[0074] When the prediction accuracy and robustness of the parameter prediction model are high, the trained parameter prediction model can accurately predict the reconstruction reference information of the oral content image corresponding to the facial expression coefficients of different preset facial expressions. Therefore, the facial expression coefficients of the preset facial expression can be selected according to actual needs, so as to generate oral content images with various expressions.
[0075] When the electronic device executing the image generation method of the exemplary embodiment of the present disclosure is a user terminal, the client can call the image acquisition function of the user terminal to obtain the source image, or it can be a source image that is allowed to be used and obtained from a resource server through the network. In response to a selection operation of a preset facial expression, the facial expression coefficient of the preset facial expression can be determined, and then the facial morphology information included in the source image is extracted from the source image. Based on the facial morphology information included in the source image and the facial expression coefficient of the preset facial expression, a facial rendering image is determined. At the same time, the facial expression coefficient of the preset facial expression can be input into a parameter prediction model to predict the reconstruction reference information of the oral cavity content image.
[0076] As for the dimension information of the oral cavity content image, it can be the default dimension information of the oral cavity content image, or it can be the dimension information of the oral cavity content image required by the client in response to a selection operation of the dimension information. Finally, based on the reconstruction reference information of the oral cavity content image and the dimension information of the oral cavity content image, a target image is generated.
[0077] When the electronic device executing the image generation method of the exemplary embodiment of the present disclosure is a cloud server, the client can, in response to a user image generation instruction, send an image generation request to the cloud server. The image generation request can carry the description information of the source image and other additional information, such as the preset facial expression and the dimension information of the oral cavity content image, etc. The cloud server can parse the image generation request and generate a target image based on the preset facial expression, the dimension information of the oral cavity content image, and the source image according to the relevant description above.
[0078] In one example, the client can call the image acquisition function of the user terminal to obtain the source image, and when sending the image generation request to the cloud server, encapsulate the source image in the image generation request and send it to the cloud server. In another example, the user image generation instruction can include the network address of the source image, and the client can encapsulate the link of the source image in the image generation request. In yet another example, the user image generation instruction can include the requirement information of the source image, such as the source role attributes included in the source image, such as gender, skin color, etc., and the client can encapsulate the requirement information of the source image in the image generation request and send it to the cloud server.
[0079] The method for training the parameter prediction model in the exemplary embodiment of the present disclosure may be a supervised training method. The training set in the training phase may include: the facial expression sample coefficients and the reconstruction reference samples corresponding to the facial expression sample coefficients. The reconstruction reference samples corresponding to the facial expression sample coefficients may be used as supervision information to supervise the training process of the parameter prediction model. The facial expression sample coefficients may include the facial expression sample coefficients in the open - mouth state. In this case, the trained parameter prediction model may predict the reconstruction reference samples corresponding to the facial expression coefficients in the open - mouth state, and by using the reconstruction reference information and the dimensional information of the oral cavity content image, the oral cavity content image in the open - mouth state may be constructed.
[0080] The following will be described by way of example in conjunction with Figure 4 the schematic diagram of the training set acquisition process of the parameter prediction model in the exemplary embodiment of the present disclosure shown below.
[0081] As Figure 4 shown, the method for obtaining the training set of the parameter prediction model in the exemplary embodiment of the present disclosure in the training phase may include:
[0082] Step 401: Transfer the mouth sample texture corresponding to the facial expression coefficient sample to the geometric template of the mouth sample to obtain the first mouth image. Here, the mouth sample image may be the mouth sample image in the open - mouth state, and the geometric template of the mouth sample may be the geometric template of the mouth sample in the open - mouth state, so as to ensure that the trained parameter prediction model can predict the reconstruction reference information of the oral cavity content image adapted to the open - mouth state.
[0083] In practical applications, the exemplary embodiment of the present disclosure may collect facial images under various mouth expressions through the image acquisition function of user terminals such as VR, AR, and mobile phones, and upload these facial images to the cloud server. The cloud server may extract the facial expression coefficient samples under different mouth expressions and obtain the facial image samples through a basic single - image driving model. On this basis, key - point detection may be performed on the facial sample images to obtain the mouth sample images.
[0084] To ensure the clarity of the mouth sample texture and the subsequent model training effect, quality enhancement may be performed on the obtained facial image samples. For example, various machine models or resolution enhancement algorithms may be used to perform super - resolution processing on the facial image samples, so as to obtain clear facial image samples, ensure the clarity of the mouth texture, and improve the parameter prediction accuracy of the subsequent parameter prediction model.
[0085] Exemplarily, the process of migrating the above-mentioned mouth sample texture to the geometric template of the mouth sample is essentially a process of mapping the mouth sample texture to the geometric template of the mouth sample. When the geometric template of the mouth sample includes multiple sub-regions, the first mouth image is determined by the geometric template of the mouth sample and the texture images located in each sub-region. Here, the texture image of each sub-region is determined by the sub-mouth texture corresponding to the sub-region in the mouth sample image. For example: First, the sub-mouth texture corresponding to each sub-region in the mouth sample texture can be obtained, and then an affine transformation is performed on the sub-mouth texture corresponding to each sub-region in the mouth sample texture to obtain the first mouth image.
[0086] Exemplarily, when the exemplary embodiments of the present disclosure divide the contour region of the mouth sample, the contour region of the mouth sample can be divided into multiple sub-regions with multiple contour key points of multiple mouth samples as vertices, and each sub-region can be determined by at least three contour key points. For example: When the triangulation algorithm is used to divide the contour region of the mouth sample, the obtained geometric template of the mouth sample can be as Figure 5 shown. At this time, each sub-region is a triangular sub-region and is determined by three contour key points.
[0087] Step 402: Based on the dimensionality reduction strategy of the first mouth image, determine the reconstruction reference sample corresponding to the facial expression sample coefficient. Here, according to the different dimensionality reduction strategies of the first mouth image, the specific forms of the dimensionality reduction image reconstruction reference information and the dimensionality information of the dimensionality reduction image are different.
[0088] Taking PCA dimensionality reduction as an example, the reconstruction reference sample corresponding to the facial expression sample coefficient can represent the eigenvalues of the image data covariance matrix of the first mouth image in k dimensions, that is, the PCA coefficients of the first mouth image. Correspondingly, the eigenvectors of the image data covariance matrix of the first mouth image in k dimensions can also be obtained and saved in advance as the dimensionality information of the oral cavity content image for generating the oral cavity content image in the inference stage.
[0089] Considering that there is a one-to-one correspondence between the facial expression sample coefficient, the eigenvalues of the image data covariance matrix in k dimensions, and the eigenvectors of the image data covariance matrix in k dimensions, therefore, when saving the eigenvectors of the image data covariance matrix in k dimensions as the dimensionality information of the oral cavity content image in advance, the facial expression sample coefficient can be saved at the same time, and the mapping relationship between the eigenvectors of the image data covariance matrix in k dimensions and the facial expression sample coefficient can be established. In this case, when selecting a preset facial expression during the image generation process, the dimensionality information of the oral cavity content image can also be indirectly determined, so as to ensure that the reconstructed oral cavity content image meets the requirements.
[0090] Exemplarily, the n-dimensional image data of the first mouth image can be decentralized according to actual needs, then the covariance matrix of the n-dimensional image data of the first mouth image is calculated, and then the eigenvectors and eigenvalues of the covariance matrix of the n-dimensional image data are obtained. The eigenvectors corresponding to the k largest eigenvalues are selected as the principal components. The k largest eigenvalues are used as the PCA coefficients of the first mouth image. Both n and k represent positive integers, and n > k.
[0091] When the reduced-dimensional image reconstruction reference information includes the eigenvalues of the covariance matrix of the image data of the first mouth image in k dimensions, and the reduced-dimensional image dimension information includes the eigenvectors of the covariance matrix of the image data of the first mouth image in k dimensions, the eigenvectors of the covariance matrix of the image data of the first mouth image in k dimensions can be combined based on the eigenvalues of the covariance matrix of the image data of the first mouth image in k dimensions, so as to restore the first oral content image sample.
[0092] Compared with the first mouth image, the amount of data required for the first oral content image sample is relatively small. Therefore, using the PCA coefficients of the first mouth image as the supervision information to train the parameter prediction model can, on the one hand, reduce the computational amount of the parameter prediction model, so that the parameter prediction model can converge quickly during the training stage, and on the other hand, it can also reduce the requirements for the architecture complexity of the parameter prediction model, thereby making the parameter prediction model lightweight. For example: A model composed of a multi-layer perceptron can meet the data processing and training accuracy requirements of the parameter prediction model.
[0093] Finally, when training the parameter prediction model, by obtaining a large number of facial expression coefficient samples of different mouth expressions and generating corresponding facial image samples, the amount of data included in the training set of the parameter prediction model during the training stage can be increased, and the prediction accuracy and robustness of the trained parameter prediction model can be improved, so that the parameter prediction model can be used to generate corresponding reconstruction reference information personalized for different mouth expression coefficients.
[0094] In the exemplary embodiment of the present disclosure, based on multiple contour key points of the mouth sample, the contour area of the mouth sample can be determined, and then the contour area of the mouth sample is divided into regions to determine the geometric template of the mouth sample. For example: Any character image can be selected, and the mouth contour key points of the character image are detected. If the character image is a character video, the mouth key points can be detected for each frame of the character image included in the character video to obtain the position information of multiple candidate contour key points of the mouth sample and determine the attributes of each candidate contour key point. Then, the position information of the candidate contour key points with the same attribute in different frames of the character image is averaged to obtain multiple contour key points of the mouth sample.
[0095] To ensure the fineness of mapping the geometric template of the mouth sample and improve the authenticity of the obtained first mouth image, key point interpolation can be performed based on multiple contour key points of the mouth sample to increase the number of contour key points of the mouth sample. In this case, when using multiple contour key points of the mouth sample as vertices to divide the contour area of the mouth sample into multiple sub-regions, it can ensure that the number of generated sub-regions is as large as possible, thereby achieving the purpose of refined texture mapping and improving the clarity and authenticity of the generated first mouth image.
[0096] The regional division method of the geometric template of the mouth sample can be referred to to perform regional division on each mouth sample image, and multiple sub-mouth textures of the mouth sample image can be obtained. The layout method of the sub-mouth texture in the mouth sample image is the same as the layout method of the sub-region in the geometric template of the mouth sample.
[0097] For example: The same regional division strategy is used to perform regional division on the mouth sample image and the contour area of the mouth sample. The layout method of the generated sub-mouth texture in the mouth sample image is the same as the layout method of the sub-region in the geometric template of the mouth sample. On this basis, an affine transformation relationship between the two can be established through the contour key points with the same attributes of the mouth sample image and the geometric template of the mouth sample.
[0098] After determining the affine transformation relationship between the mouth sample image and the geometric template of the mouth sample, the sub-mouth texture corresponding to each sub-region in the mouth sample texture can be obtained first, and then the sub-mouth texture corresponding to each sub-region in the mouth sample texture is subjected to affine transformation to obtain the first mouth image.
[0099] As a possible implementation, the training set of the image-driven model in the training stage of the exemplary embodiment of the present disclosure includes source image samples, the facial three-dimensional pose information of the source image samples, the facial three-dimensional pose information of the driving image samples, and expression sample control information.
[0100] Figure 6 The flowchart of obtaining the expression sample control information of the exemplary embodiment of the present disclosure is shown. As Figure 6 shown, the method for obtaining the expression sample control information of the exemplary embodiment of the present disclosure may include:
[0101] Step 601: Based on the mouth texture sample information included in the driving image sample and the geometric template of the mouth sample, determine the second oral cavity content image sample. The geometric template of the mouth sample here can be generated with reference to the relevant description above and will not be elaborated here.
[0102] In practical applications, when determining the second oral content image sample, the exemplary embodiment of the present disclosure can migrate the mouth texture sample of the driving image sample to the geometric template of the mouth sample to obtain the second mouth image, and then determine the second oral content image sample based on the second mouth image.
[0103] Exemplarily, when the geometric template of the mouth sample includes multiple sub-regions, the second mouth image includes the geometric template of the mouth sample and a texture image located in each sub-region, and the texture image of each sub-region is determined by the sub-mouth texture corresponding to the sub-region in the driving image sample.
[0104] Exemplarily, the driving image sample of the exemplary embodiment of the present disclosure can be a video containing an arbitrary character (the arbitrary character is defined as a reference character). For each frame of the video image contained in the video, a mouth key detection can be performed. If a mouth contour key point is detected in a certain frame of the video image, it means that the frame of the video image includes the reference character. When the mouth key point detection of all frames of the video image is completed, the mouth contour key points included in each frame of the video image can be obtained. At this time, the mouth contour key points included in each frame of the video image can be attributed, and then the mouth contour key points with the same attributes included in different frames of the video image can be aligned.
[0105] When the mouth contour key points of the same attribute included in each frame of video image are obtained, the reference character mouth texture area included in each frame of video image can be determined. Therefore, referring to the method of mapping the geometric template of the mouth sample using the mouth sample texture in the previous text, the reference character mouth texture included in each frame of video image is used to map the geometric template of the mouth sample, thereby obtaining the second oral content image sample corresponding to each frame of video image. In order to simplify the text, the process of obtaining the second oral content image sample corresponding to each frame of video image will not be described in detail here.
[0106] Step 602: Obtain a facial rendering sample based on the facial morphology information of the source image sample and the facial expression information of the driving image sample. The facial morphology information of the source image here can refer to the relevant description above, which will not be repeated here. It should be understood that the source image sample of the exemplary embodiment of the present disclosure can be any frame in a video, and the driving image sample can be any frame in the video except the source image sample.
[0107] The facial expression information of the driving image sample may be a facial expression coefficient of a reference character included in the driving image sample. When the facial morphology information included in the source image is determined by the facial texture features of the source image and the facial three-dimensional information of the source image, the facial expression coefficient of the reference character may be used to drive the facial three-dimensional information of the source image, so that the facial expression included in the obtained facial rendering sample is related to the facial expression of the reference character included in the driving image.
[0108] When the driving image sample is a video containing a reference character, the facial expression coefficients of each frame of the video image can be obtained. Then, based on the facial morphology information of the source image sample and the facial expression coefficients of each frame of the video image, the facial rendering sample corresponding to each frame of the video image can be determined. This facial rendering sample is essentially an added facial rendering image sample, whose included character appearance is the same as the source character appearance included in the source image sample, but the facial expression of the character it contains is related to the facial expression of the reference character included in the driving image.
[0109] Step 603: Obtain the expression sample control information based on the facial rendering sample and the second oral content image sample. Here, it can be to fuse the character facial rendering sample and the second oral content image sample. The obtained expression sample control information is essentially a facial fusion image formed by the character facial rendering sample and the second oral content image sample. For example: when Render2 represents the facial rendering sample and mouth area2 represents the second oral content image sample, the expression sample control information Control img2 = Render2 + mouth area2.
[0110] To improve the generation accuracy of the target image in the inference stage, whether it is the geometric template of the mouth sample used in the process of obtaining the expression sample control information or the geometric template of the mouth sample used in the process of training the parameter prediction model in the inference stage, it is the geometric template of the same mouth sample. By setting this geometric template of the mouth sample, it can ensure the alignment of the expression sample control information input in the training stage and the expression control information input in the inference stage of the image driving model, thereby improving the accuracy of the target character included in the target image.
[0111] To lightweight the image driving model and improve the convergence speed of the image driving model, the second oral content image sample in the exemplary embodiment of the present disclosure can represent the dimensionality-reduced image of the second mouth image. That is to say, the second mouth image can be dimensionally reduced to reduce the data volume of the second oral content image sample.
[0112] In practical applications, based on the dimensionality reduction strategy of the second mouth image, the dimensionality-reduced image reconstruction reference information and the dimensionality-reduced image dimension information of the second mouth image can be determined, and then the second oral content image sample can be obtained based on the dimensionality-reduced image reconstruction reference information and the dimensionality-reduced image dimension information of the second mouth image.
[0113] The dimensionality reduction strategy of the second mouth image in the exemplary embodiments of the present disclosure can be set according to actual needs. Taking PCA dimensionality reduction as an example, the n-dimensional image data of the second mouth image can be centralized, and then the covariance matrix of the n-dimensional image data of the second mouth image is calculated. Next, the eigenvectors and eigenvalues of the n-dimensional image data covariance matrix are obtained, and the eigenvectors corresponding to the k largest eigenvalues are selected as the k-dimensional principal components of the second mouth image. And the k largest eigenvalues are used as the PCA coefficients of the second mouth image. Both n and k represent positive integers, and n > k.
[0114] On this basis, the k-dimensional PCA coefficients of the second mouth image can be used to combine the k-dimensional principal components of the second mouth image, so as to obtain the second oral content image sample. This can not only reduce the data volume of the second oral content image sample, but also remove noise data and reduce the influence of noise data on the training accuracy. In addition, while reducing the data volume, it can also reduce the model complexity requirements for the image-driven model and accelerate the convergence speed of the image-driven model to complete the model training as soon as possible.
[0115] One or more technical solutions provided in the exemplary embodiments of the present disclosure determine a facial rendering image based on the facial expression coefficient of a preset facial expression and the facial morphology information included in the source image, so that the facial features of the facial rendering image have both the character facial texture of the source image and the preset facial expression. When the preset facial expression includes at least the mouth expression in the open-mouth state, the expression control information can be determined based on the facial rendering image and the oral content image, so that the expression control information has the dual control functions of facial expression guidance and oral content generation. Based on this, the exemplary embodiments of the present disclosure can, under the control of the expression control information, not only control the character facial expression of the target image to present the preset facial expression, but also improve the clarity of the character oral content of the target image, thereby enhancing the authenticity of the character oral state of the target image.
[0116] In the related art, the blurring problem that occurs in oral content generation in the single-image driving problem is essentially a control problem for oral content generation. To address this problem, the exemplary embodiments of the present disclosure adopt the same concept in the training stage and the testing stage of the image-driven model. That is, in the training stage, by adding a second oral content image sample to the facial rendering sample, and in the inference stage, by adding an oral content image to the expression control information, it is ensured that the control information input to the image-driven model all contains an oral content image, so as to drive the source image sample or the source image to generate oral content, such as a target image with clear oral texture, using the control information (the expression sample control information in the training stage and the expression control information in the inference stage).
[0117] In addition, during the training phase of the image-driven model, PCA analysis is performed on the second oral content image samples corresponding to the driving image samples to obtain k-dimensional PCA coefficients and k-dimensional principal components. The second oral content image samples are reconstructed using the k-dimensional PCA coefficients and k-dimensional principal components, thereby achieving the purpose of dimensionality reduction for the second oral content image samples. During the testing phase of the image-driven model, image enhancement can be performed on the first mouth image corresponding to the facial expression coefficient samples to obtain a high-resolution first mouth image. Then, PCA analysis is performed on the first mouth image to obtain the PCA coefficients corresponding to the facial expression sample coefficients, which are used as the supervision information for the parameter prediction model, enabling the trained parameter prediction model to predict the PCA coefficients based on the facial expression coefficients. In this way, during the testing phase of the image-driven model, the trained parameter prediction model can be used to predict the PCA coefficients corresponding to the preset facial expressions, thereby reconstructing the oral content image and adding it as the expression control information to the facial rendering image, improving the clarity of the image-driven model and the stability of the oral content.
[0118] The above mainly introduces the solution provided by the embodiments of the present disclosure from the perspective of the electronic device. It can be understood that in order for the electronic device to implement the above functions, it includes the corresponding hardware structure and / or software module for executing each function. Those skilled in the art should easily realize that, combined with the units and algorithm steps of each example described in the embodiments disclosed herein, the present disclosure can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the form of hardware or computer software driving the hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present disclosure.
[0119] The embodiments of the present disclosure can divide the functional units of the electronic device according to the above method examples. For example, each functional module can be divided corresponding to each function, or two or more functions can be integrated into one processing module. The above integrated modules can be implemented in the form of hardware or in the form of software functional modules. It should be noted that the division of modules in the embodiments of the present disclosure is illustrative, only a logical function division, and there may be other division methods in actual implementation.
[0120] In the case of dividing each functional module corresponding to each function, the exemplary embodiments of the present disclosure provide an image generation device, which can be an electronic device or a chip applied to an electronic device. Figure 7 The schematic block diagram of the functional modules of the image generation device according to the exemplary embodiments of the present disclosure is shown. As Figure 7 shown, the image generation device 700 includes:
[0121] A determination module 701, configured to determine a facial rendering image based on a facial expression coefficient of a preset facial expression and facial morphology information included in a source image, and determine expression control information based on the facial rendering image and an oral content image, where the preset facial expression at least includes a mouth expression in an open - mouth state;
[0122] A generation module 702, configured to generate a target image based on the expression control information, the source image, three - dimensional facial information of the source image, and three - dimensional facial information corresponding to the preset facial expression.
[0123] In a possible implementation manner, the determination module 701 is further configured to determine reconstruction reference information of the oral content image based on the facial expression coefficient of the preset facial expression, and determine the oral content image based on the reconstruction reference information of the oral content image and dimension information of the oral content image.
[0124] In a possible implementation manner, the determination module 701 is configured to obtain a preset mouth expression coefficient from the facial expression coefficient of the preset facial expression, and determine the reconstruction reference information of the oral content image based on the preset mouth expression coefficient.
[0125] In a possible implementation manner, the oral content image represents a dimensionality - reduced image of a mouth expression image corresponding to the preset mouth expression coefficient, and the dimensionality - reduced image is determined by dimensional parameters and coupling parameters in multiple dimensions;
[0126] The dimension information of the oral content image includes the dimensional parameters of the dimensionality - reduced image in multiple dimensions, and the reconstruction reference information of the oral content image includes the coupling parameters of the dimensionality - reduced image in multiple dimensions.
[0127] In a possible implementation manner, the reconstruction reference information of the oral content image is obtained by a parameter prediction model based on the facial expression coefficient of the preset facial expression. The training set of the parameter prediction model in the training stage includes: facial expression sample coefficients and reconstruction reference samples corresponding to the facial expression sample coefficients, and the facial expression sample coefficients include mouth expression coefficients in an open - mouth state.
[0128] In a possible implementation manner, the apparatus further includes an acquisition module 703, configured to transfer a mouth sample texture corresponding to a facial expression coefficient sample to a geometric template of the mouth sample to obtain a first mouth image, and determine the reconstruction reference sample corresponding to the facial expression sample coefficient based on a dimensionality - reduction strategy of the first mouth image.
[0129] In a possible implementation, the geometric template of the mouth sample includes a plurality of sub-regions. The first mouth image includes the geometric template of the mouth sample and the texture images located in each of the sub-regions. The texture image of each sub-region is determined by the sub-mouth texture corresponding to the sub-region in the mouth sample image.
[0130] In a possible implementation, the target image is generated by an image-driven model. The image-driven model includes a motion estimation network and an image generation network. The generation module 702 is configured to input the facial three-dimensional pose information of the source image, the facial three-dimensional pose information corresponding to the preset facial expression, and the source image into the motion estimation network to obtain the expression change estimation information, and input the expression change estimation information, the expression control information, and the source image into the image generation network to obtain the target image.
[0131] In a possible implementation, the training set of the image-driven model in the training phase includes source image samples, the facial three-dimensional pose information of the source image samples, the facial three-dimensional pose information of the driving image samples, and expression sample control information. The device further includes an acquisition module 703, configured to determine a second oral content image sample based on the mouth texture sample information included in the driving image sample and the geometric template of the mouth sample, obtain a character facial rendering sample based on the facial morphology information of the source image sample and the facial expression information of the driving image sample, and obtain the expression sample control information based on the character facial rendering sample and the second oral content image sample.
[0132] In a possible implementation, the acquisition module 703 is configured to transfer the mouth texture sample of the driving image sample to the geometric template of the mouth sample to obtain the second mouth image, and determine the second oral content image sample based on the second mouth image. The second oral content image sample represents the dimensionality-reduced image of the second mouth image.
[0133] In a possible implementation, the geometric template of the mouth sample includes a plurality of sub-regions. The second mouth image includes the geometric template of the mouth sample and the texture images located in each of the sub-regions. The texture image of each sub-region is determined by the sub-mouth texture corresponding to the sub-region in the driving image sample.
[0134] In a possible implementation, the acquisition module 703 is configured to extract the dimensionality-reduced image reconstruction reference information and the dimensionality-reduced image dimension information of the second mouth image, and obtain the second oral content image sample based on the dimensionality-reduced image reconstruction reference information and the dimensionality-reduced image dimension information of the second mouth image.
[0135] Figure 8A schematic block diagram of a chip according to an exemplary embodiment of the present disclosure is shown. As Figure 8 shown, the chip 800 includes one or more (including two) processors 801 and a communication interface 802. The communication interface 802 can support the electronic device to perform the data sending and receiving steps in the above method, and the processor 801 can support the electronic device to perform the data processing steps in the above method.
[0136] Optionally, as Figure 8 shown, the chip 800 further includes a memory 803. The memory 803 can include a read-only memory and a random access memory, and provide operation instructions and data to the processor. A part of the memory can also include a non-volatile random access memory (NVRAM).
[0137] In some embodiments, as Figure 8 shown, the processor 801 executes corresponding operations by calling the operation instructions stored in the memory (the operation instructions can be stored in the operating system). The processor 801 controls the processing operations of any one of the terminal devices, and the processor can also be referred to as a central processing unit (CPU). The memory 803 can include a read-only memory and a random access memory, and provide instructions and data to the processor 801. A part of the memory 803 can also include NVRAM. For example, in the application, the memory, the communication interface, and the memory are coupled together through a bus system. The bus system can include a power bus, a control bus, a status signal bus, etc. in addition to the data bus. However, for the sake of clear illustration, in Figure 8 all kinds of buses are labeled as the bus system 804.
[0138] The method disclosed in the embodiments of the present disclosure can be applied to a processor or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. In the implementation process, the steps of the above method can be completed by the integrated logic circuit in the hardware of the processor or instructions in the form of software. The above-mentioned processor may be a general-purpose processor, a digital signal processor (DSP), an ASIC, a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present disclosure. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present disclosure can be directly embodied as being executed and completed by a hardware decoding processor, or executed and completed by a combination of hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as random access memory, flash memory, read-only memory, programmable read-only memory or electrically erasable programmable memory, registers, etc. The storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method.
[0139] An exemplary embodiment of the present disclosure further provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor. The memory stores a computer program capable of being executed by the at least one processor, and when the computer program is executed by the at least one processor, it is used to cause the electronic device to execute the method according to the embodiments of the present disclosure.
[0140] An exemplary embodiment of the present disclosure further provides a non-transitory computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor of a computer, it is used to cause the computer to execute the method according to the embodiments of the present disclosure.
[0141] An exemplary embodiment of the present disclosure further provides a computer program product, including a computer program, wherein when the computer program is executed by a processor of a computer, it is used to cause the computer to execute the method according to the embodiments of the present disclosure.
[0142] Reference Figure 9, a block diagram of an electronic device 900 that can be a server or a client of the present disclosure will now be described. It is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0143] As Figure 9 shown, the electronic device 900 includes a computing unit 901, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 902 or a computer program loaded from a storage unit 908 into a random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the device 900 can also be stored. The computing unit 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0144] As Figure 9 shown, a plurality of components in the electronic device 900 are connected to the I / O interface 905, including: an input unit 906, an output unit 907, a storage unit 908, and a communication unit 909. The input unit 906 can be any type of device that can input information into the electronic device 900. The input unit 906 can receive input digital or character information, and generate key signal inputs related to the user settings and / or function controls of the electronic device. The output unit 907 can be any type of device that can present information, and can include but is not limited to a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 908 can include but is not limited to a magnetic disk, an optical disk. The communication unit 909 allows the electronic device 900 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks, and can include but is not limited to a modem, a network card, an infrared communication device, a wireless communication transceiver, and / or a chipset, such as a BluetoothTM device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.
[0145] As Figure 9As shown, the computing unit 901 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 901 executes the various methods and processes described above. For example, in some embodiments, the method of the exemplary embodiments of the present disclosure can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 900 via the ROM 902 and / or the communication unit 909. In some embodiments, the computing unit 901 can be configured to execute the method of the exemplary embodiments of the present disclosure by any other suitable means (e.g., by means of firmware).
[0146] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0147] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0148] As used in this disclosure, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, apparatus, and / or device that provides machine instructions and / or data to a programmable processor (e.g., a magnetic disk, an optical disk, a memory, a programmable logic device (PLD)), including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal that provides machine instructions and / or data to a programmable processor.
[0149] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic, speech, or tactile input).
[0150] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0151] A computer system can include clients and servers. The clients and servers are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.
[0152] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present disclosure are executed in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, a terminal, a user device, or other programmable devices. The computer program or instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer program or instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired or wireless manner. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center integrating one or more available media. The available medium can be a magnetic medium, such as a floppy disk, a hard disk, or a magnetic tape; it can also be an optical medium, such as a digital video disc (DVD); or it can be a semiconductor medium, such as a solid state drive (SSD).
[0153] Although the present disclosure has been described in connection with specific features and their embodiments, it is obvious that various modifications and combinations can be made without departing from the spirit and scope of the present disclosure. Accordingly, this specification and the drawings are merely exemplary illustrations of the present disclosure defined by the appended claims, and are considered to cover any and all modifications, variations, combinations, or equivalents within the scope of the present disclosure. Obviously, those skilled in the art can make various changes and modifications to the present disclosure without departing from the spirit and scope of the present disclosure. Thus, if these modifications and variations of the present disclosure fall within the scope of the claims of the present disclosure and their equivalent technologies, the present disclosure also intends to include these changes and modifications therein.
Claims
1. An image generation method, characterized in that, Including: Determine a facial rendering image based on a facial expression coefficient of a preset facial expression and facial morphology information included in a source image; Determine expression control information based on the facial rendering image and an oral content image, where the preset facial expression includes at least a mouth expression in an open - mouth state; Generate a target image based on the expression control information, the source image, the facial three - dimensional information of the source image, and the facial three - dimensional information corresponding to the preset facial expression.
2. The method according to claim 1, wherein The method further includes: Determine reconstruction reference information of the oral content image based on the facial expression coefficient of the preset facial expression; Determine the oral content image based on the reconstruction reference information of the oral content image and the dimensional information of the oral content image.
3. The method according to claim 2, wherein The oral content image represents a dimensionality - reduced image of a mouth expression image corresponding to the preset facial expression, and the dimensionality - reduced image is determined by dimensional parameters and coupling parameters in multiple dimensions of the dimensionality - reduced image; The dimensional information of the oral content image includes the dimensional parameters of the dimensionality - reduced image in multiple dimensions, and the reconstruction reference information of the oral content image includes the coupling parameters of the dimensionality - reduced image in multiple dimensions.
4. The method according to claim 2, wherein The reconstruction reference information of the oral content image is obtained by a parameter prediction model based on the facial expression coefficient of the preset facial expression. The training set of the parameter prediction model in the training stage includes: facial expression sample coefficients and reconstruction reference samples corresponding to the facial expression sample coefficients, and the facial expression sample coefficients include mouth expression coefficients in an open - mouth state.
5. The method according to claim 4, wherein The method for obtaining the training set of the parameter prediction model in the training stage includes: Transfer the mouth sample texture corresponding to the facial expression coefficient sample to the geometric template of the mouth sample to obtain a first mouth image; Determine the reconstruction reference sample corresponding to the facial expression sample coefficient based on the dimensionality - reduction strategy of the first mouth image.
6. The method according to claim 5, characterized in that, The geometric template of the mouth sample includes multiple sub - regions. The first mouth image is determined by the geometric template of the mouth sample and texture images located in each sub - region, and the texture image of each sub - region is determined by the sub - mouth texture corresponding to the sub - region in the mouth sample texture.
7. The method according to any one of claims 1 to 6, characterized in that, The target image is generated by an image - driving model. The image - driving model includes a motion estimation network and an image generation network. Generating the target image based on the expression control information, the source image, the facial three - dimensional information of the source image, and the facial three - dimensional information corresponding to the preset facial expression includes: Input the facial three - dimensional pose information of the source image, the facial three - dimensional pose information corresponding to the preset facial expression, and the source image into the motion estimation network to obtain expression change estimation information; Input the expression change estimation information, the expression control information, and the source image into the image generation network to obtain the target image.
8. The method according to claim 7, wherein The training set of the image - driving model in the training stage includes source image samples, facial three - dimensional pose information of source image samples, facial three - dimensional pose information of driving image samples, and expression sample control information. The method for obtaining the expression sample control information includes: Determine a second oral content image sample based on the mouth texture sample information and the geometric template of the mouth sample included in the driving image sample; Obtain a character facial rendering sample based on the facial morphology information of the source image sample and the facial expression information of the driving image sample; Obtain expression sample control information based on the character facial rendering sample and the second oral content image sample.
9. The method according to claim 8, characterized in that The determining the second oral content image sample based on the mouth texture sample information and the geometric template of the mouth sample included in the driving image sample includes: Transfer the mouth texture sample of the driving image sample to the geometric template of the mouth sample to obtain a second mouth image; Based on the second mouth image, determine the second oral content image sample, and the second oral content image sample represents a dimensionality-reduced image of the second mouth image.
10. The method according to claim 8, wherein The geometric template of the mouth sample includes a plurality of sub-regions, and the second mouth image is determined by the geometric template of the mouth sample and the texture images located in each of the sub-regions, and the texture image of each sub-region is determined by the corresponding sub-mouth texture of the sub-region in the driving image sample.
11. The method according to claim 8, wherein The determining the second oral content image sample based on the second mouth image includes: Extract the dimensionality-reduced image reconstruction reference information and the dimensionality-reduced image dimension information of the second mouth image; Based on the dimensionality-reduced image reconstruction reference information and the dimensionality-reduced image dimension information of the second mouth image, obtain the second oral content image sample.
12. An image generation device, characterized in that, Includes: A determining module, configured to determine a facial rendering image based on the facial expression coefficient of a preset facial expression and the facial morphology information included in the source image, and determine expression control information based on the facial rendering image and the oral content image, where the preset facial expression includes at least a mouth expression in an open-mouth state; A generating module, configured to generate a target image based on the expression control information, the source image, the three-dimensional facial information of the source image, and the three-dimensional facial information corresponding to the preset facial expression.
13. An electronic device, characterized in that, Includes: A processor; And, A memory storing a program; Wherein, the program includes instructions, and when the instructions are executed by the processor, the processor executes the method according to any one of claims 1 to 11.
14. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium stores computer instructions for causing the computer to execute the method according to any one of claims 1 to 11.