Image generation method and device, equipment, medium and product
By combining the key point processing of source images and driving information, special attention is paid to the postures of non-face areas, and images that maintain the appearance and posture are generated are generated, which solves the problem of inconsistency in image generation in the prior art and improves the quality of image generation.
Patent Information
- Application Number
- CN202410177955.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-02-08
- Publication Date
- 2025-08-08
AI Technical Summary
The prior art is difficult to effectively generate and maintain the consistency of the appearance information of the source image and the attitude of the driving information in a single-image driving scenario, especially in the posture performance of the non-face area.
By acquiring source image and driving information, image generation processing is performed using implicit key points and face key points of source image, as well as implicit key points and face key points of driving information, special attention is paid to the posture situation of non-face areas, and combined with preset image sequences and deformation processing, image results that maintain the consistent appearance and posture are generated.
In the automatic generation process under driver information, we fully pay attention to the posture conditions of each area, avoid inconsistency defects, and improve image generation quality.
Smart Images

Figure CN120451032A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to an image generation method, apparatus, device, medium, and product. Background Art
[0002] For some application scenarios, such as single-image driving scenarios, these scenarios have the following requirements: after a source image is given, the source image can be driven and processed according to the driving information input by the user to obtain a driving result, such as one or more images, so that the driving result can present the object appearance information described by the source image and the changes described by the driving information. Summary of the Invention
[0003] The present application provides an image generation method, apparatus, device, medium, and product, which are conducive to improving the quality of image generation.
[0004] In order to achieve the above objectives, the technical solutions provided by this application are as follows:
[0005] The present application provides an image generation method, the method comprising:
[0006] Obtain source image and driving information;
[0007] Image generation processing is performed based on the source image, implicit key points of the source image, facial key points of the source image, implicit key points of the driving information, and facial key points of the driving information to obtain an image generation result corresponding to the driving information; the facial key points are used to describe the facial area; and the implicit key points are used to supplement key points in other areas except at least one sub-area of the facial area.
[0008] In a possible implementation manner, the image generation result is determined based on part or all of the implicit key points.
[0009] In one possible implementation, the implicit key points include implicit key points in at least two areas; the at least two areas include the facial area and at least one non-facial area; and the image generation result is determined based on the implicit key points in the at least one non-facial area.
[0010] In a possible implementation manner, a target key point among the implicit key points partially or completely matches the facial region; and the image generation result is determined based on other key points among the implicit key points except the target key point.
[0011] In a possible implementation, the target key point is matched with the at least one sub-region; the at least one sub-region includes at least one of an eye region and a mouth region.
[0012] In a possible implementation manner, the facial key points include facial key points within the at least one sub-region; and the target key points are determined based on a distance between the implicit key points and the facial key points within the at least one sub-region.
[0013] In one possible implementation, the target key point is determined based on a point identifier of at least one position change key point;
[0014] The process of determining the at least one position change key point includes:
[0015] Acquire a preset image sequence, where the preset image sequence is used to describe a change in the at least one sub-region;
[0016] Performing key point position change analysis based on implicit key points of at least two images in the preset image sequence to obtain an analysis result;
[0017] According to the analysis result, the at least one position-changing key point is determined from the implicit key points of the at least two images.
[0018] In a possible implementation manner, the driving information includes N frames of information, where N is a positive integer;
[0019] The process of determining the image generation result corresponding to the n-th frame information includes:
[0020] Determining deformation reference data corresponding to the n-th frame information based on the implicit key points of the source image, the facial key points of the source image, the implicit key points of the n-th frame information, and the facial key points of the n-th frame information; n is a positive integer, n≤N;
[0021] Performing deformation processing on the source image according to the deformation reference data corresponding to the n-th frame information to obtain a deformed image corresponding to the n-th frame information;
[0022] Performing a completion process on the deformed image corresponding to the n-th frame information to obtain a completed image corresponding to the n-th frame information and completion position representation data corresponding to the n-th frame information;
[0023] Based on the completed position representation data corresponding to the n-th frame information, the reference image corresponding to the n-th frame information and the completed image corresponding to the n-th frame information are fused to obtain the image generation result corresponding to the n-th frame information; the reference image is determined based on the source image and the deformation reference data corresponding to the n-th frame information.
[0024] In a possible implementation manner, the reference image is the deformed image corresponding to the n-th frame information;
[0025] or,
[0026] The process of obtaining the reference image includes:
[0027] Performing super-resolution processing on the source image to obtain a super-resolution image;
[0028] Performing deformation processing on the super-resolution image according to the deformation reference data corresponding to the n-th frame information to obtain a super-resolution deformation result corresponding to the n-th frame information;
[0029] A reference image corresponding to the n-th frame information is determined according to the super-fractional deformation result corresponding to the n-th frame information.
[0030] In a possible implementation manner, if n≥2, the process of determining the deformation reference data corresponding to the n-th frame information includes:
[0031] Determining predicted deformation parameters corresponding to the n-th frame information based on implicit key points of the source image, facial key points of the source image, implicit key points of the n-th frame information, and facial key points of the n-th frame information;
[0032] Based on the weight corresponding to at least one area in the source image, the predicted deformation parameters corresponding to the n-th frame information and the deformation reference data corresponding to the n-1-th frame information are weighted to obtain the deformation reference data corresponding to the n-th frame information; the at least one area is determined based on the area analysis result of the source image; the area analysis result is used to describe the position of the at least one area in the source image.
[0033] In one possible implementation, the implicit key points are determined using an implicit key point detection model; and the image generation process is implemented using a preset model.
[0034] After obtaining the image generation result corresponding to the driving information, the method further includes:
[0035] The implicit key point detection model and the preset model are updated according to the image generation result and the label information corresponding to the image generation result; the label information is determined according to the driving information.
[0036] In a possible implementation manner, the driving information and the source image are both images extracted from the same sample video; and the label information is the driving information.
[0037] In a possible implementation manner, the process of acquiring the driving information includes:
[0038] After receiving the text data, converting the text data into voice data;
[0039] The driving information is constructed based on at least one head posture template and the predicted expression coefficient of each frame data in the voice data.
[0040] In one possible implementation, the driving information includes N frames of information, where N is a positive integer; the nth frame information is determined based on the nth frame data in the voice data and the head posture template matching the nth frame data; the at least one head posture template includes the head posture template matching the nth frame data; n is a positive integer, n≤N.
[0041] In a possible implementation, the source image is determined based on a pre-constructed virtual image; the image generation result includes a generated image corresponding to each frame of data in the voice data;
[0042] After obtaining the image generation result corresponding to the driving information, the method further includes:
[0043] A video corresponding to the virtual image is constructed based on the voice data and the image generation result.
[0044] The present application provides an image generation device, comprising:
[0045] A data acquisition unit, used for acquiring source images and driving information;
[0046] An image generation unit is configured to perform image generation processing based on the source image, the implicit key points of the source image, the facial key points of the source image, the implicit key points of the driving information, and the facial key points of the driving information to obtain an image generation result corresponding to the driving information; the facial key points are used to describe the facial area; and the implicit key points are used to supplement the key points in other areas except for at least one sub-area of the facial area.
[0047] The present application provides an electronic device, the device comprising: a processor and a memory;
[0048] The memory is used to store instructions or computer programs;
[0049] The processor is used to execute the instructions or computer programs in the memory so that the electronic device executes the image generation method provided in this application.
[0050] The present application provides a computer-readable medium, which stores instructions or computer programs. When the instructions or computer programs are executed on a device, the device executes the image generation method provided by the present application.
[0051] The present application provides a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains program code for executing the image generation method provided by the present application.
[0052] Compared with the related art, this application has at least the following advantages:
[0053] In the technical solution provided in the present application, after obtaining the source image and driving information, image generation processing is performed based on the source image, the implicit key points of the source image, the facial key points of the source image, the implicit key points of the driving information, and the facial key points of the driving information to obtain an image generation result corresponding to the driving information, so that the appearance information presented by the image generation result is consistent with the appearance information presented by the source image, and the posture presented by the image generation result is consistent with the posture represented by the driving information, so that the driving result of the source image under the driving information can be automatically generated. Among them, because the facial key points are used to describe the facial area, and the implicit key points are used to supplement the key points in other areas except the facial area, when performing image generation processing based on the facial key points and the implicit key points, not only the posture of the facial area, such as the expression state, etc., can be paid attention to, but also some non-facial areas, such as the hair area, the hat area and other areas. The posture conditions can be paid attention to, so that the image generation process can pay attention to the posture conditions of each area as comprehensively as possible, and then the final image generation result can present the various posture conditions represented by the driving information as comprehensively as possible, so that defects caused by only paying attention to the facial state, such as incoordination, can be effectively avoided, which is conducive to improving the image generation quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only some embodiments recorded in this application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0055] Figure 1 A flowchart of an image generation method provided in an embodiment of the present application;
[0056] Figure 2A schematic diagram of an image generation process in a model update scenario provided in an embodiment of the present application;
[0057] Figure 3 A schematic diagram of an image generation process in an image sequence generation scenario provided in an embodiment of the present application;
[0058] Figure 4 A schematic structural diagram of an image generating device provided in an embodiment of the present application;
[0059] Figure 5 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0060] In order to help those skilled in the art better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of this application.
[0061] In order to better understand the technical solution provided by this application, the image generation method provided by this application is described below with reference to some drawings. Figure 1 As shown, the image generation method provided by the embodiment of the present application includes the following S1-S4. Figure 1 A flowchart of an image generation method provided in an embodiment of the present application.
[0062] S1: Get source image and driving information.
[0063] The source image is used to provide other information in addition to some information such as expression information during the image generation process, such as appearance information of an object. The object refers to the foreground described by the source image; and this application does not limit the implementation of the object.
[0064] In addition, the present application does not limit the implementation method of the source image. For ease of understanding, the following description combines two cases.
[0065] Case 1: For some scenarios, such as model training scenarios, the source image mentioned above can refer to a frame of video image randomly extracted from a sample video, such as Figure 2The source image shown. The sample video refers to the video required for the model training process; and this application does not limit the method for obtaining the sample video. For example, it can be implemented using any existing or future sample video acquisition method. It can be seen that in one possible implementation, when the image generation method provided by this application is applied to the model update process, such as Figure 2 During the model training process shown in FIG, the source image may be an image extracted from a sample video.
[0066] Case 2: For some scenarios, such as performing a certain image generation task, the source image mentioned above may refer to an image input by the user in some way, such as an image taken by a camera, an image uploaded manually, an image selected by a selection operation, etc. It can be seen that in one possible implementation, when the image generation method provided by this application is applied to an image generation task, such as Figure 3 When performing the image generation task shown in FIG, the source image may refer to an image provided by the user, such as Figure 3 Source image shown.
[0067] The driving information is used to provide some information during the image generation process, such as information on facial expression status, head posture, etc.; and the present application does not limit the driving information. For example, the driving information may include some facial expression coefficients. For another example, the driving information may include one or more images. It can be seen that under one possible implementation, the driving information may include N frames of information, where N is a positive integer. Among them, the nth frame information refers to the information in the nth arrangement position in the driving information; and the present application does not limit the implementation method of the nth frame information. For example, the nth frame information can be implemented using facial expression coefficients or images. n is a positive integer, n≤N.
[0068] In addition, the present application does not limit the implementation method of the driving information. For ease of understanding, the following description is combined with two cases.
[0069] Case 1: For some scenarios, such as model training scenarios, the above-mentioned driving information may refer to the driving image corresponding to the above-mentioned source image. The driving image refers to an image that is required to provide guidance information when performing model update processing based on the source image; and this application does not limit the driving image. For example, when the source image refers to a frame of image randomly extracted from a sample video, the driving image corresponding to the source image may refer to another frame of image randomly extracted from the sample video, such as Figure 2As shown in the driving image, in one possible implementation, the driving information and the source image are both extracted from the same sample video, so that the training process based on the source image and the driving image is a self-supervised task of video reconstruction. This makes the training process not involve cross-object and cross-domain training datasets, thereby reducing the difficulty of obtaining training data.
[0070] In case 2, for some scenarios, such as performing an image generation task, the driving information can be determined based on the data involved in the image generation task, such as text, video, or audio data, used to provide the driving signal. This allows the driving information to be used to provide posture information such as facial expression and / or head posture, such as changes in facial expression. For ease of understanding, the following example is used to illustrate.
[0071] As an example, in some application scenarios, the above process of obtaining the driving information may include the following steps 11 and 12.
[0072] Step 11: After receiving the text data, convert the text data into voice data.
[0073] The text data is used to describe the constraint information required for reference when performing image generation processing, such as expression constraints, etc.; and this application does not limit the text data. For example, the text data can be used Figure 3 The following text is implemented:
[0074] In addition, this application does not limit the method for obtaining the above text data. For example, the text data may refer to text content input by the user. For another example, the text data may refer to text obtained and output by a model by processing user input content. The user input content refers to content provided by the user to the model, such as text, voice, video, image, etc. The model is used to convert the user input content into text.
[0075] In addition, the present application does not limit the implementation of the above step 11. For example, it can be implemented with the help of any existing or future method that can convert text into speech, such as text to speech (TTS).
[0076] Based on the relevant content of step 11 above, it can be seen that after obtaining the above text data, such as Figure 3 After a piece of text content is shown, the text data can be converted into speech data, such as Figure 3 The voice data shown is used to enable the voice data to express the text data in an audio manner, so that the posture information corresponding to the text data, such as the mouth state, can be determined based on the voice data.
[0077] Step 12: Construct driving information based on at least one head posture template and the predicted expression coefficients of each frame of the above speech data.
[0078] The head posture template refers to a pre-set template for describing the head posture; and the head posture template can be pre-set according to an actual application scenario.
[0079] In addition, for the voice data converted from the above text data, the voice data may include N frames of data. The nth frame of data refers to the data at the nth arrangement position in the voice data. The predicted expression coefficient of the nth frame of data refers to the expression coefficient corresponding to the nth frame of data, such as Figure 3 The coefficient n shown is used so that the predicted expression coefficient of the n-th frame data can describe the posture information of the n-th frame data, such as the state of the mouth, etc., so that the predicted expression coefficient of the n-th frame data can describe the posture information corresponding to the text content expressed by the n-th frame data, such as expression information, etc. In addition, the predicted expression coefficient of the n-th frame data can be obtained by performing expression coefficient prediction processing on the n-th frame data; and the present application does not limit the implementation method of the expression coefficient prediction processing. For example, it can be implemented using any existing or future method that can perform expression coefficient prediction processing on a frame of data, such as audio2bs and other methods. n is a positive integer, n≤N, and N is a positive integer.
[0080] In addition, this application does not limit the implementation method of the above step 12. For ease of understanding, two examples are used below to illustrate.
[0081] Example 1: When the at least one head posture template above includes only one head posture template, and the head posture template includes N frames of head action description data arranged in sequence, the above step 12 may specifically include the following steps 121-122.
[0082] Step 121: Based on the head movement description data of the nth frame in the above head posture template and the predicted expression coefficient of the nth frame data in the above voice data, construct the posture description information of the nth frame, so that the posture description information of the nth frame includes the posture information described by the head movement description data of the nth frame and the posture information described by the predicted expression coefficient of the nth frame data, where n is a positive integer and n≤N.
[0083] Among them, the head action description data of the nth frame in the above head posture template is used to describe the head posture pre-set for the nth frame; and the present application does not limit the implementation method of the head action description data of the nth frame. For example, it can be implemented in the form of a template, text, or image, etc., n is a positive integer, n≤N. It should be noted that the present application does not limit the implementation method of the head posture template. For example, the head posture template may refer to an action sequence pre-set according to the application scenario. For another example, the head posture template may be an action sequence extracted from a video provided by the user. For another example, the head posture template may be a template found in a template library that best matches the above source image. Among them, the template library is used to provide a variety of head posture templates, and each head posture template is an action sequence.
[0084] The nth frame posture description information is used to describe some postures corresponding to the nth frame data in the above voice data, such as facial expression status and / or head posture, etc., and this application does not limit the implementation method of the nth frame posture description information. For example, it can be implemented using images.
[0085] In addition, the above n-frame posture description information is determined based on the n-frame head movement description data and the predicted expression coefficient of the n-frame data, so that the n-frame posture description information includes the posture information described by the n-frame head movement description data and the posture information described by the predicted expression coefficient of the n-frame data.
[0086] In addition, the present application does not limit the implementation method of the above step 121. For example, it can be specifically: image rendering processing is performed based on the head action description data of the nth frame in the above head posture template and the predicted expression coefficient of the nth frame data in the above voice data to obtain the nth frame posture description information, so that the nth frame posture description information can describe the various postures corresponding to the nth frame data in the form of images, such as head posture, mouth status, etc.
[0087] Step 122: Construct driving information based on the first frame posture description information, the second frame posture description information, ..., and the Nth frame posture description information, so that the driving information includes the first frame posture description information, the second frame posture description information, ..., and the Nth frame posture description information.
[0088] Based on the relevant content of steps 121 to 122 above, it can be seen that for some application scenarios, if a head posture template is set in advance and the head posture template is used to describe an action sequence, the action of each frame in the head posture template can be combined with the predicted expression coefficient of each frame data in the above voice data to obtain driving information, so that the driving information includes the combination results of each frame, so that an image sequence can be generated subsequently based on the driving information.
[0089] Example 2: When the at least one head posture template includes multiple head posture templates, the above step 12 may specifically include the following steps 123 and 124.
[0090] Step 123: Based on the n-th frame data in the above speech data, find the template that best matches the n-th frame data from the above multiple head posture templates as the head posture template corresponding to the n-th frame data, where n is a positive integer, n≤N.
[0091] It should be noted that the present application does not limit the implementation method of the above step 123. For example, when each template in the above multiple head posture templates is used to describe a frame of posture, the specific implementation of step 123 can be: for the n-th frame data in the above voice data, the n-th frame data can be matched with each template in the above multiple head posture templates to obtain a matching result, and based on the matching result, the template that best matches the n-th frame data is determined as the head posture template corresponding to the n-th frame data, such as Figure 3 The template n shown is used so that the head posture template corresponding to the n-th frame data can represent the head movement that best matches the n-th frame data, which is beneficial to improving the image generation effect.
[0092] In fact, in order to better improve the image generation effect, the present application also provides another possible implementation of the above step 123. Under this implementation, when each of the above multiple head posture templates is a posture sequence, the above voice data includes at least one audio segment, and each audio segment includes at least one frame of data, the step 123 can be specifically as follows: for any audio segment, select a template that matches the audio segment from the multiple head posture templates, and determine each frame of posture data in the matching template as the head posture template corresponding to each frame of data in the audio segment. It should be noted that the present application does not limit the implementation method of the matching. For example, the matching process between the audio segment and the template can be completed based on the semantics, rhythm, strength and other information in the audio segment.
[0093] Step 124: Based on the head posture template corresponding to the n-th frame data in the above voice data and the predicted expression coefficient of the n-th frame data, construct the n-th frame posture description information, so that the n-th frame posture description information includes the posture information described by the head posture template corresponding to the n-th frame data, and the posture information described by the predicted expression coefficient of the n-th frame data, where n is a positive integer, n≤N.
[0094] It should be noted that the present application does not limit the implementation method of the above step 124. For example, the step 124 may specifically be: constructing the n-frame posture description information based on the head posture template corresponding to the n-frame data in the above voice data and the predicted expression coefficient of the n-frame data, so that the n-frame posture description information includes the head posture described by the head posture template and the predicted expression coefficient of the n-frame data. For another example, the step 124 may specifically be: performing image rendering processing based on the head posture template corresponding to the n-frame data in the above voice data and the predicted expression coefficient of the n-frame data to obtain the n-frame posture description information. For another example, the step 124 may specifically be: performing coefficient adjustment processing on the predicted expression coefficient of the n-frame data based on the head posture template corresponding to the n-frame data to obtain the n-frame posture description information, so that the n-frame posture description information simultaneously satisfies the constraints described by the head posture template corresponding to the n-frame data and the constraints described by the predicted expression coefficient of the n-frame data.
[0095] Step 125: Construct driving information based on the first frame posture description information, the second frame posture description information, ..., and the Nth frame posture description information, so that the driving information includes the first frame posture description information, the second frame posture description information, ..., and the Nth frame posture description information.
[0096] Based on the relevant content of steps 123 to 125 above, it can be seen that for some application scenarios, if many head posture templates are pre-set, and each head template is used to describe a frame of head posture, then the template that matches each frame of audio in the voice data can be determined from these templates first, so that the head posture described by the matching template and the mouth state corresponding to the audio have a relatively high adaptability; then, each frame of audio in the voice data is combined with the corresponding matching template to obtain driving information, so that the driving information includes the combination results of each frame, so that the driving information can better express the posture changes corresponding to the voice data, such as changes in expression, changes in head posture, etc., which is conducive to improving flexibility.
[0097] Based on the above content, it can be seen that in one possible implementation, the nth frame information in the above driving information can be determined based on the nth frame data in the above voice data and the head posture template matched to the nth frame data, where n is a positive integer, n≤N, and N is a positive integer. The head posture template matched to the nth frame data refers to the template 3002 determined from the at least one head posture template above that best matches the nth frame data.
[0098] Based on the relevant content of steps 11 to 12 above, it can be seen that for some application scenarios, after receiving the text data, the text data is first converted into voice data; then, the expression coefficient prediction processing is performed on each frame of the voice data; then, based on the extracted expression coefficient and the pre-set head posture template, the driving information is constructed so that the driving information can better describe the posture changes corresponding to the text data, such as expression changes, so that an image sequence that can represent the posture changes can be generated based on the driving information, such as a video.
[0099] Based on the relevant content of S1 above, it can be seen that for some application scenarios, such as model training scenarios or image generation task execution scenarios, the source image and driving information are obtained so that image generation processing can be performed based on the source image and the driving information later, so that the posture information in the final generated image is consistent with the posture information in the driving information, and other information in the final generated image except the posture information, such as appearance information, is consistent with the corresponding information in the source image.
[0100] S2: Performing image generation processing based on the source image, the implicit key points of the source image, the facial key points of the source image, the implicit key points of the driving information, and the facial key points of the driving information to obtain an image generation result corresponding to the driving information; the facial key points are used to describe the facial area; the implicit key points are used to supplement the key points in other areas except for at least one sub-area in the facial area.
[0101] The facial landmarks (LMK) of the source image are used to describe some coordinate positions in the facial region of the source image, so that the facial landmarks can describe the posture characteristics of the facial region in the source image, such as expression characteristics.
[0102] In addition, the facial key points have good controllability and interpretability, so that they can clearly represent certain areas of the face, such as the eyes, mouth, and facial contours. Obviously, the facial key points are limited to certain facial areas, making it impossible for them to describe the posture characteristics of other areas outside of this part of the face, such as the forehead, cheeks, and hairstyle. Consequently, subsequent posture adjustment processing of this part of the face can only be achieved with the help of these facial key points.
[0103] In addition, this application does not limit the method for obtaining the facial key points of the source image. For example, it can be implemented by any existing or future method that can detect facial key points of an image. For example, in some application scenarios, in order to better improve the accuracy, the process of obtaining the facial key points of the source image can be as follows: the source image is input into a pre-built facial key point detection model, such as Figure 2 or Figure 3 Model 2 is shown, so that the facial key point detection model can perform facial key point detection processing on the source image, obtain and output the facial key points of the source image. The facial key point detection model refers to a pre-built model with relatively good facial key point detection capabilities, so that the facial key point detection model can be used to perform facial key point detection processing on the input data of the facial key point detection model; and this application does not limit the implementation of the facial key point detection model.
[0104] The implicit key points (Neural Keypoints, NK) of the source image are used to describe some coordinate positions in the source image, so that the implicit key points are used to describe the distribution characteristics of the source image, so that the implicit key points can represent the posture characteristics of multiple areas in the source image, such as the face area, hat area, hair area, background area, etc., and thus the implicit key points can represent the position information in the source image as comprehensively as possible, such as the background position, hair position, hat position, etc., so that the implicit key points can supplement the position information of other areas other than at least one sub-area in the face area, such as the hat, hair accessories, etc. as comprehensively as possible. The sub-area refers to a part of the face area; and the present application does not limit the implementation method of the at least one sub-area. For example, the at least one sub-area can include at least one of the eye area and the mouth area.
[0105] In addition, for the implicit key points mentioned above, since the implicit key points have not been artificially standardized and defined, the coordinate positions described by the implicit key points do not have precise semantic meanings, which makes the implicit key points uncontrollable, and thus makes the implicit key points more flexible, so that the implicit key points are more generalizable, making the implicit key points applicable to various data, and thus making the implicit key points able to better complete the posture information of other areas other than at least one sub-area in the facial area, such as hats, hair accessories and other areas.
[0106] Based on the above two paragraphs, it can be seen that in one possible implementation, the implicit key points of the above source image can be used to supplement the key points in other areas except at least one sub-area in the facial area, so that the implicit key points of the source image can complete the position information of other areas except at least one sub-area in the facial area, which is conducive to improving the generalization ability of the image generation process.
[0107] In addition, for some application scenarios, such as scenarios that focus on the posture of various parts of the body, in order to better improve efficiency, the implicit key points of the above source image can be used to describe the posture characteristics of each body part area in the source image, so that the implicit key points of the source image can be used to supplement the key points in other body part areas except at least one sub-area in the face area, so that the implicit key points of the source image can complete the position information of other body part areas except at least one sub-area in the face area.
[0108] In addition, this application does not limit the method of obtaining the implicit key points of the source image above. For example, it can be specifically: inputting the source image into the implicit key point detection model, such as Figure 2 or Figure 3 The model 1 shown is used to enable the implicit key point detection model to perform implicit key point detection processing on the source image, obtain and output the implicit key points of the source image. The implicit key point detection model is used to perform implicit key point detection processing on the input data of the implicit key point detection model; and this application does not limit the implementation method of the implicit key point detection model. For example, it can be implemented using any machine learning model, such as any end-to-end machine learning model.
[0109] For the above-mentioned driving information, the facial key points of the driving information are used to describe the facial state in the driving information, such as the expression state, etc.; and when the driving information includes N frames of information, the facial key points of the driving information include the facial key points of the N frames of information. Among them, the facial key points of the n-th frame information are used to describe some coordinate positions within the facial area of the n-th frame information, so that the facial key points can describe the posture characteristics of the facial area in the n-th frame information. n is a positive integer, n≤N, and N is a positive integer. It should be noted that the implementation method of the facial key points of the n-th frame information is similar to the implementation method of the facial key points of the source image above. For the sake of brevity, it will not be repeated here.
[0110] In fact, to further improve image generation, this application also provides a possible implementation of the process for determining facial key points in the aforementioned driving information. In this implementation, when the driving information includes expression coefficients, the facial key points in the driving information can be determined by searching a first mapping relationship for facial key points corresponding to the expression coefficients, and using these as the facial key points in the driving information. The first mapping relationship is used to describe the facial key points corresponding to the expression coefficients under different expressions; and the first mapping relationship is pre-determined based on a sample image. The sample image refers to an image, such as a three-dimensional digital human, required to construct the first mapping relationship. It should be noted that this application does not limit the process for constructing the first mapping relationship. For example, the process can specifically include: first obtaining the expression coefficients and facial key points of the sample image; then, based on the correspondence between the expression coefficients and the facial key points of the sample image, constructing the first mapping relationship so that the first mapping relationship includes the correspondence.
[0111] In addition, the implicit key points of the driving information are used to characterize the positional information in the driving information, so that the implicit key points can be used to characterize the posture characteristics of the driving information, such as facial expression, hairstyle, hat posture, etc., so that the implicit key points of the driving information can be used to supplement the key points in other regions besides at least one sub-region of the facial region. Moreover, when the driving information includes N frames of information, the implicit key points of the driving information can include the implicit key points of the N frames of information. The implicit key points of the nth frame information are used to describe certain coordinate positions in the nth frame information, so that the implicit key points are used to describe the distribution characteristics of the nth frame information, so that the implicit key points can represent the posture characteristics of each region in the nth frame information, and thus, the implicit key points can represent the positional information described by the nth frame information as comprehensively as possible. n is a positive integer, n≤N, and N is a positive integer. It should be noted that the implementation method of the implicit key points of the nth frame information is similar to the implementation method of the implicit key points of the source image above, and for the sake of brevity, it will not be repeated here.
[0112] In fact, in order to better improve the image generation effect, the present application also provides a possible implementation method of the process of determining the implicit key points of the above-mentioned driving information. Under this implementation method, when the driving information includes an expression coefficient, the implicit key point of the driving information can be specifically: searching for the implicit key point corresponding to the expression coefficient from the second mapping relationship as the implicit key point of the driving information. The second mapping relationship is used to describe the implicit key points corresponding to the expression coefficient under different expressions; and the second mapping relationship is determined in advance based on the above-mentioned sample image. It should be noted that the present application does not limit the construction process of the second mapping relationship. For example, it can be specifically: first obtain the expression coefficient of the sample image and the implicit key point of the sample image; then construct the second mapping relationship based on the correspondence between the expression coefficient of the sample image and the implicit key point of the sample image, so that the second mapping relationship includes the correspondence.
[0113] Based on the above content, it can be seen that because implicit key points have good generalization ability, when combining implicit key points with facial key points for image generation processing, it can effectively overcome the problem caused by the large difference between the object described by the source image and the above sample image, thereby helping to improve the image generation effect.
[0114] For the above-mentioned driving information, the image generation result corresponding to the driving information refers to one or more images obtained by performing posture driving on the source image using the driving information, so that the posture information in the image generation result is consistent with the posture information in the driving information, and all other information in the image generation result except the posture information is consistent with the corresponding information in the source image.
[0115] In addition, for the above-mentioned driving information, if the driving information includes N frames of information, the image generation result corresponding to the driving information may include the image generation result corresponding to the N frames of information. The image generation result corresponding to the n-th frame information refers to the image obtained by performing posture driving on the source image using the n-th frame information, so that the posture information in the image generation result corresponding to the n-th frame information is consistent with the posture information in the n-th frame information, and the other information in the image generation result corresponding to the n-th frame information except the posture information is consistent with the corresponding information in the source image, so that the image generation result corresponding to the n-th frame information can represent the generated image corresponding to the n-th frame information, such as the generated image corresponding to the n-th frame data, etc. n is a positive integer, n≤N, and N is a positive integer.
[0116] Furthermore, this application does not limit the implementation method of S2 above. For example, it can be implemented with the help of a pre-built machine learning model.
[0117] In fact, in order to better improve the image generation effect, the present application also provides a possible implementation of the above S2. Under this implementation, the S2 may specifically include the following S21-S23.
[0118] S21: Determine key point information of the source image based on the implicit key points of the source image and the facial key points of the source image.
[0119] The key point information of the source image is used to represent the posture information in the source image, such as facial expression, hat posture, hairstyle posture, etc.; and the key point information is determined based on the implicit key points of the source image and the facial key points of the source image, so that the key point information includes the facial key points and includes some or all of the implicit key points, so that the key point information can represent the posture information in the source image as comprehensively as possible, such as the posture information of various body parts in the source image. It can be seen that under one possible implementation, the key point information of the source image includes some or all of the implicit key points of the source image and the facial key points of the source image.
[0120] In addition, the present application does not limit the implementation method of the above S21. For example, the S21 can specifically be: splicing the implicit key points of the above source image and the facial key points of the source image to obtain the key point information of the source image, so that the key point information includes the implicit key points and the facial key points, so that the key point information can represent the posture information in the source image as comprehensively as possible.
[0121] Research has found that facial key points can more accurately describe facial states, such as the state of the eyes and mouth. Therefore, in order to avoid the uncontrollability of implicit key points interfering with some or all facial states, the present application also provides a possible implementation method of the key point information of the above-mentioned source image. Under this implementation method, when the implicit key points of the source image include implicit key points in at least two areas, and the at least two areas include a facial area and at least one non-facial area, the key point information of the source image may include the implicit key points in the at least one non-facial area and the facial key points of the source image, so that the image generation result corresponding to the above-mentioned driving information can be determined based on the implicit key points in the at least one non-facial area. The at least one non-facial area refers to other areas in the source image other than the facial area, such as other body parts.
[0122] Based on the content of the above paragraph, in order to better improve the accuracy, the present application also provides a possible implementation method of the above key point information. Under this implementation method, when the target key point in the above implicit key point matches part or all of the facial area, the key point information can be determined based on other key points in the implicit key point except the target key point, so that the key point information includes other key points in the implicit key point except the target key point, so that the image generation result corresponding to the above driving information can be determined based on the other key points. Based on this, it can be seen that under a possible implementation method, the image generation result is determined based on other key points in the implicit key point except the target key point. For better understanding, a possible implementation method of S21 above is described below as an example.
[0123] As an example, in a possible implementation, the above S21 may specifically include the following steps 21 and 22.
[0124] Step 21: Determine the target key points corresponding to the source image from the implicit key points of the source image, so that the target key points are partially or completely matched with the facial region.
[0125] Among them, the target key point corresponding to the source image refers to the key point existing in the implicit key points of the source image and matching part or all of the facial area, so that the target key point can represent the implicit key point falling within part or all of the facial area.
[0126] In addition, the present application does not limit the implementation method of the target key points corresponding to the above source image. For example, in some application scenarios, the target key points corresponding to the source image may refer to the key points existing in the implicit key points of the source image and matching the facial area, so that the target key points can represent the implicit key points falling within the facial area.
[0127] Research has found that using facial key points to adjust the posture of the eyes and / or mouth is more conducive to improving the quality of image generation. Therefore, in order to better avoid the uncontrollability of implicit key points from interfering with the eye state and mouth state, the present application also provides a possible implementation method of the target key points corresponding to the above source image. Under this implementation method, the target key points corresponding to the source image can be matched with at least one sub-region in the facial area, and the at least one sub-region includes at least one of the eye area and the mouth area, so that the target key point can refer to a key point that exists in the implicit key points of the source image and falls into at least one of the eye area and the mouth area.
[0128] In addition, this application does not limit the method for obtaining the target key points corresponding to the above source image. For ease of understanding, the following is an explanation with examples.
[0129] As an example, in some application scenarios, in order to improve flexibility, when the facial key points of the above source image include facial key points within at least one sub-region, the target key points corresponding to the source image can be determined based on the distance between the implicit key points of the source image and the facial key points within the at least one sub-region, so that the target key points can represent the implicit key points that fall within the at least one sub-region, such as the eye region and / or the mouth region. It should be noted that this application does not limit the method for obtaining the distance. For example, it can be implemented using any existing or future distance calculation method, such as Euclidean distance or cosine distance.
[0130] It can be seen that in one possible implementation, for the above source image, after obtaining the implicit key points of the source image and the facial key points of the source image, the distance between each implicit key point and each facial key point can be calculated first; then, based on the distance, the implicit key points falling within the eye area and mouth area described by the facial key points are determined as the target key points corresponding to the source image.
[0131] In practice, for the implicit key points described above, although they lack precise semantic meaning, the implicit key point identification process can determine a point identifier for each implicit key point, such as a serial number, so that the implicit key points with that point identifier converge within a certain range, such as the mouth corner area. Because different implicit key points have different point identifiers, they converge within different ranges, allowing the point identifiers to be used to subsequently determine which implicit key points fall within the eye and mouth areas. Based on this, the present application also provides a possible implementation of the target key points corresponding to the source image described above. In this implementation, the target key point is determined based on the point identifier of at least one position-varying key point. The at least one position-varying key point refers to a predetermined key point that can characterize the characteristics of at least one sub-region described above. Furthermore, for any point identifier of a position-varying key point, the point identifier is used to identify the position-varying key point. Furthermore, the present application does not limit the process for determining the at least one position-varying key point, which may specifically include steps 211 through 213 below.
[0132] Step 211: Acquire a preset image sequence, where the preset image sequence is used to describe the change of at least one sub-region.
[0133] The preset image sequence refers to an image sequence required to be used when determining the point identifiers of the implicit key points to be deleted, such as a video.
[0134] In addition, the present application does not limit the implementation method of the above preset image sequence. For example, when the above at least one sub-area includes at least one of the eye area and the mouth area, the preset image sequence is used to describe the changes in the at least one sub-area, such as the changes in the opening and closing of the eyes and / or mouth. It can be seen that in one possible implementation method, if the at least one sub-area includes the eye area, the preset image sequence may include a pre-recorded blinking video, and the blinking video is used to describe the changes in the eyes. In another possible implementation method, if the at least one sub-area includes the mouth area, the preset image sequence may include a pre-recorded mouth movement video, and the mouth movement video is used to describe the changes in the mouth. In yet another possible implementation method, if the at least one sub-area includes the eye area and the mouth area, the preset image sequence may include a pre-recorded blinking video and a pre-recorded mouth movement video.
[0135] Step 212: performing key point position change analysis based on implicit key points of at least two images in a preset image sequence to obtain analysis results.
[0136] For the preset image sequence mentioned above, the t-th frame image in the preset image sequence refers to the image at the t-th arrangement position in the preset image sequence; and the implicit key points of the t-th frame image are used to describe some coordinate positions in the t-th frame image, so that the implicit key points are used to describe the distribution characteristics of the t-th frame image, so that the implicit key points can represent the posture characteristics of each area in the t-th frame image, and further enable the implicit key points to represent the position information described by the t-th frame image as comprehensively as possible. It should be noted that the implementation method of the implicit key points of the t-th frame image is similar to the implementation method of the implicit key points of the source image mentioned above. For the sake of brevity, it will not be repeated here. t is a positive integer, t≤T, T is a positive integer, and T represents the number of images in the preset image sequence.
[0137] The analysis results are used to indicate which implicit key points in the preset image sequence have not changed their position coordinates (or the degree of change is less than a threshold), and which implicit key points have changed their position coordinates continuously, so that the analysis results can indicate whether the position coordinates corresponding to the same point identifier in different frame images in the preset image sequence have changed.
[0138] Step 213: According to the above analysis results, determine at least one position-changed key point from the implicit key points of the at least two images.
[0139] The position-changing key points are used to represent implicit key points that are located at different positions in different frame images in a preset image sequence.
[0140] Based on the relevant content of steps 211 to 213 above, it can be seen that in some application scenarios, at least one position change key point mentioned above can be determined in advance using a preset image sequence, and the point identifiers of these position change key points can be recorded and stored. This allows the implicit key points falling within the at least one sub-region to be directly deleted using the point identifier during the subsequent image generation process, thereby improving image generation efficiency. Since the point identifiers of these position change key points are fixed and do not change with different processed data, in order to further improve efficiency, these position change key points can be pre-acquired and stored so that they can be directly read from the storage space later.
[0141] Based on the relevant content of step 21 above, it can be known that for the source image, after obtaining the implicit key points of the source image, the target key points corresponding to the source image can be determined from the implicit key points, so that the target key points can represent the implicit key points falling into part or all of the facial area, such as the eye area and the nose area, so that the target key points can represent the key points that need to be deleted.
[0142] Step 22: Determine the key point information of the source image based on the key points in the implicit key points of the source image except the target key points corresponding to the source image.
[0143] It should be noted that the present application does not limit the implementation method of the above step 22. For example, it can be specifically as follows: first delete the target key points corresponding to the source image from the implicit key points of the source image to obtain the remaining key points corresponding to the source image; then splice the remaining key points corresponding to the source image with the facial key points of the source image to obtain the key point information of the source image.
[0144] Based on the relevant content of steps 21 to 22 above, it can be known that for the above source image, after obtaining the implicit key points of the source image, the target key points corresponding to the source image are first determined from the implicit key points, so that the target key points can represent the hidden key points falling into certain areas, such as the eye area and the mouth area; then, based on the other key points in the implicit key points except the target key points, and the facial key points of the source image, the key point information of the source image is determined, so that the key point information includes the other key points and the facial key points, so that the image generation result of the above driving information can be determined based on the other key points. In this way, the posture interference caused by the target key points can be effectively avoided, so that the key point information can represent the posture characteristics of each area in the source image as accurately as possible, which is conducive to improving the image generation effect.
[0145] Based on the relevant content of S21 above, it can be known that for the above source image, after obtaining the source image, the implicit key points of the source image and the facial key points of the source image can be determined first; then, based on the implicit key points of the source image and the facial key points of the source image, the key point information of the source image can be determined, so that the key point information can represent the posture characteristics of each area in the source image as accurately as possible.
[0146] S22: Determine key point information of the driving information based on the implicit key points of the driving information and the facial key points of the driving information.
[0147] The key point information of the driving information is used to characterize the posture information in the driving information; and when the driving information includes N frames of information, the key point information of the driving information may include the key point information of the N frames of information. The key point information of the nth frame information is used to describe the posture information in the nth frame information, and the key point information of the nth frame information is determined based on the implicit key points of the nth frame information and the facial key points of the nth frame information, so that the key point information of the nth frame information includes the facial key points and the key point information of the nth frame information includes some or all of the implicit key points, so that the key point information of the nth frame information can represent the posture characteristics of the nth frame information as comprehensively as possible, such as the posture characteristics of various body parts in the nth frame information. n is a positive integer, n≤N, and N is a positive integer. It should be noted that the implementation method of the key point information of the nth frame information is similar to the implementation method of the key point information of the source image above. For the sake of brevity, it will not be repeated here.
[0148] Thus, in one possible implementation, for the nth frame information above, when the implicit key points of the nth frame information include implicit key points within at least two regions, and the at least two regions include a facial region and at least one non-facial region, the key point information of the nth frame information includes the implicit key points within the at least one non-facial region and the facial key points of the nth frame information. This effectively avoids posture interference caused by certain key points in the nth frame information's implicit key points, thereby improving image generation. n is a positive integer, n≤N, and N is a positive integer.
[0149] In addition, the implementation of S22 is similar to that of S21, and for the sake of brevity, it will not be repeated here. It can be seen that in a possible implementation, if the driving information includes N frame information, then S22 may specifically include the following steps 31 and 32.
[0150] Step 31: Determine the target key point corresponding to the n-th frame information from the implicit key points of the n-th frame information, so that the target key point matches part or all of the face area, where n is a positive integer, n≤N, and N is a positive integer.
[0151] The target key points corresponding to the n-th frame information refer to some or all of the implicit key points existing in the n-th frame information and falling within the facial region.
[0152] In addition, the implementation of the target key point corresponding to the n-th frame information is similar to the implementation of the target key point corresponding to the source image described above, and for the sake of brevity, it is not further described here. For example, the target key point corresponding to the n-th frame information can be matched with at least one sub-region within the face region, and the at least one sub-region includes at least one of the eye region and the mouth region, so that the target key point corresponding to the n-th frame information can be an implicit key point existing in the implicit key points of the n-th frame information that falls within at least one of the eye region and the mouth region.
[0153] In addition, the method for obtaining the target key points corresponding to the n-th frame information above is similar to the method for obtaining the target key points corresponding to the source image above. For the sake of brevity, it will not be repeated here. For example, when the facial key points of the n-th frame information include facial key points within at least one sub-region, the target key points corresponding to the n-th frame information can be determined based on the distance between the implicit key points of the n-th frame information and the facial key points within the at least one sub-region, so that the target key points can represent the implicit key points falling within the at least one sub-region. For another example, the target key points corresponding to the n-th frame information are determined based on the point identifier of the at least one position change key point above. It should be noted that the relevant content of the at least one sub-region and the at least one position change key point can be found above.
[0154] Based on the relevant content of step 31 above, for the nth frame information above, after obtaining the implicit key points of the nth frame information, the target key points corresponding to the nth frame information can be determined from the implicit key points, so that the target key points can represent the implicit key points falling within part or all of the facial area, such as the eye area and the nose area, so that the target key points can indicate the key points that need to be deleted. n is a positive integer, n≤N, and N is a positive integer.
[0155] Step 32: Determine the key point information of the nth frame information based on the key points in the nth frame information except the target key point corresponding to the nth frame information, where n is a positive integer, n≤N, and N is a positive integer.
[0156] It should be noted that the present application does not limit the implementation of step 32 above. For example, it may specifically include: first, deleting the target key point corresponding to the n-th frame information from the implicit key points of the n-th frame information to obtain the remaining key points corresponding to the n-th frame information; then, concatenating the remaining key points corresponding to the n-th frame information with the facial key points of the n-th frame information to obtain the key point information of the n-th frame information. n is a positive integer, n≤N, and N is a positive integer.
[0157] Based on the relevant contents of steps 31 to 32 above, it can be seen that for the nth frame information in the above driving information, after obtaining the implicit key points of the nth frame information, the target key points corresponding to the nth frame information are first determined from the implicit key points, so that the target key points can represent the hidden key points falling within certain areas, such as the eye area and the mouth area; then, based on the other key points in the implicit key points except the target key points corresponding to the nth frame information and the facial key points of the nth frame information, the key point information of the nth frame information is determined, so that the key point information includes the other key points and the facial key points, so that the image generation result of the above driving information can be determined based on the other key points. In this way, the posture interference caused by the target key points can be effectively avoided, so that the key point information can represent the posture characteristics of each area in the nth frame information as accurately as possible, thereby improving the image generation effect. n is a positive integer, n≤N, and N is a positive integer.
[0158] In addition, this application does not limit the correlation between the execution time of S22 and the execution time of S21. For example, the former is earlier than the latter. Another example is that the latter is earlier than the former. Another example is that the two are the same.
[0159] Based on the relevant content of S22 above, it can be seen that for the above driving information, when the driving information includes the first frame information, the second frame information, ..., and the Nth frame information, after obtaining the nth frame information, the implicit key points of the nth frame information and the facial key points of the nth frame information can be first determined; then, based on the implicit key points of the nth frame information and the facial key points of the nth frame information, the key point information of the nth frame information can be determined so that the key point information can accurately represent the posture characteristics of each region in the nth frame information as much as possible. n is a positive integer, n≤N, and N is a positive integer.
[0160] S23: performing image generation processing according to the source image, the key point information of the source image, and the key point information of the driving information to obtain an image generation result corresponding to the driving information.
[0161] It should be noted that this application does not limit the implementation of the above S23.
[0162] In addition, in order to better improve the image generation effect, the present application also provides a possible implementation method of the process of determining the image generation result corresponding to the nth frame information above. Under this implementation method, the process of determining the image generation result corresponding to the nth frame information may include the following steps 41-44.
[0163] Step 41: Determine deformation reference data corresponding to the n-th frame information based on the key point information of the source image and the key point information of the n-th frame information.
[0164] Among them, the deformation reference data corresponding to the nth frame information refers to the parameters required to perform posture adjustment processing on the source image based on the nth frame information, such as expression adjustment processing and / or head posture adjustment processing, such as posture offset + area that cannot be generated by deformation, etc.
[0165] In addition, this application does not limit the implementation method of the deformation reference data corresponding to the n-th frame information above. For example, it may include a two-dimensional geometric deformation field (Deformations) and an occlusion map (Occlusion Map). The two-dimensional geometric deformation field is used to describe the posture offset of the area that can be generated by deformation, that is, the posture offset of the deformable area. The occlusion map is used to describe the position of the area that cannot be generated by deformation, that is, the position of the non-deformable area.
[0166] In addition, the present application does not limit the implementation method of the above step 41. For example, it can be implemented using any existing or future method that can calculate the offset based on two posture data. For another example, step 41 can be implemented using a deformation network (DMN). It can be seen that under one possible implementation method, step 41 can specifically be: inputting the key point information of the source image and the key point information of the n-th frame information into the DMN, obtaining the deformations and occlusion maps output by the DMN as the deformation reference data corresponding to the n-th frame information, so that the deformation reference data includes the deformations and occlusion maps.
[0167] It can be seen that for some application scenarios, after obtaining the key point information of the source image and the key point information of the n-th frame information, the key point information of the source image and the key point information of the n-th frame information can be input into the DMN to obtain the two-dimensional geometric deformation field and occlusion map output by the DMN as the deformation reference data corresponding to the n-th frame information, so that the deformation reference data can represent the parameters required to perform deformation processing on the source image.
[0168] Research has found that for some image sequence generation scenarios, the posture information, such as facial expression status, between different frame images in the generated image sequence needs to show relatively large changes, but the changes in other information, such as background, clothing, etc. are relatively small or even no changes. Therefore, in order to better improve the generation effect of the image sequence, this application also provides a possible implementation method of the process of determining the deformation reference data corresponding to the nth frame information above. Under this implementation method, if n≥2, the process of determining the deformation reference data corresponding to the nth frame information includes the following steps 411-412.
[0169] Step 411: Determine the predicted deformation parameters corresponding to the n-th frame information based on the key point information of the source image and the key point information of the n-th frame information.
[0170] The predicted deformation parameters corresponding to the n-th frame information refer to the parameters predicted to be used when performing posture adjustment processing on the source image based on the n-th frame information, such as the posture offset + the area that cannot be generated by deformation.
[0171] In addition, the present application does not limit the implementation method of the above step 411. For example, it can be specifically: inputting the key point information of the source image and the key point information of the n-th frame information into the DMN, obtaining the deformations and occlusion map output by the DMN as the predicted deformation parameters corresponding to the n-th frame information, so that the predicted deformation parameters include the deformations and occlusion map.
[0172] Based on the relevant content of step 411 above, it can be known that in some application scenarios, the predicted deformation parameters corresponding to the n-th frame information are determined based on the implicit key points of the source image, the facial key points of the source image, the implicit key points of the n-th frame information, and the facial key points of the n-th frame information, so that the predicted deformation parameters refer to the parameters predicted to be used when the source image is subjected to posture adjustment processing based on the n-th frame information.
[0173] Step 412: Based on the weight corresponding to at least one region in the source image, the predicted deformation parameter corresponding to the n-th frame information and the deformation reference data corresponding to the n-1-th frame information are weighted to obtain the deformation reference data corresponding to the n-th frame information; the at least one region is determined based on the regional analysis result of the source image; the regional analysis result is used to describe the position of the at least one region in the source image.
[0174] Among them, the regional analysis result of the source image refers to the result obtained by analyzing the source image, so that the regional analysis result is used to describe the positions of different regions in the source image, such as the position of the background in the source image, the positions of different body parts in the source image, the positions of different facial parts in the source image, etc.
[0175] In addition, the present application does not limit the process of obtaining the above-mentioned regional analysis results. For example, it can be implemented using any existing or future method that can divide an image into different regions, such as a body part recognition method, a facial part recognition method, or an image segmentation method.
[0176] In addition, for the above source image, at least one area in the source image refers to the area described by the regional analysis result of the source image, so that the at least one area can represent some areas analyzed from the source image; and the weight corresponding to the at least one area refers to the weight required to be used when smoothing different areas, so that the weight corresponding to the at least one area can represent the degree of smoothing performed on different areas, such as performing a greater degree of smoothing on the background, clothing and other areas. This is conducive to ensuring the stability of different frame images in these areas, thereby effectively avoiding the impact caused by the instability of these areas, thereby helping to improve the generation effect of the image sequence.
[0177] It can be seen that in one possible implementation, if the region parsing results of the source image described above are used to describe the locations of K regions in the source image, then at least one region in the source image may include K regions. The weight corresponding to the kth region is used to characterize the degree of smoothing for the kth region. Furthermore, this application does not limit the implementation of the weight corresponding to the kth region; for example, it may include the weight value of a historical frame and the weight value of a current frame. The weight value of the historical frame refers to the degree of influence of the historical frame when smoothing the kth region. The weight value of the current frame refers to the degree of influence of the current frame when smoothing the kth region. Furthermore, this application does not limit the process for determining the weight corresponding to the kth region; for example, it may specifically be: based on the category of the kth region, searching a pre-established mapping relationship for the smoothing weight corresponding to the category, and using this as the weight corresponding to the kth region. The category of the kth region indicates the identity of the kth region, such as eyes, hair, etc. This mapping relationship is used to record the smoothing weights corresponding to multiple categories. k is a positive integer, k≤K, and K is a positive integer.
[0178] The deformation reference data corresponding to the n-1th frame information refers to the historical information required to be referenced when smoothing the deformation reference data corresponding to the nth frame information above; and the determination process of the deformation reference data corresponding to the n-1th frame information is similar to the determination process of the deformation reference data corresponding to the nth frame information.
[0179] In addition, the present application does not limit the implementation method of the above step 412. For example, it can be specifically: based on the weight corresponding to at least one area in the source image, the predicted deformation parameters corresponding to the n-th frame information and the deformation reference data corresponding to the n-1-th frame information are weighted to obtain the deformation reference data corresponding to the n-th frame information, so that the deformation parameters of some areas in the deformation reference data corresponding to the n-th frame information, such as the background, clothes, etc., are as consistent as possible with the deformation parameters of the corresponding areas in the deformation reference data corresponding to the n-1-th frame information, and the deformation parameters of other areas in the deformation reference data corresponding to the n-th frame information, such as the eyes, mouth, etc., are as consistent as possible with the deformation parameters of the corresponding areas in the predicted deformation parameters corresponding to the n-1-th frame information. This is conducive to ensuring that areas such as the background and clothes remain stable in images generated from different frames, and ensuring that areas such as the eyes and mouth change flexibly in images generated from different frames, thereby helping to improve the generation quality of the image sequence.
[0180] Research has found that some pose information in Deformations requires smoothing, such as clothing and background, but almost no pose information in the Occlusion Map requires smoothing. Therefore, to further improve efficiency, this application also provides a possible implementation of step 412 above. In this implementation, when the predicted deformation parameters corresponding to the n-th frame information include Deformations and the Occlusion Map, step 412 may specifically include: performing weighted processing on the Deformations in the predicted deformation parameters corresponding to the n-th frame information and the Deformations in the deformation reference data corresponding to the (n-1)-th frame information according to the weight corresponding to at least one region in the source image to obtain the Deformations in the deformation reference data corresponding to the n-th frame information. In this way, pose adjustment is performed only on the Deformations, which effectively avoids the resource overhead caused by pose adjustment on the Occlusion Map, thereby improving efficiency.
[0181] Based on the relevant contents of steps 411 to 412 above, it can be seen that for some application scenarios, after obtaining the key point information of the source image and the key point information of the n-th frame information, the key point information of the source image and the key point information of the n-th frame information can be first input into the DMN to obtain deformations and an occlusion map output by the DMN. Then, based on the regional analysis result of the source image and the deformations in the deformation reference data corresponding to the n-1-th frame information above, the deformations output by the DMN are smoothed. Subsequently, the smoothed deformations and the occlusion map output by the DMN can be used as the deformation reference data corresponding to the n-1-th frame information, so that the deformation reference data can better represent the parameters required when performing posture adjustment processing on the source image based on the n-1-th frame information, such as the posture offset + the area that cannot be generated by deformation.
[0182] In addition, to further improve accuracy, the present application also provides a possible implementation of step 41 above. In this implementation, step 41 may specifically include: determining the deformation reference data corresponding to the nth frame information based on the source image, key point information of the source image, at least one reference feature of the source image, and key point information of the nth frame information. The source image can provide some reference image information for the deformation reference data determination process, thereby making the deformation reference data determined based on the source image more accurate, thereby facilitating improved accuracy.
[0183] Based on the relevant content of step 41 above, it can be known that for some application scenarios, for the nth frame information, the deformation reference data corresponding to the nth frame information can be determined based on the implicit key points of the source image, the facial key points of the source image, the implicit key points of the nth frame information, and the facial key points of the nth frame information, so that the deformation reference data can better represent the parameters required to perform posture adjustment processing on the source image based on the nth frame information, thereby making the image generated based on the deformation reference data more accurate, which is beneficial to improving the image generation effect.
[0184] Step 42: Perform deformation processing on the source image according to the deformation reference data corresponding to the n-th frame information to obtain a deformed image corresponding to the n-th frame information.
[0185] Among them, the deformed image corresponding to the nth frame information refers to the result obtained by deforming the source image based on the deformation reference data corresponding to the nth frame information, so that the posture information in the deformable area of the deformed image corresponding to the nth frame information is consistent with the posture information in the deformable area of the nth frame information, and the posture information in the non-deformable area of the deformed image corresponding to the nth frame information is consistent with the posture information in the non-deformable area of the source image.
[0186] Based on the relevant content of step 42 above, it can be seen that for some application scenarios, after obtaining the deformation reference data corresponding to the n-th frame information above, such as the Deformations+Occlusion Map, the Deformations can be used to perform pixel-by-pixel geometric deformation on the source image, and the Occlusion Map can be used to occlude areas that cannot be generated by deformation, thereby obtaining the deformed image corresponding to the n-th frame information.
[0187] Step 43: Completing the deformed image corresponding to the n-th frame information to obtain the completed image corresponding to the n-th frame information and the completed position representation data corresponding to the n-th frame information.
[0188] Among them, the completed image corresponding to the nth frame information refers to the image obtained by completing the non-deformable area in the deformed image corresponding to the nth frame information, so that the posture information of the non-deformable area in the completed image corresponding to the nth frame information is consistent with the posture information of the non-deformable area in the nth frame information.
[0189] The complement position representation data corresponding to the n-th frame information is used to describe the position to be complemented when the deformed image corresponding to the n-th frame information is complemented, such as the position of the non-deformable area; and the present application does not limit the implementation method of the complement position representation data corresponding to the n-th frame information, for example, it can be represented by a mask image, such as by Figure 3 The alpha channel shown is implemented.
[0190] Furthermore, this application does not limit the implementation of step 43 above; for example, it can be implemented using a completion network. The completion network is used to perform the completion process; and this application does not limit the implementation of the completion network; for example, it can be implemented using any existing or future image generation network, such as an end-to-end P2P network. Thus, in one possible implementation, the completion network can be implemented using a generator network.
[0191] Based on the relevant content of step 43 above, it can be known that for some application scenarios, after obtaining the deformed image corresponding to the n-th frame information, the deformed image corresponding to the n-th frame information can be input into P2P to obtain the predicted image and α channel output by the P2P, and the predicted image is used as the completed image corresponding to the n-th frame information, and the α channel is used as the completed position representation data corresponding to the n-th frame information.
[0192] Step 44: Based on the completed position representation data corresponding to the n-th frame information, the reference image corresponding to the n-th frame information and the completed image corresponding to the n-th frame information are fused to obtain the image generation result corresponding to the n-th frame information; the reference image is determined based on the source image and the deformation reference data corresponding to the n-th frame information.
[0193] The reference image corresponding to the n-th frame information is used to influence the image information within the deformable area; and the reference image is determined based on the source image and the deformation reference data corresponding to the n-th frame information.
[0194] In addition, the present application does not limit the implementation method of the reference image corresponding to the n-th frame information above. For example, in order to improve efficiency, the reference image corresponding to the n-th frame information can be implemented using the deformed image corresponding to the n-th frame information.
[0195] For another example, in order to better improve the image generation quality, the present application also provides a process for obtaining the reference image corresponding to the nth frame information above, which can be specifically referred to as steps 441 to 443 below.
[0196] Step 441: Perform super-resolution processing on the source image to obtain a super-resolution image.
[0197] Among them, super-resolution processing is used to improve the image quality of an image; and this application does not limit the implementation method of the super-resolution processing. For example, it can be implemented using any existing or future super-resolution implementation method.
[0198] Based on the relevant content of step 441 above, it can be known that for the above source image, after obtaining the source image, super-resolution processing can be performed on the source image to obtain a super-resolution image, so that the super-resolution image can have better image quality, so that the super-resolution image can better describe the image information in the source image, such as appearance information.
[0199] Step 442: Perform deformation processing on the super-resolution image according to the deformation reference data corresponding to the n-th frame information to obtain the super-resolution deformation result corresponding to the n-th frame information.
[0200] Among them, the super-resolution deformation result corresponding to the n-th frame information refers to the image obtained by deforming the super-resolution image based on the deformation reference data corresponding to the n-th frame information, so that the posture information of the deformable area in the super-resolution deformation result is consistent with the posture information of the deformable area in the n-th frame information.
[0201] In addition, the present application does not limit the implementation of the above step 442. For example, when the deformation reference data corresponding to the above n-th frame information includes Deformations and Occlusion Map, the step 442 can be specifically as follows: deforming the super-resolved image according to the Deformations to obtain the super-resolved deformation result corresponding to the n-th frame information, such as Figure 2 or Figure 3 Deformation 2 shown.
[0202] Step 443: Determine a reference image corresponding to the n-th frame information according to the super-fractional deformation result corresponding to the n-th frame information.
[0203] It should be noted that the present application does not limit the implementation of the above step 443. For example, it can specifically be: directly determining the super-fractional deformation result corresponding to the n-th frame information as the reference image corresponding to the n-th frame information.
[0204] Based on the relevant content of steps 441 to 443 above, it can be seen that in some application scenarios, the reference image corresponding to the nth frame information above can be determined based on the super-resolution processing result of the source image and the deformation reference data corresponding to the nth frame information, so that the reference image can provide more accurate image information within the deformable area, thereby making the image generation result corresponding to the nth frame information determined based on the reference image have better quality, which is conducive to improving the image generation quality.
[0205] In addition, the present application does not limit the implementation of the above step 44. For ease of understanding, some examples are provided below for illustration.
[0206] Example 1. In a possible implementation, the above step 44 can be specifically as follows: first, based on the completed position representation data corresponding to the n-th frame information, delete the pixel points in the non-deformable area from the reference image corresponding to the n-th frame information, and based on the completed position representation data corresponding to the n-th frame information, extract the pixel points in the non-deformable area from the completed image corresponding to the n-th frame information; then, fill the extracted pixel points into the image after the pixel points are deleted to obtain the image generation result corresponding to the n-th frame information, so that the pixel points in the non-deformable area in the image generation result corresponding to the n-th frame information come from the completed image corresponding to the n-th frame information, and the pixel points in the deformable area in the image generation result corresponding to the n-th frame information come from the reference image. In this way, the influence of the completion processing on the deformable area can be effectively avoided, thereby helping to improve the image generation quality.
[0207] Example 2. In one possible implementation, step 44 above may specifically be: first, based on the completed position representation data corresponding to the n-th frame information, determine the weights corresponding to each pixel point in the reference image corresponding to the n-th frame information and the weights corresponding to each pixel point in the completed image corresponding to the n-th frame information; then, based on these weights, perform weighted processing on the reference image and the completed image to obtain the image generation result corresponding to the n-th frame information, so as to realize the fusion of the pixel points in the non-deformable area of the completed image into the corresponding area in the reference image, which is conducive to improving the image generation quality.
[0208] Based on the relevant contents of steps 41 to 44 above, it can be known that for some application scenarios, after obtaining the key point information of the source image and the key point information of the n-th frame information, these two data can be first input into the DMN to obtain the Deformations and Occlusion Map predicted by the DMN, and according to the regional analysis results of the source image and the Deformations corresponding to the n-1-th frame information, different regions in the Deformations corresponding to the n-th frame information are smoothed to different degrees to obtain smoothed Deformations, so that the smoothed Deformations are used to perform pixel-by-pixel geometric deformation on the source image, and obtain a deformation result whose posture is close to the posture information in the n-th frame information, and make the Occlusion Map is used to block some areas that cannot be generated by deformation, such as eyes, mouths, etc.; then, after P2P receives the deformed source image that blocks the area to be generated, the P2P performs a refined completion process on the deformed source image, obtains and outputs a predicted image that completes the corresponding area and an alpha channel used to describe the position of the completed area, so that the completed area in the predicted image can be fused to a certain deformation result of the source image based on the alpha channel. Figure 3 In the deformation 1 or deformation 2 shown, the image generation result corresponding to the n-th frame information is obtained, which is beneficial to improving the image quality.
[0209] In addition, in some application scenarios, the above S23 can be implemented with the help of a preset model. Based on this, it can be seen that in one possible implementation, the S23 can be specifically: the preset model performs image generation processing based on the source image, the key point information of the source image, and the key point information of the driving information, and obtains and outputs the image generation result corresponding to the driving information. Among them, the preset model is used to perform image generation processing on the input data of the preset model, and this application does not limit the implementation method of the preset model. For example, the preset model can include a deformation network and a completion network.
[0210] Based on the relevant content of S23 above, it can be known that in some application scenarios, after obtaining the source image and driving information, first, the key point information of the source image is determined based on the implicit key points of the source image and the facial key points of the source image, so that the key point information of the source image can describe the posture presented by the source image as comprehensively as possible, such as facial state, hair posture, hat posture, etc., and the key point information of the driving information is determined based on the implicit key points of the driving information and the facial key points of the driving information, so that the key point information of the driving information can describe the posture represented by the driving information as much as possible; then, image generation processing is performed based on the source image, the key point information of the source image and the key point information of the driving information to obtain the image generation result corresponding to the driving information, so that the appearance information presented by the image generation result is consistent with the appearance information presented by the source image, and the posture presented by the image generation result is consistent with the posture represented by the driving information, so that the driving result of the source image under the driving information can be automatically generated. Among them, because the facial key point is used to describe the facial area, and the implicit key point is used to supplement the key points in other areas except the facial area, the key point information determined based on the facial key point and the implicit key point can not only better represent the posture of the facial area, but also better represent the posture of some non-facial areas, such as the hair area, the hat area and other areas, so that the key point information can represent the posture of each area as comprehensively as possible, and then the image generation result determined based on the key point information can present the various postures represented by the driving information as comprehensively as possible, so that it can effectively avoid defects caused by only focusing on the facial state, such as incoordination, etc., which is conducive to improving the image generation quality.
[0211] Based on the relevant content of the image generation result corresponding to the driving information above, it can be known that the image generation result can be determined based on part or all of the implicit key points, so that in the process of determining the image generation result, attention can be paid to some non-face areas, such as the hair area, hat area and other areas. The posture conditions, so that the image generation process can pay attention to the posture conditions of each area as comprehensively as possible, and then the final image generation result can present the various posture conditions represented by the driving information as comprehensively as possible, which is conducive to improving the image generation quality.
[0212] Based on the relevant contents of S1 to S2 above, it can be seen that for the image generation method provided in the embodiment of the present application, after obtaining the source image and driving information, image generation processing is performed based on the source image, the implicit key points of the source image, the facial key points of the source image, the implicit key points of the driving information, and the facial key points of the driving information to obtain the image generation result corresponding to the driving information, so that the appearance information presented by the image generation result is consistent with the appearance information presented by the source image, and the posture presented by the image generation result is consistent with the posture represented by the driving information, so that the driving result of the source image under the driving information can be automatically generated. Among them, because the facial key points are used to describe the facial area, and the implicit key points are used to supplement the key points in other areas except the facial area, when performing image generation processing based on the facial key points and the implicit key points, not only the posture of the facial area can be paid attention to, but also the posture of some non-facial areas, such as the hair area, the hat area and other areas can be paid attention to, so that the image generation process can pay attention to the posture of each area as comprehensively as possible, and then the final image generation result can present the various postures represented by the driving information as comprehensively as possible, so that defects caused by only paying attention to the facial state, such as incoordination, can be effectively avoided, which is conducive to improving the image generation quality.
[0213] In addition, the present application does not limit the execution subject of the image generation method provided in the embodiment of the present application. For example, the image generation method provided in the embodiment of the present application can be applied to a terminal device. For another example, the image generation method provided in the embodiment of the present application can also be implemented with the help of a data interaction process between a terminal device and a server. Among them, the terminal device can be a smart phone, a computer, a personal digital assistant (PDA), a tablet computer, etc. The server can be a stand-alone server, a cluster server, or a cloud server.
[0214] In addition, this application does not limit the application scenarios of the above image generation method. For ease of understanding, the following description is combined with three scenarios.
[0215] Scenario 1: The image generation method provided in this application can be applied to some model update scenarios, such as online or offline model update scenarios. Based on this, it can be seen that this application also provides a model update process, which can specifically include the following steps 51-53.
[0216] Step 51: Obtain source image and driving information.
[0217] It should be noted that the relevant contents of step 51 can be referred to the relevant contents of S1 above. For example, step 51 may specifically be: randomly extracting two frames of images from the sample video, one frame of image being used as the source image and the other frame of image being used as the driving information.
[0218] Step 52: A preset model performs image generation processing based on the source image, the implicit key points of the source image, the facial key points of the source image, the implicit key points of the driving information, and the facial key points of the driving information to obtain an image generation result corresponding to the driving information; the facial key points are used to describe the facial area; the implicit key points are used to supplement the key points in other areas except for at least one sub-area in the facial area; the implicit key points are determined using the implicit key point detection model.
[0219] The implicit key point detection model is used to perform implicit key point detection on input data of the implicit key point detection model. Thus, in one possible implementation, the implicit key points of the source image described above are obtained by performing implicit key point detection on the source image using the implicit key point detection model, and the implicit key points of the driving information described above can be obtained by performing implicit key point detection on the driving information using the implicit key point detection model.
[0220] In addition, please see above for relevant content of preset models.
[0221] Step 53: Based on the image generation result corresponding to the above driving information and the label information corresponding to the image generation result, update the implicit key point detection model and the preset model, and return to continue executing the above step 51 and its subsequent steps until the preset stop condition is reached; the label information is determined based on the driving information.
[0222] Among them, for the image generation result corresponding to the above-mentioned driving information, the label information corresponding to the image generation result is used to represent the true value of the image corresponding to the driving information, so that the label information can serve as guidance information corresponding to the image generation result; and this application does not limit the implementation method of the label information, for example, the label information can be implemented using the driving information.
[0223] In addition, the present application does not limit the implementation method of the above step 53. For example, it can be specifically as follows: first, determine the model loss based on the similarity between the image generation result corresponding to the above driving information and the label information corresponding to the image generation result, so that the model loss can represent the model performance; then update the implicit key point detection model and the preset model based on the model loss, and return to continue executing the above step 51 and its subsequent steps until the preset stop condition is reached.
[0224] Furthermore, the present application does not limit the implementation of the preset stop condition. For example, the preset stop condition may specifically include: the model loss is lower than a preset loss threshold. For another example, the preset stop condition may include: the rate of change of the model loss is lower than a preset rate of change threshold. For another example, the preset stop condition may include: the number of model updates is higher than a preset number threshold.
[0225] Based on the relevant content of steps 51 to 53 above, it can be seen that for some model update scenarios, two different frames of images are first selected from the same video as the source image and driving image of the current round; then the implicit key point detection model, the preset model, the source image and the driving image are used to obtain the image generation result; then, based on the difference between the image generation result and the driving image, the implicit key point detection model and the preset model are updated so that the updated implicit key point detection model and the preset model have better performance, and the next round of update process is performed based on the updated implicit key point detection model and the preset model, and the iterative cycle is repeated until the preset stop condition is reached, which is conducive to improving model performance.
[0226] Scenario 2: The image generation method provided in this application can be applied to perform some image generation tasks, such as the generation task of a single image or the generation task of an image sequence.
[0227] In fact, in order to better reduce the resource pressure of the client, the image generation method provided by this application can be completed by the server and the client in a collaborative manner. For ease of understanding, the following description is made in conjunction with the generation process of the image sequence.
[0228] As an example, the image sequence generation process provided in this application may specifically include the following steps 61 to 64.
[0229] Step 61: After the server obtains the source image, the server obtains the implicit key points of the source image and the facial key points of the source image.
[0230] Step 62: After the server obtains the nth frame information in the driving information, the server obtains the implicit key points of the nth frame information and the facial key points of the nth frame information, where n is a positive integer and n≤N.
[0231] Step 63: After the client receives the implicit key points of the source image, the facial key points of the source image, the implicit key points of the n-th frame information, and the facial key points of the n-th frame information sent by the server, the client performs image generation processing based on the source image, the implicit key points of the source image, the facial key points of the source image, the implicit key points of the n-th frame information, and the facial key points of the n-th frame information to obtain an image generation result of the n-th frame information, such as Figure 3 The nth frame is shown as the generated image, where n is a positive integer, n≤N.
[0232] Based on the relevant contents of steps 61 to 63 above, it can be seen that for some execution scenarios of generation tasks, the image generation method provided by this application can be implemented in a collaborative manner with the help of the server and the client. Among them, because the source image corresponding to different frame information in a task is the same image, the server can be used to complete the relevant processing for the source image, so that the relevant content of the source image can be provided to the client in a one-time manner. This can avoid the resource overhead caused by the client completing the relevant processing for the source image, thereby helping to reduce the resource pressure on the client. In addition, because the relevant processing of each frame information in the driving information has a large resource overhead, the server can be used to complete the relevant processing for each frame information, so that the server can subsequently continuously send the relevant content of different frame information to the client. This can avoid the resource overhead caused by the client completing the relevant processing for each frame information, thereby helping to reduce the resource pressure on the client. In addition, because sending images requires a large amount of communication resources, in order to balance the resource overhead as much as possible, the client can perform image generation processing based on multiple key point information, which is conducive to improving the image generation effect.
[0233] Scenario three: the image generation method provided in this application can be used to generate a video of a virtual image; and the video generation process can include the following steps 71 to 73.
[0234] Step 71: Determine the source image based on the pre-built virtual image.
[0235] The term "avatar" refers to a pre-created virtual image for a user. This application does not limit the implementation of this virtual image; for example, it could be a three-dimensional digital human. Furthermore, this application does not limit the method for obtaining this virtual image; for example, it could be determined based on user-provided information, such as an image. Alternatively, the virtual image could be selected by the user from a selection of candidate virtual images.
[0236] In addition, the present application does not limit the implementation method of the above step 71. For example, it can specifically be: taking a photo of the virtual image to obtain the source image.
[0237] Step 72: After acquiring the above text data, after receiving the text data, convert the text data into voice data, and determine the driving information based on the voice data.
[0238] It should be noted that, for the relevant content of step 72, please refer to the above.
[0239] Step 73: Perform image generation processing based on the source image, the implicit key points of the source image, the facial key points of the source image, the implicit key points of the driving information, and the facial key points of the driving information to obtain an image generation result corresponding to the driving information, so that the image generation result includes a generated image corresponding to each frame of data in the voice data.
[0240] Step 74: Based on the above voice data and the above image generation result, construct a video corresponding to the above virtual image, so that the video is used to describe the state of the virtual image in different frames.
[0241] It should be noted that this application does not limit the implementation method of the above step 74.
[0242] Based on the relevant content of steps 71 to 74 above, it can be seen that in some application scenarios, the image generation method provided in this application can be used to determine a driving video for a virtual character. In particular, because the image generation method has good performance, the resulting driving video can better meet user needs, which is conducive to improving the user experience.
[0243] Based on the image generation method provided in the embodiment of the present application, the embodiment of the present application also provides an image generation device. Figure 4 Explain and illustrate. Figure 4 This is a schematic diagram of the structure of an image generation device provided in an embodiment of the present application. It should be noted that for the technical details of the image generation device provided in an embodiment of the present application, please refer to the relevant content of the image generation method above.
[0244] like Figure 4 As shown, the image generation device 400 provided in the embodiment of the present application includes:
[0245] The data acquisition unit 401 is used to acquire source images and driving information;
[0246] An image generation unit 402 is configured to perform image generation processing based on the source image, implicit key points of the source image, facial key points of the source image, implicit key points of the driving information, and facial key points of the driving information to obtain an image generation result corresponding to the driving information; the facial key points are used to describe the facial area; and the implicit key points are used to supplement key points in other areas except for at least one sub-area of the facial area.
[0247] In a possible implementation manner, the image generation result is determined based on part or all of the implicit key points.
[0248] In one possible implementation, the implicit key points include implicit key points in at least two areas; the at least two areas include the facial area and at least one non-facial area; and the image generation result is determined based on the implicit key points in the at least one non-facial area.
[0249] In a possible implementation manner, a target key point among the implicit key points partially or completely matches the facial region; and the image generation result is determined based on other key points among the implicit key points except the target key point.
[0250] In a possible implementation, the target key point is matched with the at least one sub-region; the at least one sub-region includes at least one of an eye region and a mouth region.
[0251] In a possible implementation manner, the facial key points include facial key points within the at least one sub-region; and the target key points are determined based on a distance between the implicit key points and the facial key points within the at least one sub-region.
[0252] In one possible implementation, the target key point is determined based on a point identifier of at least one position change key point;
[0253] The process of determining the at least one position change key point includes: obtaining a preset image sequence, which is used to describe the changes in the at least one sub-area; performing key point position change analysis based on the implicit key points of at least two images in the preset image sequence to obtain analysis results; and determining the at least one position change key point from the implicit key points of the at least two images based on the analysis results.
[0254] In a possible implementation manner, the driving information includes N frames of information, where N is a positive integer;
[0255] The image generation unit 403 is specifically used to: determine the deformation reference data corresponding to the n-th frame information based on the implicit key points of the source image, the facial key points of the source image, the implicit key points of the n-th frame information, and the facial key points of the n-th frame information; n is a positive integer, n≤N; based on the deformation reference data corresponding to the n-th frame information, deform the source image to obtain the deformed image corresponding to the n-th frame information; perform complement processing on the deformed image corresponding to the n-th frame information to obtain the complemented image corresponding to the n-th frame information and the complemented position representation data corresponding to the n-th frame information; based on the complemented position representation data corresponding to the n-th frame information, fuse the reference image corresponding to the n-th frame information with the complemented image corresponding to the n-th frame information to obtain the image generation result corresponding to the n-th frame information; the reference image is determined based on the source image and the deformation reference data corresponding to the n-th frame information.
[0256] In a possible implementation manner, the reference image is the deformed image corresponding to the n-th frame information.
[0257] In one possible implementation, the process of acquiring the reference image includes: performing super-resolution processing on the source image to obtain a super-resolution image; performing deformation processing on the super-resolution image based on the deformation reference data corresponding to the n-th frame information to obtain a super-resolution deformation result corresponding to the n-th frame information; and determining the reference image corresponding to the n-th frame information based on the super-resolution deformation result corresponding to the n-th frame information.
[0258] In one possible implementation, the image generation unit 403 is specifically used to: if n≥2, determine the predicted deformation parameters corresponding to the n-th frame information based on the implicit key points of the source image, the facial key points of the source image, the implicit key points of the n-th frame information, and the facial key points of the n-th frame information; based on the weight corresponding to at least one area in the source image, perform weighted processing on the predicted deformation parameters corresponding to the n-th frame information and the deformation reference data corresponding to the n-1-th frame information to obtain the deformation reference data corresponding to the n-th frame information; the at least one area is determined based on the area analysis result of the source image; the area analysis result is used to describe the position of the at least one area in the source image.
[0259] In one possible implementation, the implicit key points are determined using an implicit key point detection model; and the image generation process is implemented using a preset model.
[0260] The image generating device 400 further includes:
[0261] A model updating unit is used to update the implicit key point detection model and the preset model based on the image generation result and the label information corresponding to the image generation result; the label information is determined based on the driving information.
[0262] In a possible implementation manner, the driving information and the source image are both images extracted from the same sample video; and the label information is the driving information.
[0263] In one possible implementation, the data acquisition unit 401 is specifically used to: after receiving the text data, convert the text data into voice data; and construct the driving information based on at least one head posture template and the predicted expression coefficient of each frame data in the voice data.
[0264] In one possible implementation, the driving information includes N frames of information, where N is a positive integer; the nth frame information is determined based on the nth frame data in the voice data and the head posture template matching the nth frame data; the at least one head posture template includes the head posture template matching the nth frame data; n is a positive integer, n≤N.
[0265] In a possible implementation, the source image is determined based on a pre-constructed virtual image; the image generation result includes a generated image corresponding to each frame of data in the voice data;
[0266] The image generating device 400 further includes:
[0267] A video construction unit is used to construct a video corresponding to the virtual image based on the voice data and the image generation result.
[0268] Based on the relevant content of the above-mentioned image generating device 400, it can be known that for the image generating device 400 provided in the embodiment of the present application, after obtaining the source image and driving information, image generation processing is performed based on the source image, the implicit key points of the source image, the facial key points of the source image, the implicit key points of the driving information, and the facial key points of the driving information to obtain the image generation result corresponding to the driving information, so that the appearance information presented by the image generation result is consistent with the appearance information presented by the source image, and the posture presented by the image generation result is consistent with the posture represented by the driving information, so that the driving result of the source image under the driving information can be automatically generated. Among them, because the facial key points are used to describe the facial area, and the implicit key points are used to supplement the key points in other areas except the facial area, when performing image generation processing based on the facial key points and the implicit key points, not only the posture of the facial area can be paid attention to, but also the posture of some non-facial areas, such as the hair area, the hat area and other areas can be paid attention to, so that the image generation process can pay attention to the posture of each area as comprehensively as possible, and then the final image generation result can present the various postures represented by the driving information as comprehensively as possible, so that defects caused by only paying attention to the facial state, such as incoordination, can be effectively avoided, which is conducive to improving the image generation quality.
[0269] In addition, an embodiment of the present application also provides an electronic device, which includes a processor and a memory: the memory is used to store instructions or computer programs; the processor is used to execute the instructions or computer programs in the memory, so that the electronic device executes any implementation of the image generation method provided in the embodiment of the present application.
[0270] See also Figure 5 , which shows a schematic structural diagram of an electronic device 500 suitable for implementing the embodiments of the present disclosure. The terminal devices in the embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 5 The electronic device shown is only an example and should not limit the functions and scope of use of the embodiments of the present disclosure.
[0271] like Figure 5As shown, the electronic device 500 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage device 508 into a random access memory (RAM) 503. Various programs and data required for the operation of the electronic device 500 are also stored in the RAM 503. The processing device 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0272] Typically, the following devices may be connected to the I / O interface 505: an input device 506 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 507 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 508 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 509. The communication device 509 may allow the electronic device 500 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 5 The electronic device 500 is shown with various devices, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead.
[0273] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 509, or installed from the storage device 508, or installed from the ROM 502. When the computer program is executed by the processing device 501, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.
[0274] The electronic device provided by the embodiment of the present disclosure and the method provided by the above embodiment belong to the same inventive concept. For technical details not fully described in this embodiment, please refer to the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.
[0275] An embodiment of the present application further provides a computer-readable medium, in which instructions or computer programs are stored. When the instructions or computer programs are executed on a device, the device executes any implementation of the image generation method provided in the embodiment of the present application.
[0276] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.
[0277] In some embodiments, the client and server can communicate using any currently known or future developed network protocol, such as HTTP (Hypertext Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.
[0278] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0279] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device can perform the method.
[0280] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including, but not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0281] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0282] The units involved in the embodiments described in this disclosure may be implemented in software or hardware, wherein the name of a unit / module does not, in some cases, limit the unit itself.
[0283] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0284] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0285] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the common and similar parts between the various embodiments. For the systems or devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method description.
[0286] It should be understood that in this application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0287] It should also be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.
[0288] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0289] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not limited to the embodiments shown herein, but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. An image generation method, characterized in that: The method comprises: Obtain source image and driving information; Image generation processing is performed based on the source image, implicit key points of the source image, facial key points of the source image, implicit key points of the driving information, and facial key points of the driving information to obtain an image generation result corresponding to the driving information; the facial key points are used to describe the facial area; and the implicit key points are used to supplement key points in other areas except at least one sub-area of the facial area.
2. The method according to claim 1, characterized in that The image generation result is determined based on part or all of the implicit key points.
3. The method according to claim 1, characterized in that The implicit key points include implicit key points in at least two areas; The at least two regions include the face region and at least one non-face region; The image generation result is determined based on implicit key points in the at least one non-face area.
4. The method according to claim 1, wherein The target key point in the implicit key point matches partially or completely with the facial region; The image generation result is determined based on other key points in the implicit key points except the target key point.
5. The method according to claim 4, characterized in that The target key point is matched with the at least one sub-region; The at least one sub-region includes at least one of an eye region and a mouth region.
6. The method according to claim 5, characterized in that The facial key points include facial key points within the at least one sub-region; The target key point is determined according to the distance between the implicit key point and the facial key point in the at least one sub-region.
7. The method according to claim 4, characterized in that The target key point is determined based on the point identification of at least one position change key point; The process of determining the at least one position change key point includes: Acquire a preset image sequence, where the preset image sequence is used to describe a change in the at least one sub-region; Performing key point position change analysis based on implicit key points of at least two images in the preset image sequence to obtain an analysis result; According to the analysis result, the at least one position-changing key point is determined from the implicit key points of the at least two images.
8. The method according to claim 1, characterized in that The driving information includes N frames of information, where N is a positive integer; The process of determining the image generation result corresponding to the n-th frame information includes: Determining deformation reference data corresponding to the n-th frame information based on the implicit key points of the source image, the facial key points of the source image, the implicit key points of the n-th frame information, and the facial key points of the n-th frame information; n is a positive integer, n≤N; Performing deformation processing on the source image according to the deformation reference data corresponding to the n-th frame information to obtain a deformed image corresponding to the n-th frame information; Performing a completion process on the deformed image corresponding to the n-th frame information to obtain a completed image corresponding to the n-th frame information and completion position representation data corresponding to the n-th frame information; Based on the completed position representation data corresponding to the n-th frame information, the reference image corresponding to the n-th frame information and the completed image corresponding to the n-th frame information are fused to obtain the image generation result corresponding to the n-th frame information; the reference image is determined based on the source image and the deformation reference data corresponding to the n-th frame information.
9. The method according to claim 8, characterized in that The reference image is the deformed image corresponding to the n-th frame information; or, The process of obtaining the reference image includes: Performing super-resolution processing on the source image to obtain a super-resolution image; Performing deformation processing on the super-resolution image according to the deformation reference data corresponding to the n-th frame information to obtain a super-resolution deformation result corresponding to the n-th frame information; A reference image corresponding to the n-th frame information is determined according to the super-fractional deformation result corresponding to the n-th frame information.
10. The method according to claim 8, characterized in that If n≥2, the process of determining the deformation reference data corresponding to the n-th frame information includes: Determining predicted deformation parameters corresponding to the n-th frame information based on implicit key points of the source image, facial key points of the source image, implicit key points of the n-th frame information, and facial key points of the n-th frame information; Based on the weight corresponding to at least one area in the source image, the predicted deformation parameters corresponding to the n-th frame information and the deformation reference data corresponding to the n-1-th frame information are weighted to obtain the deformation reference data corresponding to the n-th frame information; the at least one area is determined based on the area analysis result of the source image; the area analysis result is used to describe the position of the at least one area in the source image.
11. The method according to claim 1, wherein The implicit key points are determined using an implicit key point detection model; The image generation process is achieved by using a preset model; After obtaining the image generation result corresponding to the driving information, the method further includes: Updating the implicit key point detection model and the preset model according to the image generation result and the label information corresponding to the image generation result; The tag information is determined according to the driving information.
12. The method according to claim 11, characterized in that The driving information and the source image are both images extracted from the same sample video; The tag information is the driving information.
13. The method according to claim 1, wherein The process of obtaining the driving information includes: After receiving the text data, converting the text data into voice data; The driving information is constructed based on at least one head posture template and the predicted expression coefficient of each frame data in the voice data.
14. The method according to claim 13, characterized in that The driving information includes N frames of information, where N is a positive integer; The nth frame information is determined based on the nth frame data in the speech data and the head posture template matched by the nth frame data; the at least one head posture template includes the head posture template matched by the nth frame data; n is a positive integer, n≤N.
15. The method according to claim 13, characterized in that The source image is determined based on a pre-constructed virtual image; The image generation result includes a generated image corresponding to each frame of data in the voice data; After obtaining the image generation result corresponding to the driving information, the method further includes: A video corresponding to the virtual image is constructed based on the voice data and the image generation result.
16. An image generating device, characterized in that: include: A data acquisition unit, used for acquiring source images and driving information; an image generation unit, configured to perform image generation processing based on the source image, the implicit key points of the source image, the facial key points of the source image, the implicit key points of the driving information, and the facial key points of the driving information, to obtain an image generation result corresponding to the driving information; The facial key points are used to describe the facial area; The implicit key points are used to supplement key points in other regions except at least one sub-region in the face region.
17. An electronic device, characterized in that: The device includes: a processor and a memory; The memory is used to store instructions or computer programs; The processor is configured to execute the instructions or computer program in the memory, so that the electronic device executes the method according to any one of claims 1 to 15.
18. A computer-readable medium, characterized in that The computer-readable medium stores instructions or a computer program, and when the instructions or the computer program are executed on a device, the device is caused to execute the method according to any one of claims 1 to 15.
19. A computer program product, characterized in that The method comprises a computer program carried on a non-transitory computer-readable medium, the computer program comprising a program code for executing the method according to any one of claims 1 to 15.