Image generation method and device, equipment, medium and product
By acquiring source image and driving information, using key point information and reference features for image generation processing, the problem of inconsistent image generation results in a single image driver scene is solved, and high-quality image generation results are achieved.
Patent Information
- Application Number
- CN202410178179.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-02-08
- Publication Date
- 2025-08-08
AI Technical Summary
The prior art is difficult to effectively generate images in a single-image driving scenario, so that the image generation results are consistent with the appearance information of the source image and the attitude conditions of the driving information, especially inadequate synchronization in posture changes such as expression states and head postures.
By acquiring the source image and driving information, the image generation process is performed using the key point information of the source image, reference features such as face contour detection results and foreground segmentation results, and the key point information of the driving information, the generation results maintain the consistency of appearance and posture, and the image generation is generated and updated using a preset model, and a video or image sequence is generated.
Automatic generation of source images under driver information is realized, the image generation results are consistent with the appearance information of the source image, and the posture conditions such as expression status and head posture are consistent with the driving information, improving the quality and consistency of image generation.
Smart Images

Figure CN120451033A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to an image generation method, apparatus, device, medium, and product. Background Art
[0002] For some application scenarios, such as single-image driving scenarios, these scenarios have the following requirements: after a source image is given, the source image can be driven and processed according to the driving information input by the user to obtain a driving result, such as one or more images, so that the driving result can present the object appearance information described by the source image and the changes described by the driving information. Summary of the Invention
[0003] The present application provides an image generation method, apparatus, device, medium, and product, which are conducive to improving the quality of image generation.
[0004] In order to achieve the above objectives, the technical solutions provided by this application are as follows:
[0005] The present application provides an image generation method, the method comprising:
[0006] Obtain source image and driving information;
[0007] Image generation processing is performed based on the source image, key point information of the source image, at least one reference feature of the source image, and key point information of the driving information to obtain an image generation result corresponding to the driving information; the at least one reference feature includes one or more of a facial contour detection result, a foreground segmentation result, and an image feature.
[0008] In a possible implementation manner, the driving information includes N frames of information, where N is a positive integer;
[0009] The process of determining the image generation result corresponding to the n-th frame information includes:
[0010] Determining deformation reference data corresponding to the n-th frame information based on key point information of the source image, at least one reference feature of the source image, and key point information of the n-th frame information, where n is a positive integer, n≤N;
[0011] Performing deformation processing on the source image according to the deformation reference data corresponding to the n-th frame information to obtain a deformed image corresponding to the n-th frame information;
[0012] An image generation result corresponding to the n-th frame information is determined based on the deformed image corresponding to the n-th frame information.
[0013] In a possible implementation manner, determining the image generation result corresponding to the n-th frame information based on the deformed image corresponding to the n-th frame information includes:
[0014] An image generation result corresponding to the n-th frame information is determined based on the at least one reference feature and the deformed image corresponding to the n-th frame information.
[0015] In one possible implementation, the at least one reference feature includes at least one feature to be deformed;
[0016] The method further comprises:
[0017] For any of the features to be deformed, deformation processing is performed on the feature to be deformed according to the deformation reference data corresponding to the n-th frame information to obtain a deformation result of the feature to be deformed;
[0018] The determining, based on the deformed image corresponding to the n-th frame information, an image generation result corresponding to the n-th frame information includes:
[0019] An image generation result corresponding to the n-th frame information is determined based on the deformation result of the at least one feature to be deformed and the deformed image corresponding to the n-th frame information.
[0020] In one possible implementation, the at least one reference feature includes a facial contour detection result, a foreground segmentation result, and an image feature; the at least one feature to be deformed includes a facial contour detection result and a foreground segmentation result; and the image generation result corresponding to the n-th frame information is determined based on the deformation result of the at least one feature to be deformed, the image feature, and the deformed image corresponding to the n-th frame information.
[0021] In a possible implementation manner, if n≥2, the process of determining the deformation reference data corresponding to the n-th frame information includes:
[0022] determining a predicted deformation parameter corresponding to the n-th frame information based on key point information of the source image, at least one reference feature of the source image, and key point information of the n-th frame information;
[0023] Based on the regional analysis result of the source image and the deformation reference data corresponding to the n-1 frame information, the predicted deformation parameters corresponding to the n frame information are adjusted to obtain the deformation reference data corresponding to the n frame information; the regional analysis result is used to describe the position of at least one region in the source image.
[0024] In one possible implementation, the process of acquiring the source image includes:
[0025] After acquiring the input image, the input image is subjected to preset processing to obtain the source image; the preset processing includes at least one of expression adjustment processing and quality increase processing; the expression adjustment processing is used to adjust the image expression state to the target expression state; the quality increase processing is used to enhance the image quality.
[0026] In a possible implementation manner, the quality enhancement process includes at least one of a beautification process and a clarity enhancement process; the beautification process is used to enhance the aesthetics of the image; and the clarity enhancement process is used to enhance the clarity of the image.
[0027] In one possible implementation, the image generation process is implemented using a preset model;
[0028] The preset model is specifically configured to perform image generation processing based on the source image, key point information of the source image, at least one reference feature of the source image, and key point information of the driving information, to obtain an image generation result corresponding to the driving information, a predicted foreground segmentation corresponding to the driving information, and a predicted facial contour corresponding to the driving information;
[0029] The method further comprises:
[0030] The preset model is updated based on the image generation result, the label information corresponding to the image generation result, the predicted foreground segmentation, the label information corresponding to the predicted foreground segmentation, the predicted facial contour, and the label information corresponding to the predicted facial contour; the label information is determined based on the driving information.
[0031] In one possible implementation method, the source image is determined based on a pre-built virtual image; the driving information is determined based on voice data; the voice data is converted from text data; and the image generation result includes a generated image corresponding to each frame of data in the voice data;
[0032] After obtaining the image generation result corresponding to the driving information, the method further includes:
[0033] A video corresponding to the virtual image is constructed based on the voice data and the image generation result.
[0034] The present application provides an image generation device, characterized by comprising:
[0035] An acquisition unit, configured to acquire source images and driving information;
[0036] A generation unit is configured to perform image generation processing based on the source image, key point information of the source image, at least one reference feature of the source image, and key point information of the driving information to obtain an image generation result corresponding to the driving information; the at least one reference feature includes one or more of a facial contour detection result, a foreground segmentation result, and an image feature.
[0037] The present application provides an electronic device, characterized in that the device includes: a processor and a memory;
[0038] The memory is used to store instructions or computer programs;
[0039] The processor is used to execute the instructions or computer programs in the memory so that the electronic device executes the image generation method provided in this application.
[0040] The present application provides a computer-readable medium, characterized in that instructions or computer programs are stored in the computer-readable medium. When the instructions or computer programs are executed on a device, the device executes the image generation method provided by the present application.
[0041] The present application provides a computer program product, characterized in that it includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains program code for executing the image generation method provided by the present application.
[0042] Compared with the related art, this application has at least the following advantages:
[0043] In the technical solution provided by the present application, a source image and driving information are first obtained; then, image generation processing is performed based on the source image, key point information of the source image, at least one reference feature of the source image, and key point information of the driving information to obtain an image generation result corresponding to the driving information, so that the appearance information presented by the image generation result is consistent with the appearance information presented by the source image, and the posture conditions presented by the image generation result, such as facial expression state, head posture, etc., are consistent with the posture conditions represented by the driving information, thereby automatically generating the driving result of the source image under the driving information. In particular, because the at least one reference feature includes one or more of the facial contour detection result, the foreground segmentation result, and the image feature, the at least one reference feature can better represent the image information of the source image, such as posture information, appearance information, etc., so that the at least one reference feature can supplement some image information that cannot be represented by the key point information of the source image, thereby making the multi-condition guided image generation processing based on the key point information of the source image and the at least one reference feature have better performance, which is conducive to improving the image generation quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only some embodiments recorded in this application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0045] Figure 1 A flowchart of an image generation method provided in an embodiment of the present application;
[0046] Figure 2 A schematic diagram of a model update process provided in an embodiment of the present application;
[0047] Figure 3 A schematic diagram of an image sequence generation process provided in an embodiment of the present application;
[0048] Figure 4 A schematic diagram of a multi-terminal collaboration method provided in an embodiment of the present application;
[0049] Figure 5 A schematic structural diagram of an image generating device provided in an embodiment of the present application;
[0050] Figure 6 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0051] In order to help those skilled in the art better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of this application.
[0052] In order to better understand the technical solution provided by this application, the image generation method provided by this application is described below with reference to some drawings. Figure 1 As shown, the image generation method provided by the embodiment of the present application includes the following S1-S2. Figure 1 A flowchart of an image generation method provided in an embodiment of the present application.
[0053] S1: Get source image and driving information.
[0054] The source image is used to provide other information in the image generation process, such as the appearance information of an object, in addition to some information such as expression information. The object refers to the foreground described by the source image; and this application does not limit the implementation of the object.
[0055] In addition, the present application does not limit the implementation method of the source image. For ease of understanding, the following description combines two cases.
[0056] Case 1: For some scenarios, such as model training scenarios, the source image mentioned above can refer to a frame of video image randomly extracted from a sample video, such as Figure 2 The source image shown. The sample video refers to the video required for the model training process; and this application does not limit the method for obtaining the sample video. For example, it can be implemented using any existing or future sample video acquisition method. It can be seen that in one possible implementation, when the image generation method provided by this application is applied to the model update process, such as Figure 2 During the model training process shown in FIG, the source image may be an image extracted from a sample video.
[0057] Case 2: For some scenarios, such as performing a certain image generation task, the source image mentioned above may refer to an image input by the user in some way, such as an image taken by a camera, an image uploaded manually, an image selected by a selection operation, etc. It can be seen that in one possible implementation, when the image generation method provided by this application is applied to an image generation task, such as Figure 3 When performing the image generation task shown in FIG, the source image may refer to an image provided by the user, such as Figure 3 Source image shown.
[0058] After research, it was found that the image provided by the user may have some defects, such as low clarity. Therefore, in order to avoid the impact caused by these defects, the present application also provides a possible implementation method of the above source image. Under this implementation method, the acquisition process of the source image can be specifically as follows: after acquiring the input image, the input image is subjected to preset processing, such as Figure 4 The preset processing shown in FIG. 4 is used to obtain a source image, so that the source image can overcome the defects of the input image itself, thereby making the source image better able to represent the image information in the input image. The input image refers to an image input by a user in some way, such as Figure 4 The preset process is used to perform at least one optimization process on the input image.
[0059] In addition, this application does not limit the implementation of the preset processing in the above paragraph. For example, the preset processing may include at least one of expression adjustment processing and quality enhancement processing. The expression adjustment processing is used to adjust the image's expression state to a predetermined target expression state, such as an open-eyes + open-mouth state. Furthermore, this application does not limit the implementation of the expression adjustment processing. For example, it can be implemented using any existing or future method capable of performing expression state adjustment processing on an image, such as using a pre-built convolutional neural network with expression state adjustment capabilities. The quality enhancement processing is used to enhance image quality. Furthermore, this application does not limit the implementation of the quality enhancement processing. For example, it can be implemented using any existing or future method capable of performing image quality enhancement processing on an image. For another example, in some application scenarios, the quality enhancement processing may include at least one of beautification processing and clarity enhancement processing. Furthermore, this beautification processing is used to enhance the image's aesthetics. Furthermore, this application does not limit the implementation of the beautification processing. For example, it can be implemented using any existing or future method capable of performing image aesthetics enhancement processing. The clarity enhancement processing is used to enhance the clarity of the image, and this application does not limit the implementation method of the clarity enhancement processing. For example, it can be implemented using any existing or future method that can increase the clarity of an image, such as a super-resolution method.
[0060] Based on the above two paragraphs, it can be seen that in one possible implementation, the acquisition process of the above source image can be specifically as follows: after obtaining the input image, the input image is first subjected to expression adjustment processing to obtain an adjusted image, so that the expression state in the adjusted image is in the target expression state, such as open eyes + open mouth state; then the adjusted image is beautified to obtain a beautified image, so that the beauty of the beautified image is higher than the beauty of the adjusted image; then, the beautified image is super-resolved to obtain the source image, so that the clarity of the source image is higher than the clarity of the beautified image. Among them, because the expression adjustment processing can adjust the expression state in any image to the target expression state, the present application can use the expression adjustment processing to realize expression normalization processing of input images provided by different users, so that the source images determined based on different input images are all in the same expression state, which is conducive to maintaining the consistency of the eyes and the inside of the mouth, such as teeth, thereby effectively avoiding the quality problems caused by inconsistent eyes or teeth generated in different frames when using non-normalized images for image generation processing, thereby helping to improve the image generation quality. Furthermore, because the beautification process is used to enhance the aesthetics of the image, the source image determined based on the beautification process has a higher aesthetics, thereby effectively avoiding the quality impact caused by low image aesthetics, thereby facilitating improved image generation quality. Furthermore, because the super-resolution process is used to enhance the clarity of the image, thereby ensuring that the source image determined based on the super-resolution process has a higher clarity, thereby effectively avoiding the quality impact caused by low image clarity, thereby facilitating improved image generation quality.
[0061] The driving information is used to provide some information during the image generation process, such as information on facial expression status, head posture, etc.; and the present application does not limit the driving information. For example, the driving information may include some facial expression coefficients. For another example, the driving information may include one or more images. It can be seen that under one possible implementation, the driving information may include N frames of information, where N is a positive integer. Among them, the nth frame information refers to the information in the nth arrangement position in the driving information; and the present application does not limit the implementation method of the nth frame information. For example, the nth frame information can be implemented using facial expression coefficients or images. n is a positive integer, n≤N.
[0062] In addition, the present application does not limit the implementation method of the driving information. For ease of understanding, the following description is combined with two cases.
[0063] Case 1: For some scenarios, such as model training scenarios, the above-mentioned driving information may refer to the driving image corresponding to the above-mentioned source image. The driving image refers to an image that is required to provide guidance information when performing model update processing based on the source image; and this application does not limit the driving image. For example, when the source image refers to a frame of image randomly extracted from a sample video, the driving image corresponding to the source image may refer to another frame of image randomly extracted from the sample video, such as Figure 2 As shown in the driving image, in one possible implementation, the driving information and the source image are both extracted from the same sample video, so that the training process based on the source image and the driving image is a self-supervised task of video reconstruction. This makes the training process not involve cross-object and cross-domain training datasets, thereby reducing the difficulty of obtaining training data.
[0064] In case 2, for some scenarios, such as performing an image generation task, the driving information can be determined based on the data involved in the image generation task, such as text, video, or audio data, used to provide the driving signal. This allows the driving information to be used to provide posture information such as facial expression and / or head posture, such as changes in facial expression. For ease of understanding, the following example is used to illustrate.
[0065] As an example, in some application scenarios, the above process of obtaining the driving information may include the following steps 11 and 12.
[0066] Step 11: After receiving the text data, convert the text data into voice data.
[0067] The text data is used to describe the constraint information required for reference when performing image generation processing, such as expression constraints, etc.; and this application does not limit the text data. For example, the text data can be Figure 3 The following text is implemented:
[0068] In addition, this application does not limit the method for obtaining the above text data. For example, the text data may refer to text content input by the user. For another example, the text data may refer to text obtained and output by a model by processing user input content. The user input content refers to content provided by the user to the model, such as text, voice, video, image, etc. The model is used to convert the user input content into text.
[0069] In addition, the present application does not limit the implementation of the above step 11. For example, it can be implemented with the help of any existing or future method that can convert text into speech, such as text to speech (TTS).
[0070] Based on the relevant content of step 11 above, it can be seen that after obtaining the above text data, such as Figure 3 After a piece of text content is shown, the text data can be converted into speech data, such as Figure 3 The voice data shown is used to enable the voice data to express the text data in an audio manner, so that the posture information corresponding to the text data, such as the mouth state, can be determined based on the voice data.
[0071] Step 12: Construct driving information based on at least one head posture template and the predicted expression coefficients of each frame of data in the above voice data.
[0072] The head posture template refers to a pre-set template for describing the head posture; and the head posture template can be pre-set according to an actual application scenario.
[0073] In addition, for the voice data converted from the above text data, the voice data may include N frames of data. The nth frame of data refers to the data at the nth arrangement position in the voice data. The predicted expression coefficient of the nth frame of data refers to the expression coefficient corresponding to the nth frame of data, such as Figure 3 The coefficient n shown is used so that the predicted expression coefficient of the n-th frame data can describe the posture information of the n-th frame data, such as the state of the mouth, etc., so that the predicted expression coefficient of the n-th frame data can describe the posture information corresponding to the text content expressed by the n-th frame data, such as expression information, etc. In addition, the predicted expression coefficient of the n-th frame data can be obtained by performing expression coefficient prediction processing on the n-th frame data; and the present application does not limit the implementation method of the expression coefficient prediction processing. For example, it can be implemented using any existing or future method that can perform expression coefficient prediction processing on a frame of data, such as audio2bs and other methods. n is a positive integer, n≤N, and N is a positive integer.
[0074] In addition, this application does not limit the implementation method of the above step 12. For ease of understanding, two examples are used below for illustration.
[0075] Example 1: When the at least one head posture template above includes only one head posture template, and the head posture template includes N frames of head action description data arranged in sequence, the above step 12 may specifically include the following steps 121-122.
[0076] Step 121: Based on the head movement description data of the nth frame in the above head posture template and the predicted expression coefficient of the nth frame data in the above voice data, construct the posture description information of the nth frame, so that the posture description information of the nth frame includes the posture information described by the head movement description data of the nth frame and the posture information described by the predicted expression coefficient of the nth frame data, where n is a positive integer and n≤N.
[0077] Among them, the head action description data of the nth frame in the above head posture template is used to describe the head posture pre-set for the nth frame; and the present application does not limit the implementation method of the head action description data of the nth frame. For example, it can be implemented in the form of a template, text, or image, etc., n is a positive integer, n≤N. It should be noted that the present application does not limit the implementation method of the head posture template. For example, the head posture template may refer to an action sequence pre-set according to the application scenario. For another example, the head posture template may be an action sequence extracted from a video provided by the user. For another example, the head posture template may be a template found in a template library that best matches the above source image. Among them, the template library is used to provide a variety of head posture templates, and each head posture template is an action sequence.
[0078] The nth frame posture description information is used to describe some postures corresponding to the nth frame data in the above voice data, such as facial expression status and / or head posture, etc., and this application does not limit the implementation method of the nth frame posture description information. For example, it can be implemented using images.
[0079] In addition, the above n-frame posture description information is determined based on the n-frame head movement description data and the predicted expression coefficient of the n-frame data, so that the n-frame posture description information includes the posture information described by the n-frame head movement description data and the posture information described by the predicted expression coefficient of the n-frame data.
[0080] In addition, the present application does not limit the implementation method of the above step 121. For example, it can be specifically: image rendering processing is performed based on the head action description data of the nth frame in the above head posture template and the predicted expression coefficient of the nth frame data in the above voice data to obtain the nth frame posture description information, so that the nth frame posture description information can describe the various postures corresponding to the nth frame data in the form of images, such as head posture, mouth status, etc.
[0081] Step 122: Construct driving information based on the first frame posture description information, the second frame posture description information, ..., and the Nth frame posture description information, so that the driving information includes the first frame posture description information, the second frame posture description information, ..., and the Nth frame posture description information.
[0082] Based on the relevant content of steps 121 to 122 above, it can be seen that for some application scenarios, if a head posture template is set in advance and the head posture template is used to describe an action sequence, the action of each frame in the head posture template can be combined with the predicted expression coefficient of each frame data in the above voice data to obtain driving information, so that the driving information includes the combination results of each frame, so that an image sequence can be generated subsequently based on the driving information.
[0083] Example 2: When the at least one head posture template includes multiple head posture templates, the above step 12 may specifically include the following steps 123 and 124.
[0084] Step 123: Based on the n-th frame data in the above speech data, find the template that best matches the n-th frame data from the above multiple head posture templates as the head posture template corresponding to the n-th frame data, where n is a positive integer, n≤N.
[0085] It should be noted that the present application does not limit the implementation method of the above step 123. For example, when each template in the above multiple head posture templates is used to describe a frame of posture, the specific implementation of step 123 can be: for the n-th frame data in the above voice data, the n-th frame data can be matched with each template in the above multiple head posture templates to obtain a matching result, and based on the matching result, the template that best matches the n-th frame data is determined as the head posture template corresponding to the n-th frame data, such as Figure 3 The template n shown is used so that the head posture template corresponding to the n-th frame data can represent the head movement that best matches the n-th frame data, which is beneficial to improving the image generation effect.
[0086] In fact, in order to better improve the image generation effect, the present application also provides another possible implementation of the above step 123. Under this implementation, when each of the above multiple head posture templates is a posture sequence, the above voice data includes at least one audio segment, and each audio segment includes at least one frame of data, the step 123 can be specifically as follows: for any audio segment, select a template that matches the audio segment from the multiple head posture templates, and determine each frame of posture data in the matching template as the head posture template corresponding to each frame of data in the audio segment. It should be noted that the present application does not limit the implementation method of the matching. For example, the matching process between the audio segment and the template can be completed based on the semantics, rhythm, strength and other information in the audio segment.
[0087] Step 124: Based on the head posture template corresponding to the n-th frame data in the above voice data and the predicted expression coefficient of the n-th frame data, construct the n-th frame posture description information, so that the n-th frame posture description information includes the posture information described by the head posture template corresponding to the n-th frame data, and the posture information described by the predicted expression coefficient of the n-th frame data, where n is a positive integer, n≤N.
[0088] It should be noted that the present application does not limit the implementation method of the above step 124. For example, the step 124 may specifically be: constructing the n-frame posture description information based on the head posture template corresponding to the n-frame data in the above voice data and the predicted expression coefficient of the n-frame data, so that the n-frame posture description information includes the head posture described by the head posture template and the predicted expression coefficient of the n-frame data. For another example, the step 124 may specifically be: performing image rendering processing based on the head posture template corresponding to the n-frame data in the above voice data and the predicted expression coefficient of the n-frame data to obtain the n-frame posture description information. For another example, the step 124 may specifically be: performing coefficient adjustment processing on the predicted expression coefficient of the n-frame data based on the head posture template corresponding to the n-frame data to obtain the n-frame posture description information, so that the n-frame posture description information simultaneously satisfies the constraints described by the head posture template corresponding to the n-frame data and the constraints described by the predicted expression coefficient of the n-frame data.
[0089] Step 125: Construct driving information based on the first frame posture description information, the second frame posture description information, ..., and the Nth frame posture description information, so that the driving information includes the first frame posture description information, the second frame posture description information, ..., and the Nth frame posture description information.
[0090] Based on the relevant content of steps 123 to 125 above, it can be seen that for some application scenarios, if many head posture templates are pre-set, and each head template is used to describe a frame of head posture, then the template that matches each frame of audio in the voice data can be determined from these templates first, so that the head posture described by the matching template and the mouth state corresponding to the audio have a relatively high adaptability; then, each frame of audio in the voice data is combined with the corresponding matching template to obtain driving information, so that the driving information includes the combination results of each frame, so that the driving information can better express the posture changes corresponding to the voice data, such as changes in expression, changes in head posture, etc., which is conducive to improving flexibility.
[0091] Based on the above, in one possible implementation, the nth frame information in the driving information may be determined based on the nth frame data in the voice data and a head posture template that matches the nth frame data, where n is a positive integer, n≤N, and N is a positive integer. The head posture template that matches the nth frame data is the template that is determined from the at least one head posture template and that best matches the nth frame data.
[0092] Based on the relevant content of steps 11 to 12 above, it can be seen that for some application scenarios, after receiving the text data, the text data is first converted into voice data; then, the expression coefficient prediction processing is performed on each frame of the voice data; then, based on the extracted expression coefficient and the pre-set head posture template, the driving information is constructed so that the driving information can better describe the posture changes corresponding to the text data, such as expression changes, so that an image sequence that can represent the posture changes can be generated based on the driving information, such as a video.
[0093] Based on the relevant content of S1 above, it can be seen that for some application scenarios, such as model training scenarios or image generation task execution scenarios, the source image and driving information are obtained so that image generation processing can be performed based on the source image and the driving information later, so that the posture information in the final generated image is consistent with the posture information in the driving information, and other information in the final generated image except the posture information, such as appearance information, is consistent with the corresponding information in the source image.
[0094] S2: Performing image generation processing based on the source image, key point information of the source image, at least one reference feature of the source image, and key point information of the driving information to obtain an image generation result corresponding to the driving information; the at least one reference feature includes one or more of a facial contour detection result, a foreground segmentation result, and an image feature.
[0095] The key point information of the source image is used to represent the position information in the source image, so that the key point information can represent the posture information in the source image, such as the expression state, hat posture, hairstyle posture, etc.
[0096] In addition, the present application does not limit the implementation method of the key point information of the above source image. For example, the key point information of the source image can be determined based on at least one of the implicit key points (Neural Keypoints, NK) of the source image and the facial key points (Landmark, LMK) of the source image.
[0097] For the facial key points of the source image above, the facial key points are used to describe some coordinate positions in the facial area of the source image, so that the facial key points can describe the posture characteristics of the facial area in the source image, such as expression characteristics.
[0098] It should be noted that facial key points have relatively good controllability and interpretability, allowing them to clearly represent certain areas of the face, such as the eyes, mouth, and facial contours. Obviously, these facial key points are limited to certain facial areas, making them unable to describe the posture characteristics of other areas outside of this part of the face, such as the forehead, cheeks, and hairstyle. Consequently, subsequent posture adjustment processing of this part of the face can only be achieved with the help of these facial key points.
[0099] In addition, this application does not limit the method for obtaining the facial key points of the source image. For example, it can be implemented by any existing or future method that can detect facial key points of an image. For example, in some application scenarios, in order to better improve the accuracy, the process of obtaining the facial key points of the source image can be as follows: the source image is input into a pre-built facial key point detection model, such as Figure 2 or Figure 3 Model 2 is shown, so that the facial key point detection model can perform facial key point detection processing on the source image, obtain and output the facial key points of the source image. The facial key point detection model refers to a pre-built model with relatively good facial key point detection capabilities, so that the facial key point detection model can be used to perform facial key point detection processing on the input data of the facial key point detection model; and this application does not limit the implementation of the facial key point detection model.
[0100] For the implicit key points of the source image above, the implicit key points are used to describe some coordinate positions in the source image, so that the implicit key points are used to describe the distribution characteristics of the source image, so that the implicit key points can represent the posture characteristics of multiple areas in the source image, such as the face area, hat area, hair area, background area, etc., and thus the implicit key points can represent the position information in the source image as comprehensively as possible, such as the background position, hair position, hat position, etc., so that the implicit key points can be used to supplement the position information of other areas other than at least one sub-area in the face area, such as the hat, hair accessories, etc. as comprehensively as possible. The sub-area refers to a part of the face area; and the present application does not limit the implementation method of the at least one sub-area. For example, the at least one sub-area can include at least one of the eye area and the mouth area.
[0101] In addition, for the implicit key points mentioned above, since the implicit key points have not been artificially standardized and defined, the coordinate positions described by the implicit key points do not have precise semantic meanings, which makes the implicit key points uncontrollable, and thus makes the implicit key points more flexible, so that the implicit key points are more generalizable, making the implicit key points applicable to various data, and thus making the implicit key points able to better complete the posture information of other areas other than at least one sub-area in the facial area, such as hats, hair accessories and other areas.
[0102] Based on the above two paragraphs, it can be seen that in one possible implementation, the implicit key points of the above source image can be used to supplement the key points in other areas except at least one sub-area in the facial area, so that the implicit key points of the source image can complete the position information of other areas except at least one sub-area in the facial area, which is conducive to improving the generalization ability of the image generation process.
[0103] In addition, for some application scenarios, such as scenarios that focus on the posture of various parts of the body, in order to better improve efficiency, the implicit key points can be used to describe the posture characteristics of each body part area in the source image, so that the implicit key points of the source image can be used to supplement the key points in other body part areas except at least one sub-area in the face area, so that the implicit key points of the source image can complete the position information of other body part areas except at least one sub-area in the face area.
[0104] In addition, this application does not limit the method of obtaining the implicit key points of the source image above. For example, it can be specifically: inputting the source image into the implicit key point detection model, such as Figure 2 or Figure 3 The model 1 shown is used to enable the implicit key point detection model to perform implicit key point detection processing on the source image, obtain and output the implicit key points of the source image. The implicit key point detection model is used to perform implicit key point detection processing on the input data of the implicit key point detection model; and this application does not limit the implementation method of the implicit key point detection model. For example, it can be implemented using any machine learning model, such as any end-to-end machine learning model.
[0105] Based on the relevant content of the above source image, it can be seen that this application provides three possible implementation methods of the key point information of the above source image. For ease of understanding, the following is an explanation with examples.
[0106] Example 1: In some application scenarios, such as those focusing on facial expressions, the key point information of the source image described above can be determined based on the facial key points of the source image, so that the key point information includes the facial key points, thereby enabling the key point information to represent the facial state, such as the facial expression, in the source image. Because the facial key points have relatively good controllability and interpretability, they can clearly represent parts of the facial region, such as the eyes, mouth, and facial contours. Therefore, the key point information determined based on the facial key points can relatively accurately represent the facial expression presented in the source image.
[0107] Example 2: In some application scenarios, such as those focusing on comprehensive posture, the key point information of the source image described above can be determined based on the implicit key points of the source image, so that the key point information includes the implicit key points, thereby enabling the key point information to represent the posture information in the source image as comprehensively as possible, such as facial condition, hair posture, hat posture, etc. In particular, because the implicit key points are highly generalizable, they are applicable to various data, thereby enabling the implicit key points to describe the posture information of multiple regions in the source image. In this way, the key point information determined based on the implicit key points can relatively comprehensively represent the posture information in the source image.
[0108] Example 3: In some application scenarios, such as those focusing on overall posture, in order to better improve accuracy, this application also provides a method for obtaining the key point information of the above-mentioned source image, which can be specifically: based on the implicit key points of the source image and the facial key points of the source image, determine the key point information of the source image, so that the key point information of the source image includes part or all of the implicit key points of the source image and the facial key points of the source image. Among them, because the facial key points are used to describe the facial area, and the implicit key points are used to supplement the key points in other areas except at least one sub-area of the facial area, the key point information determined based on these two key points can represent the position information in the source image as comprehensively and accurately as possible, such as the position information of each body part in the source image.
[0109] In addition, the present application does not limit the implementation method of the step of "determining the key point information of the source image based on the implicit key points of the source image and the facial key points of the source image" in the previous paragraph. For example, it can be specifically: splicing the implicit key points of the source image and the facial key points of the source image to obtain the key point information of the source image, so that the key point information includes the implicit key points and the facial key points, so that the key point information can represent the posture information in the source image as comprehensively as possible.
[0110] Research has found that facial key points can more accurately describe facial states, such as the state of the eyes and mouth. Therefore, in order to avoid the uncontrollability of implicit key points interfering with some or all facial states, this application also provides a possible implementation of the key point information of the above-mentioned source image. In this implementation, when the implicit key points of the source image include implicit key points in at least two regions, and the at least two regions include a facial region and at least one non-facial region, the key point information of the source image can include at least the implicit key points in the at least one non-facial region and the facial key points of the source image, so that the image generation result corresponding to the above-mentioned driving information can be determined based on the implicit key points in the at least one non-facial region. It can be seen that in one possible implementation, the key point information of the source image can include the key points in the at least one non-facial region and the facial key points of the source image. The at least one non-facial region refers to other regions in the source image other than the facial region, such as other body parts.
[0111] Based on the content of the above paragraph, in order to better improve the accuracy, the present application also provides a possible implementation method of the key point information acquisition process. Under this implementation method, when the target key point in the implicit key point above matches part or all of the facial area, the key point information can be determined based on other key points in the implicit key point except the target key point, so that the key point information includes other key points in the implicit key point except the target key point, so that the image generation result corresponding to the above driving information can be determined based on the other key points. Based on this, it can be seen that under a possible implementation method, the image generation result is determined based on other key points in the implicit key point except the target key point. For better understanding, a possible implementation method of the key point information acquisition process of the above source image is described below as an example.
[0112] As an example, in a possible implementation manner, the process of obtaining key point information of the above source image may specifically include the following steps 21 and 22.
[0113] Step 21: Determine the target key points corresponding to the source image from the implicit key points of the source image, so that the target key points are partially or completely matched with the facial region.
[0114] Among them, the target key point corresponding to the source image refers to the key point existing in the implicit key points of the source image and matching some or all of the key points in the facial area, so that the target key point can represent the implicit key point falling within some or all of the facial area.
[0115] In addition, the present application does not limit the implementation method of the target key points corresponding to the above source image. For example, in some application scenarios, the target key points corresponding to the source image may refer to the key points existing in the implicit key points of the source image and matching the key points of the facial area, so that the target key points can represent the implicit key points falling within the facial area.
[0116] Research has found that using facial key points to adjust the posture of the eyes and / or mouth is more conducive to improving the quality of image generation. Therefore, in order to better avoid the uncontrollability of implicit key points from interfering with the eye state and mouth state, the present application also provides a possible implementation method of the target key points corresponding to the above source image. Under this implementation method, the target key points corresponding to the source image can be matched with at least one sub-region in the facial area, and the at least one sub-region includes at least one of the eye area and the mouth area, so that the target key point can refer to a key point that exists in the implicit key points of the source image and falls into at least one of the eye area and the mouth area.
[0117] In addition, this application does not limit the method for obtaining the target key points corresponding to the above source image. For ease of understanding, the following is an explanation with examples.
[0118] As an example, in some application scenarios, in order to improve flexibility, when the facial key points of the above source image include facial key points within at least one sub-region, the target key points corresponding to the source image can be determined based on the distance between the implicit key points of the source image and the facial key points within the at least one sub-region, so that the target key points can represent the implicit key points falling within the at least one sub-region, such as the eye region and / or the mouth region. It should be noted that this application does not limit the method for obtaining the distance. For example, it can be implemented using any existing or future distance calculation method, such as Euclidean distance or cosine distance.
[0119] It can be seen that in one possible implementation, for the above source image, after obtaining the implicit key points of the source image and the facial key points of the source image, the distance between each implicit key point and each facial key point can be calculated first; then, based on the distance, the implicit key points falling within the eye area and mouth area described by the facial key points are determined as the target key points corresponding to the source image.
[0120] In practice, for the implicit key points described above, although they lack precise semantic meaning, the implicit key point identification process can determine a point identifier for each implicit key point, such as a serial number, so that the implicit key points with that point identifier converge within a certain range, such as the mouth corner area. Because different implicit key points have different point identifiers, they converge within different ranges, allowing the point identifiers to be used to subsequently determine which implicit key points fall within the eye and mouth areas. Based on this, the present application also provides a possible implementation of the target key points corresponding to the source image described above. In this implementation, the target key point is determined based on the point identifier of at least one position-varying key point. The at least one position-varying key point refers to a predetermined key point that can characterize the characteristics of at least one sub-region described above. Furthermore, for any point identifier of a position-varying key point, the point identifier is used to identify the position-varying key point. Furthermore, the present application does not limit the process for determining the at least one position-varying key point, which may specifically include steps 211 through 213 below.
[0121] Step 211: Acquire a preset image sequence, where the preset image sequence is used to describe the change of at least one sub-region.
[0122] The preset image sequence refers to an image sequence required to be used when determining the point identifiers of the implicit key points to be deleted, such as a video.
[0123] In addition, the present application does not limit the implementation method of the above preset image sequence. For example, when the above at least one sub-area includes at least one of the eye area and the mouth area, the preset image sequence is used to describe the changes in the at least one sub-area, such as the changes in the opening and closing of the eyes and / or mouth. It can be seen that in one possible implementation method, if the at least one sub-area includes the eye area, the preset image sequence may include a pre-recorded blinking video, and the blinking video is used to describe the changes in the eyes. In another possible implementation method, if the at least one sub-area includes the mouth area, the preset image sequence may include a pre-recorded mouth movement video, and the mouth movement video is used to describe the changes in the mouth. In yet another possible implementation method, if the at least one sub-area includes the eye area and the mouth area, the preset image sequence may include a pre-recorded blinking video and a pre-recorded mouth movement video.
[0124] Step 212: performing key point position change analysis based on implicit key points of at least two images in a preset image sequence to obtain analysis results.
[0125] For the preset image sequence mentioned above, the t-th frame image in the preset image sequence refers to the image at the t-th arrangement position in the preset image sequence; and the implicit key points of the t-th frame image are used to describe some coordinate positions in the t-th frame image, so that the implicit key points are used to describe the distribution characteristics of the t-th frame image, so that the implicit key points can represent the posture characteristics of each area in the t-th frame image, and further enable the implicit key points to represent the position information described by the t-th frame image as comprehensively as possible. It should be noted that the implementation method of the implicit key points of the t-th frame image is similar to the implementation method of the implicit key points of the source image mentioned above. For the sake of brevity, it will not be repeated here. t is a positive integer, t≤T, T is a positive integer, and T represents the number of images in the preset image sequence.
[0126] The analysis results are used to indicate which implicit key points in the preset image sequence have not changed their position coordinates (or the degree of change is less than a threshold), and which implicit key points have changed their position coordinates continuously, so that the analysis results can indicate whether the position coordinates corresponding to the same point identifier in different frame images in the preset image sequence have changed.
[0127] Step 213: According to the above analysis results, determine at least one position-changed key point from the implicit key points of the at least two images.
[0128] The position-changing key points are used to represent implicit key points that are located at different positions in different frame images in a preset image sequence.
[0129] Based on the relevant content of steps 211 to 213 above, it can be seen that in some application scenarios, at least one position change key point mentioned above can be determined in advance using a preset image sequence, and the point identifiers of these position change key points can be recorded and stored. This allows the implicit key points falling within the at least one sub-region to be directly deleted using the point identifier during the subsequent image generation process, thereby improving image generation efficiency. Since the point identifiers of these position change key points are fixed and do not change with different processed data, in order to further improve efficiency, these position change key points can be pre-acquired and stored so that they can be directly read from the storage space later.
[0130] Based on the relevant content of step 21 above, it can be known that for the source image, after obtaining the implicit key points of the source image, the target key points corresponding to the source image can be determined from the implicit key points, so that the target key points can represent the implicit key points falling into part or all of the facial area, such as the eye area and the nose area, so that the target key points can represent the key points that need to be deleted.
[0131] Step 22: Determine the key point information of the source image based on the key points in the implicit key points of the source image except the target key points corresponding to the source image.
[0132] It should be noted that the present application does not limit the implementation method of the above step 22. For example, it can be specifically as follows: first delete the target key points corresponding to the source image from the implicit key points of the source image to obtain the remaining key points corresponding to the source image; then splice the remaining key points corresponding to the source image with the facial key points of the source image to obtain the key point information of the source image.
[0133] Based on the relevant content of steps 21 to 22 above, it can be known that for the above source image, after obtaining the implicit key points of the source image, the target key points corresponding to the source image are first determined from the implicit key points, so that the target key points can represent the hidden key points falling into certain areas, such as the eye area and the mouth area; then, based on the other key points in the implicit key points except the target key points, and the facial key points of the source image, the key point information of the source image is determined, so that the key point information includes the other key points and the facial key points, so that the image generation result of the above driving information can be determined based on the other key points. In this way, the posture interference caused by the target key points can be effectively avoided, so that the key point information can represent the posture characteristics of each area in the source image as accurately as possible, which is conducive to improving the image generation effect.
[0134] Based on the relevant content of the key point information of the source image above, it can be seen that in one possible implementation method, the key point information of the source image can be obtained by combining the implicit key points of the source image and the facial key points of the source image, so that the key point information can represent the posture characteristics of each area in the source image as comprehensively and accurately as possible, which is conducive to improving the image generation effect.
[0135] The at least one reference feature of the source image refers to a feature required to be referenced when performing image generation processing based on the source image; and the at least one reference feature can relatively comprehensively and accurately represent the image information in the source image, such as facial expression state, head posture, hat state, appearance information, etc.
[0136] In addition, the present application does not limit the at least one reference feature of the above source image. For example, the at least one reference feature may include one or more of a face contour detection result, a foreground segmentation result, and an image feature.
[0137] For the face contour detection results of the source image above, such as Figure 2 The outline shown is 1 or Figure 3For the outline shown, the facial contour detection result is used to describe the facial contour in the source image, so that the facial contour detection result can represent the characteristics of some areas in the source image, such as the chin area, so that the image generated based on the facial contour detection result can better present these areas, which can effectively avoid the pixel sticking in these areas, thereby helping to improve the image quality; and the present application does not limit the method for obtaining the facial contour detection result. For example, it can be implemented using any existing or future method that can determine the facial contour of an image. For example, in order to better improve efficiency, the present application also provides a method for obtaining the facial contour detection result, which can specifically be: after obtaining the facial key points of the source image, the facial key points are used to perform facial contour drawing processing to obtain the facial contour detection result of the source image.
[0138] For the foreground segmentation results of the source image above, such as Figure 2 Mask1 shown or Figure 3 For the Mask shown, the foreground segmentation result is used to describe the location of the foreground in the source image, such as the location of the object, so that the foreground segmentation result can represent the characteristics of other areas in the source image, such as the foreground area, so that the image generated based on the foreground segmentation result can better present these areas, thereby effectively avoiding the pixel sticking feeling in these areas, which is beneficial to improving the image quality; and the present application does not limit the implementation method of the foreground segmentation result, for example, it can be implemented using a mask map (Mask). In addition, the present application does not limit the acquisition process of the foreground segmentation result, for example, it can be implemented using any existing or future method that can perform foreground segmentation processing on an image.
[0139] For the image features of the source image above, such as Figure 2 or Figure 3As for the features obtained by image encoding shown, the image features are used to describe the image information carried by the source image, such as object description information, so that the image features can represent some information in the source image, such as appearance information, so that the image features can better supplement the other information in the source image except the posture information, such as appearance information, and thus make the other information in the image generated based on the image features have a higher consistency with the corresponding information in the source image, which is conducive to improving the image generation quality; and the present application does not limit the implementation method of the image features, for example, it can be implemented using embedded feature vectors (embedding). In addition, the present application does not limit the way to obtain the image features, for example, it can adopt any existing or future method that can extract features from images, such as by using an encoder with image feature extraction function.
[0140] The keypoint information of the driving information described above is used to represent the posture information in the driving information, such as facial expression. Furthermore, the implementation of the keypoint information of the driving information is similar to the implementation of the keypoint information of the source image described above. Therefore, in one possible implementation, the keypoint information of the driving information is determined based on at least one of the implicit keypoints of the driving information and the facial keypoints of the driving information.
[0141] With respect to the facial key points of the driving information above, the facial key points are used to describe the facial state, such as the expression state, in the driving information; and when the driving information includes N frames of information, the facial key points of the driving information include the facial key points of the N frames of information. The facial key points of the n-th frame information are used to describe some coordinate positions within the facial region of the n-th frame information, so that the facial key points can describe the posture characteristics of the facial region in the n-th frame information. n is a positive integer, n≤N, and N is a positive integer. It should be noted that the implementation method of the facial key points of the n-th frame information is similar to that of the facial key points of the source image above. For the sake of brevity, they will not be described here.
[0142] In fact, to further improve image generation, this application also provides a possible implementation of the process for determining facial key points in the aforementioned driving information. In this implementation, when the driving information includes expression coefficients, the facial key points in the driving information can be specifically determined by searching a first mapping relationship for facial key points corresponding to the expression coefficients, and using these as the facial key points in the driving information. The first mapping relationship is used to describe the facial key points corresponding to the expression coefficients under different expressions; and the first mapping relationship is pre-determined based on a sample image. The sample image refers to a virtual image, such as a three-dimensional digital human, required to construct the first mapping relationship. It should be noted that this application does not limit the process for constructing the first mapping relationship. For example, the process can specifically be: first, obtaining the expression coefficients and facial key points of the sample image; then, based on the correspondence between the expression coefficients of the sample image and the facial key points of the sample image, constructing the first mapping relationship so that the first mapping relationship includes the correspondence.
[0143] Regarding the implicit key points in the driving information described above, these implicit key points are used to characterize posture information, etc., in the driving information. These implicit key points are used to characterize posture characteristics in the driving information, such as facial expression, hairstyle, and hat posture. This allows the implicit key points in the driving information to supplement key points in regions other than at least one sub-region within the facial region. Furthermore, when the driving information includes N frames of information, the implicit key points in the driving information can include the implicit key points of all N frames. The implicit key points in the nth frame of information are used to describe certain coordinate positions within the nth frame of information, so that the implicit key points describe the distribution characteristics of the nth frame of information. This allows the implicit key points to represent the posture characteristics of each region within the nth frame of information, thereby enabling the implicit key points to represent the position information described by the nth frame of information as comprehensively as possible. n is a positive integer, n≤N, and N is a positive integer. It should be noted that the implementation of the implicit key points in the nth frame of information is similar to that of the implicit key points in the source image described above and will not be further described here for the sake of brevity.
[0144] In fact, in order to better improve the image generation effect, the present application also provides a possible implementation method of the process of determining the implicit key points of the above-mentioned driving information. Under this implementation method, when the driving information includes an expression coefficient, the implicit key point of the driving information can be specifically: searching for the implicit key point corresponding to the expression coefficient from the second mapping relationship as the implicit key point of the driving information. The second mapping relationship is used to describe the implicit key points corresponding to the expression coefficient under different expressions; and the second mapping relationship is determined in advance based on the above-mentioned sample image. It should be noted that the present application does not limit the construction process of the second mapping relationship. For example, it can be specifically: first obtain the expression coefficient of the sample image and the implicit key point of the sample image; then construct the second mapping relationship based on the correspondence between the expression coefficient of the sample image and the implicit key point of the sample image, so that the second mapping relationship includes the correspondence.
[0145] Based on the above content, it can be seen that because implicit key points have good generalization ability, when combining implicit key points with facial key points for image generation processing, it can effectively overcome the problem caused by the large difference between the object described by the source image and the above sample image, thereby helping to improve the image generation effect.
[0146] Based on the relevant content of the driving information above, it can be known that in one possible implementation, when the driving information includes N frames of information, the key point information of the driving information can include the key point information of the N frames of information. The key point information of the n-th frame information is used to describe the posture information in the n-th frame information, such as expression state information and / or head posture information, and the key point information of the n-th frame information is determined based on the implicit key points of the n-th frame information and the facial key points of the n-th frame information, so that the key point information of the n-th frame information includes the facial key points, and the key point information of the n-th frame information includes part or all of the implicit key points, so that the key point information of the n-th frame information can represent the posture characteristics in the n-th frame information as comprehensively as possible, such as the posture characteristics of each body part in the n-th frame information. n is a positive integer, n≤N, and N is a positive integer. It should be noted that the implementation method of the key point information of the n-th frame information is similar to the implementation method of the key point information of the source image above. For the sake of brevity, it will not be repeated here.
[0147] Thus, in one possible implementation, for the nth frame information above, when the implicit key points of the nth frame information include implicit key points within at least two regions, and the at least two regions include a facial region and at least one non-facial region, the key point information of the nth frame information may include the implicit key points within the at least one non-facial region, as well as the facial key points of the nth frame information. This effectively avoids posture interference caused by certain key points in the nth frame information's implicit key points, thereby improving image generation. n is a positive integer, n≤N, and N is a positive integer.
[0148] Furthermore, the method for obtaining key point information for each frame of the driving information is similar to the method for obtaining key point information for the source image, and for the sake of brevity, it will not be repeated here. Thus, in one possible implementation, if the driving information includes N frames of information, the key point information for the driving information may specifically include steps 31 and 32 below.
[0149] Step 31: Determine the target key point corresponding to the n-th frame information from the implicit key points of the n-th frame information, so that the target key point matches part or all of the face area, where n is a positive integer, n≤N, and N is a positive integer.
[0150] The target key points corresponding to the n-th frame information refer to some or all of the implicit key points existing in the n-th frame information and falling within the facial region.
[0151] In addition, the implementation of the target key point corresponding to the n-th frame information is similar to the implementation of the target key point corresponding to the source image described above, and for the sake of brevity, it is not further described here. For example, the target key point corresponding to the n-th frame information can be matched with at least one sub-region within the face region, and the at least one sub-region includes at least one of the eye region and the mouth region, so that the target key point corresponding to the n-th frame information can be an implicit key point existing in the implicit key points of the n-th frame information that falls within at least one of the eye region and the mouth region.
[0152] In addition, the method for obtaining the target key point corresponding to the n-th frame information above is similar to the method for obtaining the target key point corresponding to the source image above, and for the sake of brevity, it will not be repeated here. For example, when the facial key points of the n-th frame information include facial key points within at least one sub-region, the target key point corresponding to the n-th frame information can be determined based on the distance between the implicit key point of the n-th frame information and the facial key point within the at least one sub-region, so that the target key point can represent the implicit key point falling within the at least one sub-region. For another example, the target key point corresponding to the n-th frame information is determined based on the point identifier of the at least one position-changing key point above, so that the target key point can represent the key point falling within the at least one sub-region. It should be noted that, for the relevant content of the at least one sub-region and the at least one position-changing key point, please refer to the above.
[0153] Based on the relevant content of step 31 above, for the nth frame information above, after obtaining the implicit key points of the nth frame information, the target key points corresponding to the nth frame information can be determined from the implicit key points, so that the target key points can represent the implicit key points falling within part or all of the facial area, such as the eye area and the nose area, so that the target key points can indicate the key points that need to be deleted. n is a positive integer, n≤N, and N is a positive integer.
[0154] Step 32: Determine the key point information of the nth frame information based on the key points in the nth frame information except the target key point corresponding to the nth frame information, where n is a positive integer, n≤N, and N is a positive integer.
[0155] It should be noted that the present application does not limit the implementation of step 32 above. For example, it may specifically include: first, deleting the target key point corresponding to the n-th frame information from the implicit key points of the n-th frame information to obtain the remaining key points corresponding to the n-th frame information; then, concatenating the remaining key points corresponding to the n-th frame information with the facial key points of the n-th frame information to obtain the key point information of the n-th frame information. n is a positive integer, n≤N, and N is a positive integer.
[0156] Based on the relevant contents of steps 31 to 32 above, it can be seen that for the nth frame information in the above driving information, after obtaining the implicit key points of the nth frame information, the target key points corresponding to the nth frame information are first determined from the implicit key points, so that the target key points can represent the hidden key points falling within certain areas, such as the eye area and the mouth area; then, based on the other key points in the implicit key points except the target key points corresponding to the nth frame information and the facial key points of the nth frame information, the key point information of the nth frame information is determined, so that the key point information includes the other key points and the facial key points, so that the image generation result of the above driving information can be determined based on the other key points. In this way, the posture interference caused by the target key points can be effectively avoided, so that the key point information can represent the posture characteristics of each area in the nth frame information as accurately as possible, thereby improving the image generation effect. n is a positive integer, n≤N, and N is a positive integer.
[0157] Based on the key point information of the driving information described above, when the driving information includes the first frame, the second frame, ..., and the Nth frame, after obtaining the nth frame, the implicit key points of the nth frame and the facial key points of the nth frame can be first determined; then, based on the implicit key points of the nth frame and the facial key points of the nth frame, the key point information of the nth frame can be determined so that the key point information can accurately represent the posture characteristics of each region in the nth frame as much as possible. n is a positive integer, n≤N, and N is a positive integer.
[0158] The image generation result corresponding to the driving information refers to one or more images obtained by using the driving information to perform posture driving on the source image, so that the posture information in the image generation result is consistent with the posture information in the driving information, and all other information in the image generation result except the posture information is consistent with the corresponding information in the source image.
[0159] In addition, for the above-mentioned driving information, if the driving information includes N frames of information, the image generation result corresponding to the driving information may include the image generation result corresponding to the N frames of information. The image generation result corresponding to the n-th frame information refers to the image obtained by performing posture driving on the source image using the n-th frame information, so that the posture information in the image generation result corresponding to the n-th frame information is consistent with the posture information in the n-th frame information, and the other information in the image generation result corresponding to the n-th frame information except the posture information is consistent with the corresponding information in the source image, so that the image generation result corresponding to the n-th frame information can represent the generated image corresponding to the n-th frame information, such as the generated image corresponding to the n-th frame data, etc. n is a positive integer, n≤N, and N is a positive integer.
[0160] In addition, in order to better improve the image generation effect, the present application also provides a possible implementation method of the process of determining the image generation result corresponding to the nth frame information above. Under this implementation method, the process of determining the image generation result corresponding to the nth frame information may include the following steps 41-43.
[0161] Step 41: Determine deformation reference data corresponding to the n-th frame information based on key point information of a source image, at least one reference feature of the source image, and key point information of the n-th frame information.
[0162] Among them, the deformation reference data corresponding to the nth frame information refers to the parameters required to perform posture adjustment processing on the source image based on the nth frame information, such as expression adjustment processing and / or head posture adjustment processing, such as posture offset + area that cannot be generated by deformation, etc.
[0163] In addition, this application does not limit the implementation of the deformation reference data corresponding to the n-th frame information. For example, in some application scenarios, such as scenarios where posture adjustment of areas such as the eyes and mouth is not required, the deformation reference data corresponding to the n-th frame information may include a two-dimensional geometric deformation field (Deformations). The two-dimensional geometric deformation field is used to describe the posture offset of certain areas.
[0164] After research, it was found that for some application scenarios, such as scenarios where posture adjustment is required for areas such as eyes and mouth, the posture adjustment effect of deformation processing on some areas in the image, such as the mouth and eyes, may not be very good. Therefore, in order to better improve the image generation effect, the present application also provides a possible implementation method of the deformation reference data corresponding to the n-th frame information above. Under this implementation method, the deformation reference data corresponding to the n-th frame information may include a two-dimensional geometric deformation field and an occlusion map (Occlusion Map). Among them, the two-dimensional geometric deformation field is used to describe the posture offset of the area that can be generated by deformation, that is, the posture offset of the deformable area. The deformable area refers to the area that can be better posture-adjusted with the help of deformation processing, such as other areas other than the eyes and mouth. The occlusion map is used to describe the location of the area that cannot be generated by deformation, that is, the location of the non-deformable area. The non-deformable area refers to the area that cannot be better posture-adjusted with the help of deformation processing, such as the eyes, mouth and other areas.
[0165] In addition, the present application does not limit the implementation method of the above step 41. For example, the step 41 can be implemented using a deformation network (DMN). It can be seen that under one possible implementation method, the step 41 can specifically be: inputting the key point information of the source image, at least one reference feature of the source image, and the key point information of the n-th frame information into the DMN, obtaining the deformations and occlusion map output by the DMN as the deformation reference data corresponding to the n-th frame information, so that the deformation reference data includes the deformations and occlusion map. Among them, because the at least one reference feature can supplement some information in the source image, the key point information of the source image and the at least one reference feature of the source image can provide richer image features, thereby making the deformation reference data determined based on these image features more accurate, which is conducive to improving the image generation effect.
[0166] It can be seen that for some application scenarios, after obtaining the key point information of the source image, at least one reference feature of the source image, and the key point information of the n-th frame information, these data can be input into the DMN to obtain the two-dimensional geometric deformation field and occlusion map output by the DMN as the deformation reference data corresponding to the n-th frame information, so that the deformation reference data can more accurately represent the parameters required to perform deformation processing on the source image, which is conducive to improving the image generation effect.
[0167] In addition, to further improve accuracy, the present application also provides a possible implementation of step 41 above. In this implementation, step 41 may specifically include: determining deformation reference data corresponding to the n-th frame information based on the source image, key point information of the source image, at least one reference feature of the source image, and key point information of the n-th frame information. Thus, in one possible implementation, step 41 may specifically include: inputting the source image, key point information of the source image, at least one reference feature of the source image, and key point information of the n-th frame information into the DMN, obtaining deformations and an occlusion map output by the DMN as deformation reference data corresponding to the n-th frame information, so that the deformation reference data includes the deformations and occlusion map. Because the source image can provide some reference image information for the deformation reference data determination process, the deformation reference data determined based on the source image is more accurate, thereby facilitating improved accuracy.
[0168] Research has found that for some image sequence generation scenarios, the posture information between different frame images in the image sequence to be generated shows relatively large changes, but the changes in other information, such as background, clothing, etc. are relatively small or even no changes. Therefore, in order to better improve the generation effect of the image sequence, this application also provides a possible implementation method of the process of determining the deformation reference data corresponding to the nth frame information above. Under this implementation method, if n≥2, the process of determining the deformation reference data corresponding to the nth frame information includes the following steps 411-412.
[0169] Step 411: Determine a predicted deformation parameter corresponding to the n-th frame information based on key point information of a source image, at least one reference feature of the source image, and key point information of the n-th frame information.
[0170] The predicted deformation parameters corresponding to the n-th frame information refer to the parameters predicted to be used when performing posture adjustment processing on the source image based on the n-th frame information, such as the posture offset + the area that cannot be generated by deformation.
[0171] In addition, the present application does not limit the implementation method of the above step 411. For example, the specific implementation method of step 411 can be: inputting the key point information of the source image, at least one reference feature of the source image, and the key point information of the n-th frame information into the DMN, obtaining the deformations and occlusion map output by the DMN as the predicted deformation parameters corresponding to the n-th frame information, so that the predicted deformation parameters include the deformations and occlusion map.
[0172] For another example, to further improve accuracy, step 411 above may specifically include: determining the predicted deformation parameters corresponding to the n-th frame information based on the source image, key point information of the source image, at least one reference feature of the source image, and key point information of the n-th frame information. Thus, in one possible implementation, step 411 may specifically include: inputting the source image, key point information of the source image, at least one reference feature of the source image, and key point information of the n-th frame information into the DMN, obtaining deformations and an occlusion map output by the DMN as the predicted deformation parameters corresponding to the n-th frame information, such that the predicted deformation parameters include the deformations and occlusion map.
[0173] Step 412: Based on the regional analysis result of the source image and the deformation reference data corresponding to the (n-1)th frame information, the predicted deformation parameters corresponding to the (n)th frame information are adjusted to obtain the deformation reference data corresponding to the (n)th frame information; the regional analysis result is used to describe the position of at least one region in the source image.
[0174] Among them, the deformation reference data corresponding to the n-1th frame information refers to the historical information required to be referenced when smoothing the deformation reference data corresponding to the nth frame information above; and the determination process of the deformation reference data corresponding to the n-1th frame information is similar to the determination process of the deformation reference data corresponding to the nth frame information. For the sake of brevity, it will not be repeated here.
[0175] The region parsing result of a source image refers to the result obtained by parsing the source image, so that the region parsing result is used to describe the locations of different regions in the source image, such as the location of the background region in the source image, the locations of different body region areas in the source image, the locations of different facial region areas in the source image, etc. Thus, in one possible implementation, the region parsing result is used to describe the location of at least one region in the source image.
[0176] In addition, the present application does not limit the process of obtaining the above-mentioned regional analysis results. For example, it can be implemented using any existing or future method that can divide an image into different regions, such as a body part recognition method, a facial part recognition method, or an image segmentation method.
[0177] In addition, the present application does not limit the implementation method of the above step 412. For example, it can be specifically as follows: first, based on the regional analysis result of the source image, determine the weight corresponding to at least one region in the source image, so that these weights can represent the smoothness of different regions; then, based on the weight corresponding to the at least one region, perform weighted processing on the predicted deformation parameters corresponding to the n-th frame information and the deformation reference data corresponding to the n-1-th frame information to obtain the deformation reference data corresponding to the n-th frame information, so that the deformation parameters of some regions in the deformation reference data corresponding to the n-th frame information, such as the background, clothes, etc., are as consistent as possible with the deformation parameters of the corresponding regions in the deformation reference data corresponding to the n-1-th frame information, and the deformation parameters of other regions in the deformation reference data corresponding to the n-th frame information, such as the eyes, mouth, etc., are as consistent as possible with the deformation parameters of the corresponding regions in the predicted deformation parameters corresponding to the n-1-th frame information. This is conducive to ensuring that regions such as the background and clothes remain stable in the generated images of different frames, and ensuring that regions such as the eyes and mouth change flexibly in the generated images of different frames, thereby helping to improve the generation quality of the image sequence.
[0178] For the at least one area in the source image above, the at least one area refers to the area described by the regional analysis result of the source image, so that the at least one area can represent some areas analyzed from the source image; and the weight corresponding to the at least one area refers to the weight required to be used when smoothing different areas, so that the weight corresponding to the at least one area can represent the degree of smoothing performed on different areas, such as performing a greater degree of smoothing on the background, clothing and other areas. This is conducive to ensuring the stability of different frame images in these areas, thereby effectively avoiding the impact caused by the instability of these areas, and thus helping to improve the generation effect of the image sequence.
[0179] It can be seen that in one possible implementation, if the region parsing results of the source image described above are used to describe the locations of K regions in the source image, then at least one region in the source image may include K regions. The weight corresponding to the kth region is used to characterize the degree of smoothing for the kth region. Furthermore, this application does not limit the implementation of the weight corresponding to the kth region; for example, it may include the weight value of a historical frame and the weight value of a current frame. The weight value of the historical frame refers to the degree of influence of the historical frame when smoothing the kth region. The weight value of the current frame refers to the degree of influence of the current frame when smoothing the kth region. Furthermore, this application does not limit the process for determining the weight corresponding to the kth region; for example, it may specifically be: based on the category of the kth region, searching a pre-established mapping relationship for the smoothing weight corresponding to the category, and using this as the weight corresponding to the kth region. The category of the kth region indicates the identity of the kth region, such as eyes, hair, etc. This mapping relationship is used to record the smoothing weights corresponding to multiple categories. k is a positive integer, k≤K, and K is a positive integer.
[0180] Research has found that some pose information in the Deformations requires smoothing, but almost no pose information in the Occlusion Map requires smoothing. Therefore, to further improve efficiency, the present application further provides a possible implementation of step 412 above. In this implementation, when the predicted deformation parameters corresponding to the n-th frame information include deformations and an occlusion map, step 412 may specifically include: first, determining a weight corresponding to at least one region in the source image based on the region analysis result of the source image; then, based on the weight corresponding to the at least one region, performing a weighted addition on the deformations in the predicted deformation parameters corresponding to the n-th frame information and the deformations in the deformation reference data corresponding to the n-1-th frame information to obtain the deformations in the deformation reference data corresponding to the n-th frame information, so that the occlusion map in the deformation reference data corresponding to the n-th frame information is consistent with the occlusion map in the predicted deformation parameters corresponding to the n-th frame information. In this way, pose adjustment is performed only on the deformations, effectively avoiding resource overhead caused by pose adjustment on the occlusion map, thereby improving efficiency.
[0181] Based on the relevant contents of steps 411 to 412 above, it can be seen that for some application scenarios, after obtaining the key point information of the source image and the key point information of the n-th frame information, the key point information of the source image and the key point information of the n-th frame information can be first input into the DMN to obtain deformations and an occlusion map output by the DMN. Then, based on the regional analysis result of the source image and the deformations in the deformation reference data corresponding to the n-1-th frame information above, the deformations output by the DMN are smoothed. Subsequently, the smoothed deformations and the occlusion map output by the DMN can be used as the deformation reference data corresponding to the n-1-th frame information, so that the deformation reference data can better represent the parameters required when performing posture adjustment processing on the source image based on the n-1-th frame information, such as the posture offset + the area that cannot be generated by deformation.
[0182] Based on the relevant content of step 41 above, it can be known that for some application scenarios, after obtaining the key point information of the source image, at least one reference feature of the source image, and the key point information of the nth frame information, the deformation reference data corresponding to the nth frame information can be determined based on the source image, the key point information of the source image, at least one reference feature of the source image, the key point information of the nth frame information, the regional analysis result of the source image, and the deformation reference data corresponding to the n-1th frame information above, so that the deformation reference data can better represent the parameters required to perform posture adjustment processing on the source image based on the nth frame information, thereby making the image generated based on the deformation reference data more accurate, which is beneficial to improving the image generation effect.
[0183] Step 42: Perform deformation processing on the source image according to the deformation reference data corresponding to the n-th frame information to obtain a deformed image corresponding to the n-th frame information.
[0184] Among them, the deformed image corresponding to the nth frame information refers to the result obtained by deforming the source image based on the deformation reference data corresponding to the nth frame information, so that the posture information in some areas of the deformed image corresponding to the nth frame information is consistent with the posture information in the corresponding areas in the nth frame information.
[0185] In addition, the present application does not limit the implementation method of the deformed image corresponding to the nth frame information above. For example, in some application scenarios, such as scenarios where there is no need to perform posture adjustment on areas such as the eyes and mouth, the deformed image corresponding to the nth frame information may refer to the result of deforming some areas in the source image based on the deformation reference data corresponding to the nth frame information, so that the posture information in these areas in the deformed image corresponding to the nth frame information is consistent with the posture information in the corresponding areas in the nth frame information.
[0186] For example, in some application scenarios, such as scenarios where posture adjustments need to be made to areas such as the eyes and mouth, in order to better improve the image generation effect, the deformed image corresponding to the nth frame information may refer to the result of deforming the deformable area in the source image based on the deformation reference data corresponding to the nth frame information, so that the posture information in the deformable area in the deformed image corresponding to the nth frame information is consistent with the posture information in the deformable area in the nth frame information, and the posture information in the non-deformable area in the deformed image corresponding to the nth frame information is consistent with the posture information in the non-deformable area in the source image.
[0187] Based on the relevant content of step 42 above, it can be seen that for some application scenarios, after obtaining the deformation reference data corresponding to the n-th frame information above, such as the Deformations+Occlusion Map, the Deformations can be used to perform pixel-by-pixel geometric deformation on the source image, and the Occlusion Map can be used to occlude areas that cannot be generated by deformation, thereby obtaining the deformed image corresponding to the n-th frame information.
[0188] Step 43: Determine the image generation result corresponding to the n-th frame information based on the deformed image corresponding to the n-th frame information.
[0189] It should be noted that the present application does not limit the implementation method of the above step 43. For example, in some application scenarios, such as scenarios where there is no need to adjust the posture of areas such as the eyes and mouth, the step 43 can specifically be: directly determining the deformed image corresponding to the n-th frame information as the image generation result corresponding to the n-th frame information.
[0190] For example, in some application scenarios, such as scenarios where posture adjustment is required for areas such as the eyes and mouth, in order to better improve the image generation effect, the above step 43 can specifically be: performing a complement process on the deformed image corresponding to the nth frame information to obtain the image generation result corresponding to the nth frame information. The complement process is used to perform posture complement on some areas in the deformed image, such as the eyes, mouth and other non-deformable areas, so that the posture information of the non-deformable areas in the image generation result generated based on the complement process is consistent with the posture information of the non-deformable areas in the nth frame information. This can effectively overcome the impact caused by the relatively poor posture adjustment effect of the deformation process on the non-deformable areas, thereby helping to improve the image generation quality. In addition, the present application does not limit the implementation method of the complement process. For example, the complement process can be implemented using a complement network. The complement network is used to perform the complement process; and the present application does not limit the implementation method of the complement network. For example, it can be implemented using any existing or future image generation network, such as an end-to-end P2P network. It can be seen that in one possible implementation, the completion network can be implemented using a Generator Network.
[0191] After research, it was found that for the above-mentioned completion processing, when the completion processing adopts an image generation mechanism, such as the image generation mechanism described by the Generator Network, the completion processing will not only perform generation processing on the non-deformable area, but also on the deformable area. This may cause some defects in the deformable area of the image obtained based on the completion processing, such as decreased clarity. Therefore, in order to better improve the image generation effect, the embodiment of the present application also provides a possible implementation method of the above step 43. Under this implementation method, the step 43 can specifically include the following steps 431-432.
[0192] Step 431: performing completion processing on the deformed image corresponding to the n-th frame information to obtain the completed image corresponding to the n-th frame information and the completed position representation data corresponding to the n-th frame information.
[0193] Among them, the completed image corresponding to the nth frame information refers to the image obtained by completing the non-deformable area in the deformed image corresponding to the nth frame information, so that the posture information of the non-deformable area in the completed image corresponding to the nth frame information is consistent with the posture information of the non-deformable area in the nth frame information.
[0194] The complement position representation data corresponding to the n-th frame information is used to describe the position to be complemented when the deformed image corresponding to the n-th frame information is complemented, such as the position of the non-deformable area; and the present application does not limit the implementation method of the complement position representation data corresponding to the n-th frame information, for example, it can be represented by a mask image, such as by Figure 3 The alpha channel shown is implemented.
[0195] In addition, the present application does not limit the implementation method of the above step 431. For example, it can be specifically: inputting the deformed image corresponding to the n-th frame information into the completion network, so that the completion network performs completion processing on the deformed image, and obtains and outputs the completed image corresponding to the n-th frame information and the completion position representation data corresponding to the n-th frame information.
[0196] Based on the relevant content of step 431 above, it can be known that for some application scenarios, after obtaining the deformed image corresponding to the n-th frame information, the deformed image corresponding to the n-th frame information can be input into P2P to obtain the predicted image and α channel output by the P2P, and the predicted image is used as the completed image corresponding to the n-th frame information, and the α channel is used as the completed position representation data corresponding to the n-th frame information.
[0197] Step 432: Based on the completed position representation data corresponding to the n-th frame information, the reference image corresponding to the n-th frame information and the completed image corresponding to the n-th frame information are fused to obtain the image generation result corresponding to the n-th frame information; the reference image is determined based on the source image and the deformation reference data corresponding to the n-th frame information.
[0198] The reference image corresponding to the n-th frame information is used to influence the image information within the deformable area; and the reference image is determined based on the source image and the deformation reference data corresponding to the n-th frame information.
[0199] In addition, the present application does not limit the implementation method of the reference image corresponding to the n-th frame information above. For example, in order to improve efficiency, the reference image corresponding to the n-th frame information can directly use the deformed image corresponding to the n-th frame information, such as Figure 2 or Figure 3 The deformation 1 shown is implemented.
[0200] For another example, in order to better improve the image generation quality, the present application also provides a process for obtaining the reference image corresponding to the nth frame information above, which can be specifically referred to as steps 4321 to 4323 below.
[0201] Step 4321: Perform super-resolution processing on the source image to obtain a super-resolution image.
[0202] Among them, super-resolution processing is used to improve the image quality of an image; and this application does not limit the implementation method of the super-resolution processing. For example, it can be implemented using any existing or future super-resolution implementation method.
[0203] Based on the relevant content of step 4321 above, it can be known that for the above source image, after obtaining the source image, super-resolution processing can be performed on the source image to obtain a super-resolution image, so that the super-resolution image can have better image quality, thereby enabling the super-resolution image to better describe the image information in the source image, such as appearance information.
[0204] Step 4322: Perform deformation processing on the super-resolution image based on the deformation reference data corresponding to the n-th frame information to obtain the super-resolution deformation result corresponding to the n-th frame information.
[0205] Among them, the super-resolution deformation result corresponding to the n-th frame information refers to the image obtained by deforming the super-resolution image based on the deformation reference data corresponding to the n-th frame information, so that the posture information of the deformable area in the super-resolution deformation result is consistent with the posture information of the deformable area in the n-th frame information.
[0206] In addition, the present application does not limit the implementation of the above step 4322. For example, when the deformation reference data corresponding to the above n-th frame information includes Deformations and Occlusion Map, the step 4322 can be specifically as follows: deforming the super-resolved image according to the Deformations to obtain the super-resolved deformation result corresponding to the n-th frame information, such as Figure 2 or Figure 3 Deformation 2 shown.
[0207] Step 4323: Determine a reference image corresponding to the n-th frame information based on the super-fractional deformation result corresponding to the n-th frame information.
[0208] It should be noted that the present application does not limit the implementation method of the above step 4323. For example, it can specifically be: directly determining the super-fractional deformation result corresponding to the n-th frame information as the reference image corresponding to the n-th frame information.
[0209] Based on the relevant content of steps 4321 to 4323 above, it can be known that in some application scenarios, the reference image corresponding to the nth frame information above can be determined based on the super-resolution processing result of the source image and the deformation reference data corresponding to the nth frame information, so that the reference image can provide more accurate image information within the deformable area, thereby making the image generation result corresponding to the nth frame information determined based on the reference image have better quality, which is conducive to improving the image generation quality.
[0210] In addition, this application does not limit the implementation method of the above step 432. For ease of understanding, some examples are used below for explanation.
[0211] Example 1. In one possible implementation, the above step 432 may specifically be: first, based on the completed position representation data corresponding to the n-th frame information, delete the pixel points in the non-deformable area from the reference image corresponding to the n-th frame information, and based on the completed position representation data corresponding to the n-th frame information, extract the pixel points in the non-deformable area from the completed image corresponding to the n-th frame information; then, fill the extracted pixel points into the image after the pixel points are deleted to obtain the image generation result corresponding to the n-th frame information, so that the pixel points in the non-deformable area in the image generation result corresponding to the n-th frame information come from the completed image corresponding to the n-th frame information, and the pixel points in the deformable area in the image generation result corresponding to the n-th frame information come from the reference image. This can effectively avoid the influence of the completion processing on the deformable area, thereby helping to improve the image generation quality.
[0212] Example 2. In one possible implementation, step 432 above may specifically be: first, based on the completed position representation data corresponding to the n-th frame information, determine the weights corresponding to each pixel point in the reference image corresponding to the n-th frame information and the weights corresponding to each pixel point in the completed image corresponding to the n-th frame information; then, based on these weights, perform weighted processing on the reference image and the completed image to obtain the image generation result corresponding to the n-th frame information, so as to realize the fusion of the pixel points in the non-deformable area of the completed image into the corresponding area in the reference image, which is conducive to improving the image generation quality.
[0213] Based on the relevant content of steps 431 to 432 above, it can be known that for the nth frame information, after obtaining the deformed image corresponding to the nth frame information, the deformed image is first complemented to obtain the complemented image corresponding to the nth frame information and its complemented position representation data; then, based on the complemented position representation data, the complemented image and the reference image corresponding to the nth frame information, the image generation result corresponding to the nth frame information is determined, so that the deformable area in the image generation result comes from the deformable area in the reference image, and the non-deformable area in the image generation result comes from the deformable area in the complemented image. This can effectively overcome the impact of the complementation processing on the deformable area, thereby helping to improve the image generation quality.
[0214] Research has found that, for the at least one reference feature of the source image described above, the at least one reference feature can represent some image information in the source image, such as appearance information. Therefore, in order to better improve the image generation effect, this application also provides a possible implementation of step 43 described above. Under this implementation, step 43 can specifically be: based on the at least one reference feature of the source image and the deformed image corresponding to the n-th frame information, determine the image generation result corresponding to the n-th frame information. In particular, because the at least one reference feature can provide richer image information for the image generation process, the final image generation result is more accurate, which is conducive to improving the image generation effect.
[0215] Furthermore, this application is not limited to the step of "determining an image generation result corresponding to the nth frame information based on at least one reference feature of the source image and the deformed image corresponding to the nth frame information" in the preceding paragraph. For example, this step can be implemented using a machine learning model. For another example, this step can specifically include steps 433 and 434 below.
[0216] Step 433: Based on at least one reference feature of the source image, the deformed image corresponding to the n-th frame information is completed to obtain a completion result corresponding to the n-th frame information.
[0217] The completion result corresponding to the n-th frame information refers to data obtained by completing the deformed image corresponding to the n-th frame information; and this application does not limit the completion result. For example, in some application scenarios, the completion result may include the completed image corresponding to the n-th frame information. For another example, in other application scenarios, the completion result may include the completed image corresponding to the n-th frame information and the completed position representation data corresponding to the n-th frame information.
[0218] In addition, the present application does not limit the implementation method of the above step 433. For example, it can be specifically: inputting at least one reference feature of the source image and the deformed image corresponding to the n-th frame information into the completion network, and determining the completion result corresponding to the n-th frame information based on the output data of the completion network.
[0219] Step 434: Determine the image generation result corresponding to the n-th frame information according to the completion result corresponding to the n-th frame information.
[0220] It should be noted that the present application does not limit the implementation method of the above step 434. For example, if the completion result corresponding to the n-th frame information includes the completed image corresponding to the n-th frame information, then the step 434 can specifically be: determining the completed image as the image generation result corresponding to the n-th frame information.
[0221] For example, if the completion result corresponding to the n-th frame information includes the completed image corresponding to the n-th frame information and the completed position representation data corresponding to the n-th frame information, then the above step 434 can be specifically: based on the completed position representation data corresponding to the n-th frame information, the reference image corresponding to the n-th frame information and the completed image corresponding to the n-th frame information are fused to obtain the image generation result corresponding to the n-th frame information.
[0222] Based on the relevant content of steps 433 to 434 above, it can be seen that for some application scenarios, after obtaining the deformed image corresponding to the n-th frame information, the deformed image corresponding to the n-th frame information can first be supplemented based on at least one reference feature of the source image to obtain a supplemented result corresponding to the n-th frame information; and then the image generation result corresponding to the n-th frame information is determined based on the supplemented result. In particular, because the at least one reference feature can represent some image information in the source image, the at least one reference feature can provide relatively rich image information for the supplementation process, thereby making the supplementation result obtained based on the supplementation process more accurate, and further making the image generation result obtained based on the supplementation result more accurate, which is conducive to improving the quality of image generation.
[0223] Research has found that for at least one reference feature of the above source image, the at least one reference feature can represent some posture information in the source image. Therefore, in order to better improve the image generation effect, deformation processing can be performed on part or all of the at least one reference feature. Based on this, the present application also provides a possible implementation method for determining the image generation result corresponding to the n-th frame information above. Under this implementation method, when the at least one reference feature includes at least one feature to be deformed, the process of determining the image generation result corresponding to the n-th frame information can include the following steps 435-436.
[0224] Step 435: For any feature to be deformed, deformation processing is performed on the feature to be deformed according to the deformation reference data corresponding to the n-th frame information to obtain a deformation result of the feature to be deformed.
[0225] The feature to be deformed refers to a reference feature that exists in at least one reference feature of the source image and needs to be subjected to posture adjustment processing.
[0226] Furthermore, this application does not limit the implementation of the at least one feature to be deformed. For example, the at least one feature to be deformed may include features present in the at least one reference feature of the source image that can represent posture information. For ease of understanding, the following two examples are provided for illustration.
[0227] In Example 1, when the at least one reference feature of the source image includes the source image's facial contour detection result, the source image's foreground segmentation result, and the source image's image features, because the facial contour detection result, the foreground segmentation result, and the image features can all represent some pose information in the source image, to further improve accuracy, the at least one feature to be deformed may include the facial contour detection result, the foreground segmentation result, and the image features. Therefore, in one possible implementation, the at least one feature to be deformed may include all of the at least one reference feature.
[0228] Example 2: When at least one reference feature of the above source image includes the facial contour detection result of the source image, the foreground segmentation result of the source image, and the image feature of the source image, although the facial contour detection result, the foreground segmentation result and the image feature can all represent some posture information in the source image, the image feature can also comprehensively represent other information in the source image in addition to the posture information, such as appearance information, etc., so in order to avoid the posture adjustment process affecting other information represented by the image feature, the image feature may not be subjected to posture adjustment processing. Based on this, it can be seen that the above at least one feature to be deformed can include the facial contour detection result and the foreground segmentation result. It can be seen that, under a possible implementation, the at least one feature to be deformed can include part of the at least one reference feature.
[0229] Furthermore, for the gth feature to be deformed, the deformation result of the gth feature to be deformed refers to the deformation result obtained by deforming the gth feature to be deformed based on the deformation reference data corresponding to the nth frame information, so that the posture information represented by the deformation result is consistent with the corresponding posture information in the nth frame information. Here, g is a positive integer, g ≤ G, G is a positive integer, and G represents the number of features in the at least one feature to be deformed.
[0230] Based on the relevant content of step 435 above, it can be known that for at least one reference feature of the above source image, if the at least one reference feature includes at least one feature to be deformed, then after obtaining the deformation reference data corresponding to the above n-th frame information, each feature to be deformed can be deformed according to the deformation reference data to obtain the deformation results of each feature to be deformed, so that some posture information represented by the deformation result is consistent with the corresponding posture information in the n-th frame information.
[0231] Step 436 : Determine an image generation result corresponding to the n-th frame information based on the deformation result of the at least one feature to be deformed and the deformed image corresponding to the n-th frame information.
[0232] It should be noted that this application does not limit the implementation method of the above step 436. For ease of understanding, two examples are used below to illustrate.
[0233] Example 1: When the at least one feature to be deformed includes all of the at least one reference feature, the step 436 may specifically include the following steps 4361 and 4362.
[0234] Step 4361: Based on the deformation result of at least one feature to be deformed, the deformed image corresponding to the n-th frame information is completed to obtain the completion result corresponding to the n-th frame information.
[0235] It should be noted that the implementation of step 4361 is similar to the implementation of step 433 above. For the sake of brevity, it will not be repeated here.
[0236] Step 4362: Determine the image generation result corresponding to the n-th frame information based on the completion result corresponding to the n-th frame information.
[0237] It should be noted that for the relevant content of step 4362, please refer to step 434 above.
[0238] Based on the relevant contents of steps 4361 to 4362 above, it can be known that for some application scenarios, if at least one of the features to be deformed includes all of the at least one reference feature above, the deformed image corresponding to the n-th frame information can be directly complemented based on the deformation results of all the features to be deformed to obtain the complement result corresponding to the n-th frame information; and then the image generation result corresponding to the n-th frame information is determined based on the complement result. Among them, because the posture information represented by the deformation results of these features to be deformed is consistent with the corresponding posture information in the n-th frame information, and some posture information in the deformed image is also consistent with the corresponding posture information in the n-th frame information, so that there is no posture conflict between the deformation results of these features to be deformed and the deformed image, this can effectively avoid the impact caused by the existence of posture conflicts, thereby helping to improve the image generation quality.
[0239] Example 2: When at least one reference feature of the above source image includes the facial contour detection result of the source image, the foreground segmentation result of the source image, and the image feature of the source image, and at least one feature to be deformed includes the facial contour detection result and the foreground segmentation result, the above step 436 can specifically be: determining the image generation result corresponding to the nth frame information based on the deformation result of the at least one feature to be deformed, the image feature, and the deformed image corresponding to the nth frame information.
[0240] It should be noted that this application does not limit the implementation method of the step in the previous paragraph "determining the image generation result corresponding to the n-th frame information based on the deformation result of the at least one feature to be deformed, the image feature and the deformed image corresponding to the n-th frame information". For example, it may specifically include the following steps 4363-4364.
[0241] Step 4363: Based on the deformation result of the at least one feature to be deformed and the image feature, the deformed image corresponding to the n-th frame information is completed to obtain the completion result corresponding to the n-th frame information.
[0242] It should be noted that the present application does not limit the implementation method of step 4363. For example, it can specifically be: inputting the deformation result of at least one feature to be deformed of the source image, the image feature of the source image, and the deformed image corresponding to the n-th frame information into the completion network, and determining the completion result corresponding to the n-th frame information based on the output data of the completion network.
[0243] Step 4364: Determine the image generation result corresponding to the n-th frame information based on the completion result corresponding to the n-th frame information.
[0244] It should be noted that for the relevant content of step 4362, please refer to step 434 above.
[0245] Based on the relevant contents of steps 4363 to 4364 above, it can be known that for some application scenarios, if the at least one feature to be deformed includes a portion of the at least one reference feature, the deformed image corresponding to the n-th frame information can be directly complemented based on the deformation result of this portion and the remaining portion of the at least one reference feature to obtain the complement result corresponding to the n-th frame information; and then the image generation result corresponding to the n-th frame information is determined based on the complement result. In particular, because a portion of the at least one reference feature has been deformed, but the other portion has not been deformed, when the complement process is performed based on these two portions, not only relatively accurate posture information can be obtained from the former portion, but also relatively accurate other information, such as appearance information, can be obtained from the latter portion, thereby making the image generation result determined based on the complement process more accurate, which is beneficial to improving the image generation quality.
[0246] Based on the relevant content of steps 435 to 436 above, it can be seen that for some application scenarios, after obtaining the deformed image corresponding to the n-th frame information, the deformed image corresponding to the n-th frame information can first be complemented based on part or all of the deformation results in at least one reference feature of the source image to obtain the complement result corresponding to the n-th frame information; and then the image generation result corresponding to the n-th frame information is determined based on the complement result. In particular, because the posture information represented by the deformation result is consistent with the corresponding posture information in the n-th frame information, the deformation result can provide relatively accurate posture information for the complement process, thereby making the image generation result determined based on the complement process more accurate, which is conducive to improving the image generation quality.
[0247] Based on the relevant contents of steps 41 to 43 above, it can be known that for some application scenarios, after obtaining the key point information of the source image, at least one reference feature of the source image, and the key point information of the n-th frame information, these three data and the source image can be input into the DMN to obtain the Deformations and Occlusion Map predicted by the DMN, and according to the regional analysis results of the source image and the Deformations corresponding to the n-1-th frame information, different regions in the Deformations corresponding to the n-th frame information are smoothed to different degrees to obtain smoothed Deformations, so that the smoothed Deformations are used to perform pixel-by-pixel geometric deformation on the source image, and obtain a deformation result whose posture is close to the posture information in the n-th frame information, and make the Occlusion Map is used to block some areas that cannot be generated by deformation, such as eyes, mouths, etc.; then, after P2P receives the deformed source image that blocks the area to be generated, the P2P performs a refined completion process on the deformed source image, obtains and outputs a predicted image that completes the corresponding area and an alpha channel used to describe the position of the completed area, so that the completed area in the predicted image can be fused to a certain deformation result of the source image based on the alpha channel. Figure 3 In the deformation 1 or deformation 2 shown, the image generation result corresponding to the n-th frame information is obtained, which is beneficial to improving the image quality.
[0248] In addition, in some application scenarios, the above S2 can be implemented with the help of a preset model. Based on this, it can be seen that in one possible implementation, the S2 can be specifically: the preset model performs image generation processing based on the source image, the key point information of the source image, at least one reference feature of the source image, and the key point information of the driving information, and obtains and outputs the image generation result corresponding to the driving information. Among them, the preset model is used to perform image generation processing on the input data of the preset model, and this application does not limit the implementation method of the preset model. For example, the preset model can include a deformation network and a completion network.
[0249] Based on the relevant contents of S1 to S2 above, it can be seen that for the image generation method provided in the embodiment of the present application, a source image and driving information are first obtained; then, image generation processing is performed based on the source image, the key point information of the source image, at least one reference feature of the source image, and the key point information of the driving information to obtain an image generation result corresponding to the driving information, so that the appearance information presented by the image generation result is consistent with the appearance information presented by the source image, and the posture presented by the image generation result is consistent with the posture represented by the driving information, so that the driving result of the source image under the driving information can be automatically generated. Among them, because the at least one reference feature includes one or more of the facial contour detection result, the foreground segmentation result and the image feature, so that the at least one reference feature can better represent the image information of the source image, such as posture information, appearance information, etc., so that the at least one reference feature can supplement some image information that cannot be represented by the key point information of the source image, thereby making the multi-condition guided image generation processing based on the key point information of the source image and the at least one reference feature have better performance, which is conducive to improving the image generation quality.
[0250] In addition, the present application does not limit the execution subject of the image generation method provided in the embodiment of the present application. For example, the image generation method provided in the embodiment of the present application can be applied to a terminal device. For another example, the image generation method provided in the embodiment of the present application can also be implemented with the help of a data interaction process between a terminal device and a server. Among them, the terminal device can be a smart phone, a computer, a personal digital assistant (PDA), a tablet computer, etc. The server can be a stand-alone server, a cluster server, or a cloud server.
[0251] In addition, this application does not limit the application scenarios of the above image generation method. For ease of understanding, the following description is combined with three scenarios.
[0252] Scenario 1: The image generation method provided in this application can be applied to some model update scenarios, such as online or offline model update scenarios. Based on this, it can be seen that this application also provides a model update process, which can specifically include the following steps 51-53.
[0253] Step 51: Obtain source image and driving information.
[0254] It should be noted that the relevant contents of step 51 can be referred to the relevant contents of S1 above. For example, step 51 may specifically be: randomly extracting two frames of images from the sample video, one frame of image being used as the source image and the other frame of image being used as the driving information.
[0255] Step 52: A preset model performs image generation processing based on the source image, key point information of the source image, at least one reference feature of the source image, and key point information of the driving information to obtain an image generation result corresponding to the driving information, a predicted foreground segmentation corresponding to the driving information, and a predicted facial contour corresponding to the driving information.
[0256] For details about the image generation result corresponding to the driving information, please refer to the relevant content in S2 above.
[0257] The predicted foreground segmentation corresponding to the driving information is used to describe the location of the foreground prediction in the driving information.
[0258] The predicted facial contour corresponding to the driving information is used to describe the predicted position of the facial contour in the driving information.
[0259] In addition, this application does not limit the implementation of the above step 52. For example, the step 52 may be implemented as follows: Figure 2 It can be seen that in one possible implementation, when the preset model includes a deformable network and a completion network, step 52 may specifically include: the deformable network processes the key point information of the source image, at least one reference feature of the source image, and the key point information of the driving information to obtain deformable reference data corresponding to the driving information, predicted foreground segmentation corresponding to the driving information, and predicted facial contour corresponding to the driving information; then, the completion network and the deformable reference data corresponding to the driving information are used to determine the image generation result corresponding to the driving information.
[0260] Step 53: Based on the image generation result corresponding to the above driving information, the label information corresponding to the image generation result, the predicted foreground segmentation corresponding to the driving information, the label information corresponding to the predicted foreground segmentation, the predicted facial contour corresponding to the driving information, and the label information corresponding to the predicted facial contour, update the preset model, and return to continue executing the above step 51 and subsequent steps until the preset stop condition is reached; the label information is determined based on the driving information.
[0261] Among them, for the image generation result corresponding to the above driving information, the label information corresponding to the image generation result is used to represent the true value of the image corresponding to the driving information, so that the label information can be used as guidance information corresponding to the image generation result; and this application does not limit the implementation method of the label information. For example, the label information can adopt the driving information, such as Figure 2 The drive image shown is implemented.
[0262] In addition, for the predicted foreground segmentation corresponding to the above driving information, such as Figure 2For Mask2 shown, the label information corresponding to the predicted foreground segmentation is used to represent the true value of the foreground segmentation of the driving information, so that the label information can serve as guidance information corresponding to the predicted foreground segmentation; and this application does not limit the implementation method of the label information. For example, the label information can be implemented using the foreground segmentation result of the driving information or the foreground segmentation information manually labeled for the driving information in advance. The foreground segmentation result of the driving information refers to the result obtained by performing foreground segmentation processing on the driving information; and the method for obtaining the foreground segmentation result of the driving information is similar to the method for obtaining the foreground segmentation result of the source image mentioned above.
[0263] In addition, for the predicted facial contour corresponding to the above driving information, such as Figure 2 For the contour 2 shown in FIG. 2 , the label information corresponding to the predicted facial contour is used to represent the true value of the facial contour of the driving information, so that the label information can serve as guidance information corresponding to the predicted facial contour. Moreover, this application does not limit the implementation method of the label information. For example, the label information can be implemented using the facial contour detection result of the driving information or the facial contour information manually annotated for the driving information in advance. The facial contour detection result of the driving information refers to the result obtained by performing facial contour detection processing on the driving information. The facial contour detection result of the driving information is obtained in a manner similar to the method for obtaining the facial contour detection result of the source image mentioned above.
[0264] In addition, the present application does not limit the implementation method of the above step 53. For example, it can be specifically as follows: first, based on the similarity between the above image generation result and the label information corresponding to the image generation result, the similarity between the above predicted foreground segmentation and the label information corresponding to the predicted foreground segmentation, and the similarity between the above predicted facial contour and the label information corresponding to the predicted facial contour, determine the model loss so that the model loss can represent the model performance; then, update the preset model based on the model loss, and return to continue executing the above step 51 and its subsequent steps until the preset stop condition is reached.
[0265] For example, in some application scenarios, when the above key point information is determined with the help of an implicit key point detection model, in order to better improve the model performance, the present application also provides a possible implementation method of the above step 53. Under this implementation method, the step 53 can specifically be: based on the above image generation result, the label information corresponding to the image generation result, the above predicted foreground segmentation, the label information corresponding to the predicted foreground segmentation, the above predicted facial contour, and the label information corresponding to the predicted facial contour, update the implicit key point detection model and the preset model, and return to continue executing the above step 51 and its subsequent steps until the preset stop condition is reached.
[0266] Furthermore, the present application does not limit the implementation of the above-mentioned preset stop condition. For example, the preset stop condition may specifically include: the model loss is lower than a preset loss threshold. For another example, the preset stop condition may include: the rate of change of the model loss is lower than a preset rate of change threshold. For another example, the preset stop condition may include: the number of model updates is higher than a preset number threshold.
[0267] Based on the relevant content of steps 51 to 53 above, it can be seen that for some model update scenarios, two different frames of images are first selected from the same video as the source image and driving image of the current round; then the implicit key point detection model, the preset model, the source image and the driving image are used to obtain the image generation result, the predicted foreground segmentation and the predicted facial contour corresponding to the driving image; then, based on the image generation result, the predicted foreground segmentation and the predicted facial contour, the implicit key point detection model and the preset model are updated so that the updated implicit key point detection model and the preset model have better performance, and the next round of update process is performed based on the updated implicit key point detection model and the preset model, and the iterative cycle is repeated until the preset stop condition is reached, which is conducive to improving model performance. It can be seen that during the updating process, the present application not only learns the image generation capability, but also learns the foreground segmentation capability and the facial contour detection capability, so that the finally trained model has better image generation performance, foreground segmentation performance and facial contour detection performance. In this way, the pixel sticking feeling generated by driving some areas such as the background boundary area, the chin area, etc. can be effectively reduced during the image generation process, thereby helping to improve the image generation quality.
[0268] Scenario 2: The image generation method provided in this application can be applied to perform some image generation tasks, such as the generation task of a single image or the generation task of an image sequence.
[0269] In fact, in order to better reduce the resource pressure of the client, the image generation method provided by this application can be completed by the server and the client in a collaborative manner. For ease of understanding, the following description is made in conjunction with the generation process of the image sequence.
[0270] As an example, the image sequence generation process provided in this application may specifically include the following steps 61 to 64.
[0271] Step 61: Get the input image on the server, such as Figure 4 After the input image is shown, the server performs some processing on the input image to obtain the source image, the key point information of the source image and at least one reference feature, such as Figure 4 Multiple features shown.
[0272] Step 62: After the server obtains the nth frame information in the driving information, the server first obtains the key point information of the nth frame information, where n is a positive integer and n≤N.
[0273] Step 63: After the client receives the source image, key point information of the source image, at least one reference feature of the source image, and key point information of the n-th frame information sent by the server, the client performs image generation processing based on the source image, key point information of the source image, at least one reference feature of the source image, and key point information of the n-th frame information to obtain an image generation result of the n-th frame information, such as Figure 3 The nth frame is shown as the generated image, where n is a positive integer, n≤N.
[0274] Based on the relevant contents of steps 61 to 63 above, it can be seen that for some execution scenarios of generated tasks, the server and the client can be used in a collaborative manner, such as Figure 4 The collaborative method shown realizes the image generation method provided by the present application. Among them, because the source image corresponding to different frame information in one task is the same image, the server can be used to complete the relevant processing for the source image, so that the relevant content of the source image can be provided to the client in a one-time manner. This can avoid the resource overhead caused by the client completing the relevant processing for the source image, which is beneficial to reducing the resource pressure of the client. Also, because the relevant processing of each frame information in the driving information has a large resource overhead, the server can be used to complete the relevant processing for each frame information, so that the server can continuously send the relevant content of different frame information to the client later. This can avoid the resource overhead caused by the client completing the relevant processing for each frame information, which is beneficial to reducing the resource pressure of the client. In addition, because more communication resources need to be consumed when sending images, in order to balance the resource overhead as much as possible, the client can perform image generation processing based on multiple key point information, which is beneficial to improving the image generation effect.
[0275] Scenario three: the image generation method provided in this application can be used to generate a video of a virtual image; and the video generation process can include the following steps 71 to 73.
[0276] Step 71: Determine the source image based on the pre-built virtual image.
[0277] The term "avatar" refers to a pre-created virtual image for a user. This application does not limit the implementation of this virtual image; for example, it could be a three-dimensional digital human. Furthermore, this application does not limit the method for obtaining this virtual image; for example, it could be determined based on user-provided information, such as an image. Alternatively, the virtual image could be selected by the user from a selection of candidate virtual images.
[0278] In addition, the present application does not limit the implementation method of the above step 71. For example, it can specifically be: taking a photo of the virtual image to obtain the source image.
[0279] Step 72: After acquiring the above text data, after receiving the text data, convert the text data into voice data, and determine the driving information based on the voice data.
[0280] It should be noted that, for the relevant content of step 72, please refer to the above.
[0281] Step 73: Perform image generation processing based on the source image, key point information of the source image, at least one reference feature of the source image, and key point information of the driving information, so that the image generation result includes a generated image corresponding to each frame data in the voice data.
[0282] It should be noted that for the relevant content of step 73, please refer to the relevant content in S2 above.
[0283] Step 74: Based on the above voice data and the above image generation result, construct a video corresponding to the above virtual image, so that the video is used to describe the state of the virtual image in different frames.
[0284] It should be noted that this application does not limit the implementation method of the above step 74.
[0285] Based on the relevant content of steps 71 to 74 above, it can be seen that in some application scenarios, the image generation method provided in this application can be used to determine a driving video for a virtual character. In particular, because the image generation method has good performance, the resulting driving video can better meet user needs, which is conducive to improving the user experience.
[0286] Based on the image generation method provided in the embodiment of the present application, the embodiment of the present application also provides an image generation device. Figure 5 Explain and illustrate. Figure 5 This is a schematic diagram of the structure of an image generation device provided in an embodiment of the present application. It should be noted that for the technical details of the image generation device provided in an embodiment of the present application, please refer to the relevant content of the image generation method above.
[0287] like Figure 5 As shown, the image generation device 500 provided in the embodiment of the present application includes:
[0288] An acquisition unit 501 is used to acquire a source image and driving information;
[0289] A generation unit 502 is configured to perform image generation processing based on the source image, key point information of the source image, at least one reference feature of the source image, and key point information of the driving information to obtain an image generation result corresponding to the driving information; the at least one reference feature includes one or more of a facial contour detection result, a foreground segmentation result, and an image feature.
[0290] In a possible implementation manner, the driving information includes N frames of information, where N is a positive integer;
[0291] The generation unit 502 is specifically used to: determine the deformation reference data corresponding to the n-th frame information based on the key point information of the source image, at least one reference feature of the source image, and the key point information of the n-th frame information; n is a positive integer, n≤N; based on the deformation reference data corresponding to the n-th frame information, deform the source image to obtain a deformed image corresponding to the n-th frame information; based on the deformed image corresponding to the n-th frame information, determine the image generation result corresponding to the n-th frame information.
[0292] In a possible implementation manner, the generating unit 502 is specifically configured to determine an image generation result corresponding to the n-th frame information based on the at least one reference feature and the deformed image corresponding to the n-th frame information.
[0293] In one possible implementation, the at least one reference feature includes at least one feature to be deformed;
[0294] The generating unit 502 is further configured to: for any of the features to be deformed, perform deformation processing on the feature to be deformed according to the deformation reference data corresponding to the n-th frame information to obtain a deformation result of the feature to be deformed;
[0295] The generating unit 502 is specifically configured to determine an image generation result corresponding to the n-th frame information according to the deformation result of the at least one feature to be deformed and the deformed image corresponding to the n-th frame information.
[0296] In one possible implementation, the at least one reference feature includes a facial contour detection result, a foreground segmentation result, and an image feature; the at least one feature to be deformed includes a facial contour detection result and a foreground segmentation result; and the image generation result corresponding to the n-th frame information is determined based on the deformation result of the at least one feature to be deformed, the image feature, and the deformed image corresponding to the n-th frame information.
[0297] In one possible implementation, if n≥2, the generation unit 502 is specifically used to: determine the predicted deformation parameters corresponding to the n-th frame information based on the key point information of the source image, at least one reference feature of the source image, and the key point information of the n-th frame information; adjust the predicted deformation parameters corresponding to the n-th frame information based on the regional analysis result of the source image and the deformation reference data corresponding to the n-1-th frame information to obtain the deformation reference data corresponding to the n-th frame information; the regional analysis result is used to describe the position of at least one region in the source image.
[0298] In one possible implementation, the acquisition unit 501 is specifically used to: after acquiring the input image, perform preset processing on the input image to obtain the source image; the preset processing includes at least one of expression adjustment processing and quality increase processing; the expression adjustment processing is used to adjust the image expression state to the target expression state; the quality increase processing is used to enhance the image quality.
[0299] In a possible implementation manner, the quality enhancement process includes at least one of a beautification process and a clarity enhancement process; the beautification process is used to enhance the aesthetics of the image; and the clarity enhancement process is used to enhance the clarity of the image.
[0300] In one possible implementation, the image generation process is implemented using a preset model;
[0301] The preset model is specifically configured to perform image generation processing based on the source image, key point information of the source image, at least one reference feature of the source image, and key point information of the driving information, to obtain an image generation result corresponding to the driving information, a predicted foreground segmentation corresponding to the driving information, and a predicted facial contour corresponding to the driving information;
[0302] The image generating device 500 further includes:
[0303] An updating unit is configured to update the preset model based on the image generation result, the label information corresponding to the image generation result, the predicted foreground segmentation, the label information corresponding to the predicted foreground segmentation, the predicted facial contour, and the label information corresponding to the predicted facial contour; the label information is determined based on the driving information.
[0304] In one possible implementation, the source image is determined based on a pre-built virtual image; the driving information is determined based on voice data; the voice data is converted from text data; and the image generation result includes a generated image corresponding to each frame of data in the voice data;
[0305] The image generating device 500 further includes:
[0306] A construction unit is used to construct a video corresponding to the virtual image based on the voice data and the image generation result.
[0307] Based on the relevant content of the above-mentioned image generation device 500, it can be seen that for the image generation device 500 provided in the embodiment of the present application, a source image and driving information are first obtained; then, image generation processing is performed based on the source image, the key point information of the source image, at least one reference feature of the source image, and the key point information of the driving information to obtain an image generation result corresponding to the driving information, so that the appearance information presented by the image generation result is consistent with the appearance information presented by the source image, and the posture situation presented by the image generation result is consistent with the posture situation represented by the driving information, so that the driving result of the source image under the driving information can be automatically generated. Among them, because the at least one reference feature includes one or more of the facial contour detection result, the foreground segmentation result and the image feature, so that the at least one reference feature can better represent the image information of the source image, such as posture information, appearance information, etc., so that the at least one reference feature can supplement some image information that cannot be represented by the key point information of the source image, thereby making the multi-condition guided image generation processing based on the key point information of the source image and the at least one reference feature have better performance, which is conducive to improving the image generation quality.
[0308] In addition, an embodiment of the present application also provides an electronic device, which includes a processor and a memory: the memory is used to store instructions or computer programs; the processor is used to execute the instructions or computer programs in the memory, so that the electronic device executes any implementation of the image generation method provided in the embodiment of the present application.
[0309] See also Figure 6 , which shows a schematic structural diagram of an electronic device 600 suitable for implementing the embodiments of the present disclosure. The terminal devices in the embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 6 The electronic device shown is only an example and should not limit the functions and scope of use of the embodiments of the present disclosure.
[0310] like Figure 6As shown, the electronic device 600 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 608 into a random access memory (RAM) 603. Various programs and data required for the operation of the electronic device 600 are also stored in the RAM 603. The processing device 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0311] Typically, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device 600 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 6 The electronic device 600 is shown with various devices, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead.
[0312] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 609, or installed from the storage device 608, or installed from the ROM 602. When the computer program is executed by the processing device 601, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.
[0313] The electronic device provided by the embodiment of the present disclosure and the method provided by the above embodiment belong to the same inventive concept. For technical details not fully described in this embodiment, please refer to the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.
[0314] An embodiment of the present application further provides a computer-readable medium, in which instructions or computer programs are stored. When the instructions or computer programs are executed on a device, the device executes any implementation of the image generation method provided in the embodiment of the present application.
[0315] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.
[0316] In some embodiments, the client and server can communicate using any currently known or future developed network protocol, such as HTTP (Hypertext Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.
[0317] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0318] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device can perform the method.
[0319] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including, but not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0320] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0321] The units involved in the embodiments described in this disclosure may be implemented in software or hardware, wherein the name of a unit / module does not, in some cases, limit the unit itself.
[0322] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0323] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0324] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the common and similar parts between the various embodiments. For the systems or devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the methods.
[0325] It should be understood that in this application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0326] It should also be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.
[0327] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0328] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not limited to the embodiments shown herein, but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. An image generation method, characterized in that: The method comprises: Obtain source image and driving information; Image generation processing is performed based on the source image, key point information of the source image, at least one reference feature of the source image, and key point information of the driving information to obtain an image generation result corresponding to the driving information; the at least one reference feature includes one or more of a facial contour detection result, a foreground segmentation result, and an image feature.
2. The method according to claim 1, characterized in that The driving information includes N frames of information, where N is a positive integer; The process of determining the image generation result corresponding to the n-th frame information includes: Determining deformation reference data corresponding to the n-th frame information based on key point information of the source image, at least one reference feature of the source image, and key point information of the n-th frame information, where n is a positive integer, n≤N; Performing deformation processing on the source image according to the deformation reference data corresponding to the n-th frame information to obtain a deformed image corresponding to the n-th frame information; An image generation result corresponding to the n-th frame information is determined based on the deformed image corresponding to the n-th frame information.
3. The method according to claim 2, characterized in that The determining, based on the deformed image corresponding to the n-th frame information, an image generation result corresponding to the n-th frame information includes: An image generation result corresponding to the n-th frame information is determined based on the at least one reference feature and the deformed image corresponding to the n-th frame information.
4. The method according to claim 2, characterized in that The at least one reference feature includes at least one feature to be deformed; The method further comprises: For any of the features to be deformed, deformation processing is performed on the feature to be deformed according to the deformation reference data corresponding to the n-th frame information to obtain a deformation result of the feature to be deformed; The determining, based on the deformed image corresponding to the n-th frame information, an image generation result corresponding to the n-th frame information includes: An image generation result corresponding to the n-th frame information is determined according to the deformation result of the at least one feature to be deformed and the deformed image corresponding to the n-th frame information.
5. The method according to claim 4, characterized in that The at least one reference feature includes a face contour detection result, a foreground segmentation result, and an image feature; The at least one feature to be deformed includes a facial contour detection result and a foreground segmentation result; The image generation result corresponding to the n-th frame information is determined based on the deformation result of the at least one feature to be deformed, the image feature, and the deformed image corresponding to the n-th frame information.
6. The method according to claim 2, characterized in that If n≥2, the process of determining the deformation reference data corresponding to the n-th frame information includes: determining a predicted deformation parameter corresponding to the n-th frame information based on key point information of the source image, at least one reference feature of the source image, and key point information of the n-th frame information; Based on the regional analysis result of the source image and the deformation reference data corresponding to the n-1 frame information, the predicted deformation parameters corresponding to the n frame information are adjusted to obtain the deformation reference data corresponding to the n frame information; the regional analysis result is used to describe the position of at least one region in the source image.
7. The method according to claim 1, characterized in that The process of acquiring the source image includes: After acquiring the input image, the input image is subjected to preset processing to obtain the source image; the preset processing includes at least one of expression adjustment processing and quality increase processing; the expression adjustment processing is used to adjust the image expression state to the target expression state; the quality increase processing is used to enhance the image quality.
8. The method according to claim 7, characterized in that The quality enhancement process includes at least one of a beautification process and a clarity enhancement process; The beautification process is used to enhance the aesthetics of the image; The clarity enhancement processing is used to enhance the clarity of an image.
9. The method according to claim 1, characterized in that The image generation process is achieved by using a preset model; The preset model is specifically configured to perform image generation processing based on the source image, key point information of the source image, at least one reference feature of the source image, and key point information of the driving information, to obtain an image generation result corresponding to the driving information, a predicted foreground segmentation corresponding to the driving information, and a predicted facial contour corresponding to the driving information; The method further comprises: The preset model is updated based on the image generation result, the label information corresponding to the image generation result, the predicted foreground segmentation, the label information corresponding to the predicted foreground segmentation, the predicted facial contour, and the label information corresponding to the predicted facial contour; the label information is determined based on the driving information.
10. The method according to claim 1, characterized in that The source image is determined based on a pre-constructed virtual image; The driving information is determined based on voice data; the voice data is converted from text data; The image generation result includes a generated image corresponding to each frame of data in the voice data; After obtaining the image generation result corresponding to the driving information, the method further includes: A video corresponding to the virtual image is constructed based on the voice data and the image generation result.
11. An image generating device, characterized in that: include: An acquisition unit, configured to acquire source images and driving information; a generating unit, configured to perform image generation processing based on the source image, key point information of the source image, at least one reference feature of the source image, and key point information of the driving information, to obtain an image generation result corresponding to the driving information; The at least one reference feature includes one or more of a face contour detection result, a foreground segmentation result, and an image feature.
12. An electronic device, characterized in that: The device includes: a processor and a memory; The memory is used to store instructions or computer programs; The processor is configured to execute the instructions or computer program in the memory, so that the electronic device executes the method according to any one of claims 1 to 10.
13. A computer-readable medium, characterized in that The computer-readable medium stores instructions or a computer program, and when the instructions or the computer program are executed on a device, the device is caused to execute the method according to any one of claims 1 to 10.
14. A computer program product, characterized in that The method comprises a computer program carried on a non-transitory computer-readable medium, the computer program comprising a program code for executing the method according to any one of claims 1 to 10.