Image synthesis method and device, medium and product
By stitching facial images with standard hairstyle images using head edge features as control conditions, the method addresses the challenge of uniform image synthesis across diverse styles, achieving efficient and cost-effective results.
Patent Information
- Application Number
- CN202510361537.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-15
AI Technical Summary
The existing image synthesis technology is difficult to meet users' diverse needs for different hairstyles and clothing styles, and the model training cost is high, so it is impossible to achieve the unity and consistency of images.
By obtaining facial images and standard hairstyle images for feature matching and stitching, extracting head edge features as control conditions, combining standard posture images for overall stitching, generating target images with standard hairstyles and stances, reducing dependence on the model.
It realizes the diversified ability of image synthesis, reduces the cost of model training, ensures the uniformity and consistency of images, and reduces the need for optimization and adjustment of models.
Smart Images

Figure CN120318371A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and particularly to an image synthesis method, device, medium, and product. Background Art
[0002] With the development of image synthesis technology, especially in combination with artificial intelligence models, the image synthesis ability has been significantly improved.
[0003] In some work scenarios, there are unified requirements for user photos, such as unified requirements for employees' hairstyles, clothing, poses during photo taking, etc. However, when actually taking work photos, due to the diverse presentation of hairstyles, clothing, etc. in different photo studios during photo taking, it is impossible to achieve complete unity. If an artificial intelligence model is used to synthesize photos uniformly, it is often necessary to specifically train a dedicated synthesis model. Although the synthesis effect has been improved, this synthesis model is difficult to meet the diverse needs of users for different hairstyle styles and clothing styles, and will require a higher model training cost. Summary of the Invention
[0004] The present disclosure provides an image synthesis method, device, medium, and product.
[0005] According to a first aspect of the present disclosure, an image synthesis method is provided. The method specifically includes: obtaining a facial image and a standard hairstyle image; performing feature matching on the facial image and the standard hairstyle image and then splicing them to obtain a head splicing image; extracting the head edge features included in the head splicing image; using the head edge features as control conditions for image synthesis to control the synthesis of the facial image and the standard hairstyle image in the head splicing image, so as to obtain a target head splicing image with the hairstyle effect strengthened by the control conditions.
[0006] Based on the above content, it can be seen that after obtaining the facial image, instead of directly inputting the facial image into the model for image synthesis. Instead, the facial image is first subjected to feature matching with the standard hairstyle image and then spliced to obtain a head splicing image. Thus, the spliced head splicing image has a standard hairstyle. The head edge features are included in the head splicing image. In the subsequent image synthesis process, the head edge features can be used to further strengthen the hairstyle synthesis effect in the synthesized image. Moreover, it is not necessary for the model to make excessive adjustments to the hairstyle in the head splicing image, and the synthesized and beautified image can be directly output. It can be seen that the requirements for the image synthesis ability of the model are significantly reduced, and there is no need to specifically train the model. While improving the diverse image synthesis ability of the model, the model application cost is also reduced.
[0007] In addition, the hairstyle effect of the target head composite image is defined by a standard hairstyle image, so that the finally synthesized target head composite image has good consistency (that is, the composite images synthesized by many users using this solution have good hairstyle consistency. This is because the differences among different users are only facial images, and the same standard hairstyle image is used, avoiding the occurrence of pose differentiation).
[0008] According to at least one embodiment of the present disclosure, after feature matching of the facial image and the standard hairstyle image, a head composite image is obtained, including: extracting a first facial feature and a first hair edge feature in the facial image; extracting a second hair edge feature in the standard hairstyle image; after matching and aligning the first hair edge feature and the second hair edge feature, splicing to obtain a head composite image including the first facial feature and the second hair edge feature.
[0009] According to at least one embodiment of the present disclosure, before extracting the facial feature and the first hair edge feature in the facial image, it further includes: detecting the human face feature in the facial image; adjusting the angle of the facial image according to the detection result; performing portrait segmentation and background replacement on the facial image after angle adjustment to obtain a facial image including the head.
[0010] According to at least one embodiment of the present disclosure, extracting the head edge feature included in the head composite image includes: extracting a first facial feature and a third hair edge feature in the head composite image; constructing a head edge feature by using the first facial feature and the third hair edge feature.
[0011] According to at least one embodiment of the present disclosure, it further includes: obtaining a standard pose image; after performing pose matching between the head composite image and the standard pose image, splicing to obtain an overall composite image; extracting the pose feature included in the overall composite image; using the head edge feature and the pose feature as control conditions for image synthesis to control the synthesis of the overall composite image, and obtaining a target composite image with the hairstyle effect strengthened by the control conditions.
[0012] Based on the above, after obtaining the facial image, instead of directly inputting the facial image into the model for image synthesis, the facial image is first feature-matched and then spliced with the standard hairstyle image to obtain a head spliced image, so that the head spliced image after splicing has a standard hairstyle. Then, the head spliced image can be further spliced with the standard pose image to obtain an overall spliced image. The overall spliced image contains standard pose features, and the head spliced image contains head edge features. In the subsequent image synthesis process, the head edge features can be used to further enhance the hairstyle synthesis effect in the synthesized image. Moreover, it is not necessary for the model to make excessive adjustments to the hairstyle and pose in the overall spliced image, and the synthesized and beautified image can be directly output. It can be seen that the requirement for the image synthesis ability of the model is significantly reduced, and there is no need to conduct targeted training on the model. While improving the diversified image synthesis ability of the model, the application cost of the model is also reduced.
[0013] According to at least one embodiment of the present disclosure, splicing the head spliced image and the standard pose image after pose matching to obtain an overall spliced image includes: performing matching processing on the head spliced image according to the pose in the standard pose image; splicing the head spliced image after the matching processing with the neck in the standard pose image; determining whether the length and width of the face, the width of the shoulders, and the length of the neck after splicing meet the threshold; if they meet, generating an overall spliced image.
[0014] According to at least one embodiment of the present disclosure, extracting the pose features included in the overall spliced image includes: using a key point extraction tool to identify the second facial features and neck and shoulder features in the overall spliced image; extracting the pose features jointly formed by the second facial features and neck and shoulder features.
[0015] According to at least one embodiment of the present disclosure, using the head edge features and pose features as control conditions for image synthesis to control the synthesis of the overall spliced image, and obtaining a target synthesized image with the hairstyle effect enhanced by the control conditions includes: aligning the first facial features included in the head edge features with the second facial features included in the pose features to generate control conditions including the head edge features and pose features; inputting the overall spliced image into the model, and using the control conditions to enhance the key hairstyle information and key pose information in the synthesis process of the overall spliced image by the model, and synthesizing a target synthesized image with a standard hairstyle and a standard pose.
[0016] According to at least one embodiment of the present disclosure, after using the head edge features as control conditions for image synthesis to control the synthesis of the facial image and the standard hairstyle image in the head spliced image, and obtaining a target head spliced image with the hairstyle effect enhanced by the control conditions, it further includes: extracting the facial features in the facial image; performing feature fusion on the facial features and the target head spliced image to obtain a target synthesized image with enhanced facial features.
[0017] According to a second aspect of the present disclosure, there is provided an electronic device, including: a memory that stores execution instructions; and a processor that executes the execution instructions stored in the memory, so that the processor executes the method described in the first aspect of any one of the embodiments of the present disclosure.
[0018] According to a third aspect of the present disclosure, there is provided a readable storage medium storing execution instructions, and when the execution instructions are executed by a processor, they are used to implement the method described in the first aspect of any one of the embodiments of the present disclosure.
[0019] According to a fourth aspect of the present disclosure, there is provided a computer program product including a computer program, and when the computer program is executed by a processor, it implements the method described in the first aspect of any one of the embodiments of the present disclosure. Description of the Drawings
[0020] The drawings illustrate exemplary embodiments of the present disclosure and, together with the description thereof, are used to explain the principles of the present disclosure. These drawings are included to provide a further understanding of the present disclosure and are included in this specification and form a part of this specification.
[0021] Figure 1 It is a schematic flowchart of an image synthesis method provided by the present disclosure.
[0022] Figure 2 It is a schematic diagram of the process of generating a head spliced image provided by the present disclosure.
[0023] Figure 3 It is a schematic diagram of head edge feature extraction provided by the present disclosure.
[0024] Figure 4 It is a schematic flowchart of a method for splicing with a standard pose image provided by the present disclosure.
[0025] Figure 5 It is a schematic diagram of splicing a head spliced image and a standard pose image exemplified by the present disclosure.
[0026] Figure 6 It is a schematic diagram of the key point extraction effect exemplified by the present disclosure.
[0027] Figure 7 It is a schematic diagram of the image synthesis process exemplified by the present disclosure.
[0028] Figure 8 It is a schematic block diagram of the structure of an image synthesis device according to an embodiment of the present disclosure.
[0029] Figure 9Structural schematic block diagram of an electronic device according to an embodiment of the present disclosure. Detailed implementation manners
[0030] The present disclosure will be further described in detail below with reference to the accompanying drawings and examples. It can be understood that the specific examples described herein are only used to explain the relevant content and do not limit the present disclosure. In addition, it should be noted that only the parts related to the present disclosure are shown in the drawings for the convenience of description.
[0031] It should be noted that, without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other. The technical solutions of the present disclosure will be described in detail below with reference to the accompanying drawings and embodiments.
[0032] In some scenarios, the user's work photos are presented externally through websites or display boards. The first impression left by the work photos on the user is very important. When the user takes work photos, although they will try to take photos with a unified hairstyle as required, the final presentation effects are still uneven. For example, the hairstyles are not unified and the hair colors are not unified. In order to ensure complete uniformity, an artificial intelligence model is used to optimize the photos taken by the user. In order to ensure the consistency of the image optimization processing effect of the artificial intelligence model, targeted training needs to be carried out according to actual needs, so as to obtain a model that meets the requirements of the synthesis effect. Although the image synthesis effect of the model obtained through targeted training has been significantly improved, if the style or style type of the work photos generated by the user changes, the model will not be able to handle it. Therefore, there is an urgent need for a solution that can simply and efficiently meet the diverse image synthesis requirements.
[0033] For the convenience of description and to make the technical solutions of the specific implementation manners of the present disclosure easier to understand, before describing the image synthesis method implemented in the present disclosure, the technical terms involved in the specific implementation manners of the present disclosure are explained as follows.
[0034] Pose feature: The feature jointly composed of the key points obtained by feature extraction of the spliced image and the neck and shoulder key points, which is used to represent the pose of the person in the synthesized image and is also one of the pose control conditions in the image synthesis process.
[0035] Figure 1 Flow schematic diagram of an image synthesis method provided by the present disclosure. As Figure 1 shown, the method includes steps 101 to 106. Among them, the method can be executed by an electronic device such as a server (local server or cloud server).
[0036] Specifically, Figure 1 the method shown includes: Step 101: Obtain a facial image and a standard hairstyle image.
[0037] The facial image mentioned here refers to the original image obtained by photographing the user, which contains a clear facial image of the user. In order to reduce the complexity of subsequent splicing with the standard pose image, when photographing the facial image, the user should try to take the same pose as the standard pose image.
[0038] The standard hairstyle image mentioned here refers to an image that has a standard hairstyle (which can be a standard hairstyle for men, a standard hairstyle for women, or a specific hairstyle), and may contain a face. In this image, there is a clear hairstyle contour and a hairstyle color that meets the standard. The standard hairstyle image can be selected according to needs. For example, a standard shoulder-length hair, a standard short hair, a standard crew cut, etc. Different standard hairstyle images are selected according to the image synthesis requirements of different users.
[0039] In an alternative solution, the standard hairstyle image can have no face and neck and shoulders, only the hairstyle image. When splicing the standard hairstyle image without a face, the relevant features of the user's facial image can be completely retained.
[0040] Step 102: Splice the facial image and the standard hairstyle image after feature matching to obtain a head spliced image.
[0041] In practical applications, an image containing a clear face of the user can be taken by oneself or by a professional photographer. In order to make the facial image and the standard hairstyle image be spliced more accurately, the collected facial image can be simply processed first. For example, the background content in the facial image is eliminated, or an image with only facial content is segmented. Through the above preliminary processing, the obtained facial image and the standard hairstyle image are more accurate when splicing.
[0042] When splicing. If the standard hairstyle image only has a hairstyle contour and does not include a face and neck and shoulders, the standard hairstyle image can be directly covered on the facial image. When covering, ensure that the positions of the eyes, nose, and mouth match the standard hairstyle image.
[0043] In an alternative solution, in order to make the facial image and the standard hairstyle image match more accurately, a reference point or reference area for the eyes, nose, and mouth can be set for the standard hairstyle image (that is, the allowable matching area range for the eyes, the allowable matching area range for the nose, and the allowable matching area range for the mouth). When matching, the eyes, nose, and mouth in the facial image can be aligned with the reference point or reference area for the eyes, nose, and mouth in the standard hairstyle image. In this way, in the head spliced image obtained by splicing, only the face is the content provided by the user, and for other hairstyles, hair accessories, hair colors, etc., they are provided by the standard hairstyle image. Therefore, the user does not need to change the current hairstyle, hair color, etc. to make a qualified image. This reduces the burden on the user to generate an image that meets the requirements of the standard hairstyle.
[0044] Step 103: Extract the head edge features included in the head stitching image.
[0045] The head edge features mentioned here include the overall contour features formed by the hairstyle and the face, the hairline contour features, and the facial features (such as eye, nose, and mouth features). Further, a head edge control graph can be constructed using the head edge features.
[0046] In practical applications, the facial image and the standard hairstyle in the head stitching image are combined into a whole. In order to better enhance the hairstyle effect when the subsequent model performs image synthesis, the face and the hairstyle can be taken as a whole to extract the head edge features. The specific feature extraction process will be described in detail in the following embodiments and will not be repeated here.
[0047] Step 104: Use the head edge features as control conditions for image synthesis to control the synthesis of the facial image and the standard hairstyle image in the head stitching image, and obtain a target head stitching image with the hairstyle effect enhanced by the control conditions.
[0048] In the head stitching image, different users can use the same standard hairstyle image, and each user only needs to provide their own facial image. In this way, all contents other than the face in the finally obtained synthesized image can be synthesized according to a unified standard. In other words, the use of the standard hairstyle image enables the unification and standardization of diverse hairstyles, hair colors, and other features. And the faces, which are supposed to remain diverse, continue to be diverse.
[0049] During synthesis, a pre-trained model can be used for synthesis. However, it should be noted that when the image synthesis model is trained, the selected training samples are diverse. It is not necessary to specify a single hairstyle as the training sample. Many hairstyles can be selected as training samples (that is, no targeted training is required), so that the trained image synthesis model has a good general effect and can meet various image synthesis requirements. To avoid the weakening of features such as the hairstyle during the synthesis process, the head edge features can be input into the synthesis model as control conditions, so as to obtain a target head stitching image with the hairstyle effect enhanced.
[0050] With the above solution, after obtaining the facial image, instead of directly inputting the facial image into the model for image synthesis, the facial image is first spliced with an image having a standard hairstyle to obtain a head spliced image. When performing subsequent image synthesis, there is no need for the model to optimize and adjust the hairstyle, and the synthesized and beautified image can be directly output. It can be seen that the requirement for the image synthesis ability of the model is significantly reduced, which means that there is no need to perform targeted training on the model. While improving the diversified image synthesis ability of the model, the training cost of model application is also reduced.
[0051] In addition, in some cases, the synthesis state of the synthesized image is defined by a standard pose image, so that the poses of the finally synthesized images are highly consistent.
[0052] In one or more embodiments of the present disclosure, splicing the facial image and the standard hairstyle image after feature matching to obtain a head spliced image includes: extracting a first facial feature and a first hair edge feature in the facial image; extracting a second hair edge feature in the standard hairstyle image; after matching and aligning the first hair edge feature and the second hair edge feature, splicing to obtain a head spliced image including the first facial feature and the second hair edge feature.
[0053] In practical applications, in order to accurately splice the facial image and the standard hairstyle image, feature extraction is required. Specifically, the first facial feature included in the facial image is extracted, that is, the eye, nose, mouth and facial contour features of the user and the first hair edge feature are extracted. At the same time, the second hair edge feature of the standard hairstyle image also needs to be extracted.
[0054] Next, the first hair edge feature is matched and aligned with the second hair edge feature. Specifically, a first center point of the first hair contour is determined based on the first hair edge feature, and a second center point of the second hair contour is determined based on the second hair edge feature. Furthermore, the first center point is matched and aligned with the second center point. When aligning, it is ensured that there is no angular deviation between the facial image and the standard hairstyle image. If there is a deviation, the facial image needs to be adjusted in angle.
[0055] It should be noted that in order to ensure the splicing effect, it is necessary to further determine whether the position of the first hair edge feature exceeds the coverage range of the standard hairstyle image. If it exceeds the coverage range, it may cause the standard hairstyle image to not completely cover the hair in the user's facial image during image splicing. At this time, the part of the first hair edge that exceeds the coverage range can be eliminated first, or the part that exceeds the coverage range in the head spliced image can be eliminated after splicing.
[0056] In an alternative solution, the method for determining whether the position of the first hair edge feature exceeds the coverage range of the standard hairstyle image includes: comparing the coverage ranges of the first hair edge feature and the second hair edge feature. If the coverage range of the second hair edge feature is greater than that of the first hair edge feature, splicing can be directly performed. Conversely, if any part of the coverage range of the first hair edge feature (for example, the top hair or the forehead hair) exceeds the coverage range of the second hair edge feature, further elimination processing of the first hair edge is required.
[0057] Such as Figure 2 is a schematic diagram of the process for generating the head splicing image provided by the present disclosure. As can be seen from Figure 2 it, a facial image is captured. After processing the facial image, only the user's face and hair remain. Next, the first facial features (that is, the eye, nose, mouth, and facial contour features) and the first hair edge feature are extracted, and the first center point A is found. The second hair edge feature is extracted, and the second center point B is found. Then, the first center point A and the second center point B are aligned. In this way, the completed head splicing image is obtained. When splicing, if the size of the facial image does not match the size of the standard hairstyle image, the facial image needs to be scaled and adjusted.
[0058] Based on the above solution, the hairstyle in the facial image provided by the user is directly replaced with the standard hairstyle. The user only needs to provide the facial-related image content. When the user captures the facial image, there is no need to make any treatment to their own hairstyle, as long as it is ensured that the face is not blocked when capturing the image. The obtained head splicing image has complete facial features and standard hairstyle features, serving as the basis for subsequent image synthesis.
[0059] In one or more embodiments of the present disclosure, before extracting the facial features and the first hair edge feature in the facial image, it further includes: detecting the human face features in the facial image; adjusting the angle of the facial image according to the detection result; performing human portrait segmentation and background replacement on the facial image after angle adjustment, obtaining the facial image including the head.
[0060] In practical applications, in order to facilitate subsequent effective processing of the facial image, it is necessary to first perform simple preliminary processing on the facial image provided by the user to eliminate the irrelevant content therein. Specifically, first detect the human face features in the facial image, that is, the eye, nose, and mouth features. Determine whether there is an angle deviation in the current face according to the eye, nose, and mouth. If there is, perform angle adjustment. After completing the angle adjustment, further segment the human portrait from the image. At the same time, replace the background of the human portrait with white or other specified solid colors, obtaining the facial image only including the user's head content, facilitating subsequent actual splicing for human portrait recognition.
[0061] Based on the above solution, the facial image is preliminarily optimized to effectively eliminate interference factors such as the background content of the original image provided by the user, facilitating more accurate subsequent splicing.
[0062] In one or more embodiments of the present disclosure, extracting the head edge features included in the head splicing image includes: extracting the first facial feature and the third hair edge feature in the head splicing image; constructing the head edge features using the first facial feature and the third hair edge feature.
[0063] When performing feature extraction, not only the first facial feature of the head splicing image is extracted, but also the third hair edge feature after synthesis is extracted. These two parts are jointly constructed to obtain the head edge features. The head edge features extracted contain the complete features of the user's head, which include hairstyle-related features, the relative position relationship between the face and the hairstyle, and facial-related features.
[0064] When performing feature extraction, a model can be used for extraction. For example, the HairFastGAN model first performs feature extraction on the user's portrait image. The purpose of this step is to identify the key feature points of the user's head, such as the facial contour, hair edge, etc. At the same time, the model also extracts the corresponding hairstyle edge features from the standard hairstyle image, such as the edge contour of the hairstyle, hair color, texture, etc.
[0065] Such as Figure 3 is a schematic diagram of the head edge feature extraction provided by the present disclosure. From Figure 3 it can be seen that the spliced head splicing image contains complete feature content, and the head edge features that need to be used as control conditions for subsequent synthesis can be extracted therefrom, including eyes, nose, mouth, facial edge contour, hair edge contour. In some cases, it may also include hair texture features.
[0066] Through the above solution, the extracted head edge features contain facial features and hairstyle contour-related features, which are used as control conditions for strengthening hairstyle features in subsequent image synthesis.
[0067] Such as Figure 4 is a schematic diagram of the method flow for splicing with the standard pose image provided by the present disclosure. From Figure 4 it can be seen that it specifically includes the following steps. Step 401: Obtain the standard pose image. Step 402: Perform pose matching on the head splicing image and the standard pose image and then splice them to obtain the overall spliced image. Step 403: Extract the pose features included in the overall spliced image. Step 404: Use the head edge features and pose features as control conditions for controlling the synthesis of the overall spliced image to obtain the target synthesized image with the hairstyle effect strengthened by the control conditions.
[0068] The standard pose image mentioned here refers to an image with at least a standard neck and shoulder pose, or a full-body pose image. This standard pose image serves as the basis for pose synthesis during subsequent image synthesis processes and is also the basis for facial image optimization during image stitching. The pose in the standard pose image can be a front pose of the upper body or a side pose of the upper body. According to the user's type image synthesis requirements, the standard pose image can be selected by the user according to their own needs. For example, the user can choose a standard pose image of a person wearing different types of work uniforms, or a standard pose image with a front or side pose according to their needs.
[0069] In an alternative solution, the standard pose image can be an upper body front pose image without a head, only with the neck and shoulders, and the person in the standard pose image is wearing compliant clothing. In some scenarios, various decorations can also be added to the clothing, such as brooches, bows, ties, etc. It only needs to be added to the standard pose image, and it is not necessary for the user taking the photo to wear or wear them, which can effectively improve the image generation efficiency. Reduce the time cost and economic cost of the user taking photos.
[0070] The head stitched image obtained after stitching is an image with a standard hairstyle, and the facial image in this head stitched image is the real image of the user. When stitching, it is necessary to ensure that the pose of the facial image is consistent with the pose of the standard pose image, and then stitch the head stitched image with the neck in the standard pose image.
[0071] For example, if the neck and shoulder pose in the standard pose image is a front pose, the pose presented by the head stitched image also needs to be a front pose, that is, taking a photo facing the camera. The facial pose can be adjusted accordingly according to needs during stitching. If the neck and shoulder pose in the standard pose image is a side pose, then the head stitched image will also be adjusted to a side pose during stitching. The head stitched image can be scaled and rotated during the stitching process to meet the stitching requirements. Of course, if the requirements for pose unity still cannot be met after processing such as scaling and rotation, the user can be notified to re-take the facial image and tell the user how to adjust the facial pose when taking the photo, or re-stitch the head stitched image.
[0072] It can be seen from this that during stitching, the standardization of key information that needs to be unified but is uncontrollable (such as clothing type, color, neck and shoulder pose, etc.) is achieved through the standard pose image. It only requires the user to provide a photo with a facial image. When using the model for image synthesis, it is not necessary to make major changes to the model or conduct a large amount of targeted training on the model. As long as it is optimized on the basis of keeping the key information unchanged, different image synthesis requirements can be met.
[0073] After generating the overall stitched image in the manner described above, pose features can be further extracted from the overall stitched image.
[0074] The pose features mentioned here can be understood as the neck-shoulder features and facial features respectively extracted from the standard pose image and the head stitched image. Since the relative position relationship of the necks in the head stitched image and the standard pose image (including scaling size, rotation angle, etc.) has been adjusted when stitching the head stitched image and the standard pose image, when extracting pose features, the neck-shoulder features and facial features (which can also be facial feature features) are extracted from the overall stitched image at the same time.
[0075] In an alternative solution, the head edge features and pose features can be placed in different heatmaps respectively, but the facial features for alignment need to be present in both of these heatmaps, so as to ensure good consistency in the facial key point coordinates of the head edge features and pose features.
[0076] Of course, if the standard pose image is in a standing pose or contains other poses with upper limbs, when extracting pose features, other features in addition to the neck-shoulder features need to be further extracted, such as upper limb features, torso features, lower limb features, and so on.
[0077] When using the model for image synthesis, the overall stitched image is input into the model. The model can directly use the overall stitched image for synthesis. The head edge features and pose features are input into the model as control conditions of the model, so that the hairstyle contour and pose of the synthesized image can be further enhanced during the process of the model synthesizing the image.
[0078] After two image stitchings (head image stitching and overall image stitching), various key information required for image synthesis is already included in the head edge features, pose features, and the overall stitched image. When the model uses the head edge features, pose features, and the overall stitched image for image synthesis, not too many modifications are needed. Only the hairstyle and face in the original overall stitched image, and the connection between the head and the neck need to be optimized on the basis of ensuring that the key information remains unchanged. Moreover, this model does not need to be specifically trained, reducing the model usage cost. When the user's synthesis requirements change, only the required standard hairstyle image needs to be replaced, which can well meet the diverse image synthesis requirements.
[0079] In one or more embodiments of the present disclosure, an overall stitched image is obtained by performing pose matching between a head stitched image and a standard pose image and then stitching them, including: performing matching processing on the head stitched image according to the pose in the standard pose image; stitching the head stitched image after the matching processing with the neck in the standard pose image; determining whether the length and width of the face, the width of the shoulders, and the length of the neck after stitching meet the thresholds; if they meet, generating an overall stitched image.
[0080] In practical applications, when performing image stitching, the face in the head stitched image is stitched with the neck in the standard pose image. Since the head stitched image and the standard pose image are images obtained by different methods, in order to ensure that the stitched image is more realistic during stitching, the head stitched image needs to be adjusted according to the standard pose image. This includes performing scaling adjustment, rotation adjustment, etc. on the head stitched image.
[0081] Such as Figure 5 is a schematic diagram of stitching the head stitched image and the standard pose image illustrated in the present disclosure. As can be seen from Figure 5 it, the standard pose image has a neck and an upper body image wearing specified clothing, and does not contain a head image.
[0082] During stitching, the head is connected to the neck, and the head stitched image covers the neck image. In this way, the length of the neck can be changed by moving the position of the head stitched image. For example, when it is necessary to increase the neck length of the stitched image, the head stitched image can be moved upward, and when it is necessary to decrease the neck length in the stitched image, the head stitched image can be moved downward.
[0083] In addition, since the sizes of the head stitched images provided by different users are different, resulting in inconsistent relative size relationships between the head stitched images and the standard pose images, therefore, it is necessary to perform scaling adjustment on the head stitched images. When scaling the head stitched images, the length and width need to be scaled proportionally to avoid facial deformation during the scaling process. In an alternative solution, during the stitching process, the length and width of the face, the length of the neck, and the width of the shoulders can be adjusted so that the adjusted ratio of the face to the neck and shoulders is more coordinated and more in line with the conventional ratio of real people. The ratio adjustment process will be described in detail in subsequent embodiments and will not be repeated here.
[0084] After stitching is completed, the stitched image is taken as a whole, and the relative positions of the eyes, nose, mouth and the neck and shoulders are fixed. When performing feature extraction subsequently, the extraction is performed according to the adjusted relative positions.
[0085] In one or more embodiments of the present disclosure, determining whether the length and width of the spliced face, the shoulder width, and the neck length meet the thresholds includes: determining whether the first ratio of the face width to the shoulder width meets the first threshold; if not, scaling and adjusting the length and width of the face; if so, not scaling and adjusting the length and width of the face. Determining whether the second ratio of the face length to the neck length meets the second threshold; if not, adjusting the position of the face.
[0086] Generally speaking, the width ratio of a person's head to shoulders and the ratio of face length to neck length are within a certain range. When determining whether the display ratio of the head spliced image to the standard pose image in the spliced image is appropriate, the first ratio of the face width to the shoulder width and the second ratio of the face length to the neck length can be calculated respectively.
[0087] More specifically, calculate the first ratio of the face width to the shoulder width, and then determine the size relationship between the first ratio and the first threshold. This first threshold can be a fixed value (such as 0.5) or an interval range value (such as 0.3 to 0.6). When calculating, if it exceeds the interval range specified by the first threshold, the aspect ratio of the head spliced image will be adjusted, thereby changing the width of the head spliced image so that the adjusted first ratio meets the first threshold.
[0088] At the same time, calculate the second ratio of the face length to the neck length, and then determine the size relationship between the second ratio and the second threshold. This second threshold can be a fixed value or an interval range value. When calculating, if the calculation result does not meet the requirements of the second threshold, the position of the head spliced image will be adjusted, thereby changing the neck length so that the adjusted second ratio meets the second threshold.
[0089] It should be noted that when adjusting the length and width of the head spliced image, both the first ratio and the second ratio need to be considered. That is to say, when adjusting, avoid the second ratio not meeting the second threshold due to adjusting the first ratio, or avoid the first ratio not meeting the first threshold due to adjusting the second ratio.
[0090] Generally speaking, since the size and dimensions of the standard pose image are the dimensions required for the finally generated image, when performing ratio adjustment, it is the size of the head spliced image that is adjusted, and the size of the standard pose image is not adjusted.
[0091] To better understand the above solution, the following will be elaborated through a specific embodiment. Assume that the first threshold is 0.25 - 0.4 and the second threshold is 0.1 - 0.25. That is, 0.25 < face / shoulder width < 0.4, 0.1 < neck length (from the chin at the bottom of the face to the bottom of the neck) / face length < 0.25. Only when both the first threshold and the second threshold are met, the spliced image meets the proportional requirements of the actual human face and neck-shoulder. In other words, the synthesized image obtained in this way is more realistic.
[0092] Based on the above embodiment, the first ratio of the face width to the shoulder width is adjusted. At the same time, the second comparison of the face length to the neck length is adjusted. So that the pose features in the adjusted synthesized image are more in line with the actual user pose requirements. The spliced image adjusted according to the ratio can show a more reasonable pose, providing a basis for subsequent feature extraction (including facial contour features and pose features), so that the image finally synthesized by the model can meet the requirements.
[0093] In one or more embodiments of the present disclosure, the pose features included in the overall spliced image are extracted, including: using a key point extraction tool to identify the second facial features and neck-shoulder features in the overall spliced image; extracting the pose features jointly constituted by the second facial features and neck-shoulder features.
[0094] The key point extraction tool mentioned here can be control_v11p_sd15_openpose, or extraction tools such as FaceRecognition and Animetrics Face Recognition. It should be noted that OpenPose of ControlNet is a tool that uses computer vision technology to detect human poses and facial key points (that is, the second facial features). And the detection results are used as conditional inputs to accurately control the pose, actions and facial expressions of the people in the images generated by Stable Diffusion, improving the controllability and realism of the generated images.
[0095] The facial key points mentioned here can be understood as the eyes, nose and mouth in the overall spliced image. The neck-shoulder key points mentioned here can be understood as the points on the neck and shoulders collected from the overall spliced image. The pose features are jointly constituted by the neck-shoulder key points and the facial key points. The pose features mentioned here are used to represent the pose of the person that the final synthesized image needs to present. This pose feature is determined by the spliced image and is input as a control condition into the synthesis model during the image synthesis process.
[0096] The second facial features mentioned here can be understood as the key points including eyes, nose and mouth.
[0097] Such as Figure 6Schematic diagram of the key point extraction effect illustrated in the present disclosure. Figure 6 As can be seen in the figure, the overall spliced image includes the facial image provided by the user, the standard hairstyle image, and the standard posture image. When extracting features, the head edge features and the posture features including the facial key points and the neck and shoulder key points are extracted.
[0098] In the above manner, the stitched image obtained after stitching is extracted as a whole, and the influence of the user posture of the head stitched image can be eliminated. In other words, users who have image synthesis needs (for example, users) only need to consider whether the head stitched image matches the standard posture image (the matching here means that the shooting posture of the head stitched image and the neck and shoulder posture in the standard posture image are known. For example, if the neck and shoulder posture in the standard posture image is a straight posture, then the head stitched image also needs to be shot in a straight posture, and the side face cannot be shot). There is no need to consider the user's clothing and neck and shoulder posture when taking pictures, and the interference of diverse and uncontrollable human factors in image synthesis is eliminated as much as possible.
[0099] In one or more embodiments of the present disclosure, facial key points and neck and shoulder key points are extracted to jointly constitute posture features, including: determining the coordinate values corresponding to the facial key points and the neck and shoulder key points; constructing a heat map for representing the posture state based on the coordinate values; and extracting posture features from the heat map using controlNet.
[0100] In practical applications, the preprocessed spliced image is input into the controlNet (control_v11p_sd15_openpose) model, and the model will automatically detect and extract key points related to human posture. These posture-related key points usually include neck and shoulder key points, joint points (such as shoulders, elbows, knees, etc.) and other important feature points.
[0101] Similarly, the model can also extract facial key points. These key points usually include the location information of facial features such as eyes, nose, mouth, etc. When extracting facial key points, the model may use a facial detection algorithm to locate the facial area and extract key points within that area.
[0102] The specific implementation process is as follows: Get the key point coordinates. The key point coordinates are detected from the image through a specific algorithm or model. These key points usually represent important features of the human body, such as facial features (eyes, nose, mouth, etc.), shoulders, elbows, wrists, hips, knees and ankles. In order to obtain these key point coordinates, deep learning models such as convolutional neural networks (CNN) or more advanced model architectures (such as ResNet, Hourglass, etc.) are usually used. These models are trained to accurately identify key points in images.
[0103] During the training phase, the model uses a large dataset of annotated key points for learning, gradually optimizing its internal parameters to improve the accuracy of key point detection. Once the model is trained, it can perform key point detection on new input images and output the coordinate information of the key points.
[0104] After obtaining the key point coordinates, the next step is to convert these coordinates into a Pose graph. A Pose graph is a special form of image representation where key points are marked on the image in a specific way (such as dots, lines, heatmaps, etc.). The conversion process usually involves the following steps: First, the key point coordinates need to be mapped from the coordinate space output by the model to the image space. This usually involves some geometric transformations, such as translation, scaling, and rotation, to ensure that the positions of the key points on the Pose graph are consistent with their positions in the original image.
[0105] Then, mark the key points on the Pose graph according to the mapped coordinates. This can be achieved by drawing dots, lines, or other shapes at the corresponding positions. The color, size, and shape of the markings can be adjusted as needed to represent the key points more clearly.
[0106] To represent the connection relationships between key points (such as bone or joint connections), connection information such as lines or arrows can be added to the Pose graph. This connection information helps to better understand the human pose and movement.
[0107] After the Pose graph is generated, it can be input into the ControlNet framework for further processing. The ControlNet framework is a control framework for image generation and editing. It can recognize and understand the key point information in the Pose graph and construct the human pose accordingly. Specifically, the ControlNet framework first identifies the key point information in the Pose graph. This usually involves some image processing techniques, such as image segmentation, feature extraction, and matching. Through these techniques, the framework can accurately identify each key point in the Pose graph and obtain its coordinate information.
[0108] After identifying the key point information, the ControlNet framework uses this information to construct the human pose. This usually involves some geometric and kinematic algorithms, such as bone models and joint angle calculations. Through these algorithms, the framework can infer the human pose and movement and generate corresponding control conditions accordingly.
[0109] In summary, the ControlNet framework generates a series of control conditions based on the constructed human pose. These control conditions will be used to guide the subsequent image generation or editing process to ensure that the synthesized image is consistent with the original Pose graph in terms of pose and movement.
[0110] In one or more embodiments of the present disclosure, the head edge feature and the pose feature are used as control conditions for image synthesis to control the synthesis of the facial image and the standard hairstyle image in the overall spliced image, and a target synthesized image with a hairstyle effect enhanced by the control conditions is obtained, including: after aligning the first facial feature included in the head edge feature with the second facial feature included in the pose feature, generating a control condition including the head edge feature and the pose feature; inputting the overall spliced image into a model, and using the control condition to enhance the hairstyle key information and pose key information (i.e., the head edge feature and the pose feature that enhance the pose effect) during the synthesis process of the overall spliced image by the model to synthesize a target synthesized image with a standard hairstyle and a standard pose.
[0111] In practical applications, after obtaining the head edge feature and the pose feature, the first facial feature in the head edge feature and the second facial feature in the pose feature are used for alignment. Specifically, the eyes, nose, and mouth in the first facial feature are aligned with the eyes, nose, and mouth in the second facial feature. Of course, the head edge feature and the pose feature can also be integrated into an overall feature.
[0112] Such as Figure 7 is a schematic diagram of the image synthesis process illustrated in the present disclosure. As can be seen from Figure 7 After obtaining the head spliced image by splicing, the HairFastGAN model is used to extract the head edge feature and the pose feature from the head spliced image. These two features carry the facial features (i.e., eyes, nose, and mouth) for alignment. After the alignment process, they are input into the synthesis model (e.g., IP-adapter) as control conditions. The synthesis model uses the overall spliced image and the control conditions to synthesize a target synthesized image with a standard hairstyle and a standard pose.
[0113] In an alternative solution, when performing image synthesis, it is also necessary to clarify the image synthesis style, which can be achieved through the LoRA model. According to the required image style, appropriate LoRA parameter sets are selected. These parameter sets can be pre-trained or fine-tuned according to a specific style. Thus, the image synthesis model can generate a synthesized image that meets the style required by the user.
[0114] If the existing LoRA parameter sets do not fully meet the style required by the user, these parameters can be fine-tuned using a style-related dataset. The fine-tuning process can be implemented through standard machine learning or deep learning frameworks.
[0115] The overall stitched image is input into the trained image synthesis model. At the same time, the extracted head edge features and pose features are integrated into the image synthesis model as control conditions. Usually, the head edge features and pose features are introduced as additional inputs during the forward propagation of the image synthesis model. During the synthesis process, the influence of the head edge features and pose features is strengthened by adjusting the loss function of the image synthesis model or introducing additional regularization terms. This helps to ensure that the synthesized image accurately reflects the head edge features and pose features of the original image captured by the user while maintaining the target style.
[0116] In summary, the control conditions in the ControlNet framework are a series of instructions or parameters generated based on the constructed human body poses. These parameters can accurately describe the human body's posture, movements, expressions, and other detailed information, thus ensuring that the generated image is consistent with the original template or user expectations. In practical applications, these control conditions can be adjusted and optimized according to specific requirements to meet the image synthesis needs in different scenarios.
[0117] In one or more embodiments of the present disclosure, after using the head edge features as control conditions for the image synthesis to control the synthesis of the facial image and the standard hairstyle image in the head stitched image and obtaining the target head stitched image with the hairstyle effect strengthened by the control conditions, it further includes: extracting the facial features in the facial image; performing feature fusion on the facial features and the target head stitched image to obtain the target synthesized image with enhanced facial features.
[0118] In practical applications, to achieve high-quality fusion of the facial features and the synthesized standard hairstyle image, a feature fusion model is needed. The feature fusion model uses a multi-scale attribute encoder to extract the attribute features of the synthesized image, uses a pre-trained face recognition model to extract the ID features of the facial image, and then, by introducing a viable variable feature fusion structure, embeds the ID features into the attribute feature space while realizing the adaptive change of the face in the form of an optical flow field. Finally, the fusion result is real, high-fidelity, and supports the adaptive perception of the target user's face shape to a certain extent.
[0119] To obtain the ID features of the facial image, the model uses a pre-trained face recognition model (i.e., a face ID extractor). This model has been trained with a large amount of real face data and can accurately identify and extract the unique features of the face, namely the ID features, which are crucial for distinguishing different individuals.
[0120] To achieve the effective fusion of ID features and attribute features, the feature fusion model introduces a feasible variable feature fusion structure. This structure can not only embed ID features into the attribute feature space to achieve deep fusion at the feature level, but also dynamically adjust the fusion strategy according to the specific content of the input image through an adaptive learning mechanism.
[0121] It should be noted that the feature fusion model realizes the adaptive change of the face in the form of an optical flow field. The optical flow field can capture the pixel-level change information in the image. By applying it to the face image, the model can simulate the subtle dynamic changes of the face, such as expressions and postures, making the fusion result more realistic and natural. Specifically, the optical flow field constructs a complete motion field by calculating the velocity vectors of each pixel point in the image. This motion field can accurately reflect the motion direction and speed of each pixel point in the face image. When this motion field is applied to the face image, the model can simulate the subtle dynamic changes of the face, such as expressions and postures, according to the information of the optical flow field. This simulation process is not just a simple pixel position adjustment, but a high-level processing of the face image based on the motion information of the optical flow field. Through this processing, the model can generate a more realistic and natural fusion result. Because the optical flow field captures the dynamic change information in the face image, the generated fusion image will be more natural and smooth in terms of expressions and postures.
[0122] In addition, the application of the optical flow field also endows the feature fusion model with a certain degree of robustness. Since the optical flow field can capture the motion information in the image, when the input image has a certain degree of motion blur or noise, the model can still accurately simulate the facial dynamics according to the information of the optical flow field. This greatly improves the applicability and stability of the model.
[0123] After the above processing process, the fusion result output by the feature fusion model not only has high fidelity and can retain the details and texture of the original image, but also can adaptively perceive the face shape of the target user to a certain extent. This means that the model can automatically adjust the fusion strategy according to the facial features of the target user, making the finally generated image more in line with the expectations and needs of the user.
[0124] In summary, through the combination of a multi-scale attribute encoder and a pre-trained face recognition model, as well as the application of a feasible variable feature fusion structure and an optical flow field, a more realistic and natural target head stitching image can be generated (of course, it can also be used to enhance the facial features of the target synthetic image). The feature fusion model can automatically adjust the fusion strategy according to the specific content of the input image and the facial features of the target user, realizing the adaptive perception of the face shape of the target user. This feature fusion method improves the image fusion quality and enhances the adaptive ability of the feature fusion model.
[0125] For ease of understanding, the image synthesis process will be described below through specific embodiments.
[0126] First, obtain the facial image provided by the user and perform portrait segmentation on the facial image.
[0127] Specifically, use a deep learning model (such as MTCNN or RetinaFace) to detect the frontal face image uploaded by the user and obtain the coordinate box of the face area. According to the detected facial key points (such as eyes, nose, mouth, etc.), calculate the rotation angle and rotate the image to a frontal face pose. Use a portrait segmentation model (such as U-Net or DeepLabV3, which can be pre-trained in advance using corresponding training samples before use) to segment the face area from the background. The segmented image only retains the face part, and the background is transparent or white. If necessary, beauty treatments such as whitening and skin smoothing can also be performed on the face part to improve the image quality. A GAN-based beauty model (such as BeautyGAN) can be used for automated processing.
[0128] Next, use a standard hairstyle image to change the standard hairstyle of the target image. Specifically, use the open-source HairFastGAN model to apply the standard hairstyle to the user's portrait image. Since the portrait image has been segmented and the background is white, the model can process the hairstyle replacement more accurately. According to the user's face shape and hairstyle requirements, fine-tune the generated avatar to ensure that the hairstyle matches the face shape. A hairstyle adjustment algorithm based on key points can be used to further optimize the fit between the hairstyle and the face shape.
[0129] In the scenario of generating standard work photos, it is also necessary to synthesize the characters in the image with standard postures. Therefore, it is also necessary to splice the head spliced image and the posture image. Specifically, splice the head spliced image of the user after changing the hairstyle with the standard posture image, extract the key points of the face and body, and generate the overall spliced image. The overall spliced image is used to ensure the stability of the posture of the generated image.
[0130] In addition, extract the head edge features from the head spliced image of the user after changing the hairstyle to ensure that the hairstyle contour of the generated image is consistent with the requirements, and the key features such as the facial contour, eyes, nose, and mouth are consistent with the user's upload Figure 1 Align the head edge features with the key points of the posture features (that is, eyes, nose, and mouth) to ensure that the two sets of features are consistent during the generation process.
[0131] Furthermore, the aligned head edge features and posture features can be used as control conditions to be input into the ip-adapter model for image synthesis. The ip-adapter model combines the control conditions to generate images with similarity and normativity, that is, synthetic images that are similar to the user's face and meet the hairstyle standard specifications and posture standard specifications.
[0132] In some cases, in order to further improve the facial similarity of the synthesized image, facial features can be extracted again from the facial image. On the basis of generating the target synthesized image, the facial image of the user obtained previously is fused, and facial features are further extracted from the facial image. A feature point-based fusion algorithm can be used to ensure that the details of the fused image are more realistic. Post-processing operations such as smoothing and edge enhancement are performed on the fused image to improve the image quality. Image enhancement techniques (such as super-resolution reconstruction) can be used to further improve the clarity and detail performance of the image.
[0133] Based on any of the above embodiments, the present disclosure also provides an image synthesis device. Figure 8 It is a schematic block diagram of the structure of the image synthesis device according to an embodiment of the present disclosure. As Figure 8 shown, the image synthesis device includes: an acquisition module 81 for acquiring a facial image and a standard hairstyle image. A splicing module 82 for splicing the facial image and the standard hairstyle image after feature matching to obtain a head splicing image. An extraction module 83 for extracting the head edge features included in the head splicing image. A synthesis module 84 for using the head edge features as control conditions for image synthesis to control the synthesis of the facial image and the standard hairstyle image in the head splicing image, and obtaining a target head splicing image with the hairstyle effect enhanced by the control conditions.
[0134] The splicing module 82 is configured to extract the first facial features and the first hair edge features in the facial image; extract the second hair edge features in the standard hairstyle image; after matching and aligning the first hair edge features and the second hair edge features, splice them to obtain a head splicing image including the first facial features and the second hair edge features.
[0135] The splicing module 82 is configured to detect the facial features in the facial image; adjust the angle of the facial image according to the detection result; perform portrait segmentation and background replacement on the facial image after angle adjustment to obtain a facial image including the head.
[0136] The extraction module 83 is configured to extract the first facial features and the third hair edge features in the head splicing image; construct head edge features by using the first facial features and the third hair edge features.
[0137] Optionally, the acquisition module 81 is configured to acquire a standard pose image. The splicing module 82 is configured to splice the head splicing image and the standard pose image after pose matching to obtain an overall splicing image. The extraction module 83 is configured to extract the pose features included in the overall splicing image. The synthesis module 84 is configured to use the head edge features and the pose features as control conditions for image synthesis to control the synthesis of the overall splicing image, and obtain a target synthesized image with the hairstyle effect enhanced by the control conditions.
[0138] The splicing module 82 is used to perform matching processing on the head splicing image according to the pose in the standard pose image; splice the head splicing image after the matching processing with the neck in the standard pose image; judge whether the length and width of the face, the width of the shoulders, and the length of the neck after splicing meet the threshold; if they meet, generate an overall splicing image.
[0139] The extraction module 83 is used to use a key point extraction tool to identify the second facial feature and the neck and shoulder feature in the overall splicing image; extract the pose feature jointly composed of the second facial feature and the neck and shoulder feature.
[0140] The synthesis module 84 is used to align the first facial feature included in the head edge feature with the second facial feature included in the pose feature, and then generate a control condition including the head edge feature and the pose feature generation control condition; input the overall splicing image into the model, and use the control condition to strengthen the hairstyle key information and the pose key information in the process of the model synthesizing the overall splicing image, and synthesize to obtain a target synthesis image with a standard hairstyle and a standard pose.
[0141] The synthesis module 84 is also used to extract the facial features of the facial image; perform feature fusion on the facial features and the target head splicing image to obtain a target synthesis image with enhanced facial features.
[0142] For the implementation process of the functions and roles of each module in the above device, please refer to the implementation process of the corresponding steps in the above method for details, which will not be elaborated here.
[0143] The execution subject of the image synthesis method in the specific implementation manner of the present disclosure may be an electronic device such as a server (including a local server or a cloud server).
[0144] Therefore, based on any one of the above embodiments, the present disclosure further provides an electronic device, and this electronic device can execute the image synthesis method of any one of the above embodiments described in the present disclosure.
[0145] Figure 9 It is a structural schematic block diagram of an electronic device according to an embodiment of the present disclosure.
[0146] The hardware structure of the electronic device 1000 can be implemented by using a bus architecture. The bus architecture can include any number of interconnected buses and bridges, depending on the specific application and overall design constraints of the hardware. The bus 1100 connects various circuits including one or more processors 1200, a memory 1300, and / or hardware modules together. The bus 1100 can also connect various other circuits 1400 such as peripheral devices, voltage regulators, power management circuits, external antennas, etc.
[0147] The bus 1100 can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Component (EISA) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only one connection line is used in this figure, but it does not mean that there is only one bus or one type of bus.
[0148] The present disclosure also provides a readable storage medium having a computer program stored therein, and the computer program, when executed by a processor, is used to implement the above method. The "readable storage medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. More specific examples of the readable storage medium include the following: an electrical connection part (electronic device) having one or more wirings, a portable computer disk cartridge (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable read-only memory (CDROM), etc.
[0149] The present disclosure also provides a computer program product. The method of the present disclosure can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed, the processes or functions of the present disclosure are executed in whole or in part.
[0150] The computer program or instructions can be stored in a readable storage medium, or transmitted from one readable storage medium to another readable storage medium. For example, the computer program or instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired or wireless manner. The readable storage medium can be any available medium that can be accessed, or a data storage device such as a server or data center integrating one or more available mediums. The available medium can be a magnetic medium, such as a floppy disk, a hard disk, or a magnetic tape; it can also be an optical medium, such as a digital video disc; or it can be a semiconductor medium, such as a solid state drive. The computer-readable storage medium can be a volatile or non-volatile storage medium, or can include both volatile and non-volatile types of storage media.
[0151] Those skilled in the art should understand that the embodiments of the present disclosure can be provided as a method, a system, or a computer program product. Therefore, the present disclosure can take the form of an all-hardware embodiment, an all-software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present disclosure can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0152] The present disclosure is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the present disclosure. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing method devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing method devices generate means for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or combinations of blocks.
[0153] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing method device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or combinations of blocks.
[0154] These computer program instructions can also be loaded onto a computer or other programmable data processing method device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or combinations of blocks.
[0155] In the description of this specification, the descriptions with reference to the terms "one embodiment / way", "some embodiments / ways", "example", "specific example", or "some examples", etc., mean that the specific features, structures, or characteristics described in connection with the embodiment / way or example are included in at least one embodiment / way or example of the present disclosure. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment / way or example. Moreover, the specific features, structures, or characteristics described can be combined in a suitable manner in any one or more embodiments / ways or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments / ways or examples described in this specification and the features of different embodiments / ways or examples.
[0156] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of the present disclosure, the meaning of "a plurality" is at least two, such as two, three, etc., unless otherwise specifically defined.
[0157] Those skilled in the art should understand that the above embodiments are only for clearly illustrating the present disclosure and are not intended to limit the scope of the present disclosure. For those skilled in the art, other changes or modifications can be made based on the above disclosure, and these changes or modifications are still within the scope of the present disclosure.
Claims
1. An image synthesis method, characterized in that, The method includes: Obtaining a facial image and a standard hairstyle image; Performing feature matching on the facial image and the standard hairstyle image and then splicing them to obtain a head spliced image; Extracting the head edge features included in the head spliced image; Using the head edge features as control conditions for image synthesis to control the synthesis of the facial image and the standard hairstyle image in the head spliced image, and obtaining a target head spliced image with the hairstyle effect strengthened by the control conditions.
2. The method according to claim 1, characterized in that, The performing feature matching on the facial image and the standard hairstyle image and then splicing them to obtain a head spliced image includes: Extracting the first facial features and the first hair edge features in the facial image; Extracting the second hair edge features in the standard hairstyle image; After matching and aligning the first hair edge features and the second hair edge features, splicing them to obtain the head spliced image including the first facial features and the second hair edge features.
3. The method according to claim 2, wherein The extracting the head edge features included in the head spliced image includes: Extracting the first facial features and the third hair edge features in the head spliced image; Using the first facial features and the third hair edge features to construct the head edge features.
4. The method according to claim 1, wherein It further includes: Obtaining a standard pose image; Performing pose matching on the head spliced image and the standard pose image and then splicing them to obtain an overall spliced image; Extracting the pose features included in the overall spliced image; Using the head edge features and the pose features as control conditions for image synthesis to control the synthesis of the overall spliced image, and obtaining a target synthesized image with the hairstyle effect strengthened by the control conditions.
5. The method according to claim 4, wherein The performing pose matching on the head spliced image and the standard pose image and then splicing them to obtain an overall spliced image includes: Performing matching processing on the head spliced image according to the pose in the standard pose image; Splicing the head spliced image after the matching processing with the neck in the standard pose image; Judging whether the length and width of the face, the width of the shoulders, and the length of the neck after splicing meet the thresholds; If they meet the requirements, generating the overall spliced image.
6. The method according to claim 4, wherein The extracting the pose features included in the overall spliced image includes: Using a key point extraction tool to identify the second facial features and the neck and shoulder features in the overall spliced image; Extracting the second facial features and the neck and shoulder features together to form the pose features.
7. The method according to claim 4, wherein The using the head edge features and the pose features as control conditions for image synthesis to control the synthesis of the overall spliced image, and obtaining a target synthesized image with the hairstyle effect strengthened by the control conditions includes: After aligning the first facial features included in the head edge features with the second facial features included in the pose features, generating control conditions including the head edge features and the pose features; Inputting the overall spliced image into a model, and using the control conditions to strengthen the hairstyle key information and the pose key information in the process of the model synthesizing the overall spliced image, and synthesizing a target synthesized image with a standard hairstyle and a standard pose.
8. An electronic device, characterized in that, It includes: A memory, where the memory stores execution instructions; And A processor that executes the execution instructions stored in the memory, such that the processor performs the method according to any one of claims 1 to 7.
9. A readable storage medium, characterized in that, Execution instructions are stored in the readable storage medium, and when the execution instructions are executed by a processor, they are used to implement the method according to any one of claims 1 to 7.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method according to any one of claims 1 to 7.