Image synthesis method and device, medium and product

By aligning face images with standard poses and using extracted features for controlled synthesis, the method addresses the challenge of diverse user demands in image synthesis, enhancing model versatility and reducing costs while ensuring consistent output.

CN120318372APending Publication Date: 2025-07-15KE COM (BEIJING) TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510361976.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The prior art is difficult to achieve diversified and unified hairstyles and clothing styles in image synthesis, and the model training cost is high, which cannot meet the personalized needs of different users.

Method used

By acquiring facial images and standard pose images, pose matching and stitching are performed, facial contour features and pose features are extracted as control conditions, and image synthesis is used using the model to generate a unified and diverse synthetic image.

Benefits of technology

This reduces the image synthesis capability requirements for the model, reduces the cost of model training, and ensures consistency and diversity of the poses of the synthetic image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318372A_ABST
    Figure CN120318372A_ABST
Patent Text Reader

Abstract

The invention provides an image synthesis method and device, a medium and a product. The image synthesis method comprises the following steps: acquiring a face image and a standard posture image; performing posture matching on a face orientation posture in the face image and a neck and shoulder posture in the standard posture image, and splicing the face image and the standard posture image to obtain a spliced image; extracting facial contour features and posture features contained in the spliced image; and controlling the synthesis of the spliced image by taking the facial contour features and the attitude features as control conditions for image synthesis to obtain a synthesized image of which the attitude effect is enhanced by the control conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and particularly to an image synthesis method, device, medium and product. Background Art

[0002] With the development of image synthesis technology, especially in combination with artificial intelligence models, the image synthesis ability has been significantly improved.

[0003] In some work scenarios, there are unified requirements for staff photos, such as unified requirements for employees' hairstyles, clothing, poses during photo shooting, etc. However, when actually taking work photos, due to the diverse presentation of shooting effects such as hairstyles and clothing in different photo studios, it is impossible to achieve complete unity. If an artificial intelligence model is used to synthesize the photos uniformly, it is often necessary to specifically train a dedicated synthesis model. Although the synthesis effect has been improved, this synthesis model is difficult to meet the diverse needs of users for different hairstyle styles and clothing styles, and will require a higher model training cost. Summary of the Invention

[0004] The present disclosure provides an image synthesis method, device, medium and product.

[0005] According to a first aspect of the present disclosure, an image synthesis method is provided. The method specifically includes: obtaining a facial image and a standard pose image; after performing pose matching on the facial orientation pose in the facial image and the neck and shoulder pose in the standard pose image, splicing the facial image and the standard pose image to obtain a spliced image; extracting the facial contour features and pose features included in the spliced image; using the facial contour features and pose features as control conditions for image synthesis to control the synthesis of the spliced image, and obtaining a synthesized image with the pose effect strengthened by the control conditions.

[0006] Based on the above, after obtaining the facial image, instead of directly inputting the facial image into the model for image synthesis, the facial image is first pose-matched with an image having a standard pose and then spliced to obtain a spliced image. The spliced image is obtained by adjusting the facial pose according to the neck and shoulder pose in the standard pose image. During subsequent image synthesis, the model does not need to make excessive adjustments to the pose, and can directly output the synthesized and beautified image. It can be seen that the requirements for the image synthesis ability of the model are significantly reduced, and there is no need to conduct targeted training on the model. While improving the model's diverse image synthesis ability, the model application cost is also reduced. In addition, the synthesis pose of the synthesized image is limited by the standard pose image, so that the pose of the finally synthesized image has good consistency (that is, many users use this solution to synthesize images with good pose consistency, because the difference between different users is only the facial image, and the same standard pose image is used, avoiding the occurrence of pose differentiation).

[0007] According to at least one embodiment of the present disclosure, after performing pose matching between the facial orientation pose in the facial image and the neck and shoulder pose in the standard pose image, splicing the facial image and the standard pose image to obtain a spliced image, including: performing matching processing on the facial image according to the pose in the standard pose image; splicing the matched facial image with the neck in the standard pose image; determining whether the length and width of the face, the width of the shoulders, and the length of the neck after splicing meet the thresholds; if not, adjusting the size and / or position of the facial image; if so, generating a spliced image.

[0008] According to at least one embodiment of the present disclosure, determining whether the length and width of the face, the width of the shoulders, and the length of the neck after splicing meet the thresholds; if not, adjusting the size and / or position of the facial image, including: determining whether the first ratio of the facial width to the shoulder width meets the first threshold; if not, scaling and adjusting the length and width of the face of the facial image; determining whether the second ratio of the facial length to the neck length meets the second threshold; if not, adjusting the facial position of the facial image.

[0009] According to at least one embodiment of the present disclosure, extracting the facial contour features and pose features included in the spliced image, including: using a key point extraction tool to identify the facial key points, neck and shoulder key points, and facial edge key points in the spliced image; extracting the pose features jointly formed by the facial key points and the neck and shoulder key points; extracting the facial edge key points to form the facial contour features.

[0010] According to at least one embodiment of the present disclosure, extracting facial key points and neck-shoulder key points to jointly form pose features includes: determining coordinate values corresponding to the facial key points and neck-shoulder key points; constructing a Pose graph for representing the pose state based on the coordinate values; and extracting pose features from the Pose graph using controlNet.

[0011] According to at least one embodiment of the present disclosure, optimizing a facial image includes: performing beautification processing on the facial image; and scaling and rotating the beautified facial image according to the neck-shoulder pose in a standard pose image.

[0012] According to at least one embodiment of the present disclosure, using facial contour features and pose features as control conditions for image synthesis to control the synthesis process of a spliced image, and obtaining a synthesized image with the pose effect enhanced by the control conditions, includes: determining style parameters for controlling image generation; generating control conditions using the facial contour features and pose features; inputting the spliced image into a model fine-tuned with the style parameters, and using the control conditions to enhance key information in the synthesis process of the spliced image by the model, so as to obtain a synthesized image with a standard pose.

[0013] After using facial contour features and pose features as control conditions for image synthesis to control the synthesis of a spliced image and obtaining a synthesized image with the pose effect enhanced by the control conditions according to at least one embodiment of the present disclosure, it further includes: extracting facial feature features in the facial image; and performing feature fusion on the facial feature features and the synthesized image to obtain a target synthesized image with enhanced facial feature features.

[0014] According to a second aspect of the present disclosure, there is provided an electronic device, including: a memory storing execution instructions; and a processor that executes the execution instructions stored in the memory, such that the processor executes the method according to the first aspect of any one of the embodiments of the present disclosure.

[0015] According to a third aspect of the present disclosure, there is provided a readable storage medium storing execution instructions, where the execution instructions, when executed by a processor, are used to implement the method according to the first aspect of any one of the embodiments of the present disclosure.

[0016] According to a fourth aspect of the present disclosure, there is provided a computer program product including a computer program, where the computer program, when executed by a processor, implements the method according to the first aspect of any one of the embodiments of the present disclosure. Brief Description of the Drawings

[0017] The drawings illustrate exemplary embodiments of the present disclosure and, together with the description thereof, are used to explain the principles of the present disclosure. These drawings are included to provide a further understanding of the present disclosure and are included in this specification and form a part of this specification.

[0018] Figure 1 Schematic flowchart of an image synthesis method provided by the present disclosure.

[0019] Figure 2 Schematic diagram of stitching a facial image and a standard pose image illustrated by the present disclosure.

[0020] Figure 3 Schematic diagram of the key point extraction effect illustrated by the present disclosure.

[0021] Figure 4 Schematic diagram of the feature fusion process illustrated by the present disclosure.

[0022] Figure 5 Schematic diagram of the image synthesis process provided by the present disclosure.

[0023] Figure 6 Schematic block diagram of an image synthesis device according to an embodiment of the present disclosure.

[0024] Figure 7 Schematic block diagram of an electronic device according to an embodiment of the present disclosure. Detailed implementation manners

[0025] The present disclosure will be further described in detail below with reference to the accompanying drawings and examples. It can be understood that the specific examples described herein are only used to explain the relevant content and do not limit the present disclosure. Additionally, it should be noted that for the sake of description, only parts related to the present disclosure are shown in the drawings.

[0026] It should be noted that, without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other. The technical solution of the present disclosure will be described in detail below with reference to the drawings and in combination with the embodiments.

[0027] In some scenarios, the work photos of staff will be presented to the public through websites or display boards. The first impression left by the work photos on users is very important. When taking work photos, although staff will try their best to dress and take photos according to unified requirements, the final presented effects are still uneven. For example, the hairstyles are not unified and the hair colors are not unified. In order to ensure complete uniformity, an artificial intelligence model is used to optimize the photos taken by staff. To ensure the consistency of the image optimization effect of the artificial intelligence model, targeted training needs to be carried out according to actual needs, so as to obtain a model that meets the requirements of the synthesis effect. Although the image synthesis effect of the model obtained through targeted training has been significantly improved, if the style or style type of the work photos generated by staff changes, this model will not be able to handle it. Therefore, there is an urgent need for a solution that can simply and efficiently meet the diverse image synthesis requirements.

[0028] For ease of description and to make the technical solutions of the specific embodiments of the present disclosure easier to understand, before describing the image synthesis method implemented in the present disclosure, the technical terms involved in the specific embodiments of the present disclosure are explained as follows.

[0029] Pose feature: A feature jointly composed of key points obtained by feature extraction of the spliced image and neck and shoulder key points, used to represent the pose of the person in the synthesized image, and is also one of the pose control conditions in the image synthesis process.

[0030] Figure 1 The following is a schematic flowchart of an image synthesis method provided by the present disclosure. As Figure 1 shown, the method includes steps 101 to 103. Among them, the method can be executed by an electronic device such as a server (local server or cloud server).

[0031] Specifically, Figure 1 the method shown includes: Step 101: Obtain a facial image and a standard pose image.

[0032] The facial image mentioned here refers to the original image obtained by photographing the user, and the clear facial image of the user is included in the original image. In order to reduce the complexity of subsequent splicing with the standard pose image, when taking the facial image, the user should try to take the same pose as the standard pose image.

[0033] The standard pose image mentioned here refers to an image with at least a standard neck and shoulder pose, or a full-body pose image. The standard pose image serves as the basis for pose synthesis in the subsequent image synthesis process and is also the basis for optimizing the facial image in the image splicing process. The pose in the standard pose image can be a front pose of the upper body or a side pose of the upper body. According to the user's type of image synthesis requirements, the standard pose image can be selected by the user according to their own needs. For example, the user can select a standard pose image of wearing different types of work uniforms, or select a standard pose image with a front or side pose according to the need.

[0034] In an optional solution, the standard pose image can be an upper body front pose image without a head, only with a neck and shoulders, and the person in the standard pose image is wearing compliant clothing. In some scenarios, various decorations can also be added to the clothing, such as adding brooches, bows, ties, etc. It only needs to be added to the standard pose image, and it is not necessary for the staff taking the photo to wear or wear them, which can effectively improve the image generation efficiency. Reduce the time cost and economic cost of the staff taking photos.

[0035] Step 102: After performing pose matching on the facial orientation pose in the facial image and the neck and shoulder pose in the standard pose image, splice the facial image and the standard pose image to obtain a spliced image.

[0036] After obtaining the facial image, the facial image can be directly stitched with the standard pose image after simple processing. For example, the head in the facial image can be stitched with the neck in the standard pose image lacking a head.

[0037] The pose matching between the facial image and the standard pose image mentioned here can be understood as that the facial orientation pose of the person when taking the facial image should match the neck and shoulder pose in the standard pose image. Among them, the neck and shoulder pose in the standard pose image can be a front pose, a side pose, etc. The facial orientation pose can be the pose when the front face of the face faces the camera when taking the image, the pose when the side face of the face faces the camera, the pose when the face looks down at the camera, the pose when the face looks up at the camera, etc. To make the stitching effect more realistic, it is necessary to ensure that the facial orientation pose in the captured facial image is exactly the same as the neck and shoulder pose. For example, when stitching, it is necessary to ensure that the pose of the facial image is the same as the pose of the standard pose image. For example, if the neck and shoulder pose in the standard pose image is a front pose, the pose presented by the facial image also needs to be a front pose, that is, taking a photo facing the camera. The facial pose can be adjusted accordingly according to needs during stitching. If the neck and shoulder pose in the standard pose image is a side pose, the facial image will also be adjusted to a side pose before stitching during stitching. The facial image can be scaled and rotated during the stitching process to meet the stitching requirements.

[0038] Generally speaking, the standard pose image is prepared in advance, while the facial image is captured by the user. Due to problems such as the shooting angle or the user's facial angle, the facial orientation poses of the facial images are diverse. When matching, it may be necessary to find a facial image that best matches the standard pose image from multiple facial images before stitching.

[0039] Of course, if the requirements for pose unification still cannot be met after processing such as scaling and rotation, the staff can be notified to retake the facial image and tell the staff how to adjust the facial pose when shooting.

[0040] From this, it can be seen that during stitching, the key information that needs to be unified but is uncontrollable (such as clothing type, color, neck and shoulder pose, etc.) is standardized through the standard pose image. It only requires the staff to provide a photo with the facial image. When using the model for image synthesis, it is not necessary to make major changes to the model or conduct a large number of targeted trainings on the model. As long as the key information remains unchanged and optimization is carried out, different image synthesis requirements can be met.

[0041] Step 103: Extract the facial contour features and pose features included in the stitched image.

[0042] After the stitched image is generated in the manner described above, facial contour features and posture features may be further extracted from the stitched image.

[0043] The facial contour features mentioned here can be understood as key points on the edge of the range formed by the hair, chin, left and right cheeks. The facial contour features can accurately describe the unique face shape of the staff member's face.

[0044] The posture features mentioned here can be understood as the neck and shoulder features and facial features extracted from the standard posture image and the facial image respectively. Since the relative position relationship between the facial image and the neck in the standard posture image has been adjusted (including scaling, rotation angle, etc.) when the images are stitched together, when the posture features are extracted, the neck and shoulder features and facial features (also ear, nose, and mouth features) are extracted from the stitched image at the same time.

[0045] When extracting features from the spliced image, facial contour features and posture features can be extracted simultaneously to ensure that the key point coordinates of the facial contour features and the posture features can accurately correspond. In an optional solution, the facial contour features and the posture features can be put into a heat map to ensure that the key point coordinates of the facial contour features and the posture features have good consistency.

[0046] Of course, if the standard posture image is a standing posture, or contains other postures of the upper limbs, then when extracting posture features, it is necessary to further extract other features in addition to the neck and shoulder features, such as upper limb features, torso features, lower limb features, etc.

[0047] Step 104: Using facial contour features and posture features as control conditions for image synthesis to control the synthesis of the spliced images, and obtaining a synthetic image with the posture effect enhanced by the controlled conditions.

[0048] In the spliced image, different staff members can use the same standard pose image, and each staff member provides his or her own facial image. In this way, all contents except the face in the final composite image can be synthesized according to a unified standard. In other words, the use of standard pose images standardizes the diverse poses, clothing and other features. And the faces that need to be diverse can continue to be diverse.

[0049] When the model is used for image synthesis, the spliced image is input into the model. The model can directly use the spliced image for synthesis. The facial contour features and posture features are input into the model as the control conditions of the model, so that the facial contour and posture of the synthesized image can be further enhanced during the process of the model synthesizing the image.

[0050] Since the facial contour features, pose features, and the stitched image already contain all the key information required for image synthesis. When the model uses the facial contour features, pose features, and the stitched image for image synthesis, it doesn't need to make excessive modifications. It only needs to optimize the connection between the head and the neck in the original stitched image on the basis of ensuring that the key information remains unchanged. Moreover, this model doesn't need to be specifically trained, reducing the cost of using the model. When the synthesis requirements of the staff change, only the required standard pose image needs to be replaced, which can well meet the diverse image synthesis requirements.

[0051] Through the above solution, after obtaining the facial image, instead of directly inputting the facial image into the model for image synthesis. First, the facial image is stitched with an image having a standard pose to obtain a stitched image. In this stitched image, the facial pose is adjusted according to the neck and shoulder pose in the standard pose image. In subsequent image synthesis, the model doesn't need to make excessive adjustments to the pose, and can directly output the synthesized and beautified image. It can be seen that the requirement for the image synthesis ability of the model is significantly reduced, and there is no need to specifically train the model. While improving the diverse image synthesis ability of the model, it also reduces the application cost of the model. In addition, the synthesis state of the synthesized image is limited by the standard pose image, making the pose of the finally synthesized image have good consistency.

[0052] In one or more embodiments of the present disclosure, after matching the facial orientation pose in the facial image with the neck and shoulder pose in the standard pose image, the facial image and the standard pose image are stitched to obtain a stitched image, including: performing matching processing on the facial image according to the pose in the standard pose image; stitching the processed facial image with the neck in the standard pose image; determining whether the length and width of the stitched face, the width of the shoulders, and the length of the neck meet the threshold; if not, adjusting the size and / or position of the facial image; if so, generating the stitched image.

[0053] In practical applications, when performing image stitching, the face in the facial image is stitched with the neck in the standard pose image. Since the facial image and the standard pose image are obtained by different methods, in order to ensure that the stitched image is more realistic during stitching, the facial image needs to be adjusted according to the standard pose image. This includes scaling adjustment, rotation adjustment, etc. of the facial image.

[0054] Such as Figure 2 is a schematic diagram of stitching the facial image and the standard pose image illustrated in the present disclosure. As can be seen from Figure 2 it, there is a neck and an upper body image wearing specified clothing in the standard pose image, and no head image.

[0055] During splicing, the head is connected to the neck, and the facial image is placed on top of the neck image. In this way, the length of the neck can be changed by moving the position of the facial image. For example, when it is necessary to increase the neck length of the spliced image, the facial image can be moved upward; when it is necessary to decrease the neck length in the spliced image, the facial image can be moved downward.

[0056] In addition, since the sizes of the facial images provided by different staff members are different, the relative size relationship between the facial image and the standard pose image is also inconsistent. Therefore, it is necessary to perform scaling adjustment on the facial image. When scaling the facial image, the length and width need to be scaled proportionally to avoid facial deformation during the scaling process. In an alternative solution, during the splicing process, the length and width of the face, the length of the neck, and the width of the shoulders can be adjusted so that the adjusted ratio of the face to the neck and shoulders is more coordinated and more in line with the normal ratio of real people. The process of ratio adjustment will be described in detail in the following embodiments and will not be repeated here.

[0057] After splicing is completed, the spliced image is regarded as a whole, and the relative positions of the eyes, nose, and mouth to the neck and shoulders are fixed. When performing feature extraction subsequently, the extraction is carried out according to the adjusted relative positions.

[0058] In one or more embodiments of the present disclosure, it is determined whether the length and width of the face, the width of the shoulders, and the length of the neck after splicing meet the thresholds; if not, size and / or position adjustment is performed on the facial image, including: determining whether the first ratio of the facial width to the shoulder width meets the first threshold; if not, scaling adjustment is performed on the length and width of the face; if so, no scaling adjustment is performed on the length and width of the face. Determine whether the second ratio of the facial length to the neck length meets the second threshold; if not, adjust the position of the face.

[0059] Generally speaking, the width ratio of a person's head to shoulders, as well as the ratio of facial length to neck length, are within a certain range. When determining whether the display ratio of the facial image to the standard pose image in the spliced image is appropriate, the first ratio of the facial width to the shoulder width and the second ratio of the facial length to the neck length can be calculated respectively.

[0060] More specifically, calculate the first ratio of the facial width to the shoulder width, and then determine the magnitude relationship between the first ratio and the first threshold. This first threshold can be a fixed value (such as 0.5), or it can be an interval range value (such as 0.3 to 0.6). When calculating, if it exceeds the interval range specified by the first threshold, the aspect ratio of the facial image will be adjusted to change the width of the facial image so that the adjusted first ratio meets the first threshold.

[0061] Meanwhile, calculate the second ratio of the facial length to the neck length, and then determine the relationship between the second ratio and the second threshold. This second threshold can be a fixed value or an interval range value. When calculating, if the calculation result does not meet the requirements of the second threshold, the position of the facial image will be adjusted, thereby changing the neck length so that the adjusted second ratio meets the second threshold.

[0062] It should be noted that when adjusting the length and width of the facial image, both the first ratio and the second ratio should be considered simultaneously. That is, when adjusting, avoid the situation that the second ratio does not meet the second threshold due to adjusting the first ratio, or avoid the situation that the first ratio does not meet the first threshold due to adjusting the second ratio.

[0063] Generally speaking, since the size and dimensions of the standard pose image are the sizes required for the finally generated image, when adjusting the ratio, the size of the facial image is adjusted, and the size of the standard pose image is not adjusted.

[0064] To better understand the above solution, the following will be described in detail through a specific embodiment. Assume that the first threshold is 0.25 - 0.4 and the second threshold is 0.1 - 0.25. That is, 0.25 < facial width / shoulder width < 0.4, 0.1 < neck length (from the chin at the bottom of the face to the bottom of the neck) / facial length < 0.25. Only when both the first threshold and the second threshold are met, the spliced image meets the proportional requirements of the actual person's face and neck and shoulders. In other words, the synthesized image obtained in this way is more realistic.

[0065] Based on the above embodiment, adjust the first ratio of the facial width to the shoulder width, and at the same time, adjust the second ratio of the facial length to the neck length. Make the pose features in the adjusted synthesized image more in line with the actual user's pose requirements. The spliced image adjusted according to the ratio can reflect a more reasonable pose, providing a basis for subsequent feature extraction (including facial contour features and pose features), so that the image finally synthesized by the model can meet the requirements.

[0066] In one or more embodiments of the present disclosure, extract the facial contour features and pose features included in the spliced image, including: using a key point extraction tool to identify the facial key points, neck and shoulder key points, and facial edge key points in the spliced image; extract the pose features composed of the facial key points and the neck and shoulder key points; extract the facial edge key points to form the facial contour features.

[0067] The key point extraction tool mentioned here can be control_v11p_sd15_openpose, or other extraction tools such as FaceRecognition and Animetrics Face Recognition. It should be noted that OpenPose of ControlNet is a technology that uses computer vision to detect key points of human postures and facial expressions, and takes the detection results as conditional inputs to accurately control the postures, actions and facial expressions of the characters in the images generated by Stable Diffusion, improving the controllability and realism of the generated images.

[0068] The facial key points mentioned here can be understood as the eyes, nose, and mouth in the facial image. Of course, in some cases, they can also include the ears, that is, the five sense organs are used as key points. The neck and shoulder key points mentioned here can be understood as the points on the neck and shoulders collected from the stitched image. The neck and shoulder key points and the facial key points together constitute the pose feature. The pose feature mentioned here is used to represent the pose of the person that the final synthesized image needs to present. This pose feature is determined by the stitched image and is input as a control condition into the synthesis model during the image synthesis process.

[0069] The facial edge key points mentioned here can be understood as the key points on the facial contour excluding the five sense organs. Specifically, they are the points on the contour formed by the upper edge of the forehead, the two sides of the cheeks, and the chin. The facial edge key points are used to construct the facial contour feature, which is used to represent the facial appearance feature of the provided facial image and is input as a control condition into the synthesis model during the image synthesis process.

[0070] Such as Figure 3 is a schematic diagram of the key point extraction effect illustrated in this disclosure. As can be seen from Figure 3 it, the stitched image contains a facial image provided by a staff member (here, the staff member can be understood as the person who provides the facial image and has the need for image synthesis. If other people have the need for synthesized images, it can be any person who can legally provide facial images), and also contains a standard pose image. When performing feature extraction, facial edge key points, facial key points, and neck and shoulder key points are extracted. Among them, the facial key points and the neck and shoulder key points are used to construct the pose feature, and the facial edge key points are used to construct the facial contour feature.

[0071] By performing feature extraction on the stitched image as a whole, the influence of the posture of the user providing the facial image can be eliminated. In other words, users who need image synthesis (for example, staff) only need to consider whether the facial image matches the standard posture image (the matching here means that the shooting posture of the facial image and the neck and shoulder posture in the standard posture image are known. For example, if the neck and shoulder posture in the standard posture image is a straight posture, then the facial image also needs to be shot in a straight posture, and the side face cannot be shot). There is no need to consider the user's clothing and neck and shoulder posture when taking pictures, and the interference of diverse and uncontrollable human factors in image synthesis can be eliminated as much as possible.

[0072] In one or more embodiments of the present disclosure, facial key points and neck and shoulder key points are extracted to jointly constitute posture features, including: determining the coordinate values corresponding to the facial key points and the neck and shoulder key points; constructing a heat map for representing the posture state based on the coordinate values; and extracting posture features from the heat map using controlNet.

[0073] In practical applications, the preprocessed spliced image is input into the controlNet (control_v11p_sd15_openpose) model, and the model will automatically detect and extract key points related to human posture. These posture-related key points usually include neck and shoulder key points, joint points (such as shoulders, elbows, knees, etc.) and other important feature points.

[0074] Similarly, the model can also extract facial key points. These key points usually include the location information of facial features such as eyes, nose, mouth, etc. When extracting facial key points, the model may use a facial detection algorithm to locate the facial area and extract key points within that area.

[0075] The specific implementation process is as follows: Get the key point coordinates. The key point coordinates are detected from the image through a specific algorithm or model. These key points usually represent important features of the human body, such as facial features (eyes, nose, mouth, etc.), shoulders, elbows, wrists, hips, knees and ankles. In order to obtain these key point coordinates, deep learning models such as convolutional neural networks (CNN) or more advanced model architectures (such as ResNet, Hourglass, etc.) are usually used. These models are trained to accurately identify key points in images.

[0076] During the training phase, the model uses a large number of labeled key point data sets to learn and gradually optimize its internal parameters to improve the accuracy of key point detection. Once the model is trained, it can detect key points on new input images and output the coordinate information of the key points.

[0077] After obtaining the key point coordinates, the next step is to convert these coordinates into a Pose graph. A Pose graph is a special form of image representation where key points are marked on the image in a specific manner (such as dots, lines, heatmaps, etc.). The conversion process generally involves the following steps: First, the key point coordinates need to be mapped from the coordinate space output by the model to the image space. This usually involves some geometric transformations such as translation, scaling, and rotation to ensure that the positions of the key points on the Pose graph are consistent with their positions in the original image.

[0078] Then, mark the key points on the Pose graph according to the mapped coordinates. This can be achieved by drawing dots, lines, or other shapes at the corresponding positions. The color, size, and shape of the markings can be adjusted as needed to represent the key points more clearly.

[0079] To represent the connection relationships between key points (such as bone or joint connections), connection information such as lines or arrows can be added to the Pose graph. This connection information helps to better understand the pose and movement of the human body.

[0080] After the Pose graph is generated, it can be input into the ControlNet framework for further processing. The ControlNet framework is a control framework for image generation and editing. It can recognize and understand the key point information in the Pose graph and construct the human body pose accordingly. Specifically, the ControlNet framework first recognizes the key point information in the Pose graph. This usually involves some image processing techniques such as image segmentation, feature extraction, and matching. Through these techniques, the framework can accurately identify each key point in the Pose graph and obtain its coordinate information.

[0081] After recognizing the key point information, the ControlNet framework uses this information to construct the human body pose. This usually involves some geometric and kinematic algorithms such as bone models and joint angle calculations. Through these algorithms, the framework can infer the pose and movement of the human body and generate corresponding control conditions accordingly.

[0082] In summary, the ControlNet framework generates a series of control conditions based on the constructed human body pose. These control conditions will be used to guide the subsequent image generation or editing process to ensure that the synthesized image is consistent with the original Pose graph in terms of pose and movement.

[0083] In one or more embodiments of the present disclosure, the facial image is optimized, including: beautifying the facial image; scaling and rotating the beautified facial image according to the neck and shoulder pose in the standard pose image.

[0084] The beautification process mentioned here can be understood as performing beauty treatment on the original images provided by the staff. For example, simple beauty treatments such as whitening and freckle removal are involved, but there is no need to process the face shape and facial features in the original images to avoid distortion caused by excessive beautification and affecting the synthesis effect of the subsequent synthesized images.

[0085] In practical applications, although it is required that the staff take pictures in the same pose as the standard pose image, there will inevitably be a certain angular deviation. Especially when the original images obtained by the staff through self - shooting have a large difference from the standard pose. To facilitate subsequent image stitching, it is necessary to perform scaling and rotation processing on the facial images. For example, if the line where the facial image and the shoulders are located is not perpendicular, the facial image needs to be rotated by an angle. Another example is that when the distances between the staff and the camera are inconsistent during taking pictures, the facial image needs to be scaled.

[0086] Through the above - mentioned method, when optimizing the facial image, corresponding scaling and rotation adjustments need to be made according to the neck - shoulder pose in the standard pose image. Thus, the adjusted facial image is made to be consistent with the neck - shoulder pose in the standard pose image, providing pose control conditions for subsequent image synthesis.

[0087] In one or more embodiments of the present disclosure, using facial contour features and pose features as control conditions for image synthesis to control the synthesis process of the stitched image, and obtaining a synthesized image with the pose effect enhanced by the control conditions, includes: determining style parameters for controlling image generation; generating control conditions using facial contour features and pose features; inputting the stitched image into a model with fine - tuned style parameters, and using the control conditions to enhance the key information (i.e., facial contour features and pose features for enhancing the pose effect) in the synthesis process of the stitched image by the model, to obtain a synthesized image with a standard pose.

[0088] In practical applications, after obtaining the control conditions composed of facial contour features and pose features through the method described above, further, when synthesizing the image, the control conditions can be input into the model, so that the pose standard effect during image synthesis can be enhanced by using the control conditions.

[0089] Specifically, when synthesizing the image, it is also necessary to clarify the image synthesis style, which can be achieved through the LoRA model here. According to the required image style, appropriate LoRA parameter sets are selected. These parameter sets can be pre - trained or fine - tuned according to specific styles. Thus, the image synthesis model can generate synthesized images that meet the styles required by users.

[0090] If the existing LoRA parameter sets do not exactly match the style required by the staff, these parameters can be fine - tuned using style - related data sets. The fine - tuning process can be implemented through standard machine learning or deep - learning frameworks.

[0091] Input the stitched image into the trained image synthesis model. Meanwhile, integrate the extracted facial contour and pose features as control conditions into the image synthesis model. Usually, the facial contour features and pose features are introduced as additional inputs during the forward propagation of the image synthesis model. During the synthesis process, strengthen the influence of the facial contour and pose features by adjusting the loss function of the model or introducing additional regularization terms. This helps to ensure that the synthesized image accurately reflects the facial contour and pose features of the original image captured by the staff while maintaining the target style.

[0092] It should be noted that the control conditions are different in different image synthesis requirements. In the above-described embodiments, the control conditions include facial contour features and neck-shoulder pose features. In some scenarios, it may also include action features (such as action type, speed, acceleration, dynamic effects, etc.), expression features (such as the state features of eyes, mouth, eyebrows, etc. in a smiling state), In summary, the control conditions in the ControlNet framework are a series of instructions or parameters generated based on the constructed human body pose. These parameters can accurately describe the human body's pose, action, expression, and other detailed information, thereby ensuring that the generated image is consistent with the original template or user expectations. In practical applications, these control conditions can be adjusted and optimized according to specific requirements to meet the image synthesis needs in different scenarios.

[0093] In one or more embodiments of the present disclosure, after using the facial contour features and pose features as control conditions for the image synthesis to control the synthesis of the stitched image and obtaining the synthesized image with the pose effect strengthened by the control conditions, it further includes: extracting the facial feature of the facial image; performing feature fusion on the facial features and the synthesized image to obtain the target synthesized image with enhanced facial features.

[0094] In practical applications, in order to achieve high-quality fusion of the facial features and the synthesized image, it is necessary to use a feature fusion model. The feature fusion model uses a multi-scale attribute encoder to extract the attribute features of the synthesized image, uses a pre-trained face recognition model to extract the ID features of the facial image, and then by introducing a feasible variable feature fusion structure, while embedding the ID features into the attribute feature space, realizes the adaptive change of the face in the form of an optical flow field. Finally, the fusion result is real, high-fidelity, and supports the adaptive perception of the target user's face shape to a certain extent.

[0095] Such as Figure 4 is a schematic diagram of the feature fusion process illustrated for the present disclosure. From Figure 4As can be seen, a multi-scale attribute encoder is used to accurately capture various attribute features in the synthetic image. Through convolutional kernels and pooling layers of different scales, this encoder can efficiently extract detailed information (such as texture and color) and global information (such as shape and layout) in the image, thereby constructing a comprehensive and multi-level attribute feature representation.

[0096] To obtain the ID features of the facial image, the model adopts a pre-trained face recognition model (i.e., the face ID extractor). This model has been trained with a large amount of real face data and can accurately identify and extract the unique features of the face, namely the ID features, which are crucial for distinguishing different individuals.

[0097] To achieve the effective fusion of ID features and attribute features, the feature fusion model introduces a viable variable feature fusion structure. This structure can not only embed the ID features into the attribute feature space to achieve deep fusion at the feature level, but also dynamically adjust the fusion strategy according to the specific content of the input image through an adaptive learning mechanism.

[0098] It should be noted that the feature fusion model realizes the adaptive change of the face in the form of an optical flow field. The optical flow field can capture the pixel-level change information in the image. By applying it to the facial image, the model can simulate the subtle dynamic changes of the face, such as expressions and postures, thus making the fusion result more realistic and natural. Specifically, the optical flow field constructs a complete motion field by calculating the velocity vectors of each pixel point in the image. This motion field can accurately reflect the motion direction and speed of each pixel point in the facial image. When this motion field is applied to the facial image, the model can simulate the subtle dynamic changes of the face, such as expressions and postures, according to the information of the optical flow field. This simulation process is not just a simple pixel position adjustment, but a high-level processing of the facial image based on the motion information of the optical flow field. Through this processing, the model can generate a more realistic and natural fusion result. Because the optical flow field captures the dynamic change information in the facial image, the generated fusion image will be more natural and smooth in terms of expressions and postures.

[0099] In addition, the application of the optical flow field also endows the feature fusion model with a certain degree of robustness. Since the optical flow field can capture the motion information in the image, when the input image has a certain degree of motion blur or noise, the model can still accurately simulate the facial dynamics according to the information of the optical flow field. This greatly improves the applicability and stability of the model.

[0100] After the above processing flow, the fusion result output by the feature fusion model not only has high fidelity and can retain the details and texture of the original image, but also can adaptively perceive the face shape of the target staff to a certain extent. This means that the model can automatically adjust the fusion strategy according to the facial features of the target staff, so that the final generated image is more in line with the user's expectations and needs.

[0101] In summary, through the combination of multi-scale attribute encoder and pre-trained face recognition model, as well as the application of feasible variable feature fusion structure and optical flow field, more realistic and natural target synthetic images can be generated. The feature fusion model can automatically adjust the fusion strategy according to the specific content of the input image and the facial features of the target user, and realize the adaptive perception of the target user's face shape. This feature fusion method improves the image fusion quality and enhances the adaptive ability of the feature fusion model.

[0102] For ease of understanding, the image synthesis process will be described below through a specific embodiment. Figure 5 The following is a schematic diagram of the image synthesis process provided by the present disclosure. Figure 5 As can be seen in , the specific synthesis process is as follows.

[0103] The first step is image stitching. Suppose a user uploads a frontal face image. The frontal face image is input into the face detection model to obtain the coordinate frame of the portrait. The portrait is rotated and straightened based on the information of the coordinate frame. Then the face in the face image is segmented to obtain the coordinate points of the face area. The coordinate points are input into the portrait skin beautification model for whitening, freckle removal and other processing. However, there is no need to beautify the facial features. Finally, a beautified face image is obtained.

[0104] Then, prepare a standard ID photo image without a head (that is, a standard posture image). Assuming that the circumscribed rectangular coordinate frame of the face in the standard ID photo image is known, the center point and area of the face can be calculated. Scale the obtained user's facial image to the same area as the facial area in the standard ID photo image, and use the center alignment method to paste the user's facial image onto the standard ID photo image, and splice the obtained user's facial image and the headless standard posture image together. When splicing, check whether the following conditions are met: ①0.25<face width / shoulder width<0.4, ②0.1<neck length / face length<0.25.

[0105] Condition ① is used to control the size of the face area. If it is not within the range of ①, for example, face width / shoulder width < 0.25, it means that the face is scaled too small. Lock the aspect ratio and scale the face to face width = 0.25 × shoulder width.

[0106] Condition ② is used to control the neck length. If it is not within the range of ②, for example, neck length / face length > 0.25, it means that the chin of the face is too far from the clothes, and the generated neck will be very long. Move the face down to the position where the distance from the chin to the bottom of the neck = 0.25 × face length.

[0107] Through the above steps, a spliced image of the user's facial image and the standard ID photo image is obtained. Then, the above spliced image is pasted onto a gray background image to obtain the final required spliced image.

[0108] Use the control_v11p_sd15_openpose model to extract the pose features of the spliced template image and the facial contour features of the face as the control conditions for ControlNet. Train a style template for standard ID photos using standard ID photos. Input the spliced image into the image synthesis model. The control conditions including facial contour features and pose features, and the LoRA parameters controlling the generated style are input into the image synthesis model. Finally, an ID synthesis image with the same pose, facial key points as the spliced image, and the same dressing code is obtained.

[0109] Furthermore, in order to further enhance the similarity of the portrait, based on the generated synthesis image, fuse the user's original facial image, extract the facial feature points from the original facial image, and fuse the facial feature points with the synthesis image to generate the target synthesis image.

[0110] Based on any of the above embodiments, the present disclosure also provides an image synthesis device. Figure 6 It is a structural schematic block diagram of an image synthesis device according to an embodiment of the present disclosure. As Figure 6 shown, the image synthesis method device includes: an acquisition module 61 for acquiring a facial image and a standard pose image. A splicing module 62 for performing pose matching on the facial orientation pose in the facial image and the neck and shoulder pose in the standard pose image, and then splicing the facial image and the standard pose image to obtain a spliced image. An extraction module 63 for extracting the facial contour features and pose features included in the spliced image. A synthesis module 64 for using the facial contour features and pose features as control conditions for image synthesis to control the synthesis of the spliced image, and obtaining a synthesis image with the pose effect enhanced by the control conditions.

[0111] The splicing module 62 is configured to perform matching processing on the facial image according to the pose in the standard pose image; splice the facial image after the matching processing with the neck in the standard pose image; judge whether the length and width of the face, the width of the shoulders, and the length of the neck after splicing meet the threshold; if not, adjust the size and / or position of the facial image; if so, generate a spliced image.

[0112] The splicing module 62 is further configured to determine whether a first ratio of the face width to the shoulder width meets a first threshold; if not, scale and adjust the face length and width of the face image; determine whether a second ratio of the face length to the neck length meets a second threshold; if not, adjust the face position of the face image.

[0113] The extraction module 63 is configured to use a key point extraction tool to identify face key points, neck and shoulder key points, and face edge key points in the spliced image; extract the face key points and the neck and shoulder key points to jointly form a pose feature; extract the face edge key points to form a face contour feature.

[0114] The extraction module 63 is configured to determine the coordinate values corresponding to the face key points and the neck and shoulder key points; construct a Pose graph for representing the pose state based on the coordinate values; extract the pose feature from the Pose graph by using controlNet.

[0115] The extraction module 63 is configured to perform beautification processing on the face image; scale and rotate the beautified face image according to the neck and shoulder pose in the standard pose image.

[0116] The synthesis module 64 is configured to determine style parameters for controlling image generation; generate control conditions by using the face contour feature and the pose feature; input the spliced image into a model fine-tuned with the style parameters, and use the control conditions to strengthen the key information in the process of synthesizing the spliced image by the model, so as to obtain a synthesized image with a standard pose.

[0117] Optionally, it further includes a fusion module 65, which is configured to extract the facial feature of the face image; perform feature fusion on the facial feature and the synthesized image to obtain a target synthesized image with enhanced facial features.

[0118] The implementation processes of the functions and effects of each module in the above device are specifically described in detail in the implementation processes of the corresponding steps in the above method, and will not be elaborated here.

[0119] The execution subject of the image synthesis method in the specific implementation manner of the present disclosure may be an electronic device such as a server (including a local server or a cloud server).

[0120] Therefore, based on any one of the above embodiments, the present disclosure further provides an electronic device, and this electronic device can execute the image synthesis method of any one of the above embodiments described in the present disclosure.

[0121] Figure 7 It is a structural schematic block diagram of an electronic device according to an embodiment of the present disclosure.

[0122] The hardware structure of the electronic device 1000 can be implemented using a bus architecture. The bus architecture can include any number of interconnected buses and bridges, depending on the specific application of the hardware and the overall design constraints. The bus 1100 connects various circuits including one or more processors 1200, a memory 1300, and / or hardware modules together. The bus 1100 can also connect various other circuits 1400 such as peripheral devices, voltage regulators, power management circuits, external antennas, etc.

[0123] The bus 1100 can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Component (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity of representation, only one connecting line is shown in this figure, but it does not mean that there is only one bus or one type of bus.

[0124] The present disclosure also provides a readable storage medium in which a computer program is stored. When the computer program is executed by a processor, it is used to implement the above-mentioned method. The "readable storage medium" can be any device that can contain, store, communicate, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. More specific examples of the readable storage medium include the following: an electrical connection part (electronic device) having one or more wirings, a portable computer disk cartridge (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable read-only memory (CDROM), etc.

[0125] The present disclosure also provides a computer program product. The method of the present disclosure can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed, the processes or functions of the present disclosure are executed in whole or in part.

[0126] Computer programs or instructions can be stored in a readable storage medium, or transmitted from one readable storage medium to another. For example, the computer programs or instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired or wireless manner. The readable storage medium can be any available medium that can be accessed, or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium, such as a floppy disk, hard disk, or magnetic tape; it can also be an optical medium, such as a digital video disc; or it can be a semiconductor medium, such as a solid-state drive. The computer-readable storage medium can be a volatile or non-volatile storage medium, or can include both volatile and non-volatile types of storage media.

[0127] Those skilled in the art should understand that the embodiments of the present disclosure can be provided as a method, system, or computer program product. Therefore, the present disclosure can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present disclosure can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0128] The present disclosure is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the present disclosure. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing method devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing method devices generate a device for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or multiple flows and / or blocks

[0129] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing method devices to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the functions specified in Figure 1 one or more of the flows Figure 1 or multiple flows and / or blocks

[0130] These computer program instructions can also be loaded onto a computer or other programmable data processing method device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one process or multiple processes and / or blocks. Figure 1 one process or multiple processes and / or blocks Figure 1 steps for implementing the functions specified in one block or multiple blocks.

[0131] In the description of this specification, the description with reference to terms such as "one embodiment / way", "some embodiments / ways", "example", "specific example", or "some examples" means that the specific features, structures, or characteristics described in connection with the embodiment / way or example are included in at least one embodiment / way or example of the present disclosure. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment / way or example. Moreover, the specific features, structures, or characteristics described can be combined in a suitable manner in any one or more embodiments / ways or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments / ways or examples described in this specification and the features of different embodiments / ways or examples.

[0132] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present disclosure, "a plurality" means at least two, such as two, three, etc., unless otherwise clearly and specifically defined.

[0133] Those skilled in the art should understand that the above embodiments are only for clearly explaining the present disclosure and are not intended to limit the scope of the present disclosure. For those skilled in the art, other changes or variations can be made based on the above disclosure, and these changes or variations are still within the scope of the present disclosure.

Claims

1. An image synthesis method, characterized in that, The method includes: Obtaining a facial image and a standard pose image; After performing pose matching on the facial orientation pose in the facial image and the neck-shoulder pose in the standard pose image, splicing the facial image and the standard pose image to obtain a spliced image; Extracting the facial contour features and pose features included in the spliced image; Using the facial contour features and the pose features as control conditions for image synthesis to control the synthesis of the spliced image, and obtaining a synthesized image with the pose effect strengthened by the control conditions.

2. The method according to claim 1, wherein The step of, after performing pose matching on the facial orientation pose in the facial image and the neck-shoulder pose in the standard pose image, splicing the facial image and the standard pose image to obtain a spliced image, includes: Performing matching processing on the facial image according to the neck-shoulder pose in the standard pose image; Splicing the matched facial image with the neck in the standard pose image; Judging whether the length and width of the face, the width of the shoulders, and the length of the neck after splicing meet the thresholds; If not, adjusting the size and / or position of the facial image; If so, generating the spliced image.

3. The method according to claim 2, wherein The step of judging whether the length and width of the face, the width of the shoulders, and the length of the neck after splicing meet the thresholds; The step of, if not, adjusting the size and / or position of the facial image, includes: Judging whether the first ratio of the facial width to the shoulder width meets the first threshold; If not, performing scaling adjustment on the length and / or width of the face in the facial image; Judging whether the second ratio of the facial length to the neck length meets the second threshold; If not, adjusting the facial position of the facial image.

4. The method according to claim 1, wherein The step of extracting the facial contour features and pose features included in the spliced image, includes: Using a key point extraction tool to identify the facial key points, neck-shoulder key points, and facial edge key points in the spliced image; Extracting the facial key points and the neck-shoulder key points together to form the pose features; Extracting the facial edge key points to form the facial contour features.

5. The method according to claim 4, characterized in that, The step of extracting the facial key points and the neck-shoulder key points together to form the pose features, includes: Determining the coordinate values corresponding to the facial key points and the neck-shoulder key points; Constructing a Pose graph for representing the pose state based on the coordinate values; Using controlNet to extract the pose features from the Pose graph.

6. The method according to claim 1 or 5, characterized in that, The step of using the facial contour features and the pose features as control conditions for image synthesis to control the synthesis of the spliced image, and obtaining a synthesized image with the pose effect strengthened by the control conditions, includes: Determining the style parameters for controlling image generation; Generating control conditions using the facial contour features and the pose features; Inputting the spliced image into a model fine-tuned by the style parameters, and using the control conditions to strengthen the key information in the process of the model synthesizing the spliced image, to obtain a synthesized image with a standard pose.

7. The method according to claim 1, characterized in that After using the facial contour features and the pose features as control conditions for image synthesis to control the synthesis of the spliced image, and obtaining a synthesized image with the pose effect strengthened by the control conditions, it further includes: Extract the facial features in the facial image; Perform feature fusion on the facial features and the synthetic image to obtain a target synthetic image with enhanced facial features.

8. An electronic device, characterized in that, Comprising: A memory that stores execution instructions; And A processor that executes the execution instructions stored in the memory, such that the processor executes the method according to any one of claims 1 to 7.

9. A readable storage medium, characterized in that, Execution instructions are stored in the readable storage medium, and when the execution instructions are executed by a processor, they are used to implement the method according to any one of claims 1 to 7.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method according to any one of claims 1 to 7.