Image synthesis method, device and storage medium

By determining the collar point location and repairing the texture features of the neck area in the image synthesis method, the problem of image unreality caused by inaccurate collar positioning in the existing technology is solved, and a more natural and realistic image synthesis effect is achieved.

CN114820309BActive Publication Date: 2025-09-09MIGU CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210387545.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-13
Publication Date
2025-09-09
Estimated Expiration
2042-04-13

AI Technical Summary

Technical Problem

The images generated by existing image synthesis methods are relatively unrealistic, especially the poor combination of collar positioning points and materials, which leads to problems such as the neck being too short or too long, and the collar not fitting the neck well.

Method used

By obtaining the target individual's portrait image and clothing material image, the positioning point of the collar point of the clothing material image in the portrait image to be synthesized is determined, and the missing area of ​​the neck area image is determined based on the positioning point. The texture features of the neck area image are used to repair it, and finally the target image is synthesized.

Benefits of technology

The problem of the neck being too short or too long, and the collar not fitting the neck due to poor combination of collar positioning points and materials has been solved. The connection of the synthesized image is more natural and the realism is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114820309B_ABST
    Figure CN114820309B_ABST
Patent Text Reader

Abstract

The present invention discloses an image synthesis method, device, and storage medium, belonging to the field of image processing. The image synthesis method comprises: obtaining a portrait image to be synthesized and a clothing material image of a target individual, wherein the clothing material image does not include a neck portion; determining a positioning point of a collar point of the clothing material image in the portrait image to be synthesized; based on the positioning point, determining a missing region of the neck region image in the portrait image to be synthesized relative to the clothing material image; repairing the missing region based on the texture features of the neck region image to obtain a repaired neck region image; and synthesizing the head region image in the portrait image to be synthesized, the repaired neck region image, and the clothing material image based on the positioning point to obtain a target image of the target individual. The present invention can synthesize the portrait image to be synthesized and the clothing material image to obtain a more realistic user image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing technology, and in particular to an image synthesis method, device and storage medium. Background Art

[0002] In related technologies, in order to meet the needs of users to take various ID photos and modify ID photos through mobile phones anytime, anywhere, conveniently and quickly, material pictures and existing images can be combined to generate images that meet user needs, such as ID photos.

[0003] However, the images generated by existing image synthesis methods are relatively unrealistic. Summary of the Invention

[0004] The main purpose of the present invention is to provide an image synthesis method, device and storage medium, aiming to solve the problem that images generated by existing image synthesis methods are unrealistic.

[0005] To achieve the above objectives, in a first aspect, the present invention provides an image synthesis method, comprising:

[0006] Obtaining a portrait image to be synthesized and a clothing material image of a target individual, wherein the clothing material image does not include a neck portion;

[0007] Determine the location of the collar point of the clothing material image in the portrait image to be synthesized;

[0008] Determining, based on the positioning points, a missing region of the neck region image in the portrait image to be synthesized relative to the clothing material image;

[0009] repairing the missing area according to the texture features of the neck area image to obtain a repaired neck area image;

[0010] According to the positioning points, the head region image of the person in the portrait image to be synthesized, the repaired neck region image and the clothing material image are synthesized to obtain a target image of the target individual.

[0011] In one embodiment, determining, based on the positioning points, a missing region of the neck region image in the portrait image to be synthesized relative to the clothing material image includes:

[0012] Obtaining a face region mask image and a first neck region mask image according to the portrait image to be synthesized;

[0013] Obtaining a second neck region mask image according to the clothing material image;

[0014] mapping the second neck region mask image to the portrait image to be synthesized according to the positioning point, and subtracting the face region mask image to obtain a predicted neck region mask image;

[0015] The predicted neck region mask image is compared with the first neck region mask image to obtain the missing region.

[0016] In one embodiment, before comparing the predicted neck region mask image with the first neck region mask image to obtain the missing region, the method further includes:

[0017] Calculating the brightness value of each pixel in the intersection area of ​​the predicted neck region mask image and the first neck region mask image to obtain an average first brightness value of all pixels and an average second brightness value of pixels within a preset sorting range;

[0018] If the difference between the first average brightness value and the second average brightness value is less than or equal to a preset difference, the step of comparing the predicted neck region mask image with the first neck region mask image is performed to obtain the missing region.

[0019] In one embodiment, determining, based on the positioning points, a missing region of the neck region image in the portrait image to be synthesized relative to the clothing material image includes:

[0020] The predicted neck region mask image is compared with the first neck region mask image to obtain the missing region.

[0021] In one embodiment, after comparing the predicted neck region mask image with the first neck region mask image to obtain the missing region, the method further includes:

[0022] calculating a ratio of the number of mask pixels of the missing area to the number of mask pixels of the first neck area mask image;

[0023] Determining whether the ratio is less than a preset threshold;

[0024] If it is less than the preset threshold, the method of repairing the missing area according to the texture feature of the neck area image is executed to obtain a repaired neck area image.

[0025] In one embodiment, determining the location of the collar point of the clothing material image in the portrait image to be synthesized includes:

[0026] Inputting the portrait image to be synthesized into a trained chin and neck boundary connection point recognition model to obtain boundary points output by the chin and neck boundary connection point recognition model;

[0027] Determining a collar point of the clothing material image;

[0028] Determining a longitudinal offset of the collar point relative to the boundary point based on the head height ratio;

[0029] The positioning point of the collar point in the to-be-synthesized portrait image is determined according to the longitudinal offset.

[0030] In one embodiment, determining the longitudinal offset of the collar point relative to the boundary point based on the head height ratio includes:

[0031] determining a first horizontal distance between the two boundary points and a second horizontal distance between the two collar points;

[0032] obtaining a distance ratio between the first horizontal distance and the second horizontal distance;

[0033] Determine the vertical coordinate information of the left mandibular key point, the right mandibular key point, and the chin key point in the portrait image to be synthesized;

[0034] Determining a maximum value between a first difference between the left mandibular key point and the chin key point and a second difference between the second preset facial key point and the chin key point;

[0035] The longitudinal offset of the collar point relative to the boundary point is determined according to the distance ratio and the maximum value.

[0036] In one embodiment, determining the location point of the collar point in the to-be-synthesized portrait image according to the longitudinal offset includes:

[0037] Determining the maximum value of the vertical coordinates of the two boundary points;

[0038] Obtaining a target longitudinal coordinate of the collar point according to the maximum coordinate value and the longitudinal offset;

[0039] The first neck region mask image is traversed, and two boundary pixel points whose longitudinal coordinates are the same as the target longitudinal coordinates are determined as the positioning points.

[0040] In one embodiment, the repairing of the missing region based on the texture features of the neck region image to obtain a repaired neck region image includes:

[0041] Sampling the texture features of the neck region image to obtain sample texture features;

[0042] Filling the texture features of the missing area according to the sample texture features to obtain a filling result;

[0043] The filling result is fused with the neck region image to obtain a repaired neck region image.

[0044] In a second aspect, the present application also provides an image synthesis device, comprising: a memory, a processor, and an image synthesis program stored in the memory and executable on the processor, wherein the image synthesis program is configured to implement the steps of the image synthesis method described above.

[0045] In a third aspect, the present application further provides a computer-readable storage medium, on which an image synthesis program is stored, and when the image synthesis program is executed by a processor, the steps of the image synthesis method described above are implemented.

[0046] An embodiment of the present invention provides an image synthesis method, which comprises obtaining a portrait image to be synthesized and a clothing material image of a target individual; determining the positioning point of the collar point of the clothing material image in the portrait image to be synthesized; determining the missing area of ​​the neck region image in the portrait image to be synthesized relative to the clothing material image based on the positioning point; repairing the missing area based on the texture characteristics of the neck region image to obtain a repaired neck region image; and synthesizing the head region image in the portrait image to be synthesized, the repaired neck region image, and the clothing material image based on the positioning point to obtain a target image of the target individual. Thus, the present invention can solve problems such as the neck being too short or too long, or the collar not fitting well, caused by poor combination of the collar positioning point and the material, by using the positioning point of the collar point of the clothing material image in the portrait image to be synthesized as a reference for synthesis. Furthermore, the present invention can repair the missing area based on the texture characteristics of the neck region image, making the synthesized image more naturally connected, thereby obtaining a more realistic synthesized image. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 Schematic diagram of the structure of the image synthesis device of the present invention;

[0048] Figure 2 1 is a flow chart of a first embodiment of an image synthesis method according to the present invention;

[0049] Figure 3 Schematic diagram of key points of a face according to the present invention;

[0050] Figure 4 2 is a flow chart of a second embodiment of an image synthesis method according to the present invention;

[0051] Figure 5 2 is a flow chart of a third embodiment of an image synthesis method according to the present invention;

[0052] Figure 6 2 is a flow chart of a fourth embodiment of an image synthesis method according to the present invention;

[0053] Figure 7 Schematic diagram of the boundary points of the present invention;

[0054] Figure 8 Schematic diagram of the positioning point of the present invention.

[0055] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION

[0056] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0057] Existing technologies, in order to meet users' needs for conveniently and quickly taking and modifying ID photos anytime, anywhere, using mobile phones, can combine source images with existing images to generate images that meet user needs, such as ID photos. Currently, existing ID photo applications primarily employ solutions that replace the user's original neck with a source or generated neck and then add clothing source material, or replace the user's original neck with clothing source material textures.

[0058] However, the inventors of this application have discovered that the existing pictures synthesized after changing clothes have the following defects:

[0059] 1. The neck and shoulders of the portrait taken by the user exceed the range of the clothing material, resulting in the user's original clothing texture being exposed on the neck and shoulders after changing clothes, resulting in poor dressing effect.

[0060] 2. The neck in the user-captured image is short or partially obscured by clothing, which does not match the neck area exposed at the collar of the clothing material. This causes the clothing or background to be exposed at the collar, resulting in a poor dressing effect.

[0061] 3. Use generation or material replacement to replace the user's original neck. However, since the generated neck or material neck does not match the user's facial clarity, light and shadow connection, etc., it leads to the problem of "fake at first glance" and the dressing effect is poor.

[0062] 4. The combination of collar positioning points and materials is not good, resulting in problems such as the neck being too short or too long, and the collar not fitting the neck well.

[0063] That is, the images generated by existing image synthesis methods are relatively unrealistic.

[0064] In order to solve the above problems, the present application also provides an image synthesis method, which uses the collar point of the clothing material image in the positioning point of the portrait image to be synthesized as a reference for synthesis, and can solve problems such as the neck being too short or too long, and the collar not fitting the neck due to poor combination of the collar positioning point and the material. In addition, the missing area is repaired based on the texture features of the neck area image, so that the synthesized image connection is more natural, thereby obtaining a more realistic synthesized image.

[0065] The inventive concept of the present application is further described below with reference to some specific embodiments.

[0066] Reference Figure 1 , Figure 1 This is a structural diagram of an image synthesis device in the hardware operating environment involved in the embodiment of the present application.

[0067] like Figure 1 As shown, the image synthesis device may include: a processor 1001, such as a central processing unit (CPU), a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. The communication bus 1002 is used to implement connection and communication between these components. The user interface 1003 may include a display screen (Display) and an input unit such as a keyboard (Keyboard). Optionally, the user interface 1003 may also include a standard wired interface or a wireless interface. The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a wireless fidelity (WI-FI) interface). The memory 1005 may be a high-speed random access memory (RAM) memory or a stable non-volatile memory (NVM), such as a disk storage. The memory 1005 may also be a storage device independent of the aforementioned processor 1001.

[0068] Those skilled in the art will understand that Figure 1 The structure shown in the figure does not constitute a limitation on the image synthesis device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.

[0069] like Figure 1 As shown, the memory 1005 as a storage medium may include an operating system, a data storage module, a communication module, a user interface module and an image synthesis program.

[0070] exist Figure 1In the image synthesis device shown, the network interface 1004 is mainly used for data communication with the network server; the user interface 1003 is mainly used for data interaction with the user; the processor 1001 and the memory 1005 in the image sending device of the present invention can be set in the image sending device, and the image sending device calls the image synthesis program stored in the memory 1005 through the processor 1001 and executes the image synthesis method provided in the embodiment of the present application.

[0071] Based on the above hardware devices but not limited to the above hardware devices, a first embodiment of an image synthesis method of the present application is proposed. Figure 2 , Figure 2 This is a flowchart of the first embodiment of the image synthesis method of the present application.

[0072] In this embodiment, the method includes:

[0073] Step S101: obtaining a portrait image to be synthesized and a clothing material image of a target individual, wherein the clothing material image does not include a neck portion;

[0074] In this embodiment, the image synthesis method is performed by an image synthesis device, which is used to synthesize the user's desired composite image based on the input personal image and clothing material image. The image synthesis device can be a mobile terminal such as a mobile phone or tablet, or a local computer or cloud server.

[0075] Specifically, the image synthesis device may obtain the portrait image to be synthesized and the clothing material image through the network, or may retrieve the stored portrait image to be synthesized and the clothing material image from a local database. Alternatively, the image synthesis device may retrieve the stored portrait image to be synthesized from the local database and download the clothing material image from the network. This application is not limited to this. It is also understood that the portrait image to be synthesized may also be a portrait image of the target individual to be synthesized obtained by the image synthesis device in real time based on a photo taking instruction.

[0076] It is understood that the portrait image to be synthesized includes at least the face, neck, and collar. Alternatively, the portrait image to be synthesized may also include the top of the head to the shoulders, i.e., the ID photo of the target individual. It is also understood that in this embodiment, the neck portion of the portrait image to be synthesized is exposed.

[0077] The clothing material image includes at least a portrait model and a clothing image matching the portrait model, wherein the portrait model includes a preset neck area that is exposed outside the clothing image, and the preset neck area can be a blank area. It is worth mentioning that in this embodiment, the clothing image includes a collar portion.

[0078] Step S102: Determine the location of the collar point of the clothing material image in the portrait image to be synthesized.

[0079] The collar points are the left and right boundary points of the preset neck region of the clothing material image. By determining the positioning points of the collar points of the clothing material image within the portrait image to be synthesized, the neck region of the portrait image to be synthesized can be more proportionally coordinated with the clothing material image. The positioning points are the coordinate locations where the boundary points of the clothing material image need to be superimposed during the synthesis step.

[0080] Step S103: Determine, based on the positioning points, the missing area of ​​the neck region image in the portrait image to be synthesized relative to the clothing material image.

[0081] After determining the positioning point, that is, finding the reference point of the portrait image to be synthesized and the clothing material image during synthesis, the missing area of ​​the neck area image in the portrait image to be synthesized relative to the clothing material image can be determined by reference.

[0082] Understandably, when a user synthesizes an image, due to the user's captured image having a shorter neck or the neck being partially obscured by clothing, the neck region image in the portrait image to be synthesized may not match the predetermined neck region of the preset image. Consequently, directly superimposing the portrait image to be synthesized with the clothing source image may result in disproportionate proportions. Therefore, in this embodiment, the neck region image can be compared with the predetermined neck region of the clothing source image to determine the missing portion between the actual individual's neck and the ideal neck region, i.e., the missing region.

[0083] For example, the neck of the portrait uploaded by the user to be synthesized is relatively short, while the preset neck area in the clothing material image is relatively slender. At this time, the positioning point can be used as a reference point to determine the missing area of ​​the neck area image relative to the preset neck area, so that the proportions of the human body after synthesis are more coordinated, that is, more realistic.

[0084] Step S104: repair the missing area according to the texture features of the neck area image to obtain a repaired neck area image.

[0085] In this embodiment, the missing area can be repaired using the texture features of the neck area image, so that the texture features of the missing area are consistent with the texture features of a real individual's neck.

[0086] Step S105 : synthesizing the head region image of the person in the portrait image to be synthesized, the repaired neck region image, and the clothing material image according to the positioning points to obtain a target image of the target individual.

[0087] Specifically, in this step, the head region image, the repaired neck region image, and the clothing source image are synthesized using the positioning points as reference points to produce the target image of the target individual. For example, the user-selected background color is set as the canvas background color; the head region image is overlaid on the background layer. The repaired neck region image is then applied to the head region image. Finally, the collar points and positioning points in the clothing source image are mapped and overlaid on the repaired neck region image, completing the entire outfit change process.

[0088] After the change is completed, the target image can be cropped according to the pre-set size ratio based on the ID photo size selected by the user, and finally the ID photo effect image of the specified size can be output.

[0089] In this embodiment, the collar point of the clothing material image is synthesized with reference to the positioning point in the portrait image to be synthesized, which can solve the proportional size problems such as the neck being too short or too long, the collar and the neck not fitting together, etc. caused by the poor combination of the collar positioning point and the material. In addition, the missing area is repaired according to the texture features of the neck area image, so that the connection of the synthesized image is more natural, so as to avoid the appearance of fake situations such as inconsistent clarity between the face area and the neck area, poor light and shadow connection, inconsistent texture, etc., thereby obtaining a more realistic synthesized image.

[0090] In this embodiment, only the image of the human head area and the repaired neck area image are used during synthesis, and the image of the user's original clothing part in the portrait to be synthesized is not used. This can solve the problem that the neck and shoulders of the portrait taken by the user are beyond the range of the material clothes, resulting in the neck and shoulders of the user's original clothing texture being exposed after changing clothes. It can also solve the problem that the neck in the image taken by the user is short or partially blocked by clothes, which does not match the neck range exposed at the collar of the clothing material, resulting in the clothes or background being exposed at the collar.

[0091] In this embodiment, by repairing the missing area according to the texture features of the neck area image, it is possible to solve the problem of using generation or material replacement to replace the user's original neck, but the generated neck or material neck does not match the user's facial clarity, light and shadow connection, etc., and the neck texture is lost.

[0092] In this embodiment, the collar point of the clothing material image is synthesized with reference to the positioning point in the portrait image to be synthesized, which can solve the problems of the neck being too short or too long, and the collar not fitting the neck caused by the poor combination of the collar positioning point and the material.

[0093] As an embodiment, after step S101, the method further includes:

[0094] The portrait image to be synthesized is scaled to obtain a scaled image that matches the size of the target image.

[0095] Since the actual size of the portrait image to be synthesized may not match the size of the target image in actual use, a scaling operation can be performed in advance to expand the scope of application of the method of the application and further improve the user experience.

[0096] In this embodiment, the to-be-synthesized portrait image is scaled to obtain a scaled image that matches the target image size, including:

[0097] Step A10: Detect facial key point detection information of the portrait image to be synthesized to obtain the actual face width.

[0098] See Figure 3 , the actual face width can be calculated based on the straight line distance between points 0 and 32 in the face calculation, that is,

[0099] FaceW=sqrt((facepoint

[32] .x-facepoint[0].x)*(facepoint

[32] .x-facepoint[0].x)+(facepoint

[32] .y-facepoint[0].y)*(facepoint

[32] .y-facepoint[0].y)),

[0100] Where sqrt is the square root, FaceW is the actual face width, and facepoint[n] is the nth key point.

[0101] Step A20: Create a canvas of a preset size;

[0102] For example, if you create a square canvas with a side length of L, the canvas size can be adaptively modified based on the target image's clarity. For example, if the target image is a 1-inch or 2-inch ID photo, the canvas side length can be adaptively adjusted.

[0103] Step A30: Calculate the relative position information of the ideal face in the canvas.

[0104] Specifically, relative position information includes the ideal face's width within the canvas and the height from the chin to the bottom edge of the canvas. The face's width within the canvas can be calculated using the formula drawFaceW = faceWRatio * L, where drawFaceW is the width and faceWRatio is the ratio of face width to canvas length, pre-set based on facial aesthetic proportions.

[0105] You can also calculate the height of the face's chin from the canvas's bottom edge using the formula drawJawH = jawHRatio * L. drawJawH is the height of the face's chin from the canvas's bottom edge. jawHRatio is the ratio of the chin-to-bottom-edge height minus the canvas's side length, calculated based on the aesthetic proportions of the face.

[0106] Step A40: Calculate a scaling ratio between the width information and the actual face width in the portrait image to be synthesized, and scale the portrait image to be synthesized and the facial key point information according to the scaling ratio to obtain an intermediate scaled image;

[0107] Among them, according to the following formula:

[0108] scale=drawFaceW / FaceW calculates a first ratio between the width information and the actual face width in the portrait image to be synthesized, where scale is the scaling ratio.

[0109] Step A50: Using the pupil midpoint as a reference point, translate the intermediate zoomed image into the canvas to obtain a zoomed image.

[0110] Specifically, the midpoint of points 104 and 105 between the two pupil centers in the facial position is calculated and recorded as eyesCenter. The intermediate zoomed image is translated with eyesCenter as the anchor point, and eyesCenter in the intermediate zoomed image is translated to the position (L / 2, L-drawJawH) on the canvas. At the same time, the facial key point information is translated to obtain the zoomed image.

[0111] It can be understood that the subsequent operations for synthesizing the portrait image are all operations on scaling the image.

[0112] Based on the above embodiments, a second embodiment of the image synthesis method of the present application is proposed. Figure 4 , Figure 4 This is a flowchart of the second embodiment of the image synthesis method of the present application.

[0113] In this embodiment, before step S103, the method further includes:

[0114] Step S201: obtaining a face region mask image and a first neck region mask image according to the portrait image to be synthesized;

[0115] In this step, the zoomed image can be segmented to obtain a face region mask image and a head and neck region mask image. The face region mask image includes the hair and face. The head and neck region mask image includes the hair, face, and exposed neck region.

[0116] After subtracting the face region mask image from the head and neck region mask image, a first neck region mask image is obtained. That is, the first neck region mask image includes the user's actual neck region image. This first neck region mask image can reflect the user's actual texture features.

[0117] Step S203: Obtain a second neck region mask image according to the clothing material image.

[0118] In this step, by performing portrait segmentation on the clothing material image, a second neck region mask image can be obtained. The second neck region mask image is the preset neck region of the clothing material image.

[0119] Step S205: Mapping the second neck region mask image to the portrait image to be synthesized according to the positioning point, and subtracting the face region mask image to obtain a predicted neck region mask image;

[0120] In this step, the second neck region mask image is mapped onto the scaled image according to the correspondence between the left collar point C and the left positioning point E, and the right collar point D and the right positioning point F. The face region mask image obtained by the aforementioned segmentation is then subtracted to remove the chin portion to obtain the predicted neck region mask image.

[0121] Step S207 : Compare the predicted neck region mask image with the first neck region mask image to obtain a missing region.

[0122] That is, the predicted neck region mask image is subtracted from the first neck region mask image to obtain a mask image that reveals the missing neck portion required by the user for the clothing material image. The area corresponding to the mask image is the missing area.

[0123] Based on the second embodiment of the above method, a third embodiment of the image synthesis method of the present application is proposed, see Figure 5 , Figure 5 This is a flowchart of the third embodiment of the image synthesis method of the present application.

[0124] In this embodiment, before step S207, the method further includes:

[0125] Step S206: Calculate the brightness value of each pixel in the intersection area of ​​the predicted neck area mask image and the first neck area mask image to obtain the average first brightness value of all pixels and the average second brightness value of pixels in a preset sorting range.

[0126] Specifically, the brightness value can be calculated using the RGB channel values ​​of each pixel in this area.

[0127] Brightness=0.3*R+0.6*G+0.1*B, wherein Brightness is the brightness value. It is understandable that the brightness value can also be calculated based on other color spaces, which is not limited in this embodiment.

[0128] After obtaining the brightness values ​​of each pixel in the predicted neck region mask image, the difference between the average first brightness value of all pixels and the average of the top 10% of the brightness values ​​can be calculated. In this case, the preset ranking range is the top 10%. It is understood that the preset ranking range can be determined based on the user's specific skin color and lighting conditions, and is not limited here.

[0129] If the difference between the first average brightness value and the second average brightness value is less than or equal to the preset difference, step S207 is executed.

[0130] In this embodiment, if the difference between the first brightness value average and the second brightness value average is less than or equal to the preset difference, it is considered that there is no "yin and yang neck" situation, and the texture features of the user's real neck can be used to repair the missing area.

[0131] If the difference between the first and second brightness averages is greater than a predetermined difference, it is considered that a "yin-yang neck" exists. Even if the texture features of the user's actual neck are used to repair the missing area, the "yin-yang neck" will still exist, or even be aggravated. Therefore, the texture features of the user's actual neck should not be used to repair the missing area. At this point, step S208 is executed: a replacement neck region image is generated based on the clothing material image.

[0132] Step S209 : synthesizing the head region image, the replacement neck region image, and the clothing material image in the portrait image to be synthesized according to the positioning points to obtain a target image of the target individual.

[0133] That is, in this embodiment, when it is determined that the user's real image has problems such as "yin and yang neck" due to hand lighting, skin color or even photographic defects, that is, when the texture features of the user's real neck are inconsistent, an alternative neck area image can be generated, and the alternative neck area image can be used to replace the user's neck area image for synthesis to improve the realism of the synthesized image.

[0134] Based on the above embodiments, a fourth embodiment of the image synthesis method of the present application is proposed. Figure 6 , Figure 6 This is a flowchart of the fourth embodiment of the image synthesis method of the present application.

[0135] In this embodiment, after step S103, the method further includes:

[0136] Step S210: Calculate the ratio of the number of mask pixels of the missing area to the number of mask pixels of the first neck area mask image;

[0137] Step S211: determining whether the ratio is less than a preset threshold;

[0138] If it is less than the preset threshold, step S104 is executed to repair the missing area according to the texture features of the neck area image to obtain a repaired neck area image.

[0139] Specifically, in this embodiment, the ratio of the number of mask pixels lossSum of the missing neck portion, ie, the missing area, to the number of mask pixels neckSum of the real neck area, ie, the first neck area mask image, is calculated as lossSum / neckSum.

[0140] If the ratio lossSum / neckSum is less than the preset threshold, it is determined that the missing part of the current neck in the portrait image to be synthesized is not large, and the missing part can be filled with the real neck pixel sampling texture, that is, the subsequent repair step is less difficult and the quality after repair is higher.

[0141] If the ratio lossSum / neckSum is greater than or equal to the preset threshold, it is considered that the current neck is too missing, such as there is an occluded part, or the size difference between the user's neck area image and the preset neck area image is large, the neck area image is not enough to sample and fill the missing part, or the quality after repair and filling is low, and there are still problems such as unrealism. At this time, the user's real neck should not be used, and step S208 can be executed.

[0142] As an embodiment, determining the location of the collar point of the clothing material image in the portrait image to be synthesized includes:

[0143] Step B10: input the portrait image to be synthesized into a trained chin and neck boundary connection point recognition model to obtain boundary points output by the chin and neck boundary connection point recognition model.

[0144] See Figure 7 In this embodiment, the chin-neck boundary connection point recognition model can be pre-trained by selecting a large number of portrait images with marked boundary points. The portrait image to be synthesized is input into the chin-neck boundary connection point recognition model, and the chin-neck boundary connection point recognition model outputs the boundary points.

[0145] Step B20: Determine the collar point of the clothing material image.

[0146] In this step, the collar point of the clothing material image can be directly detected through the second neck region mask image in the clothing material image.

[0147] Step B30: Determine the longitudinal offset of the collar point relative to the boundary point according to the head height ratio.

[0148] In this embodiment, the head height ratio is an industry standard or industry-recommended head height ratio. Alternatively, this can be used as a basis for determining the longitudinal offset of the collar point relative to the boundary point. The head height ratio can also be a user-entered, actual head height ratio, which is not a limitation in this embodiment.

[0149] Specifically, as an option of this embodiment, the longitudinal offset of the collar point relative to the boundary point can be determined based on the size of the mandibular portion and the chin. In this case, step B30 includes:

[0150] (1) determining a first horizontal distance between the two boundary points and a second horizontal distance between the two collar points;

[0151] (2) obtaining a distance ratio between the first horizontal distance and the second horizontal distance;

[0152] The distance ratio can be determined by the following formula:

[0153] Ratio = LenAB / LenCD;

[0154] Where Ratio is the distance ratio, LenAB is the first horizontal distance between boundary point A and boundary point B, and LenCD is the second horizontal distance between collar point C and collar point D.

[0155] (3) determining the vertical coordinate information of the left mandibular key point, the right mandibular key point, and the chin key point in the portrait image to be synthesized;

[0156] (4) determining the maximum value of a first difference between the left mandibular key point and the chin key point and a second difference between the second preset facial key point and the chin key point;

[0157] Specifically, the maximum vertical distance between points 11, 21, and 16 of the facial key points is calculated, that is,

[0158] H=max((facepoint

[16] .y-facepoint

[11] .y),(facepoint

[16] .y-facepoint

[21] .y));

[0159] Among them, H is the maximum value.

[0160] (4) Determine the longitudinal offset of the collar point relative to the boundary point based on the distance ratio and the maximum value.

[0161] The vertical downward offset of collar points C and D relative to boundary points A and B is calculated using the distance ratio and H according to the following formula:

[0162] diff=H*Ratio; where diff is the longitudinal offset.

[0163] Step B40: Determine the location of the collar point in the portrait image to be synthesized based on the longitudinal offset.

[0164] After the longitudinal offset is calculated, the boundary point can be used as a reference point and the longitudinal offset can be used as an offset to determine the positioning point of the collar point in the portrait image to be synthesized.

[0165] In one embodiment, since the user's posture is inevitably asymmetrical, such as the left side of the face being larger, the right side being larger, or the left side being higher and the right side being lower, in order to further make the proportions after synthesis more realistic and natural, step B40 specifically includes:

[0166] (1) determining the maximum value of the vertical coordinates of the two boundary points;

[0167] (2) obtaining the target longitudinal coordinate of the collar point according to the maximum coordinate value and the longitudinal offset;

[0168] (3) Traversing the first neck region mask image, two boundary pixel points whose longitudinal coordinates are the same as the target longitudinal coordinates are determined as the positioning points.

[0169] Specifically, the target vertical coordinate of the collar point clothy = max(Ay, By) + diff, where max represents the maximum value, Ay is the vertical coordinate of the boundary point A, and By is the vertical coordinate of the boundary point B. Traverse the first neck region mask image and search for the horizontal coordinate xLeft of the first valid pixel and the horizontal coordinate xRight of the last valid pixel of clothy. Figure 8 , define the coordinates (xLeft, clothy) (xRight, clothy) as the left positioning point E and the right positioning point F, and then obtain the positioning point.

[0170] As an embodiment, step S104 includes:

[0171] Step C10: sampling the texture features of the neck region image to obtain sample texture features.

[0172] In this embodiment, the first neck region mask image includes texture features of the user's actual neck region image, and the first neck region mask image can be sampled, that is, the first neck region mask image is used as the region of interest of the sampler to obtain sample texture features.

[0173] Step C20: Fill in the texture features of the missing area according to the sample texture features to obtain a filling result.

[0174] By using an image restoration filling technology (inpainting), texture features are filled in the missing area according to the sample texture features to obtain a filling result.

[0175] Step C30: Fusing the filling result with the neck region image to obtain a repaired neck region image.

[0176] Specifically, the filling result can be naturally fused with the original neck area image texture through edge fusion to obtain the repaired neck area image.

[0177] In addition, an embodiment of the present invention further proposes a computer storage medium, on which a video encoding program is stored, and when the video encoding program is executed by a processor, the steps of the video encoding method as described above are implemented. Therefore, no further description will be given here. In addition, the description of the beneficial effects of adopting the same method will not be repeated. For technical details not disclosed in the computer-readable storage medium embodiment involved in this application, please refer to the description of the method embodiment of this application. As an example, the program instructions can be deployed to be executed on one computing device, or on multiple computing devices located at one location, or on multiple computing devices distributed at multiple locations and interconnected by a communication network.

[0178] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware through a computer program. The above-described program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes in the above-described method embodiments. The above-described storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).

[0179] It should also be noted that the device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed across multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment. In addition, in the drawings of the device embodiments provided by the present invention, the connection relationship between the modules indicates that there is a communication connection between them, which may be specifically implemented as one or more communication buses or signal lines. A person of ordinary skill in the art can understand and implement the present invention without inventive effort.

[0180] Through the description of the above embodiments, it is clear to those skilled in the art that the present invention can be implemented by means of software plus necessary general-purpose hardware, and of course it can also be implemented by means of dedicated hardware including application-specific integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. In general, all functions performed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structures used to implement the same function can also be various, such as analog circuits, digital circuits, or dedicated circuits, etc. However, for the present invention, software program implementation is a better embodiment in most cases. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer's floppy disk, USB flash drive, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc., and includes a number of instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods of each embodiment of the present invention.

[0181] The above are only preferred embodiments of the present invention and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention description and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.

Claims

1. An image synthesis method, characterized in that: include: Obtaining a portrait image to be synthesized and a clothing material image of a target individual, wherein the clothing material image does not include a neck portion; Determine a positioning point of the collar point of the clothing material image in the portrait image to be synthesized; the positioning point is a coordinate position where the collar point of the clothing material image needs to be superimposed; Determining, based on the positioning points, a missing region of the neck region image of the portrait image to be synthesized relative to a preset neck region of the clothing material image; repairing the missing area according to the texture features of the neck area image to obtain a repaired neck area image; According to the positioning points, the head region image of the person in the portrait image to be synthesized, the repaired neck region image and the clothing material image are synthesized to obtain a target image of the target individual.

2. The image synthesis method according to claim 1, wherein: The step of determining, based on the positioning points, a missing region of the neck region image in the portrait image to be synthesized relative to a preset neck region of the clothing material image includes: Obtaining a face region mask image and a first neck region mask image according to the portrait image to be synthesized; Obtaining a second neck region mask image according to the clothing material image; the second neck region mask image is a mask image of the preset neck region; mapping the second neck region mask image to the portrait image to be synthesized according to the positioning point, and subtracting the face region mask image to obtain a predicted neck region mask image; The predicted neck region mask image is compared with the first neck region mask image to obtain the missing region.

3. The image synthesis method according to claim 2, wherein: Before comparing the predicted neck region mask image with the first neck region mask image to obtain the missing region, the method further includes: Calculating the brightness value of each pixel in the intersection area of ​​the predicted neck region mask image and the first neck region mask image to obtain an average first brightness value of all pixels and an average second brightness value of pixels within a preset sorting range; If the difference between the first average brightness value and the second average brightness value is less than or equal to a preset difference, the step of comparing the predicted neck region mask image with the first neck region mask image is performed to obtain the missing region.

4. The image synthesis method according to claim 2, wherein: After comparing the predicted neck region mask image with the first neck region mask image to obtain the missing region, the method further includes: calculating a ratio of the number of mask pixels of the missing area to the number of mask pixels of the first neck area mask image; Determining whether the ratio is less than a preset threshold; If it is less than the preset threshold, the method of repairing the missing area according to the texture feature of the neck area image is executed to obtain a repaired neck area image.

5. The image synthesis method according to claim 2, wherein: The step of determining the positioning point of the collar point of the clothing material image in the portrait image to be synthesized includes: Inputting the portrait image to be synthesized into a trained chin and neck boundary connection point recognition model to obtain boundary points output by the chin and neck boundary connection point recognition model; Determining a collar point of the clothing material image; Determining a longitudinal offset of the collar point relative to the boundary point based on the head height ratio; The positioning point of the collar point in the to-be-synthesized portrait image is determined according to the longitudinal offset.

6. The image synthesis method according to claim 5, characterized in that: The determining of the longitudinal offset of the collar point relative to the boundary point according to the head height ratio includes: determining a first horizontal distance between the two boundary points and a second horizontal distance between the two collar points; obtaining a distance ratio between the first horizontal distance and the second horizontal distance; Determine the vertical coordinate information of the left mandibular key point, the right mandibular key point, and the chin key point in the portrait image to be synthesized; Determine a maximum value of a first difference between the left mandibular key point and the chin key point and a second difference between the right mandibular key point and the chin key point; The longitudinal offset of the collar point relative to the boundary point is determined according to the distance ratio and the maximum value.

7. The image synthesis method according to claim 6, characterized in that: Determining the positioning point of the collar point in the to-be-synthesized portrait image according to the longitudinal offset includes: Determining the maximum value of the vertical coordinates of the two boundary points; Obtaining a target longitudinal coordinate of the collar point according to the maximum coordinate value and the longitudinal offset; The first neck region mask image is traversed, and two boundary pixel points whose longitudinal coordinates are the same as the target longitudinal coordinates are determined as the positioning points.

8. The image synthesis method according to any one of claims 1 to 7, characterized in that: The step of repairing the missing region according to the texture features of the neck region image to obtain a repaired neck region image includes: Sampling the texture features of the neck region image to obtain sample texture features; Filling the texture features of the missing area according to the sample texture features to obtain a filling result; The filling result is fused with the neck region image to obtain a repaired neck region image.

9. An image synthesis device, characterized in that: include: A memory, a processor, and an image synthesis program stored in the memory and executable on the processor, wherein the image synthesis program is configured to implement the steps of the image synthesis method according to any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores an image synthesis program, which, when executed by a processor, implements the steps of the image synthesis method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Image generation method and device, electronic equipment and storage medium

    CN112562034A