Portrait photo generation method and apparatus, and electronic device

By adjusting the portrait characteristics of the to-process portrait photos based on the bone key points of the reference photo before the portrait photo is generated, the problem of facial features mismatch in portrait photo generation in the prior art is solved. The generated photo background and clothing are consistent with the reference photo, and the stability and quality of the generated photo are improved.

WO2025167078A1PCT designated stage Publication Date: 2025-08-14HUAWEI TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/116476
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-08
Filing Date
2024-09-03
Publication Date
2025-08-14

AI Technical Summary

Technical Problem

In the generation of portrait photos, due to the difference between the bone layout of the original photograph and the bone layout of the reference photo, directly replacing the facial features of the portrait can easily lead to facial defects and mismatch problems in the generated photos.

Method used

Before the portrait photo is generated, the portrait characteristics of the pending portrait photo are adjusted based on the bone key points of the reference photo, the first target portrait photo is generated, and then the portrait feature is replaced to ensure the consistency of the bone key points.

Benefits of technology

It effectively avoids mismatch problems in the replacement of portrait features. The generated photo background and clothing are consistent with the reference photos, and the portrait features are consistent with the original photos, improving the stability and quality of the generated photos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024116476_14082025_PF_FP_ABST
    Figure CN2024116476_14082025_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of computers, and particularly to a portrait photo generation method and apparatus, and an electronic device. The method comprises: acquiring a plurality of portrait photos to be processed and extracting portrait feature information in the portrait photos to be processed; acquiring a reference portrait photo and extracting skeleton key points in the reference portrait photo; generating a first target portrait photo on the basis of the portrait feature information in the portrait photos to be processed and the skeleton key points in the reference portrait photo; and generating a second target portrait photo on the basis of the first target portrait photo and the reference portrait photo. In the embodiments of the present application, the portrait feature information in the portrait photos to be processed is first adjusted to adapt to the skeleton key points in the reference portrait photo, and portrait feature replacement is performed on the first target portrait photo and the reference portrait photo, so that the problem of mismatch of facial features can be effectively eliminated, and the subjective effect and resolution of the target portrait photo are optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Portrait photo generation method, device and electronic device

[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on February 8, 2024, with application number 202410177154.1 and application name “Portrait Photo Generation Method, Device and Electronic Device”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of computer technology, and in particular to a method, device and electronic device for generating a portrait photo. Background Art

[0003] Portrait photo generation is a technology that generates diverse portraits based on original and reference photos. The resulting photos share the same composition, style, clothing, background, and specific content as the reference photos. Existing technologies typically replace the facial features of the portrait in the original photo with those in the reference photo through a series of end-to-end face processing algorithms, including face detection, face alignment, feature extraction, and feature replacement. However, the skeletal layout in the original photo may differ from that in the reference photo. Directly replacing the facial features can result in mismatches, resulting in facial defects in the final photo.

[0004] Summary of the Invention

[0005] In view of this, the present application provides a portrait photo generation method, device and electronic device, which pre-adjusts the structure of the portrait features of the portrait photo to be processed based on the skeletal key points of the reference photo, and then replaces the portrait features of the adjusted portrait photo to be processed and the reference photo to obtain the final target portrait photo.

[0006] In a first aspect, an embodiment of the present application provides a method for generating a portrait photo. First, a plurality of portrait photos to be processed are obtained and portrait feature information in the portrait photos to be processed is extracted, a reference portrait photo is obtained and skeletal key points in the reference portrait photo are extracted; then, a first target portrait photo is generated based on the portrait feature information in the portrait photos to be processed and the skeletal key points in the reference portrait photo; finally, a second target portrait photo is generated based on the first target portrait and the reference photo.

[0007] In an embodiment of the present application, before replacing portrait features of a to-be-processed portrait photo with a reference photo, the layout of the portrait feature information in the to-be-processed portrait photo is adjusted based on the skeletal key points in the reference photo, thereby generating a first target portrait photo. The structural layout of the first target portrait photo is identical to that of the reference photo, and the content of the first target portrait photo is identical to that of the to-be-processed portrait photo. When replacing portrait features of the first target portrait photo with the reference photo, no mismatch in portrait features occurs, resulting in an ideal target portrait photo.

[0008] In an optional embodiment, the portrait photos to be processed include: gaze-state portrait photos, and / or non-gaze-state portrait photos. The method for obtaining gaze-state portrait photos and / or non-gaze-state portrait photos may include: while the user is looking at the portrait capture device, using the portrait capture device to capture gaze-state portrait photos of the user at different angles; and / or while the user's gaze direction remains unchanged, using the portrait capture device to capture non-gaze-state portrait photos of the user at different angles. Portrait feature information can be determined based on the gaze-state portrait photos and / or non-gaze-state portrait photos.

[0009] In the embodiment of the present application, by collecting portrait photos of the user in the gaze state and the non-gaze state from multiple angles, the diversity of portrait features can be increased, which is conducive to adjusting the portrait features in the portrait photos to be processed.

[0010] In an optional embodiment, generating a first target portrait photo based on the portrait feature information in the to-be-processed portrait photo and the skeleton key points in the reference portrait photo includes:

[0011] A first target portrait photo in an original shooting scene is generated by taking the skeleton key points in the reference portrait photo as a structural reference and taking the portrait feature information in the portrait photo to be processed as a content reference.

[0012] In an optional embodiment, generating a first target portrait photo based on the portrait feature information in the to-be-processed portrait photo and the skeleton key points in the reference portrait photo includes:

[0013] Determining the skeleton key points in the reference portrait photo as the skeleton key points of the first target portrait photo;

[0014] The portrait feature information in the portrait photo to be processed is adaptively filled into the skeleton key points of the first target portrait photo to generate the first target portrait photo in the original shooting scene.

[0015] In this embodiment of the present application, the skeleton key points of the first target portrait photo are identical to those of the reference photo, and the portrait feature information in the first target portrait photo is adjusted based on the portrait feature information in the portrait photo to be processed, thereby generating a first target portrait photo with the same portrait structure and layout as the reference photo. The clothing and background of the first target portrait photo are identical to those of the portrait photo to be processed, which makes the generation of the first target portrait photo more stable.

[0016] In an optional embodiment, generating a second target portrait photo based on the first target portrait and the reference photo includes:

[0017] Identifying first portrait feature information of the first target portrait photo and second portrait feature information of the reference photo;

[0018] The second portrait feature information is replaced with the first portrait feature information corresponding to the same skeletal key point to generate the second target portrait photo.

[0019] In this embodiment of the present application, the portrait features of the first target portrait photo and the reference photo are replaced. Since the skeletal key points of the first target portrait photo and the reference photo are the same, the resulting second target portrait photo does not have a portrait feature mismatch. The background and clothing of the second target portrait photo are the same as those of the reference photo.

[0020] In a second aspect, an embodiment of the present application provides a device for generating a portrait photo, comprising:

[0021] A first extraction module is used to obtain a plurality of portrait photos to be processed and extract portrait feature information from the portrait photos to be processed;

[0022] A second extraction module is used to obtain a reference portrait photo and extract skeleton key points in the reference portrait photo;

[0023] A first generating module is used to generate a first target portrait photo based on the portrait feature information in the portrait photo to be processed and the skeleton key points in the reference portrait photo;

[0024] The second generating module is configured to generate a second target portrait photo based on the first target portrait and the reference photo.

[0025] In a third aspect, an embodiment of the present application provides an electronic device comprising a memory for storing computer program instructions and a processor for executing the program instructions, wherein, when the computer program instructions are executed by the processor, the electronic device is triggered to execute the method described in any one of the first or second aspects above.

[0026] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which includes a stored program, wherein when the program is running, the device where the computer-readable storage medium is located is controlled to execute the method described in any one of the first aspect or the second aspect.

[0027] In a fifth aspect, an embodiment of the present application provides a computer program product, which includes executable instructions. When the executable instructions are executed on a computer, the computer executes the method described in any one of the first aspect or the second aspect.

[0028] Using the solution provided in the embodiments of the present application, multiple portrait photos to be processed are obtained and portrait feature information is extracted from the portrait photos to be processed; a reference portrait photo is obtained and skeletal key points are extracted from the reference portrait photo; a first target portrait photo is generated based on the portrait feature information in the portrait photo to be processed and the skeletal key points in the reference portrait photo; and a second target portrait photo is generated based on the first target portrait photo and the reference photo. The structure of the portrait features of the portrait photo to be processed is pre-adjusted based on the skeletal key points of the reference photo, and then the portrait features of the adjusted portrait photo to be processed and the reference photo are replaced to obtain the final target portrait photo, thereby avoiding the problem of mismatch between the portrait features and the skeletal key points during the portrait feature replacement process. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0030] FIG1 is a schematic structural diagram of an electronic device provided in an embodiment of the present application;

[0031] FIG2 is a flow chart of a method for generating a portrait photo provided in an embodiment of the present application;

[0032] FIG3 is a schematic diagram illustrating an example of a method for generating a portrait photo provided by an embodiment of the present application;

[0033] FIG4 is a schematic diagram illustrating another method for generating a portrait photo provided by an embodiment of the present application;

[0034] FIG5 is a schematic diagram illustrating another method for generating a portrait photo provided by an embodiment of the present application;

[0035] FIG6 is a schematic diagram illustrating another method for generating a portrait photo provided by an embodiment of the present application;

[0036] FIG7 is a schematic diagram illustrating another method for generating a portrait photo provided by an embodiment of the present application;

[0037] FIG8 is a schematic diagram illustrating another method for generating a portrait photo provided by an embodiment of the present application;

[0038] FIG9 is a schematic diagram illustrating another method for generating a portrait photo provided by an embodiment of the present application;

[0039] FIG10 is a schematic diagram illustrating another method for generating a portrait photo provided by an embodiment of the present application;

[0040] FIG11 is a schematic diagram illustrating another method for generating a portrait photo provided by an embodiment of the present application;

[0041] FIG12 is a schematic diagram illustrating another method for generating a portrait photo provided by an embodiment of the present application;

[0042] FIG13 is a flow chart of a method for generating a portrait photo according to an embodiment of the present application;

[0043] FIG14 is a schematic structural diagram of a portrait photo generation device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0044] In order to better understand the technical solution of the present application, the embodiments of the present application are described in detail below with reference to the accompanying drawings.

[0045] It should be clear that the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0046] The terms used in the embodiments of the present application are for the purpose of describing specific embodiments only and are not intended to limit the present application. The singular forms "a", "an", "the" and "the" used in the embodiments of the present application and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise.

[0047] It should be understood that the term "and / or" as used herein simply describes a relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A alone, A and B together, or B alone. Furthermore, the character " / " in this document generally indicates an "or" relationship between the associated objects.

[0048] FIG1 is a schematic structural diagram of an electronic device provided in an embodiment of the present application.

[0049] The electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, an earphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, an air pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.

[0050] It should be understood that the structure illustrated in the embodiments of the present invention does not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0051] The processor 110 may include one or more processing units, for example, the processor 110 may include a central processing unit (CPU), a graphics processing unit (GPU), and a display controller, etc. The different processing units may be independent devices or integrated into one or more processors.

[0052] The controller can generate operation control signals according to the instruction operation code and timing signal to complete the control of instruction fetching and execution.

[0053] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. In some embodiments, the processor 110 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, etc.

[0054] The charging management module 140 is configured to receive charging input from a charger. The charger can be either a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 140 can receive charging input from the wired charger via the USB interface 130. In some wireless charging embodiments, the charging management module 140 can receive wireless charging input via the wireless charging coil of the electronic device 100. While charging the battery 142, the charging management module 140 can also provide power to the electronic device via the power management module 141.

[0055] The power management module 141 is used to connect the battery 142, the charging management module 140 and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140 to provide power to the processor 110, the internal memory 121, the display screen 194, etc.

[0056] The wireless communication function of the electronic device 100 can be implemented through the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modem processor and the baseband processor.

[0057] Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. The mobile communication module 150 can provide solutions for wireless communications including 2G / 3G / 4G / 5G, etc., applied to the electronic device 100. The mobile communication module 150 may include at least one filter, a switch, a power amplifier, a low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves from antenna 1, filter, amplify, and process the received electromagnetic waves, and transmit them to the modulation and demodulation processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modulation and demodulation processor and convert it into electromagnetic waves for radiation through antenna 1.

[0058] The modulation and demodulation processor may include a modulator and a demodulator. The modulator is used to modulate the low-frequency baseband signal to be transmitted into a medium-high frequency signal. The demodulator is used to demodulate the received electromagnetic wave signal into a low-frequency baseband signal.

[0059] The wireless communication module 160 receives electromagnetic waves via the antenna 2, modulates and filters the electromagnetic wave signals, and sends the processed signals to the processor 110. The wireless communication module 160 can also receive signals to be transmitted from the processor 110, modulate the frequency of the signals, amplify the signals, and convert them into electromagnetic waves for radiation via the antenna 2.

[0060] In some embodiments, antenna 1 of electronic device 100 is coupled to mobile communication module 150 , and antenna 2 is coupled to wireless communication module 160 , so that electronic device 100 can communicate with the network and other devices through wireless communication technology.

[0061] Electronic device 100 implements display functionality through a GPU, display screen 194, and the like. A GPU is a microprocessor for image processing and is connected to display screen 194. The GPU is used to perform mathematical and geometric calculations for graphics rendering. Processor 110 may include one or more GPUs that execute program instructions to generate or modify display information.

[0062] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), etc.

[0063] The electronic device 100 can realize the shooting function through the ISP, camera 193, video codec, GPU and display screen 194.

[0064] The ISP processes data fed back by camera 193. For example, when taking a photo, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, which is then passed to the ISP for processing and converted into a visible image. The ISP can also perform algorithmic optimization on image noise, brightness, and skin tone. It can also optimize parameters such as exposure and color temperature of the captured scene. In some embodiments, the ISP can be located within camera 193.

[0065] The camera 193 is used to capture still images or videos. The object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, and then passes the electrical signal to the ISP for conversion into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard RGB, YUV or other format. In some embodiments, the electronic device 100 may include 1 or N cameras 193, where N is a positive integer greater than 1.

[0066] The digital signal processor is used to process digital signals. In addition to processing digital image signals, it can also process other digital signals. For example, when the electronic device 100 selects a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy.

[0067] The video codec is used to compress or decompress digital video. The electronic device 100 may support one or more video codecs.

[0068] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100.

[0069] The internal memory 121 can be used to store computer executable program code, which includes instructions. The internal memory 121 may include a program storage area and a data storage area. The program storage area may store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc. The data storage area may store data created during the use of the electronic device 100 (such as audio data, a phone book, etc.), etc. The processor 110 executes various functional applications and data processing of the electronic device 100 by running the instructions stored in the internal memory 121 and / or the instructions stored in the memory provided in the processor.

[0070] The electronic device 100 can implement audio functions such as music playback and recording through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the headphone jack 170D, and the application processor.

[0071] The audio module 170 is used to convert digital audio information into analog audio signal output, and is also used to convert analog audio input into digital audio signals. The audio module 170 can also be used to encode and decode audio signals.

[0072] The speaker 170A, also called a "horn", is used to convert audio electrical signals into sound signals.

[0073] The receiver 170B, also called a "handset", is used to convert audio electrical signals into sound signals.

[0074] Microphone 170C, also called "microphone" or "microphone", is used to convert sound signals into electrical signals.

[0075] The headphone jack 170D is used to connect a wired headphone and can be the USB interface 130 or a 3.5mm open mobile terminal platform (OMTP) standard interface or a cellular telecommunications industry association of the USA (CTIA) standard interface.

[0076] The pressure sensor 180A is used to sense the pressure signal and convert the pressure signal into an electrical signal. In some embodiments, the pressure sensor 180A can be disposed on the display screen 194 .

[0077] The gyro sensor 180B can be used to determine the motion posture of the electronic device 100. In some embodiments, the angular velocity of the electronic device 100 around three axes (i.e., x, y, and z axes) can be determined by the gyro sensor 180B. The gyro sensor 180B can also be used for image stabilization.

[0078] The air pressure sensor 180C is used to measure air pressure. In some embodiments, the electronic device 100 calculates the altitude using the air pressure value measured by the air pressure sensor 180C to assist in positioning and navigation.

[0079] The magnetic sensor 180D includes a Hall sensor, and the electronic device 100 can use the magnetic sensor 180D to detect the opening and closing of the flip leather case.

[0080] Accelerometer 180E can detect the magnitude of acceleration of electronic device 100 in all directions (generally three axes). It can also detect the magnitude and direction of gravity when electronic device 100 is stationary. It can also be used to identify the electronic device's posture, enabling applications such as switching between landscape and portrait modes and pedometers.

[0081] The distance sensor 180F is used to measure distance. The electronic device 100 can measure distance by infrared or laser.

[0082] The proximity light sensor 180G may include, for example, a light emitting diode (LED) and a light detector, such as a photodiode. The electronic device 100 may use the proximity light sensor 180G to detect when a user holds the electronic device 100 close to their ear to talk, thereby automatically turning off the screen to save power.

[0083] The ambient light sensor 180L is used to sense the ambient light brightness, the fingerprint sensor 180H is used to collect fingerprints, and the temperature sensor 180J is used to detect temperature.

[0084] Touch sensor 180K, also known as a "touch control device," can be mounted on display screen 194. Together, touch sensor 180K and display screen 194 form a touch screen, also known as a "touch screen." Touch sensor 180K detects touch operations applied to or near the touch sensor. The touch sensor can communicate the detected touch operations to the application processor to determine the type of touch event.

[0085] Bone conduction sensor 180M can acquire vibration signals. In some embodiments, bone conduction sensor 180M can acquire vibration signals from vibrating bones in the human body. Bone conduction sensor 180M can also contact the human pulse to receive blood pressure signals.

[0086] The buttons 190 include a power button, a volume button, and the like. The buttons 190 may be mechanical buttons or touch buttons. The electronic device 100 may receive key inputs and generate key signal inputs related to user settings and function control of the electronic device 100.

[0087] Motor 191 can generate vibration prompts. Motor 191 can be used for incoming call vibration prompts or touch vibration feedback. Motor 191 can also correspond to different vibration feedback effects for touch operations on different areas of display screen 194.

[0088] The indicator 192 may be an indicator light, which may be used to indicate the charging status, power level changes, messages, missed calls, notifications, etc.

[0089] The SIM card interface 195 is used to connect a SIM card. The SIM card can be connected to or disconnected from the electronic device 100 by inserting or removing the SIM card into or from the SIM card interface 195 .

[0090] The following describes the workflow of the electronic device 100 in conjunction with specific scenarios.

[0091] Portrait photo generation is a technology that uses a camera to capture facial, head, and shoulder information from multiple perspectives for portrait photography scenarios, and then generates diverse portrait photos based on specified scene content such as composition, style, clothing, and background. Based on the needs of portrait photography clients and users, the multi-perspective acquisition system can be deployed in professional studios, using professional cameras and professional lighting conditions, or it can be deployed using acquisition devices such as mobile phones according to front-end guidance. The subsequent image processing system can adaptively store, process, and generate images for acquisition devices and conditions of varying quality. Based on the preferred photographic content, the system can reference target photographic samples to generate portrait photos with similar poses, lighting, clothing, and backgrounds, or it can reference deep learning models that have learned specific scenes, clothing, expressions, and other information. The camera captures head and shoulder images of people in both looking and non-looking states from multiple angles, extracts relevant features of the portrait (such as facial features, face shape, neck shape, etc.) from the original portrait photos taken, and uses reference photos or reference models selected by the user to generate new portrait photos. The photos are then processed by the photography client and screened by the user to further select a rich photo album for each user to meet the portrait photography needs of diverse compositions, styles, and scenes.

[0092] There are two relatively advanced portrait photo generation solutions in the existing technology in the industry: (1) Based on the end-to-end face replacement deep learning model, this type of solution replaces the facial features of the target person with the portrait in the reference photo through a series of end-to-end face processing algorithms such as face detection, face alignment, feature extraction, and feature replacement. The deep learning model relied on is mainly trained by a face-forged portrait dataset, but the generation effect is also limited by the resolution, diversity and number of parameters of the dataset. (2) Based on the image generation large model and face processing related algorithms, the portrait photo generation system of this type of solution is built with a large image generation model with a large number of parameters as the core to build a complete generation system. Thanks to the large amount and wide range of training data, the large model can gradually transform the feature distribution of the reference image into the feature distribution of the target image, and has good generalization for the conversion process of related features such as face and body. By fine-tuning the features of the target person on the large model, the model will have the ability to transform the feature distribution of the portrait in the reference image into the feature distribution of the target person. However, due to the randomness of the large image generation model, the generation process may have unstable defects and face replacement failures. Therefore, it is necessary to combine relevant face image processing algorithms to ensure the generation quality.

[0093] The effectiveness of the first method is limited by the dataset's resolution and diversity, as well as the number of model parameters. Consequently, the model can only process limited image resolution, and facial artifacts are easily visible in complex scenes. Furthermore, the facial alignment, feature extraction, and subsequent replacement process gradually lose original facial details, resulting in blurred facial features, hard edges, and poor three-dimensionality in the resulting portraits. The lighting and shadow information of the face region in the reference image is difficult to preserve, and lighting artifacts are common. The second method, which uses text and a fine-tuned facial model to control a large model for portrait generation, is highly random and prone to artifacts. The content of the images controlled by the text is unstable, making it difficult to meet diverse and high-quality requirements. Furthermore, the detailed details of clothing and background in portraits generated entirely from noisy images depend entirely on the quality and generalization of the large model, making it difficult to ensure the appropriateness of clothing, the richness of the background, and the yield rate of the generated images. Training solely on the facial region ignores sensitive features such as hairstyle, neck shape, relative size of facial contours, relative shoulder motion, and eye gaze, resulting in similarly styled portraits with few facial details and expressions.

[0094] In response to the above technical problems, an embodiment of the present application provides a portrait photo generation method. Before replacing the portrait features of the original photo and the reference photo, the portrait features of the original photo are adjusted based on the skeletal key points of the reference photo, and then the portrait features of the adjusted photo are replaced. This can effectively avoid feature mismatch and related defects caused by direct portrait feature replacement.

[0095] FIG2 is a flow chart of a method for generating a portrait photo provided by an embodiment of the present application. The method can be applied to the above-mentioned electronic device. As shown in FIG2 , the method may specifically include:

[0096] 201. Collect original portrait photos from multiple angles;

[0097] 202. Original portrait photo preprocessing;

[0098] 203. Portrait feature extraction;

[0099] 204. Skeleton key point extraction;

[0100] 205. Generate a first target portrait photo;

[0101] 206. Replace the portrait features to generate a second target portrait photo;

[0102] 207. Portrait photo restoration and beautification;

[0103] 208. Beautify flaws;

[0104] 209.Portrait photo output.

[0105] First, the electronic device captures original portrait photos from multiple preset angles. The original portrait photos may specifically include the user's head, neck, and shoulder areas. Optionally, the original portrait photos include gaze-state portrait photos and non-gaze-state portrait photos. Specifically, referring to Figure 3, when capturing gaze-state portrait photos, the user is looking at the camera lens. When the camera lens changes position, the user needs to turn their eyes or neck to keep their gaze fixed on the lens. When capturing non-gaze-state portrait photos, the user does not need to look at the camera lens. They can always look forward or at other designated areas. The camera lens captures non-gaze-state portrait photos of the user from the front and side at different positions.

[0106] After the original portrait photo is collected, the electronic device needs to pre-process the original portrait photo in step 202, which may include noise removal, image beautification, head and shoulder and face area cropping and other pre-processing operations.

[0107] In step 203, the electronic device extracts portrait features from the pre-processed original portrait photo. A face is composed of parts such as eyes, nose, mouth, and chin. The geometric description of these parts and the structural relationship between them can be used as an important feature for identifying a face.

[0108] Users can specify a reference photo (or reference model) and then replace the reference photo's features to create the desired target portrait. The reference photo determines the background, structure, pose, and expression of the target portrait, while the original photo determines the facial features of the target portrait.

[0109] If the electronic device directly replaces the portrait features of the reference photo and the original portrait photo, the target portrait photo finally generated may have defects. For example, the face in the reference photo is fatter and the user's own face is thinner, then after the face replacement, the user's facial features are difficult to adapt to the facial contours in the reference photo. Therefore, in step 204, the electronic device first extracts the skeletal key points of the reference photo. Referring to Figure 4, the skeletal key points may specifically include the positions of the ears, glasses, nose and mouth, as well as the proportions and postures of the head, neck and shoulders. The skeletal key points can represent the contours of the face, the positions of the facial features and the overall posture and expression in the reference photo. The skeletal key points in Figure 4 are only exemplary descriptions. In actual scenarios, the electronic device can determine more skeletal key points to represent more details of the portrait in the reference photo.

[0110] In step 205, the electronic device can generate a first target portrait photo based on the portrait features of the original portrait photo and the skeleton key points of the reference photo. Specifically, the electronic device determines the skeleton key points of the reference photo as the skeleton key points of the first target portrait, and then adaptively fills the portrait features of the original portrait photo to the corresponding skeleton key point positions to generate the first target portrait photo. The first target portrait photo is a portrait photo under the original shooting scene, and its background, clothing, facial features and other contents are all determined by the original portrait photo. The skeleton key points of the first target portrait photo are the same as the skeleton key points of the reference photo. The first target portrait photo can be shown in Figure 5. Its background, clothing and specific portrait are the contents of the user's original shooting scene, which is different from the reference photo in Figure 4. However, the skeleton key points of Figure 5 are the same as the skeleton key points of Figure 4. Therefore, the geometric structure of the facial features, the head, shoulders and neck proportions and other contents in the first target portrait photo are the same as those in the reference photo.

[0111] After generating the first target portrait photo, the electronic device can replace the portrait features of the first target portrait photo and the reference photo in step 206 to generate a second target portrait photo. The background, clothing, expression, and other content in the second target portrait photo are the same as those in the reference photo, while the specific appearance of the portrait is the same as the original portrait photo. Through this process, the user can replace the face in the reference photo with their own face. In specific implementation, the original portrait photo, the reference photo, and the skeleton key points can be directly input into the image to generate a large model, resulting in the final target portrait photo.

[0112] In the embodiment of the present application, the electronic device does not directly generate the target portrait photo from the reference photo and the original portrait photo. Instead, it pre-generates a first target portrait photo of the original captured scene, with the portrait posing in the same composition and posture as the reference image. The second target portrait photo is then generated by extracting the facial features, face shape, and head and shoulders posture, combined with the reference photo. The disadvantage of direct generation is that the face in the reference image is uncontrollable and may differ from the user's face in various ways, resulting in a mismatch between the various facial parts of the final generated portrait. Therefore, extracting skeletal key points can correctly guide the generation of the target portrait photo.

[0113] In step 207, the electronic device performs face detection on the second target portrait photo and segments the head and shoulders to obtain the foreground and background of the photo. The foreground image is then processed using image processing algorithms such as super-resolution and beautification before being re-applied to the background image. Refinement operations such as fusion and blemish removal, as well as facial beautification and skin enhancement, are then performed to obtain the final portrait photo.

[0114] In step 208, the electronic device uses the beauty and shaping algorithm to process and compare the several groups of portrait photos obtained in step 207 again, screens out photos with large differences before and after processing, and then uses the face similarity detection algorithm (deep learning model or similarity numerical calculation method) to screen out photos with low similarity.

[0115] After the above process is completed, the electronic device can output the final portrait photo.

[0116] The following describes the portrait photo generation method of the present application through specific embodiments.

[0117] The portrait photo generation method of the embodiment of the present application can be applied to professional and commercial use, and the specific equipment may include a camera, a server and a display device. According to the actual store layout conditions (lighting conditions, curtain conditions, equipment conditions, venue size, etc.), the camera deployment scheme can have different variations. As shown in Figure 7, a single camera can be used for multiple rounds of shooting or multiple cameras can be used for shooting, and the shooting range can cover about 30 to 90° from the front of the face to the left and right sides. Camera shooting can be divided into two processes: on the one hand, it is the process of collecting portrait photos in the gaze state, and on the other hand, it is the process of collecting portrait photos in the non-gaze state. Referring to Figure 8, the first two rows of photos in Figure 8 are portrait photos in the gaze state collected by the camera at multiple angles. During the process of changing the angle of the camera, the user does not change the direction of gaze. The last two rows of photos in Figure 8 are portrait photos in the non-gaze state collected by the camera at multiple angles. During the process of changing the angle of the camera, the user continues to look at the camera lens. By collecting the user's gaze portrait photos and non-gaze portrait photos, the diversity of the user's facial features can be enriched, which facilitates the subsequent generation of target portrait photos.

[0118] The collected user image (original portrait photo) will be uploaded to the server. The user can select a reference photo on the display device, or flexibly specify the scene, clothing, theme, composition and other content. The server generates the target portrait photo based on the user's selection and the user image collected by the camera, and sends the target portrait photo to the display device.

[0119] In an optional embodiment, the user can batch-select multiple reference photos. The server retrieves the photos shown in FIG8 and extracts the user's portrait features. For each reference photo selected by the user, the server generates a corresponding target portrait photo based on the portrait photo generation method provided in the embodiment of this application. Referring to FIG9 , the three target portrait photos in FIG9 are generated from three different reference photos. The server batch-distributes the target portrait photos to a display device, which then displays the target portrait photos to the user.

[0120] In an optional embodiment, the portrait photo generation method of the embodiment of the present application can also be implemented through a terminal device. Taking a smartphone as an example, the user collects portrait photos of himself from multiple angles through the front / rear cameras of the smartphone. As shown in Figure 10, the user keeps the lower body still and obtains portrait photos from multiple angles by rotating his arms. The collected portrait photos also include gaze-state portrait photos and non-gaze-state portrait photos. The collection process of the gaze-state portrait photos can be referred to Figure 11. During the rotation of the smartphone or the rotation of the user's head, the user continues to look at the camera. The collection process of the non-gaze-state portrait photos can be referred to Figure 12. During the rotation of the smartphone or the rotation of the user's head, the user looks forward. After the portrait photos are collected, the user can select a reference photo on the smartphone, and the smartphone outputs the target portrait photo based on the portrait photo generation method of the embodiment of the present application.

[0121] In the embodiments of the present application, by capturing both gaze- and non-gaze-state portraits from multiple angles, it is possible to capture head and neck movements in both gaze and non-gaze states, thereby improving the generalization and yield rate of portrait photo generation, particularly the robustness of sensitive features such as eye contact and head and neck posture. The portrait photo generation method of the embodiments of the present application supports diverse compositions, generating target portrait photos after the user customizes the composition, scene, clothing, expression, and other content. The subjective effect and resolution of the generated photos are significantly superior to those of photos in the prior art.

[0122] FIG13 is a flow chart of a method for generating a portrait photo provided by an embodiment of the present application. The method can be applied to electronic devices such as the above-mentioned server and smart phones. As shown in FIG13 , the method may include:

[0123] Step 1301: Acquire multiple portrait photos to be processed and extract portrait feature information from the portrait photos to be processed;

[0124] Step 1302, obtaining a reference portrait photo and extracting skeleton key points in the reference portrait photo;

[0125] Step 1303: generating a first target portrait photo based on the portrait feature information in the portrait photo to be processed and the skeleton key points in the reference portrait photo;

[0126] Step 1304 : Generate a second target portrait photo based on the first target portrait photo and the reference photo.

[0127] Other details can be found in the above description.

[0128] Figure 14 is a schematic diagram of the structure of a portrait photo generation device provided in an embodiment of the present application. The device can be deployed in the above-mentioned electronic device. As shown in Figure 14, the device may include: a first extraction module 1401, a second extraction module 1402, a first generation module 1403, and a second generation module 1404.

[0129] The first extraction module 1401 is used to obtain a plurality of portrait photos to be processed and extract portrait feature information from the portrait photos to be processed;

[0130] The second extraction module 1402 is used to obtain a reference portrait photo and extract skeleton key points in the reference portrait photo;

[0131] A first generating module 1403 is configured to generate a first target portrait photo based on the portrait feature information in the portrait photo to be processed and the skeleton key points in the reference portrait photo;

[0132] The second generating module 1404 is configured to generate a second target portrait photo based on the first target portrait and the reference photo.

[0133] An embodiment of the present application also provides a non-transitory computer-readable storage medium, which stores computer instructions. The computer instructions enable the computer to execute the portrait photo generation method provided in the embodiment of the present application.

[0134] The above-mentioned non-temporary computer-readable storage medium can adopt any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples (non-exhaustive list) of computer-readable storage media include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (Read Only Memory; hereinafter referred to as: ROM), an erasable programmable read-only memory (Erasable Programmable Read Only Memory; hereinafter referred to as: EPROM) or flash memory, optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device.

[0135] A computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including, but not limited to, electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0136] Program code embodied on a computer readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0137] Those skilled in the art will clearly understand that the techniques in the embodiments of the present invention can be implemented using software plus a necessary general-purpose hardware platform. Based on this understanding, the technical solutions in the embodiments of the present invention, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in various embodiments of the present invention or certain portions of the embodiments.

[0138] In this specification, reference can be made to the same or similar parts between the various embodiments. In particular, for the device embodiment and the terminal embodiment, since they are basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the description in the method embodiment.

Claims

1. A method for generating a portrait photo, characterized in that: include: Acquire multiple portrait photos to be processed and extract portrait feature information from the portrait photos to be processed; Obtain a reference portrait photo and extract skeleton key points in the reference portrait photo; Generate a first target portrait photo based on the portrait feature information in the portrait photo to be processed and the skeleton key points in the reference portrait photo; A second target portrait photo is generated based on the first target portrait photo and the reference photo.

2. The method according to claim 1, characterized in that The portrait photos to be processed include: portrait photos in a gaze state, and / or portrait photos in a non-gaze state.

3. The method according to claim 2, characterized in that The step of obtaining a plurality of portrait photos to be processed includes: When the user is looking at the portrait capturing device, the portrait capturing device captures the user's gaze portraits at different angles; or / and When the user's gaze direction does not change, the portrait photos of the user in a non-gaze state at different angles are collected by the portrait collection device.

4. The method according to claim 2, characterized in that The extracting of the portrait feature information from the portrait photo to be processed includes: The portrait feature information is determined according to the gaze-state portrait photo and / or the non-gaze-state portrait photo.

5. The method according to claim 1, wherein The step of generating a first target portrait photo based on the portrait feature information in the to-be-processed portrait photo and the skeleton key points in the reference portrait photo includes: A first target portrait photo in an original shooting scene is generated by taking the skeleton key points in the reference portrait photo as a structural reference and taking the portrait feature information in the portrait photo to be processed as a content reference.

6. The method according to claim 1, characterized in that The step of generating a first target portrait photo based on the portrait feature information in the to-be-processed portrait photo and the skeleton key points in the reference portrait photo includes: Determining the skeleton key points in the reference portrait photo as the skeleton key points of the first target portrait photo; The portrait feature information in the portrait photo to be processed is adaptively filled into the skeleton key points of the first target portrait photo to generate the first target portrait photo in the original shooting scene.

7. The method according to claim 6, characterized in that The generating of the second target portrait photo based on the first target portrait photo and the reference photo includes: Identifying first portrait feature information of the first target portrait photo and second portrait feature information of the reference photo; The second portrait feature information is replaced with the first portrait feature information corresponding to the same skeletal key point to generate the second target portrait photo.

8. A portrait photo generating device, characterized in that: include: A first extraction module is used to obtain a plurality of portrait photos to be processed and extract portrait feature information from the portrait photos to be processed; A second extraction module is used to obtain a reference portrait photo and extract skeleton key points in the reference portrait photo; A first generating module is used to generate a first target portrait photo based on the portrait feature information in the portrait photo to be processed and the skeleton key points in the reference portrait photo; The second generating module is configured to generate a second target portrait photo based on the first target portrait and the reference photo.

9. An electronic device, characterized in that: The electronic device comprises a memory for storing computer program instructions and a processor for executing the program instructions, wherein when the computer program instructions are executed by the processor, the electronic device executes the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a stored program, wherein when the program is executed, the device where the computer-readable storage medium is located is controlled to execute the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Face changing method and device, electronic equipment and storage medium

    CN110889381A

  • Video face changing method and device

    CN112949360A

  • Image generation method and device, electronic equipment and computer readable medium

    CN114627529A

  • Human body grid model face changing method and device

    CN116843838A

  • Video face changing method and device, storage medium and equipment

    CN116993576A