Method for Face Synthesis Based on Expanded Processing Frames and Predicted Feature Points
Patent Information
- Application Number
- KR1020260040287
- Authority / Receiving Office
- KR · KR
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-11-17
- Filing Date
- 2026-03-06
- Publication Date
- 2026-09-21
Smart Images

Figure PAT00007_ABST
Abstract
Description
Technology Field
[0001] The technical concept of the present invention relates to a processing frame expansion and predicted feature point-based face synthesis method, wherein a first image and a second image containing a target to be replaced from the first image are received as input data, a processing frame including a shooting frame and an outer frame formed outside the shooting frame is set for the first image, a plurality of feature points are obtained from the first image and the second image respectively, and a face replacement is performed by aligning and transforming a face included in the second image with the face region of the first image based on the feature points and synthesizing them, and if it is determined that at least one of the first feature points exists in the outer frame, a predicted feature point corresponding to outside the shooting frame is calculated in the processing frame to continuously perform face replacement. Background Technology
[0002] Face synthesis (face replacement) technology allows for the natural replacement of a person's face with another face within video content, and is being utilized in various fields such as entertainment, broadcasting / video editing, and non-face-to-face communication. However, conventional face synthesis technology has faced problems where the quality of the synthesis rapidly deteriorates or the face replacement is undone when feature point detection becomes unstable due to the shooting environment or the movement of the subject during the process of matching and synthesis based on facial feature points (landmarks) detected in the video.
[0003] For example, some feature points may momentarily go undetected if part of the face is obscured by the subject's hair, hands, or mask, or if the face rotates or tilts. Additionally, tracking of feature points may be interrupted if the subject moves to the edge of the screen, causing part of the face to move out of the frame, or if the subject moves away from the camera, resulting in a small face being captured. Furthermore, if detection reliability is temporarily reduced due to changes in lighting, blur, noise, etc., the updating of inter-frame matching parameters becomes unstable, which may cause the composite result to shake, the position to jump abruptly, or the composite itself to stop.
[0004] In particular, in environments where face replacement is performed in real-time or near-real-time, the replacement must be maintained continuously even if such undetected / false detections or frame out-of-frame errors occur momentarily, so that users can view the resulting video without any sense of unfamiliarity. Nevertheless, conventional technology lacks the means to reliably compensate when feature points deviate even partially from the frame or when key feature points are omitted due to occlusion, resulting in persistent inconveniences such as the alignment of the target face being disrupted or the synthesis being interrupted.
[0005] Therefore, a technology is required that can prevent rapid degradation of synthesis quality and disassembly by stably determining frame deviations and calculating / correcting predicted feature points corresponding to outside the frame to stably update matching parameters, thereby enabling continuous face replacement even when some feature points within the image are not detected or are outside the captured frame. The problem to be solved
[0006] The problem that the technical concept of the present invention aims to solve is to provide a face synthesis method that allows the face replacement to be maintained continuously without interruption during the process of synthesizing a person's face in an image by replacing it with another face, even if part of the face is obscured by hair or hands, if the face moves to the edge of the screen and moves even partially outside the shooting frame, or if the face moves away and feature point detection becomes unstable.
[0007] Furthermore, the present invention provides a face synthesis method capable of ensuring continuity of face registration and synthesis by additionally setting an outer frame formed outside the shooting frame to set a processing frame including the shooting frame and the outer frame, and calculating predicted feature points corresponding to outside the frame in the processing frame even if some feature points are omitted within the shooting frame or deviate outside the frame.
[0008] In addition, the present invention provides a face synthesis method capable of preventing the phenomenon in which face replacement is unnecessarily released or resumed due to instantaneous noise or shaking by mapping feature point coordinates to a multidimensional embedding space and calculating a distance function and a frame deviation index for the boundary of the captured frame, and determining the deviation state according to a hysteresis condition in which the deviation threshold and the return threshold are set differently, so as to stably determine the situation in which a feature point deviates from the captured frame without false detection.
[0009] Furthermore, the present invention provides a face synthesis method that enables stable continuation of matching parameter updates and face replacement even in situations of frame deviation or partial occlusion by calculating predicted coordinates using the relative positional relationship of feature points remaining within the captured frame for feature points that have deviated from the frame, verifying and correcting based on whether the calculated predicted coordinates correspond to the relative positional relationship between the remaining feature points, and then calculating predicted feature points. means of solving the problem
[0010] A face synthesis method performed by an electronic device according to an embodiment of the present disclosure may include: receiving a first image and a second image containing a target to be replaced in the first image; setting a shooting frame for the first image and additionally setting an outer frame formed outside the shooting frame to set a processing frame including the shooting frame and the outer frame; acquiring a first feature point in the first image and acquiring a second feature point from the second image; replacing and outputting a face in the first image with a face in the second image based on the first feature point and the second feature point; when it is determined that at least one of the first feature points exists in the outer frame, calculating a predicted feature point corresponding to outside the shooting frame in the processing frame based on the relative positional relationship between feature points remaining within the shooting frame; and continuously performing face replacement based on the predicted feature point.
[0011] According to one embodiment, the step of acquiring the first feature point and acquiring the second feature point from the second image may include: detecting a plurality of the first feature points corresponding to the face, neck, and shoulders of the person in the processing frame of the first image; if at least one of the first feature points is not detected or the detection reliability is less than a reference value, additionally detecting an auxiliary feature point corresponding to the upper body of the person; and detecting a plurality of second feature points corresponding to the face included in the second image.
[0012] According to one embodiment, the step of replacing the face of the first image with the face of the second image and outputting it may include: a step of calculating a first feature line by connecting at least two of the first feature points to each other and calculating a first feature surface by connecting at least three of the first feature points to each other; a step of calculating a second feature line by connecting at least two of the second feature points to each other and calculating a second feature surface by connecting at least three of the second feature points to each other; a step of calculating a matching parameter including at least one of the position, size, and direction of the face included in the second image based on the correspondence relationship between the first feature line or the first feature surface and the second feature line or the second feature surface; a step of calculating a motion vector, a velocity vector, or an acceleration vector based on the inter-frame change of the first feature points and updating the matching parameter frame by frame; and a step of synthesizing the face included in the second image, transformed according to the updated matching parameter, to the face region of the first image so that the face of the first image appears as the face included in the second image. there is.
[0013] According to one embodiment, the face synthesis method further includes a step of determining that at least one of the first feature points deviates from the shooting frame, and the step of determining that at least one of the first feature points deviates from the shooting frame may include: a step of mapping the coordinate values of the first feature point to a multidimensional embedding space; a step of calculating a distance function for the boundary of the shooting frame and calculating a frame deviation index based on the distance function; and a step of determining a deviation state according to a hysteresis condition in which, if the frame deviation index satisfies a deviation threshold, it is determined as a deviation, but the return threshold is set to be different from the deviation threshold.
[0014] According to one embodiment, the step of calculating the frame deviation index comprises: calculating the distance function for the boundary of the captured frame based on the boundary of the captured frame and the feature point distribution of the first feature point within the embedding space; normalizing the distance function to calculate the deviation reliability; and calculating the frame deviation index based on the cumulative value, moving average value, or smoothed value of the deviation reliability. The feature point distribution includes at least one of the distribution center, variance, main distribution direction of the first feature point, or the ratio of the first feature point located outside the captured frame. The distance function may be calculated as a weighted combined value of boundary distance values weighted by the feature point distribution.
[0015] According to one embodiment, the step of calculating the predicted feature point may include: when the deviation feature point among the first feature points is determined to be outside the shooting frame and is a pair of corresponding feature points, setting a symmetric feature point remaining within the shooting frame as a reference feature point and calculating the predicted coordinates for the deviation feature point based on the relative positional relationship between the reference feature point and the face center axis; verifying the predicted coordinates based on whether the predicted coordinates correspond to the relative positional relationship between the first feature points remaining within the shooting frame excluding the reference feature point, and correcting the predicted coordinates according to the verification result; and calculating the feature point corresponding to the corrected predicted coordinates as the predicted feature point.
[0016] According to one embodiment, the step of calculating the predicted feature point may include: when the deviation feature point among the first feature points determined to be outside the shooting frame is a feature point that does not correspond to a pair, setting at least one of the first feature points remaining within the shooting frame as a reference feature point, and calculating the predicted coordinates for the deviation feature point based on at least one of a relative distance, angle, or ratio relationship with the reference feature point; verifying the predicted coordinates based on whether the predicted coordinates correspond to the relative positional relationship between the first feature points remaining within the shooting frame excluding the reference feature point, and correcting the predicted coordinates according to the verification result; and calculating the feature point corresponding to the corrected predicted coordinates as the predicted feature point. Effects of the invention
[0017] The face synthesis method according to an embodiment of the present invention can improve the continuity of face synthesis and user-perceived quality by ensuring that the face replacement is not interrupted even if partial obscuration by hair or hands, rotation or tilting of the face, deviation from the shooting frame due to movement of the subject, or instability in feature point detection due to changes in distance occurs during the process of synthesizing by replacing a person's face with another face in an image.
[0018] In addition, an outer frame can be additionally set outside the shooting frame to establish a processing frame that includes the shooting frame and the outer frame, and by calculating predicted feature points corresponding outside the shooting frame to continuously perform face replacement, the phenomenon of the composite unraveling or shaking abruptly due to the loss of feature points near the edges of the frame can be mitigated.
[0019] Additionally, the coordinate values of feature points are mapped into a multidimensional embedding space, and a distance function and frame deviation index for the captured frame boundary are calculated. Since the deviation state can be determined according to a hysteresis condition in which the deviation threshold and return threshold are set to be different, it is possible to prevent the deviation state from frequently switching due to instantaneous noise, blur, or false detection, and to improve the updating of matching parameters and the stability of the synthesis results.
[0020] Furthermore, since the system can be configured to calculate predicted coordinates of feature points that have deviated from the frame by utilizing the relative positional relationship between feature points remaining within the captured frame, and to calculate predicted feature points after verification and correction based on whether the calculated predicted coordinates correspond to the relative positional relationship between the remaining feature points, the position, size, and orientation of the synthesized face are naturally maintained even in situations of frame deviation or partial occlusion, thereby effectively reducing abrupt jumps, snapbacks, or uncompositing of the synthesized result.
[0021] The effects obtainable from the exemplary embodiments of the present disclosure are not limited to those mentioned above, and other unmentioned effects can be clearly derived and understood by those skilled in the art to which the exemplary embodiments of the present disclosure belong from the description below. That is, unintended effects resulting from the implementation of the exemplary embodiments of the present disclosure can also be derived by those skilled in the art from the exemplary embodiments of the present disclosure. Brief explanation of the drawing
[0022] FIG. 1 is a block diagram illustrating a connection structure between an external terminal and an electronic device according to an embodiment of the present invention. FIG. 2 is a block diagram illustrating an electronic device according to an embodiment of the present invention. FIG. 3 is a block diagram illustrating a processor according to an embodiment of the present invention. FIG. 4 is a diagram illustrating the process of setting a processing frame according to an embodiment of the present invention. FIG. 5 is a diagram illustrating the process of obtaining a first feature point and a second feature point according to an embodiment of the present invention. FIG. 6 is a diagram illustrating the process of calculating predicted feature points according to an embodiment of the present invention. FIG. 7 is a diagram comparing the state before and after performing the substitution according to an embodiment of the present invention. FIG. 8 is a flowchart illustrating a face synthesis method according to an embodiment of the present invention. FIG. 9 is a flowchart illustrating the steps of obtaining a first feature point and a second feature point according to an embodiment of the present invention. FIG. 10 is a flowchart illustrating the step of replacing the face of a first image with the face of a second image and outputting it according to an embodiment of the present invention. FIG. 11 is a flowchart illustrating the steps of determining whether a first feature point according to an embodiment of the present invention has moved out of the shooting frame and continuously performing face replacement. FIG. 12 is a flowchart illustrating the step of determining whether a first feature point is outside the shooting frame according to an embodiment of the present invention. FIG. 13 is a flowchart illustrating the steps for calculating predicted feature points according to an embodiment of the present invention. Specific details for implementing the invention
[0023] Before specifically describing the present disclosure, the method of description in the specification and drawings is described.
[0024] First, the terms used in this specification and claims have been selected based on general terms considering their functions in the various embodiments of this disclosure. However, these terms may vary depending on the intent of those skilled in the art, legal or technical interpretations, and the emergence of new technologies. Additionally, some terms have been arbitrarily selected by the applicant. Such terms may be interpreted according to the meanings defined in this specification; in the absence of specific definitions, they may be interpreted based on the overall content of this specification and common technical knowledge in the relevant field.
[0025] In addition, the same reference numbers or symbols described in each drawing attached to this specification represent parts or components that perform substantially the same function. For convenience of explanation and understanding, the same reference numbers or symbols are used to describe different embodiments. That is, even if components having the same reference number are all depicted in multiple drawings, the multiple drawings do not imply a single embodiment.
[0026] Additionally, in this specification and claims, terms including ordinal numbers, such as "first," "second," etc., may be used to distinguish between components. These ordinal numbers are used to distinguish identical or similar components from one another, and the meaning of the terms should not be limited by the use of such ordinal numbers. For example, the order of use or arrangement of components combined with such ordinal numbers should not be restricted by the number. If necessary, each ordinal number may be used interchangeably.
[0027] In this specification, singular expressions include plural expressions unless the context clearly indicates otherwise. In this application, terms such as "comprising" or "consisting of" are intended to specify the existence of the features, numbers, steps, actions, components, parts, or combinations thereof described in the specification, and should be understood as not precluding the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.
[0028] In the embodiments of the present disclosure, terms such as "module," "unit," "part," etc. are used to refer to a component that performs at least one function or operation, and such component may be implemented in hardware or software, or in a combination of hardware and software. Additionally, a plurality of "modules," "units," "parts," etc. may be integrated into at least one module or chip and implemented as at least one processor, except where each needs to be implemented in specific individual hardware.
[0029] Furthermore, in the embodiments of the present disclosure, when a part is described as being connected to another part, this includes not only a direct connection but also an indirect connection through another medium. Additionally, the meaning that a part includes a certain component implies that, unless specifically stated otherwise, it does not exclude other components but may include additional components.
[0030] Hereinafter, prior to describing the present invention, the term "image" in this specification may be used as a concept including still images and video. Here, a still image may refer to content expressed as a single frame, such as a commonly used still image, and a video may refer to content composed of multiple frames in which the still images are arranged in chronological order.
[0031] Hereinafter, a face synthesis method according to an embodiment of the present invention will be described in detail using the attached drawings.
[0032] FIG. 1 is a block diagram illustrating a connection structure between an external terminal and an electronic device according to an embodiment of the present invention.
[0033] The electronic device (100) can transmit and receive information with an external terminal (200) through a network. The external terminal (200) may be a user terminal used by a user, may include a terminal for monitoring or outputting synthesis results in real-time or near-real-time in a broadcasting station or content production environment, and may refer to various types of terminals capable of receiving and displaying face synthesis result images from the electronic device (100). Additionally, the electronic device (100) may be implemented as at least one of a physical server, a terminal device, an edge device, or a computing system implemented in a cloud computing environment.
[0034] The electronic device (100) can receive a first image containing a target for face replacement and a second image containing a target to be used for face replacement from an external terminal (200). The electronic device (100) can perform face replacement based on the received first image and second image, and transmit a composite result image with the face replacement performed to the external terminal (200).
[0035] Additionally, the external terminal (200) can transmit control information related to changing the synthesis target, changing the synthesis option, or starting or stopping the synthesis to the electronic device (100), and the electronic device (100) can perform face replacement or change the method of providing the synthesis result image according to the control information.
[0036] That is, the electronic device (100) receives a first image and a second image from an external terminal (200), and can provide a composite result image in which face replacement is performed based on the received images to the external terminal (200) in real-time or near-real-time.
[0037] Hereinafter, an electronic device for performing a face synthesis method according to an embodiment of the present invention will be described in detail using the attached drawings.
[0038] FIG. 2 is a block diagram illustrating an electronic device according to an embodiment of the present invention.
[0039] Hereinafter, the present specification may be described with reference to FIG. 1.
[0040] The electronic device (100) may be configured as a server or electronic device that performs outputting a face of a desired image by synthesizing it with a first image received from an external terminal (200), and may include a processor (110), RAM (Random Access Memory) (120), storage (130), and a communication unit (140).
[0041] The processor (110) can control the overall operation of the electronic device (100). The processor (110) may include a single processor core or a CPU (Central Processing Unit) that includes multiple processor cores.
[0042] The electronic device (100) may include one or more processors (110). The processor (110) may process or execute programs, data, or instructions stored in storage (130).
[0043] Specifically, the processor (110) can perform face replacement processing based on a first image containing a target for face replacement received from an external terminal (200) and a second image containing a target to be used for face replacement. For example, the processor (110) can perform face replacement by extracting feature points from the first image and the second image, calculating and updating matching parameters based on the extracted feature points, and synthesizing the face included in the second image by matching and transforming it to the face region of the first image.
[0044] Additionally, the processor (110) can set a processing frame including an outer frame formed outside the shooting frame so that face replacement is continuously maintained even when part of the feature point in the first image is not detected or goes outside the shooting frame during face replacement, and calculate and correct the predicted feature point corresponding to outside the shooting frame in the processing frame to continuously perform face replacement.
[0045] Furthermore, the processor (110) can generate a composite result image in which face replacement has been performed in real-time or near-real-time and provide it to an external terminal (200), and, if necessary, can control the method of providing the composite result or the composite processing according to the composite option or control information.
[0046] RAM (120) can temporarily store programs, data, or instructions. For example, programs and / or data stored in storage (130) may be temporarily stored in RAM (120) under the control of the processor (110) or boot code. For example, RAM (120) may include DRAM, SRAM, SDRAM, etc.
[0047] Storage (130) can store an OS (Operating System), programs, and data for performing face synthesis processing (e.g., a first image containing a target for face replacement, a second image containing a target to be used for face replacement, processing frame setting information including a captured frame and an outer frame, feature point detection results and auxiliary feature point information, matching parameters and frame-by-frame update information, threshold values and hysteresis condition information for determining frame deviation, reference information for calculating predicted feature points, a synthesis result image or intermediate data for generating a synthesis result, etc.).
[0048] Storage (130) may include ROM, flash memory, PRAM, MRAM, RRAM, FRAM, etc., and may be implemented as HDD (Hard Disk Drive), SSD (Solid State Drive), etc.
[0049] The storage (130) may store programs necessary for the electronic device (100) to perform face synthesis processing, and the processor (110) may load the programs stored in the storage (130) to perform face replacement and synthesis functions.
[0050] More specifically, the storage (130) may also store setting information used to maintain the continuity of face replacement during the face synthesis process. For example, the storage (130) may store processing frame setting information including a captured frame and an outer frame, feature point detection settings and reference values, setting information for calculating matching parameters and updating per frame, threshold values and hysteresis conditions for determining frame deviation, and reference information used for calculating and verifying predicted feature points.
[0051] Additionally, the storage (130) may store intermediate processing data such as first feature points and second feature points extracted per video or per frame, auxiliary feature points, feature lines or feature planes, matching parameters and update results, frame deviation indicators and deviation states, predicted coordinates and predicted feature points, and verification and / or correction results of predicted coordinates. This intermediate processing data can be used to ensure that the matching parameters are stably updated and the face replacement is not interrupted even when some feature points are not detected or the captured frame is out of the way during the synthesis process.
[0052] Furthermore, the storage (130) may store management information related to the target face used for face replacement, and the management information may include at least one of a target identifier, a reference input of the target face, a target feature point, and / or matching reference information. Additionally, the management information may be mapped and stored in the storage (130) per processing session, thereby supporting the maintenance of consistency in the synthesis results during the process of continuously applying the same target face.
[0053] The communication unit (140) can perform data transmission and reception through communication with an external terminal (200) or an external server. For example, the communication unit (140) can transmit and receive data by various communication methods. For example, the communication unit (140) can perform communication by, for example, 3G, LTE, Wi-Fi, Bluetooth, BLE (Bluetooth Low Energy), Zigbee, NFC (Near Field Communication), and communication methods using ultrasound, and can include wired communication, wireless communication, short-range communication, and long-range communication.
[0054] The communication unit (140) can receive a first image containing a target for face replacement and a second image containing a target to be used for face replacement from an external terminal (200) and transmit them to an electronic device (100). Additionally, the communication unit (140) can transmit a composite result image in which face replacement has been performed at the electronic device (100) to the external terminal (200) so that the composite result can be checked in real-time or near-real-time at the external terminal (200).
[0055] Additionally, the communication unit (140) can receive control information related to changing the synthesis target, starting or stopping the synthesis, or changing the synthesis option from an external terminal (200) and transmit it to the processor (110), and can transmit the processing result of the processor (110) to the external terminal (200).
[0056] Hereinafter, a processor of an electronic device that performs a face synthesis method according to an embodiment of the present invention will be described in detail using the attached drawings.
[0057] FIG. 3 is a block diagram illustrating a processor according to an embodiment of the present invention.
[0058] FIG. 4 is a diagram illustrating the process of setting a processing frame according to an embodiment of the present invention.
[0059] FIG. 5 is a diagram illustrating the process of obtaining a first feature point and a second feature point according to an embodiment of the present invention.
[0060] FIG. 6 is a diagram illustrating the process of calculating predicted feature points according to an embodiment of the present invention.
[0061] FIG. 7 is a diagram comparing the state before and after performing the substitution according to an embodiment of the present invention.
[0062] Hereinafter, the present specification may be described with reference to FIGS. 1 and FIGS. 2.
[0063] A processor (110) according to an embodiment of the present invention may include a frame setting module (111) for setting a processing frame by setting a shooting frame and an outer frame, a face detection module (112) for detecting a face region in a first image and a second image, a feature point extraction module (113) for detecting a first feature point and a second feature point from the first image and the second image, a matching replacement processing module (114) for performing face replacement by matching and synthesizing the face of the second image to the face region of the first image based on the first feature point and the second feature point, an out-of-bounds determination module (115) for determining whether the first feature point is out of the shooting frame, and a predicted feature point calculation module (116) for calculating predicted coordinates based on the relative positional relationship between remaining feature points within the shooting frame for the out-of-bounds feature point, and calculating the predicted feature point after verifying and correcting the predicted coordinates.
[0064] The frame setting module (111) may mean a module that sets a shooting frame (PF) for the first image (V1) and additionally sets an outer frame (OF) to extend the processing range outside the shooting frame (PF), thereby setting a processing frame (TF) that includes the shooting frame (PF) and the outer frame (OF).
[0065] More specifically, the frame setting module (111) can set a shooting frame (PF) based on the entire screen of the first image (V1) or the area in the first image (V1) where a composite target exists. The shooting frame (PF) may represent a basic processing range in which face replacement is performed in the first image (V1), and may be set to correspond to the frame size of the input image or the terminal screen ratio.
[0066] The frame setting module (111) can set the processing frame (TF) to surround the shooting frame (PF) in order to process a wider range than the shooting frame (PF). For example, the processing frame (TF) may include a range extended outward by a preset pixel distance from the boundary of the shooting frame (PF), and may have the same or different extension widths in the up, down, left, and right directions. Additionally, the extension width of the processing frame (TF) may be determined variably according to the setting value, the resolution of the first image (V1), the size of the person, or the speed of movement.
[0067] The Outer Frame (OF) can be defined not as an independent, separate image frame, but as a regional concept referring to the area existing between the Capture Frame (PF) and the Processing Frame (TF). That is, the Outer Frame (OF) refers to the area extending from the outer edge of the Capture Frame (PF) to the boundary of the Processing Frame (TF), and can function as an auxiliary processing area for capturing the trend of feature points moving out of the Capture Frame (PF) or for calculating predicted feature points (OP) corresponding to the area outside the Capture Frame (PF).
[0068] The reason for extending and setting the processing frame (TF) separately from the shooting frame (PF) in this manner is to ensure that face replacement is not interrupted even when the face or upper body in the first image (V1) approaches the boundary of the shooting frame (PF) or temporarily moves out of the shooting frame (PF). For example, if a person moves and part of the face goes out of the shooting frame (PF), if some first feature points (P1) are obscured by hands or hair, or if the face is shifted to the edge of the screen due to a change in camera angle, and only the shooting frame (PF) is used for processing, feature points may be rapidly lost, causing the alignment to become unstable or the face replacement to stop. On the other hand, if the processing frame (TF) is set wider than the shooting frame (PF), the area corresponding to the feature points moving out of the shooting frame (PF) can be processed together in the outer frame (OF), and predicted feature points (OP) corresponding to outside the shooting frame (PF) can be calculated to support the continuous maintenance of alignment and replacement processing.
[0069] The face detection module (112) may refer to a module that detects a face region that serves as a reference for face replacement in the first image (V1) and the second image (V2), and provides the detected face region information to the feature point extraction module (113) and the matching replacement processing module (114) to be described later.
[0070] More specifically, the face detection module (112) can detect an area containing a face in the first image (V1) and generate face area information indicating the location and size of the face area, and can also detect an area containing a target face in the second image (V2) and generate face area information. For example, the face area information may include at least one of a bounding box defining the face area, a polygonal area defining the face contour, or a face center point and a representative size.
[0071] Additionally, the face detection module (112) can generate control information to specify the section where the face is located within the processing frame (TF) using the detected face region information, or to expand the target range for feature point detection to an upper body region including the face, neck, and shoulders. Accordingly, the feature point extraction module (113) can more stably detect the first feature point (P1) and the second feature point (P2) based on the face region, and the alignment replacement processing module (114) can perform a process of aligning and synthesizing the face of the second image (V2) with the face region of the first image (V1) by referring to the detected face region information.
[0072] The feature point extraction module (113) may mean a module that detects feature points corresponding to the upper body of a person from the first image (V1) and the second image (V2) to obtain the first feature point (P1) and the second feature point (P2).
[0073] More specifically, the feature point extraction module (113) can detect a plurality of first feature points (P1) corresponding to the upper body including the face of a person based on the face region detected in the first image (V1) by the face detection module (112). Additionally, the feature point extraction module (113) can detect a plurality of second feature points (P2) corresponding to the upper body including the face of a target person based on the face region detected in the second image (V2) by the face detection module (112).
[0074] The first feature point (P1) and the second feature point (P2) may be coordinate-based feature points set for the upper body, including the face, neck, and shoulders. For example, the first feature point (P1) and the second feature point (P2) may include at least one of the following points in the face area: both eyes, nose, both ends of the mouth, the center of the lips, the tip of the chin, the end point of the eyebrows, the end point of the nostrils, the protruding point of the cheekbones, the center of the forehead, and / or multiple points on the face contour. Additionally, in the neck and shoulder area, they may include at least one of the following points: the center or both ends of the clavicle, the center of the neck, the end line of both shoulders, and / or multiple points on the shoulder contour. Such feature points are not necessarily limited to the eyes, nose, and / or mouth, but may be extended and set to any point capable of stably representing the face shape or the posture of the upper body.
[0075] The feature point extraction module (113) can calculate the detection reliability along with the coordinate values for each feature point during the feature point detection process, and can additionally detect auxiliary feature points corresponding to the upper body if the detection reliability is below a reference value or if some feature points are not detected. For example, the auxiliary feature points may include at least one of additional points on the face contour, additional points on the eyebrows, the center point of the cheek, the middle point of the jawline, the contour point of the neckline, and / or additional points on the shoulder contour. Through this, the feature point extraction module (113) can support the stable maintenance of the feature point set even if part of the person is obscured or a change in the viewing angle occurs.
[0076] The alignment replacement processing module (114) may refer to a module that aligns two faces using a first feature point (P1) detected in a first image (V1) and a second feature point (P2) detected in a second image (V2), and synthesizes the face of the second image (V2) into the face area of the first image (V1) to generate a result image (V1').
[0077] More specifically, the alignment replacement processing module (114) can establish a correspondence relationship between the first feature point (P1) and the second feature point (P2) that correspond to the face. For example, it can establish a correspondence relationship between feature points that stably represent the shape of the face, such as the area around both eyes, the area around the nose, the area around the mouth, the jawline, the facial contour, the center of the forehead, and the cheekbone area, and feature points corresponding to the neck or shoulders may also be referenced together for reference stabilization when the face position shakes.
[0078] The alignment replacement processing module (114) can create a first feature line or a first feature surface by connecting at least some of the feature points with established correspondences, and can also create a corresponding second feature line or a second feature surface at the second feature point (P2). For example, it can construct a line connecting the two eyes, a line connecting the center of the nose and the mouth, a polygonal line formed along the facial contour, or a feature surface in the shape of a triangle or a square defined by some feature points of the eyes / nose / mouth.
[0079] The alignment replacement processing module (114) can calculate alignment parameters based on the correspondence between the first feature line or first feature plane and the second feature line or second feature plane. The alignment parameters may include at least one of position, size, and / or direction as transformation information for aligning the face of the second image (V2) with the face region of the first image (V1), and may be expressed as at least one of translation transformation, similarity transformation, affine transformation, and / or homography transformation depending on the implementation. Additionally, if there is a large change in face pose, the alignment parameters may be calculated by including a 3D pose estimation result.
[0080] The alignment replacement processing module (114) can calculate at least one of a motion vector, a velocity vector, and / or an acceleration vector based on the inter-frame change of the first feature point (P1), and use this to update the alignment parameters frame by frame. For example, the phenomenon of the face momentarily bouncing or shaking can be mitigated by updating by weighted combining the alignment parameters of the previous frame and the newly calculated alignment parameters in the current frame, or by applying a moving average or filtering.
[0081] That is, the matching replacement processing module (114) can transform the face of the second image (V2) according to the updated matching parameters, set an area in the first image (V1) to which the face replacement is to be applied, and synthesize the transformed face in that area. The area to which the face replacement is to be applied can be defined as a face contour mask based on the first feature point (P1), and at least one of blending, boundary smoothing, alpha mixing, color correction, and / or brightness correction can be applied for natural joining at the edge boundary. As a result, a result image (V1') can be output so that the face of the first image (V1) appears as the face of the second image (V2).
[0082] The deviation determination module (115) may mean a module that determines whether at least one of the first feature points (P1) obtained in the first image (V1) is outside the shooting frame (PF) and exists in the outer frame (OF) area, and generates information for determining the frame deviation state.
[0083] More specifically, the deviation determination module (115) can evaluate the relationship with the boundary of the shooting frame (PF) using the coordinate values of the first feature point (P1). For example, it can simply compare whether the first feature point (P1) is located inside the shooting frame (PF), or it can more reliably determine whether there is a deviation by calculating the distance from the boundary of the shooting frame (PF) or the degree of deviation outside the boundary, including whether there is a presence within the outer frame (OF) area.
[0084] The deviation determination module (115) maps the coordinate values of the first feature point (P1) into a multidimensional embedding space, calculates a distance function for the boundary of the captured frame (PF) based on the feature point distribution in the embedding space, and can calculate a frame deviation index based on the distance function. Here, the feature point distribution may include at least one of the distribution center, dispersion, main distribution direction of the first feature point (P1), and / or the ratio of the first feature point (P1) located outside the captured frame (PF), and the distance function may be calculated as a weighted combination of boundary distance values weighted by this feature point distribution.
[0085] The descent determination module (115) can determine whether to descent by comparing the calculated frame descent indicator with a threshold value, and can apply a hysteresis condition to prevent the descent state from fluctuating repeatedly near the frame boundary. For example, the descent threshold value and the return threshold value can be set differently so that the frame descent indicator satisfies the descent threshold value, the system switches to a descent state, and then maintains the descent state until the frame descent indicator satisfies the return threshold value.
[0086] The result of the determination of the deviation determination module (115) can be used as a trigger for calculating a predicted feature point (OP), and if it is determined that the deviation state is outside the shooting frame (PF) and exists in the outer frame (OF) area, it can be provided to the predicted feature point calculation module (116) to calculate a predicted feature point (OP) corresponding outside the shooting frame (PF) based on the relative positional relationship with the remaining feature point (RP) within the shooting frame (PF).
[0087] The predicted feature point calculation module (116) may mean a module that calculates a predicted feature point (OP) corresponding to a feature point determined to be out of the shooting frame (PF) among the first feature points (P1) based on the judgment result of the deviation judgment module (115).
[0088] More specifically, the prediction feature point calculation module (116) can calculate a prediction coordinate corresponding to outside the shooting frame (PF) based on the relative positional relationship between the feature points (RP) remaining within the shooting frame (PF), and generate a feature point corresponding to the calculated prediction coordinate as a prediction feature point (OP).
[0089] In one embodiment, the prediction feature point calculation module (116) may vary the method of calculating prediction coordinates depending on whether the feature point determined to be outside the shooting frame (PF) is a pair of corresponding feature points. For example, for feature points that correspond to each other due to the left-right symmetry relationship of the face, the symmetric feature point remaining within the shooting frame (PF) is set as the reference, and the prediction coordinates of the feature point determined to be outside the shooting frame (PF) can be calculated based on the relative positional relationship between the reference feature point and the face center axis.
[0090] In addition, if a feature point determined to be outside the shooting frame (PF) is a feature point that does not correspond to a pair, the prediction feature point calculation module (116) can set at least one of the feature points (RP) remaining within the shooting frame (PF) as a reference feature point and calculate prediction coordinates based on at least one of the relative distance, angle, and ratio relationship with the reference feature point.
[0091] The prediction feature point calculation module (116) can verify the prediction coordinates based on whether the calculated prediction coordinates correspond to the relative positional relationship between the remaining feature points (RP) within the shooting frame (PF) excluding the reference feature point, and can calculate the prediction feature point (OP) corresponding to the corrected prediction coordinates after correcting the prediction coordinates according to the verification result.
[0092] Finally, the prediction feature point calculation module (116) can transmit the calculated prediction feature point (OP) to the alignment replacement processing module (114), and the alignment replacement processing module (114) can utilize the prediction feature point (OP) as part of the first feature point (P1) to continuously perform calculation and updating of alignment parameters, thereby enabling face replacement to continue even if feature point deviation occurs in the first image (V1).
[0093] Below, a face synthesis method will be explained in detail using the attached drawings.
[0094] FIG. 8 is a flowchart illustrating a face synthesis method according to an embodiment of the present invention.
[0095] FIG. 9 is a flowchart illustrating the steps of obtaining a first feature point and a second feature point according to an embodiment of the present invention.
[0096] FIG. 10 is a flowchart illustrating the step of replacing the face of a first image with the face of a second image and outputting it according to an embodiment of the present invention.
[0097] FIG. 11 is a flowchart illustrating the steps of determining whether a first feature point according to an embodiment of the present invention has moved out of the shooting frame and continuously performing face replacement.
[0098] FIG. 12 is a flowchart illustrating the step of determining whether a first feature point is outside the shooting frame according to an embodiment of the present invention.
[0099] FIG. 13 is a flowchart illustrating the steps for calculating predicted feature points according to an embodiment of the present invention.
[0100] Hereinafter, the present specification may be described with reference to FIGS. 1 to 7.
[0101] A face synthesis method according to an embodiment of the present invention may include the steps of receiving a first image and a second image (S110), setting a processing frame (S120), acquiring a first feature point and a second feature point (S130), replacing the face of the first image with the face of the second image and outputting it (S140), and determining whether the first feature point has moved out of the shooting frame and continuously performing face replacement (S150).
[0102] The step of receiving the first image and the second image (S110) is a step in which the electronic device (100) receives the first image (V1) and the second image (V2) from an external terminal (200), and the electronic device (100) can receive the first image (V1) containing the target for face replacement and the second image (V2) containing the target to be used for face replacement.
[0103] An external terminal (200) can transmit a first image (V1) obtained from a shooting device or storage medium to an electronic device (100), and the electronic device (100) can use the received first image (V1) as input data for performing face replacement processing. Additionally, the external terminal (200) can transmit a second image (V2) containing a target to be used for face replacement to the electronic device (100), and the electronic device (100) can use the received second image (V2) as target data to be used for face replacement processing.
[0104] The step of setting a processing frame (S120) may mean setting a shooting frame (PF) for the received first image (V1) and setting a processing frame (TF) by extending the processing range outside the shooting frame (PF).
[0105] More specifically, the frame setting module (111) may set a shooting frame (PF) based on the entire screen of the first image (V1) or set a shooting frame (PF) based on an area in the first image (V1) where a person is determined to be present. The shooting frame (PF) may refer to an image range where face replacement processing is basically performed.
[0106] The frame setting module (111) can set the processing frame (TF) to include an area extended outward by a predetermined distance from the boundary of the shooting frame (PF) in order to process a wider range than the shooting frame (PF). The processing frame (TF) can be set to surround the shooting frame (PF) and can have the same or different extension widths in the up, down, left, and right directions, and the extension width may be variably determined according to the resolution or setting value of the first image (V1).
[0107] In this way, by setting the processing frame (TF) wider than the shooting frame (PF), even if some of the first feature points (P1) are outside the shooting frame (PF), if the corresponding feature points are within the range of the processing frame (TF), the electronic device (100) can estimate the position of the feature points outside the shooting frame (PF) or calculate the predicted feature points (OP) based on the relative positional relationship with the remaining feature points (RP) within the shooting frame (PF). Accordingly, the updating of the matching parameters and face replacement can be performed continuously in subsequent steps, so that the phenomenon of the face replacement being interrupted can be mitigated even if the person moves or the composition changes.
[0108] The step of acquiring a first feature point and a second feature point (S130) may mean detecting a face region in a first image (V1), acquiring a first feature point (P1) corresponding to an upper body including the neck and shoulders around the detected face region, and acquiring a second feature point (P2) corresponding to an upper body including the face / neck / shoulders of a target to be used for face replacement from a second image (V2).
[0109] The step of acquiring a first feature point and a second feature point (S130) may include a step of detecting a face region from a first image (S131), a step of detecting a first feature point from a first image (S132), a step of additionally detecting an auxiliary feature point (S133), and a step of detecting a second feature point from a second image (S134).
[0110] The step of detecting a face region from the first image (S131) is to detect a face region in the first image (V1) where face replacement is to be performed, and to set a search range and a reference region in the first feature point (P1) detection process described later based on the detected face region.
[0111] More specifically, the face detection module (112) can set a search area within a processing frame (TF) or a captured frame (PF) in the first image (V1), extract face candidate regions, and determine whether each candidate region is a face to confirm the face region. At this time, the face region is an area within the image that includes a face, and can be represented as at least one of a bounding rectangle-shaped region, a mask-shaped region, or a polygon-shaped region. The face detection module (112) can calculate at least one of position information, size information, and direction information corresponding to the face region, and the calculated information can be used as a standard for face structure-based operations, such as limiting the range in which the first feature point (P1) is detected in a subsequent step or estimating the face center axis.
[0112] Additionally, the face detection module (112) can compare the amount of change between the face area in the previous frame and the face area in the current frame to mitigate shaking of the detection result in a single frame, and if the amount of change exceeds a reference value, it can re-verify or correct the face area of the current frame. For example, the face detection module (112) can perform re-search by expanding or contracting the search range based on the face area of the previous frame, and can update the face area as a result of moving average or smoothing processing by reflecting the trend of position change of the face area across multiple frames. Additionally, the face detection module (112) can be configured not to confirm the face area when the detection reliability is below a reference value, or to re-detect the face area by expanding the search range within the processing frame (TF).
[0113] The step of detecting a first feature point from a first image (S132) may mean detecting a plurality of first feature points (P1) corresponding to the upper body including the face, neck, and shoulders based on the face region detected by the face detection module (112).
[0114] Specifically, the feature point extraction module (113) can set a reference area for feature point detection based on at least one of the location, size, and direction information of the detected face area, and detect a first feature point (P1) in an upper body area extended to include the neck and shoulders centered on the reference area. The first feature point (P1) may include, for example, at least one of points around both eyes, points around the nose, points at both ends of the mouth, a center point of the lips, a chin end point, an end point of the eyebrows, a center point of the forehead, a point near the cheekbones, and / or multiple points on the face contour in the face area. Additionally, in the neck and shoulder area, it may include at least one of the center or end points of the clavicle, a center point of the neck, end points of both shoulders, and / or multiple points on the shoulder contour.
[0115] The feature point extraction module (113) can generate a first feature point (P1) by calculating coordinate values for a plurality of such points, and can also calculate the detection reliability for each feature point. Through this, in a subsequent step, it is possible to determine whether there is a missing feature point or a false detection, or to determine whether to additionally detect auxiliary feature points.
[0116] The step of additionally detecting auxiliary feature points (S133) may mean a step of reinforcing the first feature points (P1) by additionally detecting auxiliary feature points corresponding to the upper body from the first image (V1) when at least some of the first feature points (P1) detected in the first image (V1) are not detected or the detection reliability is below a reference value.
[0117] Specifically, the feature point extraction module (113) can evaluate the detection result of the first feature point (P1) in the first image (V1) to determine whether there are any undetected feature points or feature points with a detection reliability below a reference value. And if the condition is satisfied, the feature point extraction module (113) can additionally detect auxiliary feature points corresponding to the upper body. The auxiliary feature points may include, for example, at least one of additional points on the face contour, additional points on the eyebrow contour, center points of the cheeks, forehead contour points, middle points of the jawline, neckline contour points, additional points around the clavicle, and / or additional points on the shoulder contour. The feature point extraction module (113) can include the additionally detected auxiliary feature points in the first feature point (P1) or assign weights to the auxiliary feature points to improve the stability of the first feature point (P1).
[0118] The step of detecting a second feature point from a second image (S134) may mean detecting a plurality of second feature points (P2) corresponding to the upper body including the face / neck / shoulder of the target to be used for face replacement from the second image (V2).
[0119] Specifically, the feature point extraction module (113) can detect a second feature point (P2) based on a face region detected in the second image (V2) by the face detection module (112). The second feature point (P2) may include, for example, at least one of a point around both eyes, a point around the nose, points at both ends of the mouth, a center point of the lips, a point at the end of the chin, an end point of the eyebrows, a center point of the forehead, a point near the cheekbone, and / or a plurality of points on the face contour in the face region, and may include at least one of a center or end point of the clavicle, a center point of the neck, a point at the end line of both shoulders, and / or a plurality of points on the shoulder contour in the neck and shoulder region. The feature point extraction module (113) can generate the second feature point (P2) by calculating coordinate values for each feature point.
[0120] The step (S140) of replacing the face in the first image with the face in the second image and outputting it may mean a step of performing face replacement and outputting it so that the face included in the first image (V1) appears as the face included in the second image (V2) based on the first feature point (P1) and the second feature point (P2) obtained in the preceding step (S130).
[0121] The step of outputting by replacing the face of the first image with the face of the second image (S140) may include the step of calculating a first feature line and a first feature plane (S141), the step of calculating a second feature line and a second feature plane (S142), the step of calculating a matching parameter (S143), the step of updating the matching parameter frame by frame (S144), and the step of outputting the face of the first image so that it appears as a face included in the second image (S145).
[0122] The step of calculating the first feature line and the first feature surface (S141) may mean a step of calculating the first feature line and the first feature surface by connecting at least some of the first feature points (P1) to each other.
[0123] Specifically, the alignment replacement processing module (114) can calculate a first feature line by selecting at least two of the first feature points (P1) obtained from the first image (V1) and connecting them to each other, and can calculate a first feature surface by selecting at least three of the first feature points (P1) and connecting them in a polygonal shape. For example, the alignment replacement processing module (114) can calculate at least one of a line connecting feature points around both eyes in the face area, a line connecting feature points around the nose and feature points around the mouth, and / or a polygonal line connecting multiple feature points on the face contour in sequence as the first feature line. In addition, the alignment replacement processing module (114) can calculate a first feature surface in the shape of a triangle or a square by connecting three or more of the feature points around the eyes, feature points around the nose, and feature points around the mouth, or can calculate a first feature surface that approximates the face shape by connecting multiple feature points on the face contour.
[0124] The alignment replacement processing module (114) can select feature points to be used for connection or determine the connection order based on the detection reliability of the first feature point (P1) or the relative positional relationship between feature points when calculating the first feature line and the first feature surface, and the calculated first feature line and the first feature surface can be used as reference information for calculating alignment parameters in the step to be described later.
[0125] The step of calculating the second feature line and the second feature surface (S142) may mean a step of calculating the second feature line and the second feature surface by connecting at least some of the second feature points (P2) to each other.
[0126] Specifically, the alignment replacement processing module (114) can calculate a second feature line by selecting at least two of the second feature points (P2) obtained from the second image (V2) and connecting them to each other, and can calculate a second feature surface by selecting at least three of the second feature points (P2) and connecting them in a polygonal shape. For example, the alignment replacement processing module (114) can calculate at least one of a line connecting feature points around both eyes in the face area, a line connecting feature points around the nose and feature points around the mouth, and / or a polygonal line connecting multiple feature points on the face contour in sequence as the second feature line. In addition, the alignment replacement processing module (114) can calculate a second feature surface in the shape of a triangle or a square by connecting three or more of the feature points around the eyes, feature points around the nose, and / or feature points around the mouth, or can calculate a second feature surface that approximates the face shape by connecting multiple feature points on the face contour.
[0127] The matching replacement processing module (114) can select a feature point to be used for connection or determine the connection order based on the detection reliability of the second feature point (P2) or the relative positional relationship between the feature points when calculating the second feature line and the second feature surface, and the calculated second feature line and the second feature surface can be used as reference information for calculating matching parameters in the step to be described later.
[0128] The step of calculating the matching parameters (S143) may mean a step of calculating matching parameters for matching the face included in the second image (V2) to the face region of the first image (V1) based on the correspondence between the first feature line and the first feature plane calculated in the first image (V1) and the second feature line and the second feature plane calculated in the second image (V2).
[0129] Specifically, the alignment replacement processing module (114) can calculate alignment parameters by setting a corresponding pair between a first feature point (P1) constituting a first feature line and a first feature plane and a second feature point (P2) constituting a second feature line and a second feature plane, and determining a transformation model such that the coordinate error for the corresponding pair is minimized. For example, the alignment replacement processing module (114) can set a transformation model such that at least one of the face center point, face width and height, distance between eyes, relative position of nose and mouth and / or face center axis direction in the first image (V1) matches the corresponding value in the second image (V2).
[0130] The matching parameter may include at least one of the position / size / orientation of the face included in the second image (V2), and depending on the implementation, may be composed of at least one of a translation parameter, a scale parameter, a rotation parameter, and / or a tilt parameter. Additionally, the matching replacement processing module (114) may calculate the matching parameter to include a transformation parameter that reflects not only a 2D transformation but also a 3D pose component, by reflecting the shape difference between the first feature plane and the second feature plane when there is a large change in face pose.
[0131] Additionally, the alignment replacement processing module (114) can calculate the alignment parameters by lowering the weight of a specific feature point and increasing the weight of a relatively stable feature point when the detection reliability of that feature point is low or temporarily fluctuates during the alignment parameter calculation process. At this time, the alignment replacement processing module (114) can calculate the alignment parameters such that the cost function is minimized by defining at least one of the length ratio of the first feature line and the second feature line, the area ratio of the first feature surface and the second feature surface, the angle difference between feature lines, and / or the degree of deformation of the feature surface as a cost function.
[0132] The step of updating the matching parameters frame by frame (S144) may mean updating the matching parameters on a frame-by-frame basis based on the change of the first feature point (P1) detected in each frame when the first image (V1) includes multiple frames.
[0133] Specifically, the alignment replacement processing module (114) can calculate the feature point displacement amount by setting the alignment parameter calculated in the preceding step (S143) as an initial value and then calculating the coordinate difference between the first feature point (P1) in the current frame of the first image (V1) and the first feature point (P1) in the previous frame. The alignment replacement processing module (114) can update the position component of the alignment parameter based on the magnitude and direction of the feature point displacement amount, update the magnitude component of the alignment parameter based on the rate of change of distance between feature points, and update the direction component of the alignment parameter based on the direction of the face center axis or the change in slope of the feature line.
[0134] Additionally, the alignment replacement processing module (114) can calculate at least one of a movement vector, a velocity vector, and / or an acceleration vector by accumulating a plurality of feature point displacement amounts or calculating the rate of change between frames, and can predict the alignment parameters of the current frame based on the calculated vector, and then update them by performing fine corrections based on the predicted alignment parameters. For example, the alignment replacement processing module (114) can calculate the error regarding how much the second feature line or the second feature plane matches the first feature line or the first feature plane when the predicted alignment parameters are applied, and update the alignment parameters by correcting at least one of the position, magnitude, and direction components so that the error is reduced.
[0135] The alignment replacement processing module (114) can update the alignment parameters to a stabilized value by applying a weighted combination of the alignment parameters of the previous frame and the alignment parameters updated in the current frame, or by applying a cumulative moving average and / or smoothing filtering, in order to mitigate the phenomenon where the screen shakes due to excessive fluctuations in the alignment parameters updated per frame.
[0136] The step (S145) of outputting the face of the first image to appear as the face included in the second image may mean the step of converting the face included in the second image (V2) according to the matching parameters updated for each frame, and synthesizing the converted face to the face area of the first image (V1) and outputting it.
[0137] Specifically, the alignment replacement processing module (114) can transform a face region included in the second image (V2) by at least one of translation, scaling, rotation, and / or tilting transformation based on the updated alignment parameters. At this time, the alignment replacement processing module (114) can set a face region to be used for synthesis in the second image (V2) based on the second feature point (P2) or the second feature plane, and generate a transformation result such that the transformed face region is positioned to match the face region specified by the first feature point (P1) in the first image (V1).
[0138] The alignment replacement processing module (114) can generate a mask corresponding to the face region and combine pixel by pixel based on the mask in order to composite the face region of the transformed second image (V2) to the face region of the first image (V1). In addition, processing to smooth the mask boundary or blend the area around the boundary can be applied to mitigate the unnatural appearance of the composite boundary, and correction processing based on the statistical value of the face region can be applied to mitigate the difference in brightness or color between the face region of the first image (V1) and the face region of the second image (V2).
[0139] The alignment replacement processing module (114) can sequentially output the composite result frames for each frame, thereby providing that the face replacement results appear continuously in multiple frames of the first image (V1).
[0140] Meanwhile, in a situation where the first feature point goes out of the shooting frame, the face replacement can be maintained through the step (S150) described later.
[0141] The step (S150) of determining whether a first feature point has moved out of the shooting frame and continuously performing face replacement may mean calculating a predicted feature point (OP) corresponding to outside the shooting frame (PF) and continuously performing face replacement based on the predicted feature point (OP) so that face replacement is not interrupted even when a situation occurs where at least one of the first feature points (P1) extracted from the first image (V1) moves out of the shooting frame (PF).
[0142] The step of determining whether the first feature point is outside the shooting frame and continuously performing face replacement (S150) may include the step of determining whether the first feature point is outside the shooting frame (S151), the step of calculating a predicted feature point (S152), and the step of continuously performing face replacement based on the predicted feature point (S153).
[0143] Meanwhile, frame out-of-frame determination is performed based on the boundaries of the captured frame (PF), and the calculation of predicted coordinates and predicted feature points (OP) can be performed within the range of the processing frame (TF).
[0144] The step (S151) of determining whether a first feature point is outside the shooting frame may mean a step of determining whether each of the plurality of first feature points (P1) is outside the shooting frame (PF) and exists in the outer frame (OF) area within the processing frame (TF).
[0145] Specifically, the electronic device (100) determines whether each of the first feature points (P1) is out of the way, and if it is determined that all of the first feature points (P1) are within the shooting frame (PF), it may not perform the steps (S152 and S153) described later and may maintain the face replacement output state from the preceding step (S140). On the other hand, the electronic device (100) may perform the steps (S152 and S153) described later to calculate predicted feature points and continue face replacement only when it is determined that at least one of the first feature points (P1) is out of the shooting frame (PF) and is in the outer frame (OF) area.
[0146] The step of determining whether the first feature point is out of the captured frame (S151) may include the step of mapping the coordinate values of the first feature point to a multidimensional embedding space (S151-1), the step of calculating a distance function and calculating a frame out-of-frame index (S151-2), and the step of determining the out-of-frame state according to a hysteresis condition (S151-3).
[0147] The step of mapping the coordinate values of the first feature point to a multidimensional embedding space (S151-1) may mean that the deviation determination module (115) constructs an input representation containing position information in a reference coordinate system defined based on a shooting frame (PF) or a processing frame (TF) for each of the plurality of first feature points (P1) acquired in the first image (V1), and converts this into an embedding value in a multidimensional embedding space.
[0148] Specifically, the deviation determination module (115) can construct a feature vector for each first feature point (P1) that includes not only position information in a reference coordinate system, but also at least one of a relative position relationship with respect to the frame boundary, a change in position relative to the previous frame, a relative position relationship with respect to a remaining feature point (RP), and identification information regarding the type of feature point. Here, the feature vector is not limited to two dimensions or three dimensions, but can be composed of a multidimensional input in which multiple components necessary for deviation determination are combined.
[0149] The deviation determination module (115) inputs the constructed feature vector into an embedding transformation function to map each first feature point (P1) to an embedding value in a multidimensional embedding space. The embedding transformation function can be implemented as at least one of a pre-trained neural network-based encoder, a combination of linear or non-linear transformations, or a rule-based transformation.
[0150] Additionally, the deviation determination module (115) can construct a feature vector using the result of accumulating, moving averaging, or smoothing the location information of the first feature point (P1) or the remaining feature point (RP) in multiple frames to mitigate the rapid change in the embedding value caused by shaking of the frame-unit detection result. Accordingly, the distribution of feature points in the embedding space is stabilized, and the distance function and frame deviation index can be calculated more stably in the step (S151-2) to be described later.
[0151] The step of calculating a distance function and a frame deviation index (S151-2) may mean a step in which a deviation judgment module (115) calculates a distance function and a frame deviation index to quantify the possibility of deviation from a captured frame (PF) based on the boundary relationship between the feature point distribution of a first feature point (P1) mapped in a multidimensional embedding space and the captured frame (PF). The step of calculating a distance function and a frame deviation index (S151-2) may include a step of calculating a distance function for the boundary of the captured frame based on the feature point distribution of the first feature point, a step of normalizing the distance function to calculate a deviation reliability, and a step of calculating a frame deviation index based on the deviation reliability.
[0152] The step of calculating a distance function for the boundary of a captured frame based on the distribution of the first feature points is to calculate the distribution of the first feature points (P1) in the embedding space, and to calculate a distance function by weightedly combining the boundary distance values for the boundary of the captured frame (PF) based on the calculated distribution.
[0153] In this case, the distance function is a boundary proximity score reflecting the feature point distribution rather than a single feature point, the out-of-bounds confidence is a continuous value normalized from it, and the frame out-of-bounds indicator may refer to a value for state judgment reflecting time-axis accumulation or smoothing.
[0154] More specifically, the feature point distribution may include at least one of the distribution center, variance, main distribution direction, and / or the proportion of the first feature points (P1) located outside the captured frame (PF). For example, the distribution center may be calculated as the average or representative position of the first feature points (P1) in the embedding space, the variance may be calculated as the degree of spread of the first feature points (P1) relative to the distribution center, and the main distribution direction may be calculated as the direction in which the variance appears largest and / or the direction in which the first feature points (P1) move in a dense manner. Additionally, the proportion of the first feature points (P1) located outside the captured frame (PF) may be calculated as the ratio of the number of feature points among the first feature points (P1) that are mapped to a position outside the boundary of the captured frame (PF) and / or a weighting ratio.
[0155] Additionally, the deviation determination module (115) can calculate a boundary distance value based on the boundary of the captured frame (PF). For example, the boundary distance value can be calculated as the distance between the distribution center and the boundary of the captured frame (PF) and / or the minimum distance between each first feature point (P1) and the boundary of the captured frame (PF). Additionally, the boundary distance value may be assigned a code to distinguish between the inner and outer directions of the boundary, and may be set so that the distance value increases or is weighted when it corresponds to the outer side of the captured frame (PF).
[0156] Additionally, the deviation determination module (115) can set weights for boundary distance values based on the feature point distribution. For example, the weight of the boundary distance value can be set to increase as the center of the distribution approaches the boundary of the captured frame (PF), and the weight of the boundary distance value for the boundary in that direction can be set to increase as the variance increases in a specific direction. Additionally, as the ratio of the first feature point (P1) outside the captured frame (PF) increases, the weight of the boundary distance value included in the distance function can be increased, or the distance value can be set to be non-linearly amplified.
[0157] Therefore, the distance function is not determined solely by a single boundary distance value, but can be calculated as a combination of at least two of the following: a boundary distance value for the center of the distribution, a boundary distance value for an individual first feature point (P1), and a weighted value set according to the outer ratio. For example, the deviation determination module (115) can calculate the boundary distance value based on the center of the distribution and the boundary distance value based on the individual feature point, respectively, then combine them by reflecting the variance or the main distribution direction as a weight, and determine the result of the combination as the distance function. In this case, the weighted combination can be implemented as at least one of a weighted sum, a weighted average, and / or a non-linear combination.
[0158] The step of calculating churn reliability by normalizing the distance function is the step of calculating churn reliability by normalizing the calculated distance function according to a reference range.
[0159] More specifically, the deviation determination module (115) may set a normalization criterion based on at least one of the setting information of the captured frame (PF) or the processing frame (TF), the degree of dispersion of the feature point distribution, and / or a preset reference value. For example, the normalization criterion may be set to an allowable distance range for the boundary of the captured frame (PF), an effective processing range in the processing frame (TF), and / or a correction range that is dynamically adjusted according to the dispersion of the feature point distribution.
[0160] Additionally, the deviation determination module (115) can normalize the distance function to generate a value that is comparable even in situations with different scales. For example, since the distribution relative to the shooting frame (PF) differs between dense and dispersed cases even with the same amount of movement, the deviation determination module (115) can apply a normalization coefficient that adjusts the influence of the distance function based on the dispersion and / or outward ratio. Additionally, normalization can be performed as linear normalization and / or non-linear normalization to reflect the extent to which the distance function is located within a reference range.
[0161] Additionally, the deviation judgment module (115) can calculate the deviation reliability using values obtained by performing cumulative, moving average, and / or smoothing processing on the distance function or normalization result to mitigate inter-frame fluctuations. For example, the normalization result from the previous multiple frames can be accumulated to reduce one-off fluctuations, a moving average can be applied to reflect a gradual moving trend, and smoothing filtering can be applied to mitigate abrupt rises or falls.
[0162] Additionally, the out-of-bounds confidence can be set as a value that expresses the probability of the captured frame (PF) out-of-bounds as a continuous value. For example, the out-of-bounds confidence can be calculated to increase as approach to the boundary of the captured frame (PF) increases, as the proportion of the outer side of the captured frame (PF) increases, or as the center of distribution moves along the boundary or the dominant direction of distribution faces the boundary. In this case, the out-of-bounds confidence is not calculated using only a distance function, but can be calculated as a normalized result that reflects at least one of the center of distribution, variance, dominant direction of distribution, and / or the proportion of the outer side derived from the feature point distribution.
[0163] The step of calculating the frame breakout indicator based on breakout reliability is a step of calculating the frame breakout indicator based on the cumulative value, moving average value, and / or smoothed value of the breakout reliability.
[0164] More specifically, the breakout judgment module (115) can calculate a frame breakout indicator by reflecting the temporal trend of breakout reliability. For example, the cumulative value of the breakout reliability can be used to reflect the breakout trend over a certain period, the moving average value can be used to prioritize reflecting continuous movement over one-time fluctuations, and the smoothing value can be used to mitigate the excessive change in the indicator caused by sudden changes.
[0165] Additionally, the breakout judgment module (115) can calculate a frame breakout indicator by reflecting whether the breakout reliability is maintained above a threshold value for a certain period and / or whether the breakout reliability has an increasing trend. For example, the frame breakout indicator may be set to increase as the number of frames in which the breakout reliability exceeds the threshold value increases, and the frame breakout indicator may be set to increase as the growth rate of the breakout reliability increases.
[0166] In addition, at this stage, the deviation judgment module (115) may correct the frame deviation index by additionally reflecting at least one of the ratio of the first feature point (P1) outside the captured frame (PF), the degree of proximity of the distribution center to the boundary, and / or the boundary orientation of the main distribution direction. For example, if the ratio outside the captured frame (PF) increases above a certain value, the frame deviation index may be corrected to increase acceleratedly, and if the distribution center is close to the boundary and the main distribution direction is directed toward the outside of the boundary, the frame deviation index may be corrected to increase preferentially.
[0167] Additionally, the deviation judgment module (115) can apply correction conditions to prevent the frame deviation index from rising excessively due to temporary spikes in a single feature point or instantaneous changes in distribution. For example, if a one-time increase in the distance function occurs while the outer ratio is low, the increase in the frame deviation index can be limited, the frame deviation index can be maintained if the movement of the distribution center is within a certain range, and the frame deviation index correction can be suppressed if the main distribution direction is calculated in a direction unrelated to the boundary.
[0168] Therefore, the frame breakout indicator is not determined solely by the distance function of a single point in time or the breakout confidence of a single frame, but can be implemented as a value calculated by combining the cumulative breakout confidence across multiple frames or information related to trends and feature point distributions.
[0169] The step of determining the out-of-bounds state according to hysteresis conditions (S151-3) determines the out-of-bounds state according to hysteresis conditions in which the frame out-of-bounds indicator satisfies the out-of-bounds threshold, and the return threshold is set to be different from the out-of-bounds threshold.
[0170] More specifically, the ditching judgment module (115) can be configured to switch to a ditching state when the frame ditching indicator is above a ditching threshold. Additionally, after switching to a ditching state, it can be configured to return to a non-ditching state only when the frame ditching indicator decreases below a return threshold. In this case, the return threshold can be set lower than the ditching threshold, and accordingly, even if the frame ditching indicator fluctuates near the threshold, the phenomenon of frequently switching between the ditching state and the non-ditching state can be suppressed.
[0171] Additionally, when determining whether to switch to a switch state or a switch back, the switch state determination module (115) may be configured to refer to at least one of a cumulative value across multiple frames, a moving average value, and / or a smoothed value, rather than relying solely on a single frame value of the frame switch indicator. For example, it may be configured to switch to a switch state only when the state above the switch threshold is maintained for a number of consecutive reference frames or longer, or to return to a non-switch state only when the state below the switch threshold is maintained for a number of consecutive reference frames or longer.
[0172] Additionally, the deviation determination module (115) can be configured to update a state flag to maintain the state when a deviation state is determined, and to determine whether to perform the calculation of predicted feature points (OP) based on the updated state flag. For example, if a non-deviation state is determined, the step (S152) described below can be controlled so that it is not performed, and if a deviation state is determined, the step (S152) described below can be controlled so that it is performed.
[0173] That is, through this process, the electronic device (100) can determine whether a plurality of first feature points (P1) exist in the outer frame (OF) area outside the shooting frame (PF) for each of the plurality of frames included in the first image (V1), and if it is not determined to be outside, it can control the device to continue performing alignment replacement. In addition, if it is determined that at least one of the first feature points (P1) exists in the outer frame (OF) area outside the shooting frame (PF), the device can control the device to calculate a predicted feature point (OP), thereby ensuring that face replacement is maintained without interruption even if a part of the first feature point (P1) moves outside the shooting frame (PF) and into the outer frame (OF) area.
[0174] The step of calculating predicted feature points (S152) may mean determining whether a deviation feature point is a paired feature point for a deviation feature point determined to be outside the shooting frame (PF) among the first feature points (P1), setting a reference feature point according to the determination result, calculating predicted coordinates based on the reference feature point, and verifying and correcting the calculated predicted coordinates to calculate the predicted feature points.
[0175] The step of calculating predicted feature points (S152) may include a step of determining whether the deviating feature points correspond to a pair of feature points (S152-1), a step of calculating predicted coordinates by setting the symmetric feature point as a reference feature point if they are a pair (S152-2Y), a step of calculating predicted coordinates by setting the remaining first feature point as a reference feature point if they are not a pair (S152-2N), a step of correcting the predicted coordinates through the remaining first feature point excluding the reference feature point (S152-3), and a step of calculating predicted feature points based on the corrected predicted coordinates (S152-4).
[0176] The step of determining whether the out-of-bounds feature point is a pair of corresponding feature points (S152-1) is a step of determining whether the out-of-bounds feature point determined to be outside the shooting frame (PF) has a corresponding feature point that is symmetrical with respect to the face center axis.
[0177] More specifically, the prediction feature point calculation module (116) can determine whether the deviation feature point corresponds to a type defined as a left-right symmetrical pair by referring to the feature point type information and symmetrical pair mapping information assigned to the first feature point (P1) obtained in step (S130). For example, symmetrical pair mapping information may be set for feature points of a type in which a left-right symmetrical relationship is established, such as eyes, eyebrows, corners of the mouth, part of the jawline, clavicle point, shoulder point, etc., and for feature points of a type that have the nature of a single reference point, such as the tip of the nose or the center of the forehead, symmetrical pair mapping information may not be set or may be defined as having no pair corresponding target. At this time, the first feature point (P1) may be classified into face landmarks, contour landmarks, neck landmarks, or shoulder landmarks, and each landmark may have symmetrical pair mapping information corresponding to the corresponding type.
[0178] Additionally, the prediction feature point calculation module (116) can estimate the face center axis based on the first feature point (RP) remaining within the captured frame (PF) and determine whether a feature point corresponding to the symmetric position of the deviation feature point remains based on the estimated face center axis. For example, the prediction feature point calculation module (116) can determine pair correspondence by clustering the feature points such that the left and right position codes relative to the center axis are opposite for a plurality of feature points among the remaining feature points (RP), and by checking whether an opposite feature point mapped to the same type as the deviation feature point exists within the cluster.
[0179] Furthermore, the prediction feature point calculation module (116) may additionally determine whether the symmetric corresponding feature point of the deviation feature point is a feature point that is tracked together in the upper body structure of the same person by using feature point location information from the previous frame or the result of estimating movement between frames. That is, even if symmetric candidates of the same type exist, the module may be configured to finally determine them as paired corresponding feature points only if the candidate feature point satisfies continuity between frames or maintains a relative positional relationship with the remaining feature points (RP).
[0180] The step of calculating predicted coordinates by setting a symmetric feature point as a reference feature point (S152-2Y) is a step in which, in the preceding step (S152-1), the deviation feature point is determined to be a feature point corresponding to a pair, and if the symmetric feature point corresponding to the deviation feature point exists as a residual feature point (RP) within the captured frame (PF), the residual symmetric feature point is set as a reference feature point to calculate the predicted coordinates of the deviation feature point.
[0181] More specifically, the prediction feature point calculation module (116) can calculate the predicted coordinates of the deviation feature point by setting the remaining symmetric feature point as a reference feature point and then reflecting or symmetrically transforming the position of the reference feature point with respect to the face center axis. For example, the prediction feature point calculation module (116) can calculate the relative distance and direction between the reference feature point and the face center axis, and determine the predicted coordinates where the deviation feature point is expected to be located by applying the same relative distance and direction to the opposite side of the face center axis.
[0182] In addition, since the position of the face center axis may change from frame to frame, the prediction feature point calculation module (116) can update the face center axis frame by frame based on the distribution of residual feature points (RP) or the center axis estimation result from the previous frame, and calculate the prediction coordinates based on the updated face center axis.
[0183] Additionally, the prediction feature point calculation module (116) can correct the prediction coordinates by reflecting at least one of the relative offset, ratio, and / or correction coefficient included in the mapping information when there is predefined symmetry pair mapping information between the out-of-place feature point and the reference feature point. For example, if rotation or tilting of the face occurs, the accuracy of the prediction coordinates can be improved by adjusting the angle or scale of the symmetry transformation based on the face orientation information estimated from the residual feature point (RP).
[0184] The step of calculating predicted coordinates by setting the remaining first feature point as a reference feature point (S152-2N) is a step of calculating predicted coordinates of the deviation feature point by setting at least one of the remaining first feature points (RP) within the captured frame (PF) as a reference feature point when the deviation feature point is determined to be a feature point that does not correspond to a pair in the preceding step (S152-1), or when the symmetric feature point corresponding to the deviation feature point is not confirmed as a remaining feature point (RP) within the captured frame (PF).
[0185] More specifically, the prediction feature point calculation module (116) can select reference feature point candidates based on type information of the deviation feature point. For example, if the deviation feature point is a feature point corresponding to the center of the face or the facial contour, the prediction feature point calculation module (116) can select residual feature points (RPs) corresponding to adjacent areas such as around the eyes, around the nose, around the mouth, or around the jawline as reference feature point candidates, and if the deviation feature point is a feature point corresponding to the neck or shoulder, residual feature points (RPs) on the clavicle, neck center point, or shoulder line can be selected as reference feature point candidates.
[0186] Additionally, the prediction feature point calculation module (116) may finally select a feature point among the reference feature point candidates that has high continuity between frames or a stable relative positional relationship as the reference feature point. For example, the prediction feature point calculation module (116) may evaluate the reference feature point candidates based on at least one of the amount of position change in the previous frame, the range of variation of the movement vector, and / or detection stability, and set a candidate whose evaluation result satisfies a reference value as the reference feature point.
[0187] The prediction feature point calculation module (116) can calculate the predicted coordinates of the deviation feature point based on the set reference feature point. For example, the prediction feature point calculation module (116) can calculate the predicted coordinates by applying a predefined relative distance, angle, or ratio relationship between the reference feature point and the deviation feature point, and can determine the predicted coordinates by projecting the relative distance or relative angle relationship maintained in the previous frame onto the reference feature point location of the current frame.
[0188] In addition, the prediction feature point calculation module (116) can improve the accuracy of the prediction coordinates by estimating the degree of change in the size or direction of the face based on the distribution of the remaining feature points (RP) and correcting the relative distance, angle, or ratio relationship by reflecting the estimated degree of change.
[0189] The step of correcting the predicted coordinates through the remaining first feature points excluding the reference feature point (S152-3) is a step of verifying whether the predicted coordinates calculated in the preceding step (S152-2Y or S152-2N) correspond to the relative positional relationship with the remaining first feature points (RP) within the captured frame (PF), and correcting the predicted coordinates according to the verification result. At this time, the verification of the predicted coordinates can be performed based on whether the relative distance, angle, and ratio relationship with the remaining feature points (RP) excluding the reference feature point is within the reference error range.
[0190] More specifically, the prediction feature point calculation module (116) may select a plurality of residual feature points (RPs), excluding residual feature points set as reference feature points, as verification reference feature points. For example, the prediction feature point calculation module (116) may select at least one of residual feature points (RPs) spatially adjacent to the prediction coordinates, residual feature points (RPs) corresponding to the same face area or the same upper body area, and / or residual feature points (RPs) with stable positional changes between frames as verification reference feature points.
[0191] The prediction feature point calculation module (116) can calculate the relative distance, angle, and / or ratio relationship between the verification reference feature points and the prediction coordinates, and verify the validity of the prediction coordinates by comparing the relationship with a predefined reference relationship or a relationship in the previous frame to determine whether it is within the reference error range. For example, the prediction feature point calculation module (116) can store the relative distance or relative angle relationship at the time when the deviation feature point existed within the shooting frame (PF) in the previous frame, and compare this with the prediction coordinates calculated in the current frame to determine whether the same relationship is maintained.
[0192] If the verification result exceeds the standard error range, the prediction feature point calculation module (116) can correct the prediction coordinates so that the error is minimized. For example, the prediction feature point calculation module (116) can move the prediction coordinates so that the sum of the distance error and / or angle error with the verification standard feature points is minimized, or fine-tune the prediction coordinates in the direction of the distribution center of the residual feature point (RP) or the face center axis. Additionally, the prediction feature point calculation module (116) can mitigate excessive correction caused by temporary spikes by determining the amount of correction by giving greater weight to the feature point with higher reliability among the multiple verification standard feature points.
[0193] Consequently, through the step (S152-3) of correcting the predicted coordinates using the remaining first feature points excluding the reference feature point, the predicted coordinates can be determined not as initial estimates calculated relying only on the reference feature point, but as coordinates verified and corrected to satisfy the relative positional relationship with the remaining feature points (RP) excluding the reference feature point.
[0194] The step of calculating predicted feature points based on corrected predicted coordinates (S152-4) involves determining predicted feature points (OP) corresponding to outside the captured frame (PF) using the predicted coordinates verified and corrected in the preceding step (S152-3), and configuring the determined predicted feature points (OP) so that they can be used for subsequent face replacement processing.
[0195] More specifically, the prediction feature point calculation module (116) can generate a prediction feature point (OP) by mapping feature point type information and identification information of the deviation feature point to the corrected prediction coordinates. For example, the prediction feature point (OP) may be configured to include at least one of the corresponding feature point type, feature point index, and / or tracking information of the same feature point in the previous frame, along with the corrected prediction coordinates.
[0196] Additionally, the prediction feature point calculation module (116) checks whether the corrected prediction coordinates are located within the processing frame (TF), and if they are located within the processing frame (TF), it can confirm the prediction feature point (OP) as a valid feature point. Conversely, if the corrected prediction coordinates extend beyond the processing frame (TF), the calculation of the prediction feature point (OP) may be withheld, or the prediction feature point (OP) may be calculated after performing clipping processing that limits the corrected prediction coordinates to the boundary of the processing frame (TF).
[0197] Additionally, the predicted feature point calculation module (116) can be configured to consistently apply a coordinate system so that the predicted feature point (OP) maintains a relative positional relationship with the remaining feature points (RP), and to include the predicted feature point (OP) in the first set of feature points of the current frame for continuous processing between frames.
[0198] Finally, the predicted feature point calculation module (116) transmits the calculated predicted feature point (OP) to the alignment replacement processing module (114) so that the predicted feature point (OP) is used as supplementary information for the first feature point (P1) together with the remaining feature point (RP), thereby allowing the face replacement to be performed continuously without interruption even if some feature points deviate from the shooting frame (PF).
[0199] The step of continuously performing face replacement based on predicted feature points (S153) is a step of processing so that face replacement is maintained even in a situation where the captured frame (PF) is out of reach, using the predicted feature points (OP) calculated in the preceding step (S152).
[0200] More specifically, the prediction feature point calculation module (116) can transmit the calculated prediction feature point (OP) to the alignment replacement processing module (114), and the alignment replacement processing module (114) can continuously perform calculation and updating of alignment parameters using the remaining feature point (RP) and the prediction feature point (OP) together, thereby ensuring that face replacement is maintained without interruption. In addition, the alignment replacement processing module (114) can process so that the face included in the second image (V2) is stably replaced and output in the face area of the first image (V1) even in the section where the prediction feature point (OP) is applied.
[0201] As described above, exemplary embodiments have been disclosed in the drawings and specification. Although specific terms have been used to describe the embodiments in this specification, they are used only for the purpose of explaining the technical concept of this disclosure and are not intended to limit the meaning or the scope of this disclosure as defined in the claims. Therefore, those skilled in the art will understand that various modifications and equivalent alternative embodiments are possible therefrom. Accordingly, the true technical scope of protection of this disclosure should be determined by the technical concept of the appended claims.
Claims
Claim 1 A face synthesis method performed by an electronic device, comprising: receiving a first image and a second image including a target to be replaced in the first image; setting a shooting frame for the first image and additionally setting an outer frame formed outside the shooting frame to set a processing frame including the shooting frame and the outer frame; acquiring a first feature point in the first image and acquiring a second feature point from the second image; replacing and outputting a face in the first image with a face in the second image based on the first feature point and the second feature point; when it is determined that at least one of the first feature points exists in the outer frame, calculating a predicted feature point in the processing frame that corresponds to outside the shooting frame based on the relative positional relationship between feature points remaining within the shooting frame; and continuously performing face replacement based on the predicted feature point. Claim 2 A face synthesis method according to claim 1, wherein the step of acquiring the first feature point and acquiring the second feature point from the second image comprises: a step of detecting a face region in the first image; a step of detecting a plurality of the first feature points corresponding to the face, neck, and shoulders based on the face region; a step of additionally detecting an auxiliary feature point corresponding to the upper body including the face when at least one of the first feature points is not detected or the detection reliability is less than a reference value; and a step of detecting a plurality of second feature points corresponding to the face included in the second image. Claim 3 In paragraph 2, the step of replacing the face of the first image with the face of the second image and outputting it comprises: a step of calculating a first feature line by connecting at least two of the first feature points to each other and calculating a first feature surface by connecting at least three of the first feature points to each other; a step of calculating a second feature line by connecting at least two of the second feature points to each other and calculating a second feature surface by connecting at least three of the second feature points to each other; a step of calculating a matching parameter including at least one of the position, size, and direction of the face included in the second image based on the correspondence relationship between the first feature line or the first feature surface and the second feature line or the second feature surface; and a step of calculating a motion vector, a velocity vector, or an acceleration vector based on the inter-frame change of the first feature points and updating the matching parameter on a frame-by-frame basis. A face synthesis method characterized by including the step of synthesizing a face included in the second image, which is transformed according to the updated matching parameters, to a face region of the first image, and outputting the face of the first image to appear as a face included in the second image. Claim 4 The face synthesis method according to claim 1 further comprises a step of determining that at least one of the first feature points deviates from the shooting frame, wherein the step of determining that at least one of the first feature points deviates from the shooting frame comprises: a step of mapping the coordinate values of the first feature points to a multidimensional embedding space; a step of calculating a distance function for the boundary of the shooting frame and calculating a frame deviation index based on the distance function; and a step of determining a deviation state according to a hysteresis condition in which, if the frame deviation index satisfies a deviation threshold, a deviation is determined, but a return threshold is set to be different from the deviation threshold. Claim 5 A face synthesis method according to claim 4, wherein the step of calculating the frame deviation index comprises: a step of calculating the distance function for the boundary of the shooting frame based on the boundary of the shooting frame and the feature point distribution of the first feature point within the embedding space; a step of normalizing the distance function to calculate the deviation reliability; and a step of calculating the frame deviation index based on the cumulative value, moving average value, or smooth value of the deviation reliability; wherein the feature point distribution includes at least one of the distribution center, variance, main distribution direction of the first feature point, or the ratio of the first feature point located outside the shooting frame, and the distance function is calculated as a weighted combined value of boundary distance values weighted by the feature point distribution. Claim 6 A face synthesis method according to claim 1, wherein the step of calculating the predicted feature point comprises: a step of, when the deviation feature point determined to be outside the shooting frame among the first feature points is a paired corresponding feature point, setting a symmetric feature point remaining within the shooting frame as a reference feature point, and calculating the predicted coordinates for the deviation feature point based on the relative positional relationship between the reference feature point and the face center axis; a step of verifying the predicted coordinates based on whether the predicted coordinates correspond to the relative positional relationship between the first feature points remaining within the shooting frame excluding the reference feature point, and correcting the predicted coordinates according to the verification result; and a step of calculating the feature point corresponding to the corrected predicted coordinates as the predicted feature point. Claim 7 A face synthesis method according to claim 1, wherein the step of calculating the predicted feature point comprises: a step of, when the deviation feature point determined to be outside the shooting frame among the first feature points is a feature point that does not correspond in pairs, setting at least one of the first feature points remaining within the shooting frame as a reference feature point, and calculating the predicted coordinates for the deviation feature point based on at least one of the relative distance, angle, or ratio relationship with the reference feature point; a step of verifying the predicted coordinates based on whether the predicted coordinates correspond to the relative positional relationship between the first feature points remaining within the shooting frame excluding the reference feature point, and correcting the predicted coordinates according to the verification result; and a step of calculating the feature point corresponding to the corrected predicted coordinates as the predicted feature point.