Method, apparatus, and device for removing wrinkles in portrait videos

The method addresses the challenge of real-time wrinkle removal in videos by using key frames, 3D reconstruction, and mapping techniques to effectively eliminate deep wrinkles, ensuring natural-looking results.

JP2025522211AActive Publication Date: 2025-07-11XIAMEN MEITUZHIJIA TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2025500329
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-01-29
Filing Date
2024-01-26
Publication Date
2025-07-11
Estimated Expiration
2044-01-26

AI Technical Summary

Technical Problem

Existing video editing software lacks effective methods to remove deep wrinkles such as crow's feet and forehead wrinkles in real-time, as time-consuming segmentation algorithms are not suitable for video processing.

Method used

A method involving key frame selection, wrinkle segmentation, 3D reconstruction, and mapping processing to accurately identify and remove wrinkles in portrait videos, utilizing 3D Morphable Model (3DMM) for face reconstruction and MVP matrices for mapping, along with texture replacement and Gaussian filtering to enhance naturalness.

Benefits of technology

Accurately removes wrinkles in real-time video frames by minimizing misprocessing and enhancing the naturalness of wrinkle removal, improving accuracy and efficiency in video editing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025522211000001_ABST
    Figure 2025522211000001_ABST
Patent Text Reader

Abstract

The present invention discloses a wrinkle removal method, apparatus, device, and storage medium for portrait videos. Among the videos to be processed, a video frame that meets a preset condition is selected as the first key frame image, and wrinkle segmentation is performed on the first key frame image to obtain a first wrinkle mask image; performing connected component detection on the first wrinkle mask image to obtain N connected components, and assigning values to the corresponding pixel values in the first wrinkle mask image according to the relationship between the centroid of each connected component and a preset range to obtain a second wrinkle mask image; performing 3D reconstruction of the human face on the first key frame image and other video frames to obtain a first mvp matrix and a second mvp matrix, performing mapping processing on the second wrinkle mask image based on the first mvp matrix and the second mvp matrix to obtain a third wrinkle mask image corresponding to the current video frame; and performing wrinkle removal on the human face of the current video frame according to the third wrinkle mask image to obtain a target video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video processing, and in particular, to a wrinkle removal method, apparatus, and device for portrait videos.

Background Art

[0002] In the video editing industry, the function of facial processing of people is a frequently used function by users. Commercially available facial processing software and clip software for people usually provide functions for removing skin defects such as smoothing the skin, removing dark circles, and removing freckles and acne. However, in the case of wrinkles, only wrinkles without dark wrinkles or folds such as nasolabial folds and marionette lines can be removed. There is no effective solution for deep wrinkles such as crow's feet and forehead wrinkles. Although the means for removing wrinkles in images is relatively mature, time-consuming segmentation algorithms are difficult to meet the real-time standard and cannot be directly applied to video editing.

Summary of the Invention

Problems to be Solved by the Invention

[0003] In view of the above, an object of the present invention is to propose a wrinkle removal method, apparatus, and device for portrait videos in order to solve the conventional problem that there is no effective solution for removing wrinkles on the face of a person in a video scene.

Means for Solving the Problems

[0004] To achieve the above object, the present invention obtains a video to be processed, selects a video frame of the face of the same person that satisfies a preset condition among the videos to be processed as a first key frame image, performs matching labeling for each of the first key frame images based on a preset expression category label, and performs wrinkle segmentation on the first key frame image to obtain a corresponding first wrinkle mask image; Performing connected component detection on the first wrinkle mask image to obtain N connected components, and assigning values to the corresponding pixel values in the first wrinkle mask image according to the relationship between the centroid of each connected component and a preset range to obtain a second wrinkle mask image, and storing the expression category label corresponding to the first key frame image and the second wrinkle mask image in a first set; Performing 3D reconstruction of the face of a person on the first key frame image to obtain a corresponding 3D face model of the person and a first MVP matrix; Performing 3D reconstruction of the face of a person on other video frames in the video to be processed to obtain a second MVP matrix, and performing mapping processing of the second wrinkle mask image based on the first MVP matrix and the second MVP matrix to obtain a third wrinkle mask image corresponding to the current video frame; Performing wrinkle removal on the face of a person on the current video frame according to the third wrinkle mask image to obtain a target video. A method for removing wrinkles for portrait videos is provided, which includes the above steps.

[0005] Preferably, the step of selecting a video frame of the face of the same person that meets a preset condition in the video to be processed as the first key frame image includes: Selecting a video frame of the face of the same person whose face is not blocked in the video to be processed as a candidate key frame image; Performing hierarchical clustering on the candidate key frame images according to a preset aggregation strategy, dividing the candidate key frame images into N categories by the same expression, and extracting, as the first key frame image, the video frame with the shortest distance from the clustering center in the same category for each category.

[0006] Preferably, the fact that the face of the person is not blocked is determined by: skin area of the face of the person / area of the polygon formed by the outer contour points of the face of the person > threshold.

[0007] Preferably, when the expression category label corresponding to the first key frame image is the first category, according to the relationship between the centroid of each connected component and a preset range, a value is assigned to the corresponding pixel value in the first wrinkle mask image to obtain a second wrinkle mask image. The steps are as follows: When the centroid is within the first circle and the maximum distance from the centroid of the pixels in the connected component > 0.75 × the first radius, the wrinkle type is determined as a laughing wrinkle, and the corresponding pixel value in the first wrinkle mask image is set to 128. Here, the first radius is the distance from the midpoint between the left nostril point and the left mouth corner point to the left nostril point, and the first circle is a circle with the midpoint between the left nostril point and the left mouth corner point as the center and the first radius as the radius. This is a step. When the centroid is within the second circle and the maximum distance from the centroid of the pixels in the connected component > 0.75 × the second radius, the wrinkle type is determined as a laughing wrinkle, and the corresponding pixel value in the first wrinkle mask image is set to 128. Here, the second radius is the distance from the midpoint between the right nostril point and the right mouth corner point to the right nostril point, and the second circle is a circle with the midpoint between the right nostril point and the right mouth corner point as the center and the second radius as the radius. This is a step. When the distance between the centroid and the left eye corner point or the right eye corner point is less than the first threshold, the wrinkle type is determined as a wrinkle at the outer corner of the eye, and the corresponding pixel value in the first wrinkle mask image is set to 128. These steps include the above.

[0008] Preferably, when the expression category label corresponding to the first key frame image is the second category, according to the relationship between the centroid of each connected component and a preset range, a value is assigned to the corresponding pixel value in the first wrinkle mask image to obtain a second wrinkle mask image. The steps are as follows: Performing Gaussian filtering on the first key frame image to obtain a filtering result image. Calculating the difference between the pixel value of the first key frame image and the pixel value of the filtering result image. When the difference is less than the second threshold, the wrinkle type is determined as a relatively shallow wrinkle, and the corresponding pixel value in the first wrinkle mask image is set to 64. These steps include the above.

[0009] Preferably, the step of performing mapping processing on the second wrinkle mask image based on the first MVP matrix and the second MVP matrix to obtain a third wrinkle mask image corresponding to the current video frame is as follows: When it is determined that the expression category label of the current video frame exists in the first set, the second wrinkle mask image is back-projected onto the 3D human face model according to the first MVP matrix to obtain an intermediate wrinkle mask image, and the intermediate wrinkle mask image is projected onto the current video frame by the second MVP matrix to obtain the third wrinkle mask image corresponding to the current video frame.

[0010] Preferably, the step of removing wrinkles on the face of a person from the current video frame according to the third wrinkle mask image to obtain a target video is as follows: Performing texture replacement by searching for skin textures in adjacent non-wrinkle regions for regions with a pixel value of 255 in the third wrinkle mask image; Performing Gaussian filtering on regions with a pixel value of 128 in the third wrinkle mask image, and then fusing the Gaussian filtering result with the corresponding first key frame image according to a degree value of 50%; Performing a curve highlight operation on regions with a pixel value of 64 in the third wrinkle mask image.

[0011] In order to achieve the above object, the present invention also provides a key frame processing unit that acquires a video to be processed, selects, as a first key frame image, video frames of the face of the same person that satisfy preset conditions in the video to be processed, performs matching and labeling for each of the first key frame images based on preset expression category labels, and performs wrinkle segmentation on the first key frame images to obtain corresponding first wrinkle mask images; Performing connected component detection on the first wrinkle mask image to obtain N connected components, and assigning values to the corresponding pixel values in the first wrinkle mask image according to the relationship between the centroid of each connected component and a preset range to obtain a second wrinkle mask image, a mask image analysis unit for storing the expression category label corresponding to the first key frame image and the second wrinkle mask image in a first set, A first reconstruction unit for performing 3D reconstruction of a person's face on the first key frame image to obtain a corresponding 3D person face model and a first MVP matrix, Performing 3D reconstruction of a person's face on other video frames in the video to be processed to obtain a second MVP matrix, and performing mapping processing of the second wrinkle mask image based on the first MVP matrix and the second MVP matrix to obtain a third wrinkle mask image corresponding to the current video frame A second reconstruction unit, A wrinkle removal unit for removing wrinkles on a person's face from the current video frame according to the third wrinkle mask image to obtain a target video, and providing a wrinkle removal device for portrait videos.

[0012] In order to achieve the above object, the present invention also includes a processor, a memory, and a computer program stored in the memory. When the computer program is executed by the processor, it realizes the steps of the wrinkle removal method for portrait videos described in the above embodiments, and proposes a wrinkle removal device for portrait videos.

[0013] In order to achieve the above object, the present invention also proposes a computer-readable storage medium storing a computer program that realizes the steps of the wrinkle removal method for portrait videos described in the above embodiments when executed by a processor.

Advantages of the Invention

[0014] The beneficial effects are as follows.

[0015] In the above form, when removing wrinkles from a video by extracting key frames and performing wrinkle segmentation, it is not necessary to apply the wrinkle segmentation algorithm to each frame. Only when the facial expression of a person changes significantly, the wrinkle segmentation and analysis are performed again, effectively utilizing the originally non-real-time detection algorithm to meet the requirements for the performance and effect of wrinkle removal in a video scene.

[0016] In the above form, when removing wrinkles from a person's face in a video by performing 3D reconstruction of the person's face for a video frame and projection based on the mvp matrix, the wrinkle segmentation mask is accurately mapped, the positions of wrinkles in different sizes and poses of the face of the same person for each video frame are identified, misprocessing of non-wrinkle areas is avoided, and the effect of wrinkle removal is improved.

[0017] In the above form, by performing processing according to the face of the same person in a video, wrinkle removal processing can be performed respectively according to the faces of different people in the video, and the accuracy of wrinkle removal from the faces of people can be improved.

Brief Description of the Drawings

[0018] To more clearly explain the technical solution means in the embodiments of the present invention or the prior art, the following briefly describes the drawings necessary for the description of the embodiments or the prior art. However, the drawings in the following description are only some of the embodiments of the present invention, and it is obvious to those skilled in the art that other drawings can be obtained based on these drawings without creative effort.

Figure 1

Figure 2

Figure 3

Figure 4

Embodiments for Carrying out the Invention

[0019] To make the object, technical solution and advantages of the embodiments of the present invention clearer, hereinafter, with reference to the drawings in the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. However, it is obvious that the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative labor belong to the protection scope of the present invention. Therefore, the following detailed description of the embodiments of the present invention shown in the drawings does not limit the scope of the claimed present invention, but only shows selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative labor belong to the protection scope of the present invention.

[0020] In the description of the present invention, the terms "first" and "second" are for illustrative purposes only and should not be understood as indicating or implying relative importance or indicating the number of the indicated technical features. Therefore, the features limited by "first" and "second" may explicitly or implicitly include one or more of those features.

[0021] Hereinafter, the content of the present invention will be described in detail based on the embodiments.

[0022] Referring to FIG. 1, a flowchart of a wrinkle removal method for portrait video according to an embodiment of the present invention is shown.

[0023] In this embodiment, the method includes the following S11 to S15.

[0024] S11: Obtain the video to be processed, select the video frames of the faces of the same person that meet the preset conditions among the videos to be processed as the first key-frame images, perform matching labeling for each of the first key-frame images based on the preset expression category labels, and perform wrinkle segmentation on the first key-frame images to obtain corresponding first wrinkle mask images.

[0025] Here, the step of selecting the video frames of the faces of the same person that meet the preset conditions among the videos to be processed as the first key-frame images includes the following S11-1 and S11-2.

[0026] S11-1: Select the video frames of the faces of the same person whose faces are not blocked among the videos to be processed as candidate key-frame images.

[0027] S11-2: Perform hierarchical clustering on the candidate key-frame images according to the preset aggregation strategy, divide the candidate key-frame images into N categories by the same expression, and extract the video frame with the shortest distance from the clustering center in the same category as the first key-frame image for each category.

[0028] In this embodiment, among the videos, video frames of the faces of the same person whose faces are not blocked are selected as candidate key frames. Here, the fact that a person's face is not blocked is determined by the criterion that the skin area of the person's face / the area of the polygon formed by the outer contour points of the person's face > the threshold value. In this embodiment, the threshold value is set to 0.95. In particular, when there are multiple portraits in the video, each face of the same person is processed, and then the data for each portrait is stored in different sets Q1, Q2, ··· respectively. For the face data of the video frames that meet the above requirements, hierarchical clustering is performed according to the "bottom-up" agglomerative strategy, divided into multiple classes, and the classes with the same expression are combined. The expression is analyzed by the face points of the person at the center of the class to obtain an expression category label, for example, 0 - neutral expression, 1 - smiling with teeth, 2 - opening the mouth, 3 - being surprised, 4 - other expressions. Then, for each class, the video frame with the shortest distance from the clustering center is extracted as the key frame. In this embodiment, the key frame is at most 5 frames (that is, one expression is one frame), and there is a key frame corresponding to the expression category label that does not exist throughout the video. By using a deep learning model (as the model, a conventional model available for wrinkle segmentation, such as a U-Net model or other mature models, may be adopted.), the wrinkle segmentation of the key frame image is performed to obtain a single-channel wrinkle mask image M. The pixel values of the wrinkle area of this wrinkle mask image are marked as 255, and the pixel values of the non-wrinkle area are marked as 0.

[0029] S12: Perform connected component detection on the first wrinkle mask image to obtain N connected components, and assign values to the corresponding pixel values in the first wrinkle mask image according to the relationship between the centroid of each connected component and a preset range to obtain a second wrinkle mask image, and store the expression category label corresponding to the first key frame image and the second wrinkle mask image in a first set.

[0030] Furthermore, when the expression category label corresponding to the first key frame image is the first category, the step of obtaining the second wrinkle mask image by assigning a value to the corresponding pixel value in the first wrinkle mask image according to the relationship between the centroid of each connected component and a preset range includes the following S12-1 to S12-3.

[0031] S12-1: When the centroid is within the first circle and the maximum distance from the centroid of the pixels in the connected component > 0.75 × the first radius, determine the wrinkle type as a laughing wrinkle, set the corresponding pixel value in the first wrinkle mask image to 128, where the first radius is the distance from the midpoint between the left nostril point and the left mouth corner point to the left nostril point, and the first circle is a circle with the midpoint between the left nostril point and the left mouth corner point as the center and the first radius as the radius.

[0032] S12-2: When the centroid is within the second circle and the maximum distance from the centroid of the pixels in the connected component > 0.75 × the second radius, determine the wrinkle type as a laughing wrinkle, set the corresponding pixel value in the first wrinkle mask image to 128, where the second radius is the distance from the midpoint between the right nostril point and the right mouth corner point to the right nostril point, and the second circle is a circle with the midpoint between the right nostril point and the right mouth corner point as the center and the second radius as the radius.

[0033] S12-3: When the distance between the centroid and the left eye corner point or the right eye corner point is less than the first threshold, determine the wrinkle type as a crow's feet wrinkle, and set the corresponding pixel value in the first wrinkle mask image to 128.

[0034] And when the expression category label corresponding to the first key frame image is the second category, the step of obtaining the second wrinkle mask image by assigning a value to the corresponding pixel value in the first wrinkle mask image according to the relationship between the centroid of each connected component and a preset range includes S12-4.

[0035] S12-4: Perform Gaussian filtering on the first key frame image to obtain a filtered result image, calculate the difference between the pixel values of the first key frame image and the filtered result image, and if the difference is less than the second threshold, determine the wrinkle type as relatively shallow wrinkles, and set the corresponding pixel value in the first wrinkle mask image to 64.

[0036] Refer to the schematic diagram of the partial processing process of the image shown in FIG. 2 (in the figure, the left is the key frame image, the middle is the first wrinkle mask image, and the right is the second wrinkle mask image). In this embodiment, laughter wrinkles and crow's feet are dynamic wrinkles caused by expressions. If they are completely removed, it will look unnatural. Therefore, laughter wrinkles or crow's feet in the wrinkle mask image are further specified. Specifically, Perform connected component detection on the wrinkle mask image to obtain N connected components, and calculate the centroid of each connected component by traversing the connected components. Here, the centroid refers to the geometric center. The geometric center or centroid of an object X in n-dimensional space is the intersection of all hyperplanes that divide X into two parts with equal moments (which may also be called the average of all points in X). When the mass of an object is uniformly distributed, the centroid is the center of gravity.

[0037] In the current scene, the calculation formula for the centroid of the connected component is as follows.

[0038]

Equation

[0039] (1) When the corresponding expression category label of the key frame image is 2 or 3, perform the following analysis processing.

[0040] a. If the center of gravity is within the circle centered at Mid_left with radius Rleft, and the maximum distance from the center of gravity of the pixels in the connected component > 0.75 × Rleft, it is determined to be a laughter wrinkle, and the corresponding pixel value in the wrinkle mask image M is set to 128.

[0041] b. If the center of gravity is within the circle centered at Mid_right with radius Rright, and the maximum distance from the center of gravity of the pixels in the connected component > 0.75 × Rright, it is determined to be a laughter wrinkle, and the corresponding pixel value in the wrinkle mask image M is set to 128.

[0042] c. If the distance from the center of gravity to the left or right eye corner point is less than the threshold A, it is determined to be a crow's foot wrinkle, and the corresponding pixel value in the wrinkle mask image M is set to 128.

[0043] Here, the points Mid_left and Mid_right are the midpoints between the left and right nostril points and the mouth corner points respectively. The radius Rleft is the distance from Mid_left to the left nostril point, and Rright is the distance from Mid_right to the right nostril point.

[0044] (2) In the case of a wrinkle connected component not belonging to any of a, b, c above, and when the corresponding expression category label of the keyframe image is 0, 1, or 4, the following analysis process is performed. Perform Gaussian filtering on the keyframe image to obtain a filtered result image. Calculate the difference in pixel values between the original image and the filtered result image. If the result of subtracting the original image from the filtered result image exceeds the threshold B for wrinkles, it is determined to be a deep wrinkle and no processing is performed. Otherwise, it is determined to be a relatively shallow wrinkle, and the corresponding pixel value in the wrinkle mask image M is set to 64.

[0045] S13: Perform 3D reconstruction of the person's face on the first keyframe image to obtain the corresponding 3D person's face model and the first mvp matrix.

[0046] In this embodiment, the obtained 3D human face reconstruction model is obtained by 3D Morphable Model (3DMM). This method regards the human face space as a linear space, creates a human face space basis based on pre-collected 3D human face data, and projects the linear combination of pre-created 3D human face data to approach the human face in the 2D image. The human face space basis includes the face of the 3D average human, the face type basis of the human that constitutes the 3D human face type model, and the expression basis that constitutes the 3D expression model. The basic formula of 3DMM is shown as follows.

[0047] [Number] represents the expression basis that constitutes the 3D expression model, e j is the coefficient of the expression basis, and n and m respectively represent the numbers of the human face type basis and the expression basis.

[0048] The MVP matrix refers to three transformation matrices including model, view, and projection. The 3D human face model is projected onto a 2D image. Its initial parameters are estimated based on the feature points of the human face space basis, including the position of the camera, the rotation angle of the image plane, each component of direct light and ambient light, the image contrast, etc. Based on the extracted human face feature points, the human face space basis, and the initial parameters of the projection matrix, fitting is performed to obtain the 3D human face model corresponding to the image. That is, based on the 3D model data with the same fixed floating-point number and topology structure, the parameters of the linear combination of the 3D model are obtained from the distance between the minimized projection of the feature points on the 3D model and the 2D feature points, and the parameters are fitted to obtain the 3D human face model corresponding to the human face image and the projection matrix. Its formula is as follows.

[0049] [Formula 3] Error = MVP × M - P 2d Formula (2) In the formula, MVP represents the projection matrix of the 3D model, P 2drepresents the facial feature points of the facial image of a person detected from the facial feature points of the person, and M represents the 3D facial model of the person. By minimizing Error in combination with Equation (1), a solution is obtained to obtain the 3D facial model of the person and the mvp matrix.

[0050] S14: Perform 3D reconstruction of the face of the person on other video frames of the video to be processed to obtain a second mvp matrix, and perform mapping processing on the second wrinkle mask image based on the first mvp matrix and the second mvp matrix to obtain a third wrinkle mask image corresponding to the current video frame.

[0051] Here, the step of performing mapping processing on the second wrinkle mask image based on the first mvp matrix and the second mvp matrix to obtain a third wrinkle mask image corresponding to the current video frame is as follows: If it is determined that the expression category label of the current video frame exists in the first set, the second wrinkle mask image is back-projected onto the 3D facial model of the person according to the first mvp matrix to obtain an intermediate wrinkle mask image, and the intermediate wrinkle mask image is projected onto the current video frame by the second mvp matrix to obtain the third wrinkle mask image corresponding to the current video frame.

[0052] In this embodiment, 3D reconstruction of the face of the person is performed on other video frames in the video to obtain an mvp matrix P'. First, it is determined whether the expression category label of the current video frame exists in set Q. (1) If it exists, the corresponding wrinkle mask image is back-projected onto the 3D facial model of the person by the mvp matrix P of the key frame image to obtain an intermediate wrinkle mask image under the uv mapping coordinates of the 3D facial model of the person, and the intermediate wrinkle mask image is projected onto the current video frame image by the mvp matrix P' to obtain the wrinkle mask M' (the third wrinkle mask image) of the current video frame. (2) If it does not exist, process the current video frame according to the above steps S11 to S13, add the obtained expression category label of the current video frame to the set Q together with the corresponding wrinkle mask image, and then perform the mapping process.

[0053] S15: Perform wrinkle removal on the face of the person in the current video frame according to the third wrinkle mask image to obtain a target video.

[0054] Here, the step of performing wrinkle removal on the face of the person in the current video frame according to the third wrinkle mask image to obtain a target video includes the following S15-1 to S15-3.

[0055] S15-1: For the region with a pixel value of 255 in the third wrinkle mask image, perform texture replacement by searching for skin texture in the adjacent non-wrinkle region.

[0056] S15-2: After performing Gaussian filtering on the region with a pixel value of 128 in the third wrinkle mask image, fuse the Gaussian filtering result and the corresponding first key frame image according to a degree value of 50%.

[0057] S15-3: Perform a curve highlight operation on the region with a pixel value of 64 in the third wrinkle mask image.

[0058] In this embodiment, the step of performing wrinkle removal on the current video frame image according to the wrinkle mask M' of the current video frame includes the following (1) to (3).

[0059] (1) Regard the region with a pixel value of 255 in the wrinkle mask M' as a deep wrinkle, and directly replace the texture by searching for skin texture in the adjacent non-wrinkle region.

[0060] (2) For the region of the wrinkle mask M' with a pixel value of 128, it is regarded as a dynamic wrinkle that can be made deeper or shallower, and Gaussian filtering is performed. The Gaussian filtering result and the original image are fused according to the degree value of 50%. Thereby, when beautifying the expression lines to a certain extent, since the expression lines are dynamic wrinkles that exist regardless of age, they are not completely removed, thereby improving the processing effect of the expression wrinkles and enhancing the overall naturalness. That is, the fused result image = 0.5 × original image + 0.5 × Gaussian filtering result image.

[0061] (3) For the region of the wrinkle mask M' with a pixel value of 64, it is regarded as a relatively shallow dark wrinkle, and a curve highlight operation is performed. Further, in order to avoid the problem of turning white, the pixel value after the curve highlight is restricted to be less than or equal to the pixel value after Gaussian filtering. Refer to the schematic diagram of the wrinkle removal process of the face of a person in another video frame in the video shown in FIG. 3 (in the figure, the left is the original image of another video frame, the middle is the third wrinkle mask image, and the right is the final wrinkle removal result image corresponding to the original image).

[0062] Refer to the structural schematic diagram of the wrinkle removal device for portrait video according to an embodiment of the present invention shown in FIG. 4.

[0063] In this embodiment, the device 40 obtains the video to be processed, selects the video frame of the face of the same person that satisfies the preset conditions in the video to be processed as the first key frame image, performs matching and labeling for each of the first key frame images based on the preset expression category label, and performs wrinkle segmentation on the first key frame image to obtain a corresponding first wrinkle mask image, a key frame processing unit 41, performs connected component detection on the first wrinkle mask image to obtain N connected components, assigns values to the corresponding pixel values in the first wrinkle mask image according to the relationship between the centroid of each connected component and the preset range, obtains a second wrinkle mask image, and stores the expression category label corresponding to the first key frame image and the second wrinkle mask image in a first set, a mask image analysis unit 42 A first reconstruction unit 43 for performing 3D reconstruction of a person's face on the first key frame image to obtain a corresponding 3D person face model and a first mvp matrix; A second reconstruction unit 44 for performing 3D reconstruction of a person's face on other video frames of the video to be processed to obtain a second mvp matrix, and performing mapping processing of the second wrinkle mask image based on the first mvp matrix and the second mvp matrix to obtain a third wrinkle mask image corresponding to the current video frame; A wrinkle removal unit 45 for removing wrinkles on a person's face from the current video frame according to the third wrinkle mask image to obtain a target video.

[0064] Since the modules of each unit of the apparatus 40 can respectively execute the corresponding steps of the above method embodiment, the modules of each unit will not be described in detail herein. For details, reference may be made to the description of the corresponding steps above.

[0065] An embodiment of the present invention also provides a wrinkle removal device for portrait videos. The device includes the above-mentioned wrinkle removal device for portrait videos. The wrinkle removal device for portrait videos may adopt the structure in the embodiment of FIG. 4. Therefore, it can execute the technical solution of the method embodiment shown in FIG. 1, and their implementation principles and technical effects are similar. For details, reference may be made to the relevant descriptions of the above embodiments, and thus will not be described in detail here.

[0066] The wrinkle removal device for portrait videos is a device having a photographing function, an image processing function, or an image display function, such as a mobile phone, a digital camera, or a tablet. The device may include components such as a memory, a processor, an input unit, a display unit, and a power supply.

[0067] The memory may be used to store software programs and modules, and the processor executes the software programs and modules stored in the memory to perform various functional applications and data processing. The memory mainly includes a program storage area and a data storage area. The program storage area may store an operating system, an application program required for at least one function (such as an image playback function, etc.). The data storage area may store data created according to the use of the device. Further, the memory may include a high-speed random access memory and may also include a non-volatile memory such as at least one magnetic disk storage device, a flash memory device, or other volatile solid storage devices. Therefore, the memory may further include a memory controller that provides access to the memory by the processor and the input unit.

[0068] The input unit can receive input numerical, character, or image information and generate keyboard, mouse, lever, optical, or trackball signal inputs related to user settings and function control. Specifically, the input unit of this embodiment may include, in addition to a camera, a touch-sensitive surface (such as a touch display screen, etc.) and other input devices.

[0069] The display unit can be used to display various graphical user interfaces of the device, which may be composed of information input by the user or information provided to the user, as well as graphics, text, icons, videos, and any combination thereof. The display unit may include a display panel, or the display panel may be configured in the form of a liquid crystal display (LCD) or an organic light-emitting diode (OLED). Further, the touch-sensitive surface may cover the display panel. When the touch-sensitive surface detects a touch operation thereon or in its vicinity, it transfers the type of touch event to the processor for determination. Then, the processor provides a visual output corresponding to the display panel according to the type of touch event.

[0070] Embodiments of the present invention further provide a computer-readable storage medium, which may be the computer-readable storage medium included in the memory in the above embodiments, or may exist alone and not be incorporated into the device. At least one instruction that is loaded and executed by the processor to implement the portrait wrinkle removal method shown in FIG. 1 is stored in this computer-readable storage medium. The computer-readable storage medium may be a read-only memory, a magnetic disk, an optical disk, etc.

[0071] It should be noted that each embodiment in this specification is described step by step. Each embodiment focuses on explaining the differences from other embodiments. For the same or similar parts between each embodiment, reference may be made to each other. The apparatus embodiments, device embodiments, and storage medium embodiments are substantially similar to the method embodiments, so their descriptions are relatively simple. For related points, reference may be made to the partial description of the method embodiments.

[0072] Also, in this specification, the terms "comprise", "include" or any other variation thereof are intended to include non-exclusive inclusion, such that a process, method, article or apparatus that comprises a series of elements includes not only those elements but also other elements not expressly listed, or elements specific to such a process, method, article or apparatus. Further, unless otherwise limited, elements defined by the phrase "comprising..." do not exclude the further presence of the same elements in a process, method, article or apparatus that comprises such elements.

[0073] The above description shows and describes preferred embodiments of the present invention, but the present invention is not limited to the forms disclosed herein, should not be considered as excluding other embodiments, and is applicable to various other combinations, modifications, and environments, and it should be understood that within the scope of the concept of this specification, changes can be made by the above teachings or the techniques and knowledge in the relevant field. Changes and variations made by those skilled in the art are intended to be included within the scope of protection of the appended claims of the present invention as long as they do not depart from the spirit and scope of the present invention.

Claims

1. A method for removing wrinkles in portrait videos, comprising: obtaining a video to be processed, selecting, as a first key-frame image, a video frame of a face of the same person satisfying a preset condition among the videos to be processed, performing matching labeling for each of the first key-frame images based on a preset expression category label, and performing wrinkle segmentation on the first key-frame images to obtain corresponding first wrinkle mask images; performing connected component detection on the first wrinkle mask images to obtain N connected components, assigning values to corresponding pixel values in the first wrinkle mask images according to the relationship between the centroid of each connected component and a preset range to obtain second wrinkle mask images, and storing the expression category labels corresponding to the first key-frame images and the second wrinkle mask images in a first set; performing 3D reconstruction of the face of the person on the first key-frame images to obtain corresponding 3D face models of the person and first mvp matrices; performing 3D reconstruction of the face of the person on other video frames among the videos to be processed to obtain second mvp matrices, performing mapping processing of the second wrinkle mask images based on the first mvp matrices and the second mvp matrices, and obtaining third wrinkle mask images corresponding to the current video frames; performing wrinkle removal of the face of the person on the current video frames according to the third wrinkle mask images to obtain target videos. A method for removing wrinkles in portrait videos, characterized by comprising the above steps.

2. The step of selecting, as a first key-frame image, a video frame of a face of the same person satisfying a preset condition among the videos to be processed includes: selecting, as candidate key-frame images, video frames of faces of the same person whose faces are not occluded among the videos to be processed; performing hierarchical clustering on the candidate key-frame images according to a preset aggregation strategy, dividing the candidate key-frame images into N categories by the same expression, and extracting, as the first key-frame images, the video frames with the shortest distance from the clustering center in the same category for each category. The method for removing wrinkles in portrait videos according to Claim 1, characterized by comprising the above steps.

3. The fact that the face of the person is not occluded is determined by the ratio of the skin area of the person's face to the area of the polygon formed by the outer contour points of the person's face > a threshold value. The method for removing wrinkles for portrait videos according to claim 2 is characterized by this.

4. When the expression category label corresponding to the first key frame image is the first category, The step of obtaining the second wrinkle mask image by assigning a value to the corresponding pixel value in the first wrinkle mask image according to the relationship between the centroid of each connected component and a preset range is as follows: When the centroid is within the first circle and the maximum distance from the centroid of the pixels within the connected component > 0.75 × the first radius, it is determined that the wrinkle type is a laughing wrinkle, and the corresponding pixel value in the first wrinkle mask image is set to 128. The first radius is the distance from the midpoint between the left nostril point and the left mouth corner point to the left nostril point. The first circle is a circle with the midpoint between the left nostril point and the left mouth corner point as the center and the first radius as the radius. This is the step. When the centroid is within the second circle and the maximum distance from the centroid of the pixels within the connected component > 0.75 × the second radius, it is determined that the wrinkle type is a laughing wrinkle, and the corresponding pixel value in the first wrinkle mask image is set to 128. The second radius is the distance from the midpoint between the right nostril point and the right mouth corner point to the right nostril point. The second circle is a circle with the midpoint between the right nostril point and the right mouth corner point as the center and the second radius as the radius. This is the step. When the distance between the centroid and the left eye corner point or the right eye corner point is less than the first threshold value, it is determined that the wrinkle type is a crow's feet wrinkle, and the corresponding pixel value in the first wrinkle mask image is set to 128. The method for removing wrinkles for portrait videos according to claim 1 includes this and is characterized by this.

5. When the expression category label corresponding to the first key frame image is the second category, The step of obtaining the second wrinkle mask image by assigning a value to the corresponding pixel value in the first wrinkle mask image according to the relationship between the centroid of each connected component and a preset range is as follows: The step of performing Gaussian filtering on the first key frame image to obtain a filtered result image. Calculate the difference between the pixel value of the first key frame image and the pixel value of the filtered result image. When the difference is less than the second threshold, determine that the wrinkle type is a relatively shallow wrinkle, and set the corresponding pixel value in the first wrinkle mask image to 64. The method for removing wrinkles for portrait video according to claim 1 or 4, characterized by including the step of

6. The step of performing mapping processing on the second wrinkle mask image based on the first mvp matrix and the second mvp matrix to obtain a third wrinkle mask image corresponding to the current video frame is When it is determined that the facial expression category label of the current video frame exists in the first set, back-project the second wrinkle mask image onto the face model of the 3D person according to the first mvp matrix to obtain an intermediate wrinkle mask image, and project the intermediate wrinkle mask image onto the current video frame by the second mvp matrix to obtain the third wrinkle mask image corresponding to the current video frame. The method for removing wrinkles for portrait video according to claim 1, characterized by including the step of

7. The step of removing wrinkles on the face of a person from the current video frame according to the third wrinkle mask image to obtain a target video is The step of performing texture replacement by searching for skin texture in an adjacent non-wrinkle area for an area where the pixel value in the third wrinkle mask image is 255 The step of performing Gaussian filtering on an area where the pixel value in the third wrinkle mask image is 128, and then fusing the Gaussian filtering result and the corresponding first key frame image according to a degree value of 50% The step of performing a curve highlight operation on an area where the pixel value in the third wrinkle mask image is 64. The method for removing wrinkles for portrait video according to claim 1, characterized by including the step of

8. A wrinkle removal device for portrait video, comprising Obtain the video to be processed, select, as the first key-frame image, video frames of the face of the same person that satisfy preset conditions among the videos to be processed, perform matching labeling for each of the first key-frame images based on a preset expression category label, and perform wrinkle segmentation on the first key-frame images to obtain corresponding first wrinkle mask images, a key-frame processing unit for perform connected component detection on the first wrinkle mask images to obtain N connected components, assign values to the corresponding pixel values in the first wrinkle mask images according to the relationship between the centroid of each connected component and a preset range to obtain second wrinkle mask images, and a mask image analysis unit for storing the expression category label corresponding to the first key-frame images and the second wrinkle mask images in a first set perform 3D reconstruction of the face of the person on the first key-frame images to obtain corresponding 3D face models of the person and first mvp matrices, a first reconstruction unit for perform 3D reconstruction of the face of the person on other video frames among the videos to be processed to obtain second mvp matrices, perform mapping processing of the second wrinkle mask images based on the first mvp matrices and the second mvp matrices, and obtain third wrinkle mask images corresponding to the current video frames, a second reconstruction unit for perform wrinkle removal of the face of the person on the current video frame according to the third wrinkle mask images to obtain a target video, a wrinkle removal unit, characterized by comprising a wrinkle removal device for portrait videos

9. A wrinkle removal device for portrait videos, comprising a processor, a memory, and a computer program stored in the memory, wherein when the computer program is executed by the processor, the steps of the wrinkle removal method for portrait videos according to any one of Claims 1 to 7 are realized, characterized by a wrinkle removal device for portrait videos

10. A computer-readable storage medium, storing a computer program that, when executed by a processor, realizes the steps of the wrinkle removal method for portrait videos according to any one of Claims 1 to 7, characterized by a computer-readable storage medium

Citation Information

Patent Citations

  • Expression generation method, device and equipment and storage medium

    CN111489426A

  • Image correction device

    JP2009111947A

  • Age conversion method based on age and environmental factors for each facial part, recording medium and apparatus for performing the same

    JP2018531449A