Face image processing methods, apparatus, computer equipment and storage media

By selecting temporally correlated image frames from multiple face image frames and determining smoothing weight parameters based on time difference to process feature points, the problem of low face feature localization accuracy in traditional methods is solved, thereby improving the accuracy of face image processing and the applicability of application scenarios.

CN116824651BActive Publication Date: 2025-10-31TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210281024.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-22
Publication Date
2025-10-31
Estimated Expiration
2042-03-22

AI Technical Summary

Technical Problem

Traditional face image processing methods suffer from low accuracy in locating key facial feature points due to the complexity and diversity of facial expressions, which affects the accuracy of facial feature contour lines.

Method used

From multiple face image frames carrying time information, consecutive image frames that are temporally correlated with the target time are selected. An initial set of feature points is obtained through face feature detection. Based on the time difference, a smoothing weight parameter is determined, and the initial feature points are smoothed to obtain the target feature points and determine the face feature contour.

Benefits of technology

It improves the accuracy of face image processing results, reduces the impact of interference points on facial feature contours, and expands the applicability of application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116824651B_ABST
    Figure CN116824651B_ABST
Patent Text Reader

Abstract

This application relates to a face image processing method, apparatus, computer device, computer-readable storage medium, and computer program product. The method includes: selecting consecutive image frames that are temporally correlated with a target time from multiple face image frames carrying time information; performing face feature detection on each image frame in the consecutive image frames to obtain an initial feature point set corresponding to each image frame; determining a smoothing weight parameter corresponding to each initial feature point set based on the time difference between the time matched by the image frame and the target time; smoothing the initial feature points in each initial feature point set according to the smoothing weight parameter to obtain target feature points matching the target time; and determining a face feature contour line matching the target time based on each target feature point. Using the above method can reduce the influence of interference points on the face feature contour line, which is beneficial to improving the accuracy of the face image processing results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to a face image processing method, apparatus, computer device, computer-readable storage medium, and computer program product. Background Technology

[0002] With the rapid development of the cultural industry, the demand for high-quality characters is increasing in various sectors such as film and television works, short videos, and games. To reconstruct a high-quality character, it is necessary to process facial image data, determine the corresponding facial feature contours, and then complete the character reconstruction based on the facial feature contours.

[0003] Traditional face image processing methods first acquire a binarized image of the face, then locate facial feature keys based on the grayscale values ​​of each pixel in the binarized image, thereby obtaining the corresponding facial feature contour. However, due to the complexity and diversity of facial expressions, the morphology of facial features varies under different expressions, which may introduce interference points, affecting the accuracy of facial feature key point localization and consequently the accuracy of the facial feature contour. Therefore, traditional face image processing methods suffer from low accuracy. Summary of the Invention

[0004] Therefore, it is necessary to provide a face image processing method, apparatus, computer equipment, computer-readable storage medium, and computer program product that can improve the accuracy of processing results in response to the above-mentioned technical problems.

[0005] Firstly, this application provides a method for processing facial images. The method includes:

[0006] From multiple face image frames carrying time information, select consecutive image frames that have a temporal correlation with the target time.

[0007] Face feature detection is performed on each image frame in the continuous image frames to obtain the initial feature point set corresponding to each image frame;

[0008] Based on the time difference between the time matched by the image frame and the target time, the smoothing weight parameters corresponding to each of the initial feature point sets are determined.

[0009] According to the smoothing weight parameters, the initial feature points in each of the initial feature point sets are smoothed to obtain target feature points that match the target time.

[0010] Based on each of the target feature points, a facial feature contour line matching the target time is determined.

[0011] Secondly, this application provides a face image processing apparatus. The apparatus includes:

[0012] The continuous image frame filtering module is used to filter out continuous image frames that are temporally correlated with the target time from multiple face image frames carrying time information.

[0013] The initial feature point set determination module is used to perform face feature detection on each image frame in the continuous image frames to obtain the initial feature point set corresponding to each image frame.

[0014] The smoothing weight parameter determination module is used to determine the smoothing weight parameters corresponding to each of the initial feature point sets based on the time difference between the time matched by the image frame and the target time.

[0015] The target feature point determination module is used to smooth the initial feature points in each of the initial feature point sets according to the smoothing weight parameters, so as to obtain target feature points that match the target time.

[0016] The facial feature contour determination module is used to determine a facial feature contour that matches the target time based on each of the target feature points.

[0017] Thirdly, this application also provides a computer device. The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to perform the following steps:

[0018] From multiple face image frames carrying time information, select consecutive image frames that have a temporal correlation with the target time.

[0019] Face feature detection is performed on each image frame in the continuous image frames to obtain the initial feature point set corresponding to each image frame;

[0020] Based on the time difference between the time matched by the image frame and the target time, the smoothing weight parameters corresponding to each of the initial feature point sets are determined.

[0021] According to the smoothing weight parameters, the initial feature points in each of the initial feature point sets are smoothed to obtain target feature points that match the target time.

[0022] Based on each of the target feature points, a facial feature contour line matching the target time is determined.

[0023] Fourthly, this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, which, when executed by a processor, performs the following steps:

[0024] From multiple face image frames carrying time information, select consecutive image frames that have a temporal correlation with the target time.

[0025] Face feature detection is performed on each image frame in the continuous image frames to obtain the initial feature point set corresponding to each image frame;

[0026] Based on the time difference between the time matched by the image frame and the target time, the smoothing weight parameters corresponding to each of the initial feature point sets are determined.

[0027] According to the smoothing weight parameters, the initial feature points in each of the initial feature point sets are smoothed to obtain target feature points that match the target time.

[0028] Based on each of the target feature points, a facial feature contour line matching the target time is determined.

[0029] Fifthly, this application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, performs the following steps:

[0030] From multiple face image frames carrying time information, select consecutive image frames that have a temporal correlation with the target time.

[0031] Face feature detection is performed on each image frame in the continuous image frames to obtain the initial feature point set corresponding to each image frame;

[0032] Based on the time difference between the time matched by the image frame and the target time, the smoothing weight parameters corresponding to each of the initial feature point sets are determined.

[0033] According to the smoothing weight parameters, the initial feature points in each of the initial feature point sets are smoothed to obtain target feature points that match the target time.

[0034] Based on each of the target feature points, a facial feature contour line matching the target time is determined.

[0035] The aforementioned face image processing method, apparatus, computer equipment, computer-readable storage medium, and computer program product first select consecutive image frames that are temporally correlated with the target time from multiple face image frames. Based on the time difference between the time matched by each image frame and the target time, a smoothing weight parameter corresponding to the initial feature point set of each image frame is determined. Then, the initial feature points in each initial feature point set are smoothed according to the smoothing weight parameter to obtain target feature points that match the target time. This allows for the determination of the face feature contour line that matches the target time. This can correct the feature points in the target face image frame corresponding to the target time, reduce the influence of interference points on the face feature contour line, and improve the accuracy of the face image processing results. Attached Figure Description

[0036] Figure 1 This is a schematic diagram illustrating an application scenario of a face image processing method in one embodiment;

[0037] Figure 2 This is a flowchart illustrating a face image processing method in one embodiment;

[0038] Figure 3 This is a flowchart illustrating a face image processing method in another embodiment;

[0039] Figure 4 This is a schematic diagram illustrating the temporal relationship between the target time and the target image frame in one embodiment;

[0040] Figure 5 This is a schematic diagram showing the location of the first set of lip feature points obtained based on a facial feature point detection model in one embodiment;

[0041] Figure 6 This is a schematic diagram illustrating the process of obtaining the second inner lip feature point set corresponding to the image frame based on a semantic segmentation algorithm in one embodiment.

[0042] Figure 7 This is a flowchart illustrating a face image processing method in yet another embodiment;

[0043] Figure 8 This is a schematic diagram of four sets of lip contour lines obtained based on a face image processing method in one embodiment;

[0044] Figure 9 This is a structural block diagram of a face image processing device in one embodiment;

[0045] Figure 10 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0046] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0047] The face image processing method provided in this application can be applied to a terminal, a server, or a system including both a terminal and a server, and is implemented through the interaction between the terminal and the server. Figure 1 In the application environment shown, during the execution of the face image processing method by computer device 102, continuous image frames with temporal correlation to the target time are selected from multiple face image frames carrying time information; face feature detection is performed on each image frame in the continuous image frames to obtain an initial feature point set corresponding to each image frame; based on the time difference between the time matched by the image frame and the target time, a smoothing weight parameter corresponding to each initial feature point set is determined; according to the smoothing weight parameter, the initial feature points in each initial feature point set are smoothed to obtain target feature points matching the target time; based on each target feature point, a face feature contour line matching the target time is determined. The computer device 102 can be a terminal or a server. The terminal can be, but is not limited to, various desktop computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, etc. Portable wearable devices can be smartwatches, smart bracelets, head-mounted devices, etc. The server can be implemented using a standalone server or a server cluster composed of multiple servers. The embodiments of the present invention can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, smart transportation, and assisted driving.

[0048] In one embodiment, such as Figure 2 As shown, a face image processing method is provided. This embodiment illustrates the application of this method to a terminal. It is understood that this method can also be applied to a server, or to a system including a terminal and a server, and can be implemented through the interaction between the terminal and the server.

[0049] In this embodiment, the method includes the following steps:

[0050] Step S201: Select consecutive image frames that are temporally correlated with the target time from multiple face image frames carrying time information.

[0051] Here, the target time refers to the time at which the desired facial feature contour line is matched. A facial image frame refers to an image frame containing a facial image. The time information carried by a facial image frame refers to the time at which the facial image frame matches within the image frame sequence. A consecutive image frame refers to multiple image frames carrying consecutive time information. For example, if the inter-frame time interval of image frames is 0.1S, then the time information carried by consecutive image frames are: t, t+0.1S, t+0.2S...t+n*0.1S. Furthermore, consecutive image frames that are temporally associated with the target time can refer to image frames whose time information carried by each image frame in the consecutive image frames has a time difference from the target time that is less than a set time. Correspondingly, the image frames that are temporally associated with the target time can include image frames among multiple facial image frames whose time information is equal to the target time, as well as image frames whose time information is before and / or after the target time.

[0052] Specifically, the terminal selects multiple consecutive face image frames that are temporally correlated with the target time from multiple face image frames carrying time information, forming a continuous image frame to be processed. For example, the terminal can select continuous image frames from multiple face image frames based on the target time and the desired position of the target time in the time information carried by each image frame in the continuous image frame.

[0053] Step S203: Perform face feature detection on each image frame in the continuous image frames to obtain the initial feature point set corresponding to each image frame.

[0054] Facial feature detection refers to the process of acquiring the coordinate information of facial feature points. These facial feature points can include boundary points of parts such as lips, eyes, eyebrows, nose, and ears. The initial feature point set corresponding to a certain image frame includes multiple initial feature points in the face image of that image frame and their corresponding coordinate information.

[0055] Specifically, the terminal can employ a preset face feature detection method to perform face feature detection on each image frame in a series of image frames, obtaining an initial set of feature points corresponding to each image frame. This face feature detection method can be based on models such as ASM (Active Shape Model) and AAM (Active Appreance Model); it can also be based on cascaded shape regression algorithms such as CPR (Cascaded Pose Regression); or it can be based on deep learning face feature detection methods, such as DCNN (Dimensional Convolutional Neural Network) and TCDCN (Tasks-Constrained Deep Convolutional Network).

[0056] Step S205: Based on the time difference between the time matched by the image frame and the target time, determine the smoothing weight parameters corresponding to each initial feature point set.

[0057] The smoothing weight parameter determines the contribution of each initial feature point in the initial feature point set to the target feature point obtained after smoothing. Specifically, based on the time difference between the time matched by the image frame and the target time, the terminal can determine the smoothing weight parameter corresponding to each initial feature point set, such that the contribution of each initial feature point in the image frame to the target feature point is inversely correlated with the time difference associated with that image frame. That is, the closer the time matched by the image frame is to the target time, the greater the contribution of the initial feature point corresponding to that image frame to the target feature point.

[0058] Step S207: According to the smoothing weight parameters, the initial feature points in each initial feature point set are smoothed to obtain target feature points that match the target time.

[0059] Smoothing refers to the process of obtaining the corresponding target feature points based on multiple initial feature points using a smoothing algorithm. This smoothing algorithm can include at least one of the following: moving average, exponential average, or Gaussian moving average.

[0060] Specifically, by smoothing the initial feature points at each location in the initial feature point set according to the smoothing weight parameters, the target feature point at that location that matches the target time can be obtained. Based on the above method, by smoothing the initial feature points at each location separately, target feature points at all locations that match the target time can be obtained.

[0061] Step S209: Determine the facial feature contour line that matches the target time based on each target feature point.

[0062] In this context, facial feature contour lines refer to the boundary lines between different objects in a facial image, as well as between objects and the background. Specifically, these facial feature contour lines can include lip contour lines, eye contour lines, hair contour lines, and so on.

[0063] Specifically, based on each target feature point, the terminal can determine a facial feature contour line that matches the target time. For example, the terminal can directly connect the target feature points into a line to form a facial feature contour line. It can be understood that the facial feature contour line matching the target time obtained in this application can provide a foundation for subsequent related processing such as role reconstruction, expression analysis, and automatic facial recognition for the target time.

[0064] In one embodiment, step S209 includes: performing polynomial curve fitting on each target feature point for a preset number of times based on the least squares method to determine the facial feature contour line that matches the target time.

[0065] Curve fitting, in particular, is a data processing method that uses a continuous curve to approximate or analogize the functional relationship between coordinates represented by a set of discrete points on a plane. Curve fitting can be achieved using the least squares method. Polynomial curve fitting means that the functional relationship can be expressed using a polynomial. Furthermore, the preset degree can be 3rd, 4th, or 5th order.

[0066] Specifically, suppose the analytical expression of the polynomial curve is:

[0067] y = a0 + a1x + ... + a k x k (1)

[0068] In the formula, a k Let be the expansion coefficient. Then the sum of the distances from each target feature point to this curve, i.e., the sum of squared deviations, is:

[0069]

[0070] To obtain the expansion coefficients that meet the conditions, we need to calculate the right-hand side of the equation. a k Simplifying the partial derivatives, we get:

[0071]

[0072] Substituting the coordinate information of each target feature point into equation (3), the matrix [a0, a1, ..., a1] can be solved. kThe polynomial coefficients are obtained, and the corresponding polynomial curve is then determined to obtain a facial feature contour line that matches the target time. Based on the least squares method, polynomial curve fitting of a preset degree is performed on each target feature point to determine the facial feature contour line that matches the target time, which can further improve the accuracy of the facial feature contour line.

[0073] The aforementioned face image processing method first selects consecutive image frames that are temporally correlated with the target time from multiple face image frames. Based on the time difference between the time matched by each image frame and the target time, it determines the smoothing weight parameter corresponding to the initial feature point set of each image frame. Then, it smooths the initial feature points in each initial feature point set according to the smoothing weight parameter to obtain the target feature points that match the target time. This allows for the determination of the face feature contour line that matches the target time. On the one hand, it can correct the feature points in the target face image frame corresponding to the target time, reducing the influence of interference points on the face feature contour line and improving the accuracy of the face image processing results. On the other hand, it is not limited to the time information carried by each face image frame and can determine the face feature contour line of any target time that is temporally correlated with multiple face image frames, which is beneficial to expanding the application scenarios of the face image processing method.

[0074] As described above, consecutive image frames that are temporally correlated with the target time can include image frames before, after, and corresponding to the target time from multiple face image frames. In one embodiment, such as Figure 3 As shown, step S201 includes:

[0075] Step S301: From multiple face image frames carrying time information, determine the target image frame whose time difference between the time information it carries and the target time is less than the inter-frame time interval.

[0076] The inter-frame time interval refers to the time difference between the temporal information carried by two adjacent image frames. The target image frame is the image frame with the strongest temporal correlation to the target time among multiple face image frames.

[0077] Specifically, if among multiple face image frames, there is an image frame whose time information is the same as the target time, then that image frame is determined as the target image frame; if among multiple face image frames, there is no image frame whose time information is the same as the target time, then the image frame whose time information is less than the target time is determined as the target image frame.

[0078] It is understandable that in multiple face image frames, if no image frame carries the same time information as the target time, there will be two target image frames. In this case, if the time difference between the time information carried by the two target image frames and the target time is different, the target image frame with the smaller time difference can be further identified as the target image frame. For example... Figure 4 As shown, A to I are multiple face image frames, t1 is the first target time, E is the first target image frame that matches t1, t2 is the second target time, and G is the second target image frame that matches t2.

[0079] Step S302: Based on the desired number of image frames and the desired position of the target image frame in the consecutive image frames, consecutive image frames are selected from multiple face image frames.

[0080] The number of image frames refers to the number of image frames contained in a series of image frames. Since changes in facial expressions take time, the shorter the time interval between frames, the smaller the changes in facial expressions within the consecutive image frames. Based on this, the number of image frames can be flexibly set according to the time interval between frames: the shorter the time interval, the more image frames are desired, and vice versa. The desired position of the target image frame within the series of image frames refers to the position of the target image frame within each image frame after sorting the image frames according to the time information they carry. This desired position can be the first frame, the last frame, or any frame between the first and last frames. For example, if the desired number of image frames is odd, the desired position of the target image frame within the series of image frames can be in the middle to ensure a strong temporal correlation between all image frames in the series and the target time.

[0081] Specifically, the terminal can select consecutive image frames from multiple face image frames based on the desired number of image frames and the desired position of the target image frame within consecutive image frames. For example... Figure 4 If the desired number of image frames is 5, and the desired position of the target image frame in the consecutive image frames is the last frame, then the consecutive image frames corresponding to the first target time t1 are A to E, and the consecutive image frames corresponding to the second target time t2 are C to G.

[0082] In the above embodiments, the target image frame is first determined from multiple face image frames. Then, based on the desired number of image frames and the desired position of the target image frame in the consecutive image frames, consecutive image frames are selected from multiple face image frames. This ensures that the consecutive image frames contain the image frames with the strongest temporal correlation with the target time, which is beneficial to improving the scientific nature of the face image processing method.

[0083] In one embodiment, facial feature detection includes lip feature detection. In this embodiment, please refer to... Figure 3Step S203 includes:

[0084] Step S303: Based on the face feature point detection model, obtain the first lip feature point set corresponding to each image frame.

[0085] The facial landmark detection model can be any of the models such as ASM, AAM, DCNN, and TCDCN. The first lip landmark set refers to the set of multiple initial lip landmarks in an image frame obtained based on the facial landmark detection model. This first lip landmark set may include a first upper lip landmark set and a first lower lip landmark set, or a first outer lip landmark set and a first inner lip landmark set.

[0086] Specifically, based on the facial feature point detection model, the terminal can obtain the set of first lip feature points corresponding to each image frame in a series of image frames.

[0087] Step S304: Determine the lip region state of the face image in the image frame based on the first lip feature point set.

[0088] The lip region state can include both open and closed states. Specifically, the terminal can extract key lip feature points from a first set of lip feature points to determine the lip region state, and determine the lip region state of the face image in the corresponding image frame based on these key lip feature points. The location of the key lip feature points is not unique; for example, it can include the highest and lowest points of the center of the outer lip, corresponding points at the outer corners of the outer lip, corresponding points at the left and right inner corners of the inner lip, corresponding points at the center of the inner lip, etc. Correspondingly, the terminal's determination of the lip region state of the face image in the corresponding image frame based on the key lip feature points is also not unique. For example, the terminal can determine the lip region state based on the slope of the line connecting the lowest point of the center of the outer lip and any point at the outer corner of the outer lip; alternatively, it can determine the lip region state based on the area of ​​the region formed by the line connecting the lowest point of the center of the outer lip and the corresponding point at the outer corner of the outer lip.

[0089] In one embodiment, the first lip feature point set includes a first outer lip feature point set and a first inner lip feature point set; in this embodiment, step S304 includes: obtaining the height difference of inner lip feature points at at least one associated position of the upper and lower lips in the first inner lip feature point set; and determining the lip region state of the face image in the image frame based on each height difference.

[0090] As mentioned earlier, the initial feature point set of an image frame includes multiple initial feature points in the face image of that image frame and their corresponding coordinate information. Correspondingly, the first inner lip feature point set of that image frame includes the coordinate information of each inner lip feature point of the upper and lower lips. Figure 5 For example, this coordinate information can specifically be the coordinates of the inner lip feature points in an orthogonal coordinate system XY. Here, the inner lip feature points at the associated positions of the upper and lower lips refer to feature points on the upper and lower lips with the same X-coordinate, for example... Figure 5 Points L and M in the diagram. Furthermore, the height of the inner lip feature point refers to its Y-axis coordinate information.

[0091] Specifically, the terminal obtains the coordinate information of the inner lip feature points at at least one associated position of the upper and lower lips in the first inner lip feature point set, calculates the height difference of the inner lip feature points at each associated position based on the coordinate information, and then determines the lip region state of the face image in the image frame based on each height difference.

[0092] It should be noted that the specific method by which the terminal determines the state of the lip region in the face image in the image frame based on the height difference is not unique.

[0093] In one embodiment, determining the lip region state of a face image in an image frame based on each height difference includes: if the height difference of at least one associated position is greater than a preset height threshold corresponding to the associated position, determining that the lip region state of the face image in the image frame is an open state; if the height difference of each associated position is less than or equal to the preset height threshold corresponding to each associated position, determining that the lip region state of the face image in the image frame is a closed state.

[0094] Specifically, when the shape of the mouth changes, the height difference of the inner lip feature points at related locations of the upper and lower lips changes. In the closed state, this height difference is relatively small, and may even be zero, while in the open state, it is relatively large. Furthermore, even in the open state, the closer the position is to the corner of the mouth, the smaller the height difference of the inner lip feature points of the upper and lower lips.

[0095] Based on this, a preset height threshold can be determined according to the relative distance between the associated positions of the upper and lower lips and the corners of the mouth: the closer the associated position is to the corners of the mouth, the smaller the preset height threshold. Then, based on the relationship between the height difference of the inner lip feature points at the associated position and the preset height threshold, the lip region state of the face image in the image frame is determined: if the height difference of at least one associated position is greater than the preset height threshold corresponding to that position, the lip region state of the face image in the image frame is determined to be open; if the height differences of each associated position are less than or equal to the preset height threshold corresponding to each associated position, the lip region state of the face image in the image frame is determined to be closed. It can be understood that the preset height threshold can be set to zero. Setting corresponding preset height thresholds for different associated positions and determining the lip region state based on the relationship between the height difference of each associated position and the corresponding preset height threshold helps improve the accuracy of the lip region state judgment results.

[0096] In another embodiment, determining the lip region state of a face image in an image frame based on the height difference includes: if the maximum value among the height differences is greater than a set height threshold, determining that the lip region state of the face image in the image frame is an open state; if the maximum value among the height differences is less than or equal to the set height threshold, determining that the lip region state of the face image in the image frame is a closed state.

[0097] Specifically, the terminal acquires the height differences of inner lip feature points at multiple associated locations and compares the maximum value of each height difference with a set height threshold: if the maximum height difference is greater than the set height threshold, the lip region of the face image in the image frame is determined to be in an open state; if the maximum height difference is less than or equal to the set height threshold, the lip region of the face image in the image frame is determined to be in a closed state. Comparing the maximum height difference at each associated location with the set height threshold to determine the lip region state improves work efficiency.

[0098] Step S305: Based on the first set of lip feature points and the state of the lip region, obtain the initial set of lip feature points corresponding to the image frame.

[0099] Specifically, the terminal can obtain the initial lip feature point set corresponding to the image frame based on the lip region state and the first lip feature point set. For example, if the lip region state is closed, the overlapping points of the inner lip in the first lip feature point set are removed to obtain the corresponding initial lip feature point set; if the lip region state is open, the first lip feature point set is determined as the corresponding initial lip feature point set.

[0100] In one embodiment, the first lip feature point set includes a first outer lip feature point set and a first inner lip feature point set. In this embodiment, step S305 includes: if the lip region is in a closed state, determining the initial lip feature point set corresponding to the image frame based on the first lip feature point set; if the lip region is in an open state, obtaining the second inner lip feature point set corresponding to the image frame based on a semantic segmentation algorithm, and combining the first outer lip feature point set and the second inner lip feature point set to determine the initial lip feature point set corresponding to the image frame.

[0101] Semantic segmentation is the process of automatically segmenting object regions from an image and identifying their content. Specifically, the semantic segmentation algorithm can be based on the Swing Transformer segmentation scheme, or it can be based on various semantic segmentation neural network models.

[0102] Specifically, if the lip area is in an open state, such as Figure 5As shown, the inner lip feature points obtained based on the facial feature point detection model may be affected by the outer contour of the teeth, resulting in significant errors. Semantic segmentation algorithms, however, can effectively segment the interior of the lips, obtaining more accurate inner lip feature points. Therefore, the terminal obtains a second set of inner lip feature points corresponding to the image frame based on the semantic segmentation algorithm, and combines this with the first set of outer lip feature points to obtain the initial lip feature point set corresponding to that image frame. Furthermore, if the lip region is in a closed state, the inner lip feature points obtained based on the facial feature point detection model are not affected by the teeth, resulting in smaller errors. In this case, the terminal can determine the initial lip feature point set corresponding to the image frame based on the first lip feature point set. For example, the terminal can directly determine the first lip feature point set as the initial lip feature point set corresponding to the image frame.

[0103] In one embodiment, obtaining a second set of inner lip feature points corresponding to an image frame based on a semantic segmentation algorithm includes: performing semantic segmentation on the image frame based on the semantic segmentation algorithm to extract the inner lip region of the image frame; performing edge detection on the inner lip region to obtain a second set of inner lip feature points corresponding to the image frame.

[0104] Specifically, the terminal uses a semantic segmentation algorithm to perform semantic segmentation on the face image in the image frame, obtaining the semantic segmentation results for each region of the face image in that image frame. This semantic segmentation result includes the semantic labels and region boundaries corresponding to each region. For example... Figure 6 In the image, 1 represents the hair region, 2 the nose region, 3 the mouth region, and so on. After obtaining the semantic segmentation results, the terminal extracts the inner lip region of the corresponding image frame based on the boundary of the mouth region, such as... Figure 6 Region 4. Finally, the terminal performs edge detection on the inner lip region to obtain the corresponding inner lip feature points, such as... Figure 6 The N points in the image frame constitute the second set of inner lip feature points.

[0105] In the above embodiments, based on the state of the lip region in the face image of each image frame, the initial lip feature point set corresponding to each image frame is determined differently, which helps to improve the accuracy of the coordinate information of the initial lip feature points, thereby improving the accuracy of the face image processing results.

[0106] In one embodiment, the smoothing process includes Gaussian smoothing iteration. In this embodiment, please refer to... Figure 3 Step S207 includes:

[0107] Step S306: Extract the initial feature points corresponding to the positions in each initial feature point set to form an associated feature point set.

[0108] The Gaussian smoothing iterative process involves weighted averaging of multiple initial feature points, with the weights of each initial feature point conforming to a Gaussian distribution. The initial feature points corresponding to positions within each set of initial feature points refer to the initial feature points whose relative positions on the face image correspond to those positions within each set. For example, the left outer corner of the mouth, the left inner corner of the mouth, the tip of the nose, the leftmost point of the nostril, and the rightmost point of the nostril in each set of initial feature points. It can be understood that the relative positions of the initial feature points within the same associated feature point set are consistent on the face image, and the number of initial feature points in the associated feature point set is consistent with the number of initial feature points in the initial feature point set; the number of initial feature points in any associated feature point set is consistent with the number of initial feature points in the initial feature point set.

[0109] Specifically, the terminal extracts the initial feature points corresponding to the positions in each initial feature point set to form a set of associated feature points. For example, the feature point set of the left outer lip corner, the feature point set of the left inner lip corner, the feature point set of the nose tip, the feature point set of the leftmost point of the nostril, and the feature point set of the rightmost point of the nostril, etc.

[0110] Step S307: According to the smoothing weight parameters corresponding to each initial feature point in the associated feature point set, Gaussian smoothing iteration is performed on the initial feature values ​​of each initial feature point in sequence.

[0111] In the Gaussian smoothing iterative process, the weights of each initial feature value are inversely correlated with the time difference corresponding to that initial feature value, and also inversely correlated with the difference between that initial feature value and the reference feature value. The reference feature value can refer to the initial feature value of the corresponding initial feature point in the target image frame, for example... Figure 4 In this context, the initial feature value refers to the initial feature value of a corresponding point in image frame E or image frame G; alternatively, it can refer to the feature value of the corresponding feature point obtained after the previous Gaussian smoothing iteration. It should be noted that when there are two target image frames whose time difference with the target time is less than the inter-frame time interval, the initial feature value of the target image frame can be obtained by weighted summing of the initial feature values ​​of the corresponding initial feature points in these two target image frames. The weight of each target image frame is inversely correlated with the time difference between that image frame and the target time. For example... Figure 4 In the process, the reference feature value for the second target time t2 can be obtained by weighted summing of the initial feature values ​​of the initial feature points at corresponding positions in image frames F and G. This initial feature value is the reference feature value that matches the target time.

[0112] Specifically, the terminal performs Gaussian smoothing iteration on the initial feature values ​​of each initial feature point according to the smoothing weight parameters corresponding to each initial feature point in the associated feature point set. This results in the weight being greater the closer the time matched by the initial feature point is to the target time, and the weight being greater the closer the initial feature value corresponding to the initial feature point is to the reference feature value.

[0113] Step S308: Combine the Gaussian smoothing iteration results of each associated feature point set to obtain the target feature points that match the target time.

[0114] Specifically, each set of associated feature points can obtain a corresponding Gaussian smoothing iterative processing result, yielding a corresponding feature value. This feature value is the coordinate information of the target feature point corresponding to that set of associated feature points. Based on this, the terminal integrates the Gaussian smoothing iterative processing results of each set of associated feature points to obtain target feature points at different locations on the face image that match the target time.

[0115] In the above embodiments, during the process of smoothing the initial feature points in the initial feature point set to obtain target feature points that match the target time, the temporal correlation and feature value correlation are comprehensively considered. This can avoid the influence of abnormal initial feature points in a certain image frame on the target feature points and further improve the accuracy.

[0116] To facilitate understanding, the following will be combined with Figures 4 to 8 This section provides a detailed explanation of the process for generating the lip contour line. In one embodiment, such as... Figure 7 As shown, the face image processing method includes the following steps:

[0117] Step S701: From multiple face image frames carrying time information, determine the target image frame with the smallest time difference between the time information it carries and the target time.

[0118] Specifically, if among multiple face image frames, there is an image frame whose time information is the same as the target time, then that image frame is determined as the target image frame; if among multiple face image frames, there is no image frame whose time information is the same as the target time, then the image frame whose time information has the smallest time difference with the target time is determined as the target image frame.

[0119] Step S702: Based on the desired number of image frames and the desired position of the target image frame in the consecutive image frames, consecutive image frames are selected from multiple face image frames.

[0120] The number of image frames refers to the number of image frames contained in a series of image frames, which can be flexibly set according to the inter-frame time interval: the shorter the inter-frame time interval, the more image frames are expected, and vice versa. The expected position of the target image frame within the series of image frames refers to the position of the target image frame within each image frame after sorting the image frames according to the time information they carry. This expected position can be the first frame, the last frame, or any frame between the first and last frames. For example, if the expected number of image frames is odd, the expected position of the target image frame within the series of image frames can be in the middle to ensure a strong temporal correlation between all image frames in the series and the target time. Alternatively, the expected number of image frames can be four, with the series of image frames including the target image frame and the three frames preceding it.

[0121] Specifically, the terminal can select consecutive image frames from multiple face image frames based on the desired number of image frames and the desired position of the target image frame within consecutive image frames. For example... Figure 4 If the desired number of image frames is 4, and the desired position of the target image frame in the consecutive image frames is the last frame, then the consecutive image frames corresponding to the first target time t1 are B to E, and the consecutive image frames corresponding to the second target time t2 are D to G.

[0122] Step S703: Based on the face feature point detection model, obtain the first lip feature point set corresponding to each image frame.

[0123] The facial landmark detection model can be any of the models such as ASM, AAM, DCNN, and TCDCN. The first lip landmark set refers to the set of multiple initial lip landmarks in an image frame obtained based on the facial landmark detection model. This first lip landmark set may include a first upper lip landmark set and a first lower lip landmark set, or a first outer lip landmark set and a first inner lip landmark set.

[0124] Specifically, based on the facial feature point detection model, the terminal can obtain the set of first lip feature points corresponding to each image frame in a series of image frames.

[0125] Step S704: Obtain the height difference between two inner lip feature points located at the middle position of the upper and lower lips in the first inner lip feature point set.

[0126] The first set of inner lip feature points includes the coordinate information of each inner lip feature point of both the upper and lower lips. Figure 5For example, this coordinate information can specifically be the coordinates of the inner lip feature points in an orthogonal coordinate system XY. The two inner lip feature points located in the middle of the upper and lower lips refer to the feature points whose X-coordinate is the median among all the inner lip feature points on the upper and lower lips, for example... Figure 5 Points L and M in the diagram. Furthermore, the height of the inner lip feature point refers to its Y-axis coordinate information.

[0127] Specifically, the terminal obtains the coordinate information of two inner lip feature points in the middle position of the upper and lower lips in the first inner lip feature point set, calculates the height difference between the upper and lower lips based on the coordinate information, and then determines the lip region state of the face image in the image frame based on the height difference.

[0128] Step S705: If the height difference is greater than the preset height threshold, determine that the lip region of the face image in the image frame is in an open state.

[0129] Specifically, when the shape of the mouth changes, the height difference of the inner lip feature point in the middle of the upper and lower lips changes. In the closed state, this height difference is relatively small, and may even be zero, while in the open state, it is relatively large. Based on this, the terminal can compare this height difference with a preset height threshold. If the height difference is greater than the preset height threshold, the lip region of the face image in the image frame is determined to be in an open state. Figure 5 For example, the preset height threshold could be 5. Since the Y-coordinate of point L is 165 and the Y-coordinate of point M is 147, the difference between them is 18, which is greater than the preset height threshold. Figure 5 The lip region of the face image is in an open state.

[0130] Step S706: Perform semantic segmentation on the image frame based on the semantic segmentation algorithm, extract the inner lip region of the image frame, and perform edge detection on the inner lip region to obtain the second inner lip feature point set corresponding to the image frame.

[0131] Specifically, the terminal uses a semantic segmentation algorithm to perform semantic segmentation on the face image in the image frame, obtaining the semantic segmentation results for each region of the face image in that image frame. This semantic segmentation result includes the semantic labels and region boundaries corresponding to each region. For example... Figure 6 In the image, 1 represents the hair region, 2 the nose region, 3 the mouth region, and so on. After obtaining the semantic segmentation results, the terminal extracts the inner lip region of the corresponding image frame based on the boundary of the mouth region, such as... Figure 6 Region 4. Finally, the terminal performs edge detection on the inner lip region to obtain the corresponding inner lip feature points, such as... Figure 6 The N points in the image frame constitute the second set of inner lip feature points.

[0132] Step S707: Combine the first set of external lip feature points and the second set of internal lip feature points to determine the initial set of lip feature points corresponding to the image frame.

[0133] Specifically, if the lip area is in an open state, such as Figure 5 As shown, the inner lip feature points obtained based on the facial feature point detection model may be affected by teeth, resulting in significant errors. Semantic segmentation algorithms, however, can effectively segment the interior of the lips to obtain more accurate inner lip feature points. Therefore, the terminal obtains a second set of inner lip feature points corresponding to the image frame based on the semantic segmentation algorithm. This set is then combined with the first set of outer lip feature points to obtain the initial set of lip feature points corresponding to that image frame.

[0134] Step S708: If the height difference is less than or equal to a preset height threshold, determine that the lip region of the face image in the image frame is in a closed state.

[0135] Specifically, the terminal can compare the height difference with a preset height threshold. If the height difference is less than or equal to the preset height threshold, the terminal determines that the lip region of the face image in the image frame is in a closed state.

[0136] Step S709: Determine the first set of lip feature points as the initial set of lip feature points corresponding to the image frame.

[0137] Specifically, if the lip region is in a closed state, the inner lip feature points obtained based on the facial feature point detection model will not be affected by the teeth and the error will be small. Therefore, the terminal can directly determine the first lip feature point set as the initial lip feature point set corresponding to the image frame.

[0138] Step S710: Based on the time difference between the time matched by the image frame and the target time, determine the smoothing weight parameters corresponding to each initial feature point set of the lips.

[0139] The smoothing weight parameter determines the contribution of each initial lip feature point in the set of initial lip feature points to the target lip feature point obtained after smoothing. Specifically, based on the time difference between the matched time of the image frame and the target time, the terminal can determine the smoothing weight parameter corresponding to each set of initial lip feature points, such that the contribution of the initial lip feature points of each image frame to the target lip feature point is inversely correlated with the time difference associated with that image frame. That is, the closer the matched time of the image frame is to the target time, the greater the contribution of the initial lip feature points of that image frame to the target feature point.

[0140] Step S711: Extract the initial lip feature points corresponding to the positions in each initial lip feature point set to form a set of lip-related feature points.

[0141] The initial lip feature points corresponding to the positions in each set of initial lip feature points refer to the initial lip feature points in each set that have a relative position on the face image. For example, the left outer corner of the mouth and the left inner corner of the mouth in each set of initial lip feature points. It can be understood that the relative positions of the initial lip feature points contained in the same set of associated lip feature points are consistent on the face image, and the number of such sets is consistent with the number of initial lip feature points contained in the initial lip feature point sets; the number of initial lip feature points contained in any set of associated lip feature points is consistent with the number of initial lip feature points in the set of initial lip feature points.

[0142] Specifically, the terminal extracts the initial lip feature points corresponding to the positions in each initial lip feature point set, forming a set of associated lip feature points. For example, the set of feature points at the corner of the left outer lip and the set of feature points at the corner of the left inner lip, etc.

[0143] Step S712: According to the smoothing weight parameters corresponding to each initial feature point of the lips in the set of lip-related feature points, the initial feature values ​​of each initial feature point of the lips are sequentially subjected to Gaussian smoothing iteration.

[0144] The Gaussian smoothing iterative processing involves a weighted average of multiple initial lip feature points, with each initial lip feature point's weight conforming to a Gaussian distribution. Specifically, the terminal performs Gaussian smoothing iterative processing on the initial lip feature values ​​of each initial lip feature point according to the smoothing weight parameters corresponding to each initial lip feature point in the set of lip-related feature points. This ensures that the closer the time matched by the initial lip feature point is to the target time, the greater its weight; and the closer the initial lip feature value corresponding to the initial lip feature point is to the lip reference feature value, the greater its weight.

[0145] Taking the case where the desired number of image frames is 7, and the desired position of the target image frame in a series of image frames is the middle frame as an example. Assuming the target image frame is i, the lip reference feature value is the coordinate value corresponding to the k-th feature point in the target image frame, that is, the lip reference feature value x_temp = lips_lmk[i][k][x]; y_temp lips_lmk[i][k][y]. Set the smoothing weight parameters weights = [0.1, 0.1, 0.2, 0.4]; weights2 = [0.1, 0.1, 0.5, 0.5], where both weights and weights2 are related to the time difference.

[0146] The Gaussian smoothing iterative process is actually an iterative process where frame j starts from frame i-3 and ends at frame i+3. For any frame j, the corresponding smoothing weight parameters are:

[0147] weight=weights(abs(ji)+1) (4)

[0148] weight2=weights2(abs(ji)+1) (5)

[0149] For example, when j = i - 3, the corresponding smoothing weight parameter weight is 0.4 and weight2 is 0.5; when j = i, the corresponding smoothing weight parameter weight is 0.1 and weight2 is 0.1.

[0150] Taking the iterative process of the x-coordinate of the lip target feature point as an example, the terminal can obtain the weight w corresponding to the x-coordinate of the j-th frame based on the smoothing weight parameter:

[0151] w=exp(-weight*(x-x_temp)^2-weight2*(abs(ji)+1)^2) (6)

[0152] As can be seen from equation (6), the closer the j-th frame is to the target image frame in the time domain, the smaller the values ​​of weight and weight2, and the larger the weight w; the closer the coordinate x of the j-th frame is to the lip reference feature value x_temp, the smaller the value of (x-x_temp), and the larger the weight w.

[0153] After obtaining the weight w corresponding to the j-th frame, the x-coordinate of that frame is then weighted and superimposed with the smoothed x-coordinate result obtained from the previous Gaussian smoothing iteration to obtain the current smoothed x-coordinate result smooth_x:

[0154] smooth_x=(smooth_x+w*x_temp) / (sum_weight+w) (7)

[0155] Here, sum_weight is the total weight obtained after the previous Gaussian smoothing iteration. The initial values ​​of smooth_x and sum_weight are both set to 0.

[0156] Following the iterative steps described above, j loops from frame i-3 to frame i+3, thus obtaining the x-coordinate value smooth_x of the lip target feature point. It can be understood that the y-coordinate value of the lip target feature point can also be obtained using the same method, thereby determining the corresponding lip target feature point location.

[0157] Step S713: Combine the Gaussian smoothing iterative processing results of each set of lip-related feature points to obtain lip target feature points that match the target time.

[0158] Specifically, each set of lip-related feature points can obtain a corresponding Gaussian smoothing iterative processing result, yielding a corresponding feature value. This feature value represents the coordinate information of the lip target feature point corresponding to that set of lip-related feature points. Based on this, the terminal integrates the Gaussian smoothing iterative processing results of each set of lip-related feature points to obtain lip target feature points at different locations on the face image that match the target time.

[0159] Step S714: Based on the least squares method, perform fourth-order polynomial curve fitting on each lip target feature point to determine the lip contour line that matches the target time.

[0160] Curve fitting, in particular, is a data processing method that uses a continuous curve to approximate or analogize the functional relationship between coordinates represented by discrete points on a plane. Curve fitting can be achieved using the least squares method. Polynomial curve fitting means that the functional relationship can be expressed using a polynomial. Specifically, the terminal's least squares method performs fourth-order polynomial curve fitting on each lip target feature point to determine the lip contour line matching the target time. For example, the terminal can first divide the lip target feature points into sets of outer lip lower contour points, outer lip upper contour points, inner lip lower contour points, and inner lip upper contour points according to coordinate information. Then, it performs fourth-order polynomial curve fitting on the lip target feature points in each set to obtain the outer lip lower contour line, outer lip upper contour line, inner lip lower contour line, and inner lip upper contour line. These lip contour lines are then superimposed on a single image for output and display. Using fourth-order polynomial curve fitting to obtain the lip contour line avoids overfitting and further improves the accuracy of the lip contour line. Figure 8 As shown, four sets of lip contour lines were obtained based on the method of this embodiment, all of which can accurately represent the changes in the lip area.

[0161] In the above embodiments, based on consecutive image frames that are temporally correlated with the target time, facial feature detection is performed to obtain initial lip feature points, and each initial lip feature point is smoothed to obtain target lip feature points, thereby determining the lip contour line matching the target time. On the one hand, this can reduce the influence of interference points in the target face image frame corresponding to the target time on the lip contour line, which is beneficial to improving the accuracy of the face image processing results. On the other hand, it is not limited to the time information carried by each face image frame, and the lip contour line of any target time that is temporally correlated with multiple face image frames can be determined, which is beneficial to expanding the application scenarios of the face image processing method. In addition, based on the lip region state of the face image in each image frame, the set of initial lip feature points corresponding to each image frame is determined differently, which is beneficial to improving the accuracy of the coordinate information of the initial lip feature points, thereby improving the accuracy of the face image processing results. In the process of smoothing the initial feature points in the initial lip feature point set to obtain the target feature points matching the target time, the temporal correlation and feature value correlation are comprehensively considered, which can avoid the influence of abnormal initial lip feature points in a certain image frame on the target lip feature points, further improving the accuracy.

[0162] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0163] Based on the same inventive concept, this application also provides a face image processing apparatus for implementing the face image processing method described above. The solution provided by this apparatus is similar to the implementation scheme described in the above method; therefore, the specific limitations in one or more face image processing apparatus embodiments provided below can be found in the limitations of the face image processing method described above, and will not be repeated here.

[0164] In one embodiment, such as Figure 9 As shown, a face image processing device 900 is provided, including: a continuous image frame filtering module 901, an initial feature point set determination module 903, a smoothing weight parameter determination module 905, a target feature point determination module 907, and a face feature contour line determination module 909, wherein:

[0165] The continuous image frame filtering module 901 is used to filter out continuous image frames that are temporally related to the target time from multiple face image frames carrying time information.

[0166] The initial feature point set determination module 903 is used to perform face feature detection on each image frame in a series of image frames to obtain the initial feature point set corresponding to each image frame.

[0167] The smoothing weight parameter determination module 905 is used to determine the smoothing weight parameters corresponding to each initial feature point set based on the time difference between the time matched by the image frame and the target time.

[0168] The target feature point determination module 907 is used to smooth the initial feature points in each initial feature point set according to the smoothing weight parameter, so as to obtain the target feature points that match the target time.

[0169] The facial feature contour determination module 909 is used to determine the facial feature contour that matches the target time based on each target feature point.

[0170] In one embodiment, the continuous image frame filtering module 901 includes: a target image frame determination unit, configured to determine, from multiple face image frames carrying time information, a target image frame whose time difference with the target time is less than the inter-frame time interval; and a continuous image frame filtering unit, configured to filter continuous image frames from multiple face image frames based on the desired number of image frames and the desired position of the target image frame in the continuous image frames.

[0171] In one embodiment, facial feature detection includes lip feature detection. In this embodiment, the initial feature point set determination module 903 includes: a first lip feature point set acquisition unit, configured to obtain a first lip feature point set corresponding to each image frame based on a facial feature point detection model; a lip region state determination unit, configured to determine the lip region state of the facial image in the image frame based on the first lip feature point set; and an initial lip feature point set determination unit, configured to obtain an initial lip feature point set corresponding to the image frame based on the first lip feature point set and the lip region state.

[0172] In one embodiment, the first lip feature point set includes a first outer lip feature point set and a first inner lip feature point set. In this embodiment, the lip region state determination unit includes: a height difference determination component, used to obtain the height difference between inner lip feature points at at least one associated position of the upper and lower lips in the first inner lip feature point set; and a lip region state determination component, used to determine the lip region state of the face image in the image frame based on each height difference.

[0173] In one embodiment, the lip region state determination component is specifically used to: determine that the lip region state of the face image in the image frame is an open state if the height difference of at least one associated position is greater than the preset height threshold corresponding to the associated position; and determine that the lip region state of the face image in the image frame is a closed state if the height difference of each associated position is less than or equal to the preset height threshold corresponding to each associated position.

[0174] In another embodiment, the lip region state determination component is specifically used to: determine that the lip region state of the face image in the image frame is open if the maximum value of each height difference is greater than a set height threshold; and determine that the lip region state of the face image in the image frame is closed if the maximum value of each height difference is less than or equal to the set height threshold.

[0175] In one embodiment, the first lip feature point set includes a first outer lip feature point set and a first inner lip feature point set. In this embodiment, the lip initial feature point set determination unit is specifically used to: if the lip region is in a closed state, determine the lip initial feature point set corresponding to the image frame based on the first lip feature point set; if the lip region is in an open state, obtain the second inner lip feature point set corresponding to the image frame based on a semantic segmentation algorithm, and combine the first outer lip feature point set and the second inner lip feature point set to determine the lip initial feature point set corresponding to the image frame.

[0176] In one embodiment, the initial lip feature point set determination unit is specifically used to: perform semantic segmentation on the image frame based on a semantic segmentation algorithm to extract the inner lip region of the image frame; perform edge detection on the inner lip region to obtain the second inner lip feature point set corresponding to the image frame.

[0177] In one embodiment, the smoothing process includes Gaussian smoothing iteration. In this embodiment, the target feature point determination module 907 is specifically used to: extract the initial feature points corresponding to the positions in each initial feature point set to form an associated feature point set; perform Gaussian smoothing iteration on the initial feature values ​​of each initial feature point according to the smoothing weight parameters corresponding to each initial feature point in the associated feature point set; the weight of each initial feature value during the Gaussian smoothing iteration process is inversely correlated with the time difference corresponding to the initial feature value, and inversely correlated with the difference between the initial feature value and the reference feature value; and synthesize the Gaussian smoothing iteration results of each associated feature point set to obtain the target feature point that matches the target time.

[0178] In one embodiment, the facial feature contour determination module 909 is specifically used to: perform polynomial curve fitting on each target feature point for a preset number of times based on the least squares method to determine the facial feature contour that matches the target time.

[0179] Each module in the aforementioned face image processing device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the operations corresponding to each module.

[0180] In one embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 10 As shown, the computer device includes a processor, memory, input / output interface, communication interface, display unit, and input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interface. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage medium. The input / output interface is used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, NFC (Near Field Communication), or other technologies. When executed by the processor, the computer program implements a facial image processing method. The display unit of the computer device is used to form a visually visible image. It can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of the computer device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads set on the casing of the computer device, or external keyboards, touchpads, or mice, etc.

[0181] Those skilled in the art will understand that Figure 10 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0182] In one embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.

[0183] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps in the above method embodiments.

[0184] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0185] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0186] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0187] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A face image processing method, characterized in that, The method includes: From multiple face image frames carrying time information, select consecutive image frames that have a temporal correlation with the target time. Face feature detection is performed on each image frame in the continuous image frames to obtain the initial feature point set corresponding to each image frame; Based on the time difference between the time matched by the image frame and the target time, a smoothing weight parameter is determined for each of the initial feature point sets; the smoothing weight parameter is used to characterize the contribution of the initial feature points in the initial feature point set; the contribution is inversely correlated with the time difference; Extract the initial feature points corresponding to the positions in each of the initial feature point sets to form an associated feature point set; According to the smoothing weight parameters corresponding to each of the initial feature points in the set of associated feature points, the initial feature values ​​of each initial feature point are sequentially subjected to Gaussian smoothing iteration processing; the weight of each initial feature value during the Gaussian smoothing iteration processing is inversely correlated with the time difference corresponding to the initial feature value, and inversely correlated with the difference between the initial feature value and the reference feature value; By combining the Gaussian smoothing iterative processing results of each set of associated feature points, target feature points that match the target time are obtained. Based on each of the target feature points, a facial feature contour line matching the target time is determined.

2. The method according to claim 1, characterized in that, The step of filtering out consecutive image frames that have a temporal correlation with a target time from multiple face image frames carrying time information includes: From multiple face image frames carrying time information, determine the target image frame whose time difference between the time information it carries and the target time is less than the inter-frame time interval; Based on the desired number of image frames and the desired position of the target image frame in the consecutive image frames, consecutive image frames are selected from the multiple face image frames.

3. The method according to claim 1, characterized in that, The facial feature detection includes lip feature detection; the step of performing facial feature detection on each image frame in the continuous image frames to obtain an initial feature point set corresponding to each image frame includes: Based on the facial feature point detection model, the first lip feature point set corresponding to each of the image frames is obtained; Based on the first set of lip feature points, determine the lip region state of the face image in the image frame; Based on the first set of lip feature points and the state of the lip region, the initial set of lip feature points corresponding to the image frame is obtained.

4. The method according to claim 3, characterized in that, The first set of lip feature points includes a first set of outer lip feature points and a first set of inner lip feature points; determining the lip region state of the face image in the image frame based on the first set of lip feature points includes: Obtain the height difference of inner lip feature points at at least one associated position of the upper and lower lips in the first inner lip feature point set; Based on the height differences, the state of the lip region in the face image of the image frame is determined.

5. The method according to claim 4, characterized in that, Determining the lip region state of the face image in the image frame based on the height differences includes: If the height difference of at least one of the associated positions is greater than the preset height threshold corresponding to the associated position, the lip region of the face image in the image frame is determined to be in an open state. If the height difference between each of the associated positions is less than or equal to the preset height threshold corresponding to each of the associated positions, the lip region of the face image in the image frame is determined to be in a closed state.

6. The method according to claim 4, characterized in that, Determining the lip region state of the face image in the image frame based on the height difference includes: If the maximum value among the height differences is greater than a set height threshold, the lip region of the face image in the image frame is determined to be in an open state. If the maximum value among the height differences is less than or equal to the set height threshold, the lip region of the face image in the image frame is determined to be in a closed state.

7. The method according to claim 3, characterized in that, The first set of lip feature points includes a first set of outer lip feature points and a first set of inner lip feature points; The step of obtaining the initial feature point set corresponding to the image frame based on the first lip feature point set and the lip region state includes: If the lip region is in a closed state, the initial lip feature point set corresponding to the image frame is determined based on the first lip feature point set. If the lip region is in an open state, the second inner lip feature point set corresponding to the image frame is obtained based on the semantic segmentation algorithm, and the first outer lip feature point set and the second inner lip feature point set are combined to determine the initial lip feature point set corresponding to the image frame.

8. The method according to claim 7, characterized in that, The method for obtaining the second inner lip feature point set corresponding to the image frame based on the semantic segmentation algorithm includes: The image frame is semantically segmented based on a semantic segmentation algorithm to extract the inner lip region of the image frame; Edge detection is performed on the inner lip region to obtain the second set of inner lip feature points corresponding to the image frame.

9. The method according to any one of claims 1 to 8, characterized in that, The step of determining the facial feature contour line matching the target time based on each of the target feature points includes: Based on the least squares method, a polynomial curve fitting of a preset number is performed on each of the target feature points to determine the facial feature contour line that matches the target time.

10. A face image processing device, characterized in that, The device includes: The continuous image frame filtering module is used to filter out continuous image frames that are temporally correlated with the target time from multiple face image frames carrying time information. The initial feature point set determination module is used to perform face feature detection on each image frame in the continuous image frames to obtain the initial feature point set corresponding to each image frame. The smoothing weight parameter determination module is used to determine the smoothing weight parameter corresponding to each of the initial feature point sets based on the time difference between the time matched by the image frame and the target time; the smoothing weight parameter is used to characterize the contribution of the initial feature points in the initial feature point set; the contribution is inversely correlated with the time difference; The target feature point determination module is used to extract the initial feature points corresponding to the positions in each of the initial feature point sets, forming an associated feature point set; according to the smoothing weight parameters corresponding to each of the initial feature points in the associated feature point set, the initial feature values ​​of each of the initial feature points are sequentially subjected to Gaussian smoothing iteration processing; the Gaussian smoothing iteration processing results of each of the associated feature point sets are combined to obtain the target feature point that matches the target time; the weight of each of the initial feature values ​​in the Gaussian smoothing iteration processing is inversely correlated with the time difference corresponding to the initial feature value, and inversely correlated with the difference between the initial feature value and the reference feature value; The facial feature contour determination module is used to determine a facial feature contour that matches the target time based on each of the target feature points.

11. The apparatus according to claim 10, characterized in that, The continuous image frame filtering module includes: The target image frame determination unit is used to determine, from multiple face image frames carrying time information, a target image frame whose time difference between the time information carried and the target time is less than the inter-frame time interval. A continuous image frame filtering unit is used to filter out continuous image frames from the plurality of face image frames based on the desired number of image frames and the desired position of the target image frame in the continuous image frames.

12. The apparatus according to claim 10, characterized in that, The facial feature detection includes lip feature detection; the initial feature point set determination module includes: The first lip feature point set acquisition unit is used to obtain the first lip feature point set corresponding to each image frame based on the face feature point detection model. The lip region state determination unit is used to determine the lip region state of the face image in the image frame based on the first set of lip feature points. The initial lip feature point set determination unit is used to obtain the initial lip feature point set corresponding to the image frame based on the first lip feature point set and the lip region state.

13. The apparatus according to claim 12, characterized in that, The first set of lip feature points includes a first set of outer lip feature points and a first set of inner lip feature points; The lip region state determination unit includes: A height difference determination component is used to obtain the height difference between inner lip feature points at at least one associated position of the upper and lower lips in the first inner lip feature point set. A lip region state determination component is used to determine the lip region state of a face image in the image frame based on the height differences.

14. The apparatus according to claim 13, characterized in that, The lip region state determination component is specifically used for: If the height difference of at least one of the associated positions is greater than the preset height threshold corresponding to the associated position, the lip region of the face image in the image frame is determined to be in an open state. If the height difference between each of the associated positions is less than or equal to the preset height threshold corresponding to each of the associated positions, the lip region of the face image in the image frame is determined to be in a closed state.

15. The apparatus according to claim 13, characterized in that, The lip region state determination component is specifically used for: If the maximum value among the height differences is greater than a set height threshold, the lip region of the face image in the image frame is determined to be in an open state. If the maximum value among the height differences is less than or equal to the set height threshold, the lip region of the face image in the image frame is determined to be in a closed state.

16. The apparatus according to claim 12, characterized in that, The first lip feature point set includes a first outer lip feature point set and a first inner lip feature point set; the initial lip feature point set determination unit is specifically used for: If the lip region is in a closed state, the initial lip feature point set corresponding to the image frame is determined based on the first lip feature point set. If the lip region is in an open state, the second inner lip feature point set corresponding to the image frame is obtained based on the semantic segmentation algorithm, and the first outer lip feature point set and the second inner lip feature point set are combined to determine the initial lip feature point set corresponding to the image frame.

17. The apparatus according to claim 16, characterized in that, The initial feature point set determination unit for lips is specifically used for: The image frame is semantically segmented based on a semantic segmentation algorithm to extract the inner lip region of the image frame; Edge detection is performed on the inner lip region to obtain the second set of inner lip feature points corresponding to the image frame.

18. The apparatus according to any one of claims 10 to 17, characterized in that, The facial feature contour determination module is specifically used for: Based on the least squares method, a polynomial curve fitting of a preset number is performed on each of the target feature points to determine the facial feature contour line that matches the target time.

19. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 9.

20. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 9.

21. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Control method, control device and system for carrying out tracking shooting for object

    CN105578034A

  • Image recognition method and apparatus, electronic device, and computer-readable storage medium

    CN109409235A