Face key point filtering method, device, electronic device and storage medium
Through the facial key point filtering method, the face key points are adjusted and filtered, and the problems of face positioning jitter and lag in the existing technology are solved, achieving a more stable face positioning and reducing misalignment.
Patent Information
- Application Number
- CN202111535811.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-15
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2041-12-15
AI Technical Summary
The existing key point detection model based on convolutional neural network has jitter problems when positioning key points of the face, resulting in lag and misalignment in face positioning.
The face key point filtering method is adopted. By obtaining the first face key point of the current frame face image, adjusting the spatial signal standard deviation of the historical face key point of the preset key point filtering model, the adjusted key point filtering model is constructed, and all frame face images in the filter window are filtered to output a stable second face key point.
It effectively overcomes the lag and jitter problems in face positioning, improves the stability and follow-up of key points of the face, and achieves more stable face positioning and reduces misalignment.
Smart Images

Figure CN114187639B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology, and in particular to a method, device, electronic equipment and storage medium for filtering key points of a human face. Background Art
[0002] In scenarios where AR technology is used to achieve a variety of special effects rendering (such as face slimming, big eyes, etc.), this is achieved by positioning the key points of the face. That is, the detected face image is used as the input of the key point detection model to obtain the position of the key points of the face.
[0003] The existing key point detection model is built based on convolutional neural networks. Due to factors such as sample size and diversity limitations, the prediction results of 106 or 240 facial key points obtained by the key point detection model based on convolutional neural networks are jittery when locating the positions of facial key points. Therefore, in scenarios where AR features are implemented based on faces, lag and jitter are prone to occur when locating faces based on these facial key points, which in turn leads to the problem of face misalignment. Summary of the invention
[0004] In view of this, the embodiments of the present invention provide a facial key point filtering method, device, electronic device and storage medium to solve the problem of face misalignment caused by lag and jitter when performing face positioning based on facial key points.
[0005] To achieve the above objectives, the embodiments of the present invention provide the following technical solutions:
[0006] A first aspect of an embodiment of the present invention provides a method for filtering key points of a face, the method comprising:
[0007] In the filtering window, a current frame face image is obtained, and the current frame face image is predicted to obtain a first face key point;
[0008] Adjusting the spatial signal standard deviation of the historical facial key points of a preset key point filtering model according to the first facial key points to obtain an adjusted key point filtering model, wherein the preset key point filtering model is constructed based on the principle of a bilateral filter;
[0009] Acquire key points obtained by predicting the current frame face image from all the frame face images within the filtering window;
[0010] All the key points are input into the adjusted key point filtering model for key point filtering, and the second face key points after filtering of the current frame face image are output, and the second face key points are used to locate the face position in the current scene.
[0011] Optionally, the process of constructing the preset key point filtering model includes:
[0012] Determine the length of the filter window, obtain the historical frame face image within the filter window, predict the historical frame face image, and obtain the historical face key points;
[0013] Obtaining a time domain signal standard deviation of the historical facial key points in the time domain, and using the time domain signal standard deviation as a time domain distance weight of a bilateral filter;
[0014] Obtaining a spatial signal standard deviation of the historical facial key points in space, and using the spatial signal standard deviation as a spatial distance weight of the bilateral filter;
[0015] A bilateral filter is constructed using the time domain distance weight and the space distance weight, and the bilateral filter is used as a key point filtering model.
[0016] Optionally, adjusting the spatial signal standard deviation of the historical facial key points in the preset key point filter model according to the first facial key points to obtain the adjusted key point filter model includes:
[0017] Acquire the current face contour area represented by the first face key points;
[0018] Calculating the square root of the ratio of the current face contour area to the reference face contour area to obtain a face change ratio;
[0019] Multiplying the face change ratio by the spatial signal standard deviation of the historical face key points in the preset key point filtering model to obtain a new spatial signal standard deviation;
[0020] The preset key point filter model is reconstructed using the new spatial signal standard deviation to obtain an adjusted key point filter model.
[0021] Optionally, adjusting the spatial signal standard deviation of the historical face key points of the preset key point filter model according to the first face key point to obtain the adjusted key point filter model includes:
[0022] Calculate the difference between the first facial key point and the facial key point at the previous moment as the displacement speed;
[0023] If the displacement speed is greater than the set value, reducing the spatial signal standard deviation of the historical face key points of the key point filtering model, and reconstructing the key point filtering model using the reduced spatial signal standard deviation to obtain an adjusted key point filtering model;
[0024] If the displacement speed is less than the set value, the spatial signal standard deviation of the historical face key points of the key point filtering model is increased, and the key point filtering model is reconstructed using the increased spatial signal standard deviation to obtain an adjusted key point filtering model.
[0025] A second aspect of an embodiment of the present invention provides a facial key point filtering device, the device comprising:
[0026] A first acquisition unit is used to acquire a current frame face image within a filtering window, predict the current frame face image, and obtain a first face key point;
[0027] an adjusting unit, configured to adjust a spatial signal standard deviation of a historical face key point of a preset key point filtering model according to the first face key point to obtain an adjusted key point filtering model, wherein the preset key point filtering model is constructed based on the principle of a bilateral filter;
[0028] A second acquisition unit is used to acquire key points obtained by predicting the current frame face image from all the frame face images within the filtering window;
[0029] The processing unit is used to input all the key points into the adjusted key point filtering model for key point filtering, output the second face key point information after filtering the current frame face image, and use the second face key point to locate the face position in the current scene.
[0030] Optionally, the device further comprises: a pre-built module;
[0031] The pre-constructed module is used to determine the length of the filtering window, obtain the historical frame face image within the filtering window, predict the historical frame face image, and obtain the historical face key points; obtain the time domain signal standard deviation of the historical face key points in the time domain, and use the time domain signal standard deviation as the time domain distance weight of the bilateral filter; obtain the spatial signal standard deviation of the historical face key points in space, and use the spatial signal standard deviation as the spatial distance weight of the bilateral filter; use the time domain distance weight and the spatial distance weight to construct a bilateral filter, and use the bilateral filter as a key point filtering model.
[0032] Optionally, the adjustment unit includes:
[0033] A first acquisition module, used to acquire the current face contour area represented by the first face key points;
[0034] A first calculation module is used to calculate the square root of the ratio of the current face contour area to the reference face contour area to obtain a face change ratio; and multiply the face change ratio by the spatial signal standard deviation of the historical face key points in the preset key point filtering model to obtain a new spatial signal standard deviation;
[0035] The first adjustment module is used to reconstruct the preset key point filter model by using the new spatial signal standard deviation to obtain an adjusted key point filter model.
[0036] Optionally, the adjustment unit includes:
[0037] A second calculation module is used to calculate the difference between the first facial key point and the facial key point at the previous moment as a displacement speed, and if the displacement speed is greater than a set value, the second adjustment module is executed; if the displacement speed is less than the set value, the third adjustment module is executed;
[0038] A second adjustment module is used to reduce the standard deviation of the spatial signals of the historical face key points of the key point filtering model, and use the reduced standard deviation of the spatial signals to reduce the preset value to obtain an adjusted key point filtering model;
[0039] The third adjustment module is used to increase the spatial signal standard deviation of the historical face key points in the key point filtering model, and use the increased spatial signal standard deviation to reconstruct the key point filtering model to obtain an adjusted key point filtering model.
[0040] A third aspect of an embodiment of the present invention discloses an electronic device, which is used to run a program, wherein the program, when running, executes the facial key point filtering method disclosed in the first aspect of the present invention.
[0041] A fourth aspect of an embodiment of the present invention discloses a storage medium, which includes a storage program, wherein when the program is running, the device where the storage medium is located is controlled to execute the facial key point filtering method disclosed in the first aspect of the present invention.
[0042] Based on the facial key point filtering method, device, electronic device and storage medium provided by the above-mentioned embodiments of the present invention, within the filtering window, a current frame facial image is obtained, and the current frame facial image is predicted using a convolutional neural network (CNN) model to obtain a first facial key point; the spatial signal standard deviation of the historical facial key points of a preset key point filtering model is adjusted according to the first facial key point to obtain an adjusted key point filtering model, wherein the key point filtering model is constructed based on a bilateral filter, the length of the filtering window, the time domain signal standard deviation of the historical facial key points and the spatial signal standard deviation of the historical facial key points; all frames of facial images within the filtering window are obtained, and the current frame facial image is predicted using a convolutional neural network (CNN) model to obtain key points; all the key points are input into the adjusted key point filtering model for key point filtering, and a second facial key point after filtering the current frame facial image is output, and the second facial key point is used to locate the face position in the current scene. In the solution provided by the embodiment of the present invention, the first face key point of the current frame face image is used to adjust the key point filtering model constructed based on the bilateral filter, the length of the filtering window, the standard deviation of the time domain signal of the historical face key points, and the standard deviation of the spatial signal of the historical face key points, and then the key points of all frame face images in the filtering window are filtered based on the adjusted key point filtering model, so as to obtain face key point information with high stability and greater randomness, and use the face key point information to locate the face, thereby overcoming the lag and jitter problems existing in the prior art when positioning through face key points, and achieving the purpose of more stable face positioning without misalignment. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.
[0044] Figure 1 A schematic diagram of a flow chart of a facial key point filtering method provided by an embodiment of the present invention;
[0045] Figure 2 A structural block diagram of a facial key point filtering device provided by an embodiment of the present invention;
[0046] Figure 3 A schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0047] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0048] In this application, the terms "comprises", "comprising" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprising a ..." does not exclude the presence of other identical elements in the process, method, article or device comprising the element.
[0049] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein.
[0050] According to the background, the existing technology has the disadvantages of lag and jitter when positioning through facial key points.
[0051] To this end, the embodiment of the present invention provides a method and device for filtering key points of a face, which uses the first key point of the face image of the current frame to adjust the key point filtering model constructed based on the bilateral filter, the length of the filtering window, the standard deviation of the time domain signal of the historical face key points, and the standard deviation of the spatial signal of the historical face key points, and then filters the key points of all the face images of the frames in the filtering window based on the adjusted key point filtering model, thereby obtaining highly stable and more random face key point information, and using the face key point information to locate the face, overcoming the lag and jitter problems existing in the prior art when locating by face key points, and achieving the purpose of more stable and non-misaligned face positioning. The following is a detailed introduction through specific embodiments.
[0052] refer to Figure 1 , shows a schematic flow chart of a method for filtering key points of a face provided by an embodiment of the present invention. The method comprises the following steps:
[0053] S101: obtaining a current face frame image within a filtering window, and predicting the current face frame image to obtain a first face key point.
[0054] In step S101, the current frame of face image refers to a frame of face image captured from the image data or video image within the filtering window.
[0055] The filter window refers to the area where the key point filter model is preset to filter the object to be filtered during the filtering process. The size of this area is the size of the filter window, which is also called the length of the filter window. It should be noted that the object to be filtered here refers to the key points of the face.
[0056] It should be noted that the value of the filter window length is an empirical value, and when setting a specific value, it can be determined according to the application requirements of the actual scene. Optionally, the filter window length is set to 15, and of course it can be other values.
[0057] It should be noted that the embodiments of the present invention are applied to scenarios where a mobile terminal implements AR special effects based on human faces. Therefore, human face data or images must exist in the captured images or the processed image data or video images.
[0058] In the specific implementation of step S101, a captured frame of facial image is input into a preset convolutional neural network (CNN) model, prediction is performed in the CNN model, and the first facial key point in the current frame of facial image is output.
[0059] Facial key points refer to the facial positions in the face area that can reflect the facial expression state, including feature points of the eyes, mouth, nose and other parts.
[0060] Here, the first facial key point output by the CNN model refers to the set of coordinate points obtained based on identifying the facial key points on the face image of the current frame.
[0061] The input of the CNN model is frames of images. In the video application scenario of realizing AR special effects based on faces, multiple frames of images can be obtained by decoding the captured video. By inputting the frames of images into the CNN model in the order of their generation, the coordinate sets of facial key points corresponding to different times can be obtained.
[0062] S102: adjusting the spatial signal standard deviation of the historical facial key points of the preset key point filter model according to the first facial key point to obtain an adjusted key point filter model.
[0063] In step S102, the preset key point filtering model is constructed based on the principle of bilateral filter.
[0064] Specifically, the preset key point filter model is a key point filter constructed based on the principle of bilateral filter by utilizing the filter window length, the time domain signal standard deviation and the space signal standard deviation.
[0065] In the embodiment of the present invention, the first facial key point to be filtered is filtered from the perspective of the time domain and the perspective of the space.
[0066] It should be noted that the difference between the first face key point at the current moment and the face key point at the previous moment (historical face key point) will affect the standard deviation of the spatial signal. Therefore, the spatial signal standard deviation of the historical face key points of the preset key point filtering model is adjusted according to the first face key point, and the adjusted key point filtering model is subsequently used to perform bilateral filtering on the face key points.
[0067] S103: Acquire key points obtained by predicting the current frame face image from all the frame face images within the filtering window.
[0068] In the specific implementation of step S103, key points are obtained by predicting the current frame face image using the frame face image corresponding to each moment in the filtering window.
[0069] The time corresponding to the filter window can be determined by the length of the filter window and the time corresponding to the current frame face image.
[0070] For ease of understanding, an example is given. If t represents the time corresponding to the acquired face image of the current frame, and H represents the acquired set filter window length, then based on the acquired face image of the current frame and the filter window length H, the time corresponding to the filter window is determined to be [tH, t].
[0071] S104: inputting all key points into the adjusted key point filtering model for key point filtering, outputting second face key points after filtering the face image of the current frame, and using the second face key points to locate the face position in the current scene.
[0072] In the specific implementation of step S104, all key points are used as the adjusted key point filtering model, key point filtering is performed in the key point filtering model, and the key point filtering model outputs the second face key point. The second face key point is then used to locate the face position in the current scene.
[0073] It should be noted that both the first facial key point and the second facial key point exist in the form of coordinate sets.
[0074] Based on the facial key point filtering method provided by the above-mentioned embodiment of the present invention, the pre-constructed key point filtering model is adjusted using the first facial key point of the current frame facial image, and then the key points of all frame facial images in the filtering window are filtered based on the adjusted key point filtering model, thereby obtaining facial key point information with high stability and greater randomness, and using the facial key point information to locate the face, thereby overcoming the problems of lag and jitter that exist in the prior art when locating through facial key points, and achieving the purpose of more stable and non-misaligned face positioning.
[0075] In addition, since the key point filtering model used is adjusted using the first face key point of the face image in the current frame each time face key point filtering is implemented, it is possible to ensure that the key point filtering model used effectively implements bilateral filtering of face key points.
[0076] For the above-mentioned embodiments of the present invention Figure 1 The specific construction process of the preset key point filtering model involved in the disclosed step S102 is described in detail below.
[0077] Bilateral filter (BF) is a nonlinear filter. When bilateral filtering is performed on an image, the bilateral filter (BF) is expressed as shown in formula (1):
[0078]
[0079] Where p represents the current pixel. q represents any pixel in the filter window corresponding to the current pixel. p Represents the grayscale value of the image at point p. q Represents the grayscale value of the image at point q. r (p) represents the pixel value weight of any pixel in the filter window corresponding to the current pixel p. s (p) represents the spatial distance weight of any pixel point in the filter window corresponding to the current pixel point p.
[0080] In the process of extracting facial key points, the technicians found that the facial key points in the previous period of time adjacent to the current moment also obey the Gaussian distribution. Based on this, the technicians constructed a key point filtering model matching the facial key points based on formula (1) to facilitate bilateral filtering of the facial key points, so as to achieve bilateral filtering of the facial key points while retaining the edge information of the facial key points, thereby improving the stability and tracking of the obtained facial key point information, and overcoming the lag and jitter shortcomings of the existing technology when positioning through facial key points.
[0081] In the embodiment of the present invention, considering the characteristics of the key points of the face and the bilateral filter, the standard deviation of the time domain signal is used as the standard deviation of the pixel points in the bilateral filter. Based on formula (1) and the filter window length, the standard deviation of the time domain signal and the standard deviation of the spatial signal, a key point filtering model is constructed, as shown in formula (2).
[0082]
[0083] Wherein, t represents the current time, which can also be understood as the time corresponding to the current frame face image. H represents the filter window length. The h moment represents any moment within the filter window determined based on the current time and the filter window length. represents the time domain distance weight of the frame image corresponding to time h, as shown in formula (3). represents the spatial distance weight of the frame image corresponding to time h, as shown in formula (4). h represents the facial key points of the frame face image corresponding to time h,
[0084]
[0085] Among them, σ time Represents the standard deviation of the time domain signal. The standard deviation of the time domain signal σ time The value of can be set according to the application requirements of the actual scene. When the value of the standard deviation of the time domain signal is larger, the degree of denoising of the image is greater. Conversely, when the value of the standard deviation of the time domain signal is smaller, the degree of denoising of the image is smaller.
[0086] Optionally, the standard deviation of the time domain signal may be set to one third of the length of the filtering window, or other values.
[0087]
[0088] Among them, σ space Represents the standard deviation of the spatial signal. space It can also be used to represent the noise standard deviation of the key point filtering model. The noise standard deviation is the statistic of the noise. When the noise increases, the noise standard deviation will also increase, that is, the spatial signal standard deviation will also increase. Conversely, when the noise decreases, the noise standard deviation will also increase, that is, the spatial signal standard deviation will also decrease.
[0089] x t Represents the coordinate set of facial key points of the frame image corresponding to time t, N represents the total number of points that make up the face contour. It should be noted that, for the sake of convenience, the present invention does not distinguish between horizontal and vertical coordinates, and all are represented by x. t Give a description.
[0090] Wt represents the set of weights of each frame of the image at the time corresponding to the filtering window, as shown in formula (5):
[0091]
[0092] In combination with the above description, the process of pre-building the key point filtering model based on the principle of the bilateral filter includes the following steps:
[0093] First, the length of the filter window is determined, and the historical frame face image within the filter window is obtained. The historical frame face image is predicted to obtain the historical face key points.
[0094] Secondly, obtain the time domain signal standard deviation σ of the historical face key points in the time domain time , the standard deviation of the time domain signal is used as the time domain distance weight of the bilateral filter
[0095] Secondly, obtain the spatial signal standard deviation σ of the historical face key points in space space , the spatial signal standard deviation is used as the spatial distance weight of the bilateral filter
[0096] Finally, using the temporal distance weight and the spatial distance weight A bilateral filter is constructed and used as a key point filtering model.
[0097] In the embodiment of the present invention, the standard deviation σ of the time domain signal is used. time and the standard deviation of the spatial signal σ space , filtering the coordinate set of the face key points to be filtered from the time domain perspective and the space perspective respectively. This makes the face key points after filtering have good tracking and stability, thereby ensuring that the face positioning using the filtered face key points is more accurate.
[0098] In the specific implementation, because the difference between the first face key point at the current moment and the historical face key point will affect the standard deviation of the spatial signal, in order to ensure that the face key points obtained by subsequent filtering have better follow-up and stability, it is necessary to make corresponding adjustments to the standard deviation of the spatial signal that differs before and after.
[0099] Based on the above Figure 1 The facial key point filtering method disclosed in the embodiment of the present invention may optionally specifically perform step S102 of adjusting the spatial signal standard deviation of the historical facial key points of the preset key point filtering model according to the first facial key points to obtain the adjusted key point filtering model, including the following steps:
[0100] S11: Obtain the current face contour area represented by the first face key point.
[0101] Generally speaking, the size of the image input into the preset CNN model is usually fixed, so the output facial key points will also be proportionally converted to the original image coordinate system according to the size of the image.
[0102] Therefore, in the specific implementation of S11, the actual face contour area on the face image of the current frame can be determined according to the first face key point.
[0103] S12: Calculate the square root of the ratio of the current face contour area to the reference face contour area to obtain the face change ratio.
[0104] In the specific implementation of S12, the face change ratio s can be obtained according to formula (6): t .
[0105]
[0106] Among them, area(x t ) represents the current face contour area, area base Indicates the baseline face contour area.
[0107] It should be noted that the reference facial contour area is the facial area obtained using standard biological parameters.
[0108] S13: multiplying the face change ratio by the spatial signal standard deviation of the historical face key points in the preset key point filtering model to obtain a new spatial signal standard deviation.
[0109] In the specific implementation of S13, a new spatial signal standard deviation can be obtained according to formula (7).
[0110]
[0111] S14: Reconstruct the preset key point filter model using the new spatial signal standard deviation to obtain an adjusted key point filter model.
[0112] In the specific implementation of S14, the standard deviation of the new spatial signal obtained based on formula (7) and formula (4) is calculated to obtain a new spatial distance weight, and then the new spatial distance weight is used to replace the old spatial distance weight in formula (2) to obtain a new key point filter model, that is, the adjusted key point filter model.
[0113] In an embodiment of the present invention, the face change ratio is determined based on the current face contour area and the benchmark face contour area, and then the face change ratio is used to adjust the spatial distance weight in the key point filtering model to obtain the adjusted key point filtering model, so as to use the adjusted key point filtering model to perform subsequent bilateral filtering processing, so as to improve the stability and followability of the obtained face key point information, and overcome the lag and jitter shortcomings existing in the prior art when positioning through face key points.
[0114] Based on the above Figure 1 The facial key point filtering method disclosed in the embodiment of the present invention may optionally specifically perform step S102 of adjusting the spatial signal standard deviation of the historical facial key points of the preset key point filtering model according to the first facial key points to obtain the adjusted key point filtering model, including the following steps:
[0115] S21: Calculate the difference between the first facial key point and the facial key point at the previous moment as the displacement speed.
[0116] S22: Determine the magnitude of the displacement speed and the set value. If the displacement speed is greater than the set value, continue to step S23. If the displacement speed is less than the set value, continue to step S24. If the displacement speed is equal to the set value, return to perform the next difference calculation.
[0117] S23: reducing the standard deviation of the spatial signals of the historical face key points of the key point filter model, and reconstructing the key point filter model using the reduced standard deviation of the spatial signals to obtain an adjusted key point filter model.
[0118] In the process of implementing S23, the preset value is reduced based on the standard deviation of the spatial signals of the historical face key points of the key point filtering model, and the key point filtering model is reconstructed using the standard deviation of the spatial signals of the historical face key points after reducing the preset value to obtain the adjusted key point filtering model.
[0119] When the displacement speed is greater than the set value, it can be regarded as a large displacement speed at this time. By reducing the standard deviation of the spatial signal in the key point filtering model by the preset value, the spatial distance weight of the frame image corresponding to the previous moment is suppressed, thereby ensuring the followability of the final facial key point information.
[0120] S24: increasing the standard deviation of the spatial signals of the historical face key points of the key point filter model, and reconstructing the key point filter model using the increased standard deviation of the spatial signals to obtain an adjusted key point filter model.
[0121] In the process of implementing S24, a preset value is added on the basis of the standard deviation of the spatial signals of the historical face key points of the key point filtering model, and the key point filtering model is reconstructed using the standard deviation of the spatial signals of the historical face key points after adding the preset value to obtain the adjusted key point filtering model.
[0122] When the displacement speed is less than the set value, it can be regarded as a small displacement speed at this time. By increasing the standard deviation of the spatial signal in the key point filtering model by the preset value, the spatial distance weight of the frame image corresponding to the previous moment is improved, thereby ensuring the stability of the final facial key point information.
[0123] In an embodiment of the present invention, the difference between the facial key points at the current moment and the facial key points at the previous moment is taken as the displacement speed, and the size of the displacement speed and the set value is judged, and different processing is performed on the reference spatial signal standard deviation in the preset key point filtering model according to the judgment result, so as to obtain an adjusted key point filtering model, so as to use the adjusted key point filtering model to perform subsequent bilateral filtering processing, so as to improve the stability and followability of the obtained facial key point information, and overcome the lag and jitter shortcomings existing in the prior art when positioning through facial key points.
[0124] Based on the above embodiments of the present invention Figure 1 The disclosed human face key point filtering method, correspondingly, the embodiment of the present invention also discloses a human face key point filtering device.
[0125] refer to Figure 2 , shows a structural block diagram of a facial key point filtering device provided by an embodiment of the present invention. The device comprises: a first acquisition unit 201, an adjustment unit 202, a second acquisition unit 203 and a processing unit 204.
[0126] The first acquisition unit 201 is used to acquire a current frame face image within a filtering window, and predict the current frame face image to obtain a first face key point.
[0127] The adjustment unit 202 is used to adjust the spatial signal standard deviation of the historical face key points of the preset key point filtering model according to the first face key point to obtain an adjusted key point filtering model, wherein the preset key point filtering model is constructed based on the principle of a bilateral filter.
[0128] The second acquisition unit 203 is used to acquire key points obtained by predicting the current frame face image from all the frame face images within the filtering window.
[0129] The processing unit 204 is used to input all the key points into the adjusted key point filtering model for key point filtering, output the second face key point information after filtering the face image of the current frame, and use the second face key point to locate the face position in the current scene. Optionally, the face key point filtering device also includes a pre-construction module.
[0130] The pre-built module is used to determine the length of the filtering window, obtain the historical frame face image within the filtering window, predict the historical frame face image, and obtain the historical face key points; obtain the time domain signal standard deviation of the historical face key points in the time domain, and use the time domain signal standard deviation as the time domain distance weight of the bilateral filter; obtain the spatial signal standard deviation of the historical face key points in space, and use the spatial signal standard deviation as the spatial distance weight of the bilateral filter; use the time domain distance weight and the spatial distance weight to construct a bilateral filter, and use the bilateral filter as a key point filtering model.
[0131] Optionally, the adjustment unit includes:
[0132] The first acquisition module is used to acquire the current face contour area represented by the first face key points.
[0133] The first calculation module is used to calculate the square root of the ratio of the current face contour area to the benchmark face contour area to obtain the face change ratio; and multiply the face change ratio by the spatial signal standard deviation of the historical face key points in the preset key point filtering model to obtain a new spatial signal standard deviation.
[0134] The first adjustment module is used to reconstruct the preset key point filter model by using the new spatial signal standard deviation to obtain an adjusted key point filter model.
[0135] Optionally, the adjustment unit includes:
[0136] The second calculation module is used to calculate the difference between the first facial key point and the facial key point at the previous moment as the displacement speed. If the displacement speed is greater than the set value, the second adjustment module is executed; if the displacement speed is less than the set value, the third adjustment module is executed.
[0137] The second adjustment module is used to reduce the spatial signal standard deviation of the historical face key points of the key point filtering model, and use the reduced spatial signal standard deviation to reduce the preset value to obtain the adjusted key point filtering model.
[0138] The third adjustment module is used to increase the spatial signal standard deviation of the historical face key points in the key point filtering model, and use the increased spatial signal standard deviation to reconstruct the key point filtering model to obtain an adjusted key point filtering model.
[0139] The specific implementation principles of each unit and module in the device disclosed in the above embodiment of the present invention can be found in the corresponding content of the method disclosed in the above embodiment of the present invention, which will not be repeated here.
[0140] Based on the human face key point filtering device provided by the above embodiment of the present invention,
[0141] The pre-built key point filtering model is adjusted by using the first face key point of the current frame face image, and then the key points of all frame face images in the filtering window are filtered based on the adjusted key point filtering model, so as to obtain face key point information with high stability and greater randomness, and the face key point information is used for face positioning, overcoming the lag and jitter problems existing in the existing technology when positioning by face key points, and achieving the purpose of more stable face positioning without misalignment.
[0142] The service development method disclosed in the embodiment of the present invention is based on the face key point filtering model disclosed in the embodiment of the present invention, and the above modules of the face key point filtering model can be implemented by a hardware device composed of a processor and a memory. Specifically, the above modules are stored in the memory as program units, and the processor executes the above program units stored in the memory to realize thread control.
[0143] The processor includes a kernel, which retrieves the corresponding program unit from the memory. One or more kernels can be set, and business system development can be achieved by adjusting kernel parameters.
[0144] An embodiment of the present disclosure provides a storage medium, which includes a business system development program, wherein when the program is executed by a processor, the business system development method disclosed in the above-mentioned embodiment of the present disclosure is implemented.
[0145] The present disclosure provides an electronic device, which may be a server, a PC, a PAD, a mobile phone, etc. Figure 3 As shown, the electronic device 30 includes at least one processor 301 , at least one memory 302 connected to the processor, and a bus 303 .
[0146] The processor 301 and the memory 302 communicate with each other via the bus 303 .
[0147] The processor 301 is used to execute the program stored in the memory.
[0148] The memory 302 is used to store a program, which is at least used to: obtain a current frame face image within a filtering window, predict the current frame face image, and obtain a first face key point; adjust the spatial signal standard deviation of the historical face key points of a preset key point filtering model according to the first face key point to obtain an adjusted key point filtering model, wherein the preset key point filtering model is constructed based on the principle of a bilateral filter; obtain key points obtained by predicting the current frame face image from all frame face images within the filtering window; input all the key points into the adjusted key point filtering model for key point filtering, output a second face key point after filtering the current frame face image, and use the second face key point to locate the face position in the current scene.
[0149] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can refer to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system or system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiment. The system and system embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without creative work.
[0150] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0151] The above description of the disclosed embodiments enables one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A facial key point filtering method, characterized in that: The method comprises: In the filtering window, a current frame face image is obtained, and the current frame face image is predicted to obtain a first face key point; adjusting the spatial signal standard deviation of the historical facial key points in the preset key point filtering model according to the first facial key points to obtain an adjusted key point filtering model, wherein the preset key point filtering model is constructed based on the principle of a bilateral filter; Acquire key points obtained by predicting the current frame face image from all the frame face images within the filtering window; Input all the key points into the adjusted key point filtering model to perform key point filtering, output the second face key points after filtering the face image of the current frame, and use the second face key points to locate the face position in the current scene; The preset key point filtering model is constructed based on the principle of bilateral filter, including: Determine the length of the filter window, obtain the historical frame face image within the filter window, predict the historical frame face image, and obtain the historical face key points; Obtaining a time domain signal standard deviation of the historical facial key points in the time domain, and using the time domain signal standard deviation as a time domain distance weight of a bilateral filter; Obtaining a spatial signal standard deviation of the historical facial key points in space, and using the spatial signal standard deviation as a spatial distance weight of the bilateral filter; A bilateral filter is constructed using the time domain distance weight and the space distance weight, and the bilateral filter is used as a key point filtering model.
2. The method according to claim 1, characterized in that: The step of adjusting the spatial signal standard deviation of the historical facial key points in the preset key point filter model according to the first facial key points to obtain an adjusted key point filter model includes: Acquire the current face contour area represented by the first face key points; Calculating the square root of the ratio of the current face contour area to the reference face contour area to obtain a face change ratio; Multiplying the face change ratio by the spatial signal standard deviation of the historical face key points in the preset key point filtering model to obtain a new spatial signal standard deviation; The preset key point filter model is reconstructed using the new spatial signal standard deviation to obtain an adjusted key point filter model.
3. The method according to claim 1, characterized in that The step of adjusting the spatial signal standard deviation of the historical facial key points in the preset key point filter model according to the first facial key points to obtain an adjusted key point filter model includes: Calculate the difference between the first facial key point and the facial key point at the previous moment as the displacement speed; If the displacement speed is greater than the set value, reducing the spatial signal standard deviation of the historical face key points in the key point filtering model, and reconstructing the key point filtering model using the reduced spatial signal standard deviation to obtain an adjusted key point filtering model; If the displacement speed is less than the set value, the spatial signal standard deviation of the historical face key points in the key point filtering model is increased, and the key point filtering model is reconstructed using the increased spatial signal standard deviation to obtain an adjusted key point filtering model.
4. A facial key point filtering device, characterized in that: The device comprises: A first acquisition unit is used to acquire a current frame face image within a filtering window, predict the current frame face image, and obtain a first face key point; an adjusting unit, configured to adjust a spatial signal standard deviation of a historical face key point in a preset key point filtering model according to the first face key point, so as to obtain an adjusted key point filtering model, wherein the preset key point filtering model is constructed based on the principle of a bilateral filter; A second acquisition unit is used to acquire key points obtained by predicting the current frame face image from all the frame face images within the filtering window; A processing unit, used for inputting all the key points into the adjusted key point filtering model for key point filtering, outputting the second face key point information after filtering the face image of the current frame, and locating the face position in the current scene by using the second face key point; A pre-construction module is used to determine the length of the filtering window, obtain the historical frame face image within the filtering window, predict the historical frame face image, and obtain the historical face key points; obtain the time domain signal standard deviation of the historical face key points in the time domain, and use the time domain signal standard deviation as the time domain distance weight of the bilateral filter; obtain the spatial signal standard deviation of the historical face key points in space, and use the spatial signal standard deviation as the spatial distance weight of the bilateral filter; use the time domain distance weight and the spatial distance weight to construct a bilateral filter, and use the bilateral filter as the key point filtering model.
5. The device according to claim 4, characterized in that The adjustment unit comprises: A first acquisition module, used to acquire the current face contour area represented by the first face key points; A first calculation module is used to calculate the square root of the ratio of the current face contour area to the reference face contour area to obtain a face change ratio; and multiply the face change ratio by the spatial signal standard deviation of the historical face key points in the preset key point filtering model to obtain a new spatial signal standard deviation; The first adjustment module is used to reconstruct the preset key point filter model by using the new spatial signal standard deviation to obtain an adjusted key point filter model.
6. The device according to claim 4, characterized in that The adjustment unit comprises: A second calculation module is used to calculate the difference between the first facial key point and the facial key point at the previous moment as a displacement speed, and if the displacement speed is greater than a set value, the second adjustment module is executed; if the displacement speed is less than the set value, the third adjustment module is executed; A second adjustment module is used to reduce the standard deviation of the spatial signals of the historical face key points in the key point filtering model, and use the reduced standard deviation of the spatial signals to reduce the preset value to obtain an adjusted key point filtering model; The third adjustment module is used to increase the spatial signal standard deviation of the historical face key points in the key point filtering model, and use the increased spatial signal standard deviation to reconstruct the key point filtering model to obtain an adjusted key point filtering model.
7. An electronic device, characterized in that: The electronic device is used to run a program, wherein the program executes the facial key point filtering method as described in any one of claims 1 to 3 when running.
8. A storage medium, characterized in that: The storage medium includes a storage program, wherein when the program is running, the device where the storage medium is located is controlled to execute the facial key point filtering method as described in any one of claims 1 to 3.
Citation Information
Patent Citations
Autonomous mobile robot trajectory planning method, system and device
CN112099493A
Position determination method and device, electronic equipment and storage medium
CN113223083A