A face localization method and related device
By dividing the face into a contour area and a five-feature area for separate positioning, the delay problem caused by large calculations in the prior art is solved, and efficient and accurate face positioning is achieved in image processing scenes with high real-time performance.
Patent Information
- Application Number
- CN202111679411.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-31
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2041-12-31
AI Technical Summary
Due to the large calculation amount of the face positioning method in the prior art, the processing delay is long, making it difficult to be applied to image processing scenarios with high real-time requirements.
By dividing the face into the contour area of the face and the facial features area for separate positioning, the contour key points and facial features are identified separately, reducing the calculation amount and improving the positioning accuracy.
It realizes that the processing delay of face positioning is significantly reduced while ensuring positioning accuracy, so that face positioning technology can be extended to image processing scenarios with high real-time performance.
Smart Images

Figure CN116416658B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing, and in particular, to a face localization method and related device. Background Art
[0002] Face localization technology can effectively localize faces in image frames, and the obtained localization results can identify the positions of faces in image frames, facial shapes, distributions of facial features, etc. Such localization results can provide data bases for various image processing technologies such as face recognition, beauty enhancement, and face reconstruction.
[0003] In related technologies, in order to ensure face localization accuracy, network models that consume a large amount of system resources are generally used, and complex algorithms are adopted to implement face localization. The algorithms of such network models have a large amount of computation, resulting in a certain processing delay in face localization, making it difficult for related technologies to be applicable to more and more image processing scenarios with high real-time requirements. Summary of the Invention
[0004] To solve the above technical problems, this application provides a face localization method and related device, which are used to reduce the processing delay of face localization while ensuring face localization accuracy, so that face localization technology can be extended to image processing scenarios with higher real-time performance.
[0005] The embodiments of this application disclose the following technical solutions:
[0006] On the one hand, the embodiments of this application provide a face localization method, and the method includes:
[0007] Obtain the i-th image frame of the video content to be recognized, where the i-th image frame is one of the consecutive N image frames included in the video content to be recognized;
[0008] Determine the face contour area and the facial feature areas of the face to be localized in the face detection box through the face detection box in the i-th image frame;
[0009] Identify contour key points according to the face contour area, and identify facial feature key points according to the facial feature areas; where the contour key points are used to identify the face contour of the face to be localized, and the facial feature key points are used to identify the facial features of the face to be localized;
[0010] Generate a face localization result for the face to be localized in the i-th image frame based on the contour key points and the facial feature key points.
[0011] On the other hand, the embodiments of this application provide a face localization device, and the device includes: an acquisition unit, a determination unit, an identification unit, and a generation unit;
[0012] The obtaining unit is configured to obtain the i-th image frame of the video content to be recognized, where the i-th image frame is one of the consecutive N image frames included in the video content to be recognized;
[0013] The determining unit is configured to determine, through the face detection box in the i-th image frame, the face contour area and the facial feature areas of the face to be located in the face detection box;
[0014] The recognizing unit is configured to recognize contour key points according to the face contour area, and recognize facial feature key points according to the facial feature areas; wherein, the contour key points are used to identify the face contour of the face to be located, and the facial feature key points are used to identify the facial features of the face to be located;
[0015] The generating unit is configured to generate a face positioning result for the face to be located in the i-th image frame based on the contour key points and the facial feature key points.
[0016] On the other hand, an embodiment of the present application provides a computer device, which includes a processor and a memory:
[0017] The memory is configured to store program codes and transmit the program codes to the processor;
[0018] The processor is configured to execute the method described in the above aspect according to the instructions in the program codes.
[0019] On the other hand, an embodiment of the present application provides a computer-readable storage medium, which is configured to store a computer program, and the computer program is used to execute the method described in the above aspect.
[0020] On the other hand, an embodiment of the present application provides a computer program product or a computer program, which includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method described in the above aspect.
[0021] As can be seen from the above technical solution, when performing face localization on the video content to be recognized including N consecutive image frames, the real-time requirement is relatively high. During the movement of the face, the change amplitudes of the face contour and facial features are generally significantly different. For example, the change amplitude of the face contour is larger than that of the facial features. Therefore, in order to reduce the processing delay in the related technology, for the i-th image frame of the video content to be recognized, the face contour area and the facial feature area of the face to be located are determined through the face detection frame. Since the change amplitudes of the face parts within the face contour area or the facial feature area are relatively more consistent, separating the localization of the face contour and facial features can actually achieve a more accurate effect.
[0022] Compared with the related technology that locates each face part with different change amplitudes as a whole, the method of separately locating the face contour and facial features decouples the change amplitude of the face, which is equivalent to splitting the complex problem in the related technology into multiple simple problems for processing. This not only significantly reduces the overall computational amount of face localization, but also can more accurately identify the contour key points indicating the face contour based on the face contour area, and identify the feature key points indicating the facial features based on the facial feature area, and obtain the face localization result in the i-th image frame according to the contour key points and the feature key points. The reduced computational amount effectively reduces the time delay of face localization, enabling the face localization technology to be extended to image processing scenarios with relatively high real-time performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0024] Figure 1 Schematic diagram of the application scenario of the face localization method provided by the embodiment of the present application;
[0025] Figure 2 Schematic flow chart of a face localization method provided by the embodiment of the present application;
[0026] Figure 3 Schematic diagram of a face detection frame provided by the embodiment of the present application;
[0027] Figure 4 Schematic diagram of determining the face contour area and the facial feature area provided by the embodiment of the present application;
[0028] Figure 5 Schematic diagram of determining the face contour area and the facial feature area provided by the embodiment of the present application;
[0029] Figure 6 A schematic diagram of different regional ranges provided by an embodiment of the present application;
[0030] Figure 7 A schematic diagram of determining the distance between a to-be-detected bounding box and an image position through GIOU provided by an embodiment of the present application;
[0031] Figure 8 A schematic diagram of adjusting a face detection bounding box provided by an embodiment of the present application;
[0032] Figure 9 A schematic diagram of target area correction provided by an embodiment of the present application;
[0033] Figure 10 A schematic diagram of an application scenario of a face positioning method provided by an embodiment of the present application;
[0034] Figure 11 A flowchart of hierarchical key point positioning provided by an embodiment of the present application;
[0035] Figure 12 A schematic diagram of a face positioning device provided by an embodiment of the present application;
[0036] Figure 13 A structural diagram of a terminal device provided by an embodiment of the present application;
[0037] Figure 14 A structural diagram of a server provided by an embodiment of the present application. Detailed implementation manners
[0038] The embodiments of the present application will be described below with reference to the accompanying drawings.
[0039] In view of the face positioning technology in the related art, algorithms with relatively large computational complexity are required to ensure the face positioning accuracy, resulting in processing delays while ensuring the accuracy, making it difficult to extend the scenarios of the related art to image processing scenarios with high real-time performance. Based on this, the embodiments of the present application provide a face positioning method and related devices, which are used to reduce the computational complexity of the algorithm while ensuring the face positioning accuracy, thereby reducing the processing delay of face positioning, so that the face positioning technology can be extended to image processing scenarios with high real-time performance.
[0040] The data processing method provided by the embodiments of this application is implemented based on artificial intelligence. Artificial Intelligence (AI) uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, and is a theory, method, technology, and application system that can perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling machines to have the functions of perception, reasoning, and decision-making.
[0041] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning, autonomous driving, and intelligent transportation.
[0042] In the embodiments of this application, the artificial intelligence technologies mainly involved include the above-mentioned directions such as machine learning / deep learning. The embodiments of this application can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, intelligent transportation, assisted driving, etc.
[0043] The face positioning method provided by this application can be applied to computer devices with face positioning capabilities, such as terminal devices and servers. Among them, the terminal device can specifically be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart voice interaction device, a smart home appliance, a vehicle-mounted terminal, etc., but is not limited thereto; the server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The terminal device and the server can be directly or indirectly connected through wired or wireless communication methods, and this application does not limit this.
[0044] The computer device may have machine learning capabilities. Machine Learning (ML) is an interdisciplinary field that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.
[0045] In the face localization method provided in the embodiments of the present application, the artificial intelligence model adopted mainly involves the application of machine learning, and determines one or more of a face detection frame, a face contour area, a face facial feature area, contour key points, facial feature key points, and a face localization result through an artificial neural network model, etc.
[0046] To facilitate the understanding of the technical solution of the present application, the face localization method provided in the embodiments of the present application will be introduced below in combination with an actual application scenario.
[0047] See Figure 1 , which is a schematic diagram of the application scenario of the face localization method provided in the embodiments of the present application. In Figure 1 the described application scenario, the aforementioned computer device is a smart phone 100, and the user uses the smart phone 100 to take a self-portrait, and the smart phone 100 detects the face in real time and performs beauty processing on the detected face. The process of face detection will be described below.
[0048] After the user turns on the self-portrait function through the smart phone 100, adjustments are generally made and then the shooting button is pressed. During the adjustment process, the smart phone 100 will obtain the video content to be recognized including N consecutive image frames, and perform face detection on each image frame of the video content to be recognized so as to perform beauty processing on each image frame. The following will take the i-th image frame as an example for illustration.
[0049] The smart phone 100 obtains the i-th image frame of the video content to be recognized and determines the face detection frame in the i-th image frame, and this face detection frame can roughly identify the area where the face is located in the i-th image frame.
[0050] During the movement of a human face, there are generally significant differences in the change amplitudes of the human face contour and facial features. For example, the change amplitude of the human face contour is larger than that of the facial features. Related technologies have not discovered the differences in the change amplitudes of different regions of the human face during face movement, and use a unified model for recognition, resulting in complex algorithms and large computational amounts, thereby causing processing delays and being unable to be applied to image processing scenarios with high real-time requirements.
[0051] Based on this, in the embodiments of the present application, the face to be located is divided into multiple regions according to different change amplitudes, the face contour region and the facial feature regions of the face to be located are determined, contour key points are recognized according to the face contour region, and facial feature key points are recognized according to the facial feature regions. Based on the contour key points and the facial feature key points, a face positioning result for the face to be located in the i-th image frame is generated.
[0052] Compared with the related technologies that perform face positioning by treating each part of the human face with different change amplitudes as a whole, the method of separately positioning the face contour and facial features realizes the decoupling of the face change amplitude, which is equivalent to splitting the complex problems in the related technologies into multiple simple problems for processing, significantly reducing the overall computational amount of face positioning. At the same time, since within the same region, such as the face contour region or the facial feature regions, the change amplitudes of the face parts are relatively more consistent, separately positioning the face contour and facial features can actually achieve a more accurate effect. Thus, while ensuring the face positioning accuracy, the processing delay of face positioning is reduced, enabling the face positioning technology to be extended to image processing scenarios with high real-time requirements.
[0053] Next, taking the aforementioned computer device as an example of a terminal device, a face positioning method provided by the embodiments of the present application will be introduced with reference to the accompanying drawings.
[0054] See Figure 2 , which is a schematic flowchart of a face positioning method provided by the embodiments of the present application. As Figure 2 shown, the face positioning method may include S201 - S204.
[0055] S201: Obtain the i-th image frame of the video content to be recognized.
[0056] Since the video content to be recognized includes N consecutive image frames with a small interval between the image frames, such as 30 image frames within 1 second, when performing face positioning on the video content to be recognized, generally, face positioning is performed on each image frame of the video content to be recognized, and the real-time requirement for the face positioning method is relatively high. Herein, N is a positive integer.
[0057] However, since the related technology uses complex algorithms to implement face localization, it is not only inapplicable to image processing scenarios with high real-time requirements due to large processing delays, but also cannot be applied to edge devices on mobile terminals, such as smartphones and various devices with low-end chips, because of the large amount of computation.
[0058] Based on this, the present application proposes a face localization method to overcome the above problems. Among them, the terminal device can obtain the complete video content to be recognized for storage, and can also continuously obtain image frames of the video content to be recognized one by one and selectively store them. For example, when taking a selfie, image frames are continuously obtained, and only when the user presses the selfie button will they be saved.
[0059] For the convenience of description, hereinafter, the i-th image frame in the video content to be recognized will be taken as an example to illustrate the face localization method provided in the embodiments of the present application.
[0060] It should be noted that the face localization method provided in the embodiments of the present application is not only applicable to image processing scenarios with high real-time requirements, but also applicable to image processing scenarios without real-time requirements, and terminal devices that cannot achieve a large amount of computation, etc. The present application does not make specific limitations on this.
[0061] It can be understood that in the specific implementation manner of the present application, the video content to be recognized involves data related to user information such as faces. When the face localization method included in the present application is applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards of relevant countries and regions.
[0062] S202: Determine the face contour area and the facial feature area of the face to be located in the face detection box through the face detection box in the i-th image frame.
[0063] After obtaining the i-th image frame, a face detection box is determined in the i-th image frame by means of a cross-platform computer vision library (Open Source Computer Vision Library, OpenCV), building an artificial neural network, etc. The face detection box can identify the area where a face may exist in the i-th image frame. As Figure 3 shown, five face detection boxes are used to respectively identify the areas where the faces of each person in the i-th image frame are located. It should be noted that the face detection box can be displayed to the user through the terminal device, or can not be displayed to the user, and is only used to identify the area where a face may exist for subsequent face localization.
[0064] Through research, it is found that during the movement of a human face, the variation amplitudes of different facial regions generally have obvious differences. For example, the variation amplitude of the facial contour is larger than that of the facial features. Related technologies use a unified model for recognition, resulting in a complex algorithm and a large amount of computation, thus causing processing delays and being unable to be applied to image processing scenarios with high real-time requirements.
[0065] Based on this, in the embodiments of the present application, the face to be located is divided into multiple regions according to different variation amplitudes. For example, the facial contour region and the facial feature region of the face to be located are determined. Thus, subsequently, recognition can be performed based on the facial contour region with a larger variation amplitude, and recognition can be performed based on the facial feature region with a smaller variation amplitude. Instead of using a unified model for recognition, the algorithm complexity is reduced. At the same time, targeted separate recognition also improves the accuracy of the face location result.
[0066] Among them, the facial contour region refers to the region formed by the facial contours outside the eyebrows, eyes, nose, and mouth. The facial feature region refers to the region where the eyebrows, eyes, nose, mouth, and ears are located. As a possible implementation manner, the region formed by the eyebrows, eyes, nose, and mouth can be used as the facial feature region for subsequent recognition, so as to reduce the amount of computation without affecting the accuracy.
[0067] The embodiments of the present application do not specifically limit the manner of determining the facial contour region and the facial feature region of the face to be located in the face detection frame. For example, through the principal component analysis (PCA), deep neural network models such as the CoarseNet model (an artificial neural network model), or support vector machine (SVM), etc. Subsequently, an example will be given in a manner (S2021 - S2023), and details will not be elaborated here.
[0068] S203: Identify the contour key points according to the facial contour region, and identify the feature key points according to the facial feature region.
[0069] Among them, the contour key points are used to identify the facial contour of the face to be located, and the feature key points are used to identify the facial features of the face to be located. It can be understood that when the facial feature region is the region formed by the eyebrows, eyes, nose, and mouth, the feature key points can be the key points corresponding to the eyebrows, eyes, nose, and mouth.
[0070] The embodiments of the present application do not specifically limit the manner of identifying key points based on the region. For example, the facial contour region is input into the ContourNet model (an artificial neural network model) to obtain the contour key points, and the facial feature region is input into the InlineNet model (an artificial neural network model) to obtain the feature key points.
[0071] Thus, compared with the related art that locates each face part with different variation ranges as a whole for face localization, the method of separately locating the face contour and facial features realizes the decoupling of the face variation range, which is equivalent to splitting the complex problem in the related art into multiple simple problems for processing, significantly reducing the overall computation amount of face localization.
[0072] S204: Generate a face localization result for the face to be located in the i-th image frame based on the contour key points and the facial feature key points.
[0073] If there is a face in the i-th image frame, the face localization result of the face to be located includes the localization result obtained based on the contour key points and the facial feature key points; if there is no face in the i-th image frame, the face localization result of the face to be located is that there is no face.
[0074] As can be seen from the above technical solution, when performing face localization on the video content to be recognized including N consecutive image frames, the real-time requirement is relatively high. Since during the face movement process, the variation ranges of the face contour and the facial features generally have obvious differences. For example, the variation range of the face contour is larger than that of the facial features. Therefore, in order to reduce the processing delay in the related art, for the i-th image frame of the video content to be recognized, determine the face contour area and the facial feature area of the face to be located through the face detection box. Since the variation ranges of the face parts within the face contour area or the facial feature area are relatively more consistent, separately locating the face contour and the facial features can actually achieve a more accurate effect. Compared with the related art that locates each face part with different variation ranges as a whole for face localization, the method of separately locating the face contour and the facial features realizes the decoupling of the face variation range, which is equivalent to splitting the complex problem in the related art into multiple simple problems for processing. Not only significantly reduces the overall computation amount of face localization, but also can more accurately identify the contour key points indicating the face contour based on the face contour area, and identify the facial feature key points indicating the facial features based on the facial feature area, and obtain the face localization result in the i-th image frame according to the contour key points and the facial feature key points. The smaller computation amount effectively reduces the time delay of face localization, enabling the face localization technology to be extended to image processing scenarios with higher real-time performance.
[0075] In a possible implementation manner, an embodiment of the present application provides an S202, that is, a specific implementation manner of determining the face contour area and the facial feature area of the face to be located in the face detection box through the face detection box in the i-th image frame, specifically including S2021 - S2023.
[0076] S2021: Determine the initial face key points corresponding to the face to be located in the face detection box through the face detection box in the i-th image frame.
[0077] Among them, the initial face key points are used to identify the key points of facial features such as the forehead, nose, eyes, etc.
[0078] See Figure 4 , which is a schematic diagram of determining the face contour area and the facial feature areas provided by the embodiment of the present application. According to the position of the face detection box in the i-th image frame, the face image of the corresponding area is cropped to obtain a face box cropped image, and the face box cropped image is input into the CoarseNet model for rough positioning to obtain the initial face key points.
[0079] S2022: Divide the initial face key points into a contour point set for identifying the initial contour of the face to be located, and an inner point set for identifying the initial facial features of the face to be located.
[0080] Since the initial face key points can clarify the facial features of the face, for example, one initial face key point is a point for identifying the eyes, and another initial face key point is a point for identifying the mouth, etc., the initial face key points can be divided into a contour point set for identifying the initial contour of the face, and an inner point set for identifying the initial facial features of the face to be located.
[0081] Continue to refer to Figure 4 , the number of initial face key points obtained by rough positioning is 16, 4 of which belong to the contour point set and 12 belong to the inner point set.
[0082] S2023: Determine the face contour area according to the contour point set, and determine the facial feature areas according to the inner point set.
[0083] Continue to refer to Figure 4 , based on the contour point set, the image corresponding to the face contour area is cropped from the face box cropped image, and based on the inner point set, the image corresponding to the facial feature areas is cropped from the face box cropped image.
[0084] It should be noted that although the image corresponding to the face contour area includes the image corresponding to the facial feature areas, in subsequent recognition, it only focuses on the face contour area and does not focus on the facial feature areas, which will not increase the calculation amount.
[0085] In the embodiment of the present application, for the face to be located in the i-th image frame, the initial face key points can be roughly identified, and then the face contour area and the facial feature areas can be determined. Thus, the calculation amount can be further reduced by the rough positioning method.
[0086] Further, in the case where the face to be located is rotated in the i-th image frame, the embodiment of the present application provides a S2023, that is, a specific implementation manner of determining the face contour area according to the contour point set and determining the face facial feature area according to the inlier set, which is as follows:
[0087] If the face to be located is rotated, the first face angle of the face to be located can be determined according to the inlier set, where the first face angle indicates the inclination degree of the face to be located in the i-th image frame, or the rotation situation of the face to be located in the i-th image frame.
[0088] The embodiment of the present application does not specifically limit the manner of determining the first face angle of the face to be located according to the inlier set. For example, determine the points used to identify the eyes of the face to be located from the inlier set, thereby determining the straight line where the two eyes are located, and further determining the first face angle. Or, determine the points used to identify the ears of the face to be located from the inlier set, thereby determining the straight line where the two ears are located, and further determining the first face angle.
[0089] See Figure 5 , which is a schematic diagram of determining the face contour area and the face facial feature area provided by the embodiment of the present application. Compared with Figure 4 , Figure 5 In the case where the face to be located shown is rotated in the i-th image frame. After obtaining the face box cutout, input the face box cutout into the CoarseNet model for rough positioning to obtain the initial face key points, and obtain the contour point set and the inlier set based on the initial face key points. Obtain the points used to identify the two eyes of the face to be located from the inlier set, so as to obtain the face roll angle (Roll) direction angle value, where the Roll direction angle value belongs to a type of first face angle and is the angle with the earth's horizontal plane, indicating left tilt or right tilt. In other words, it is the angle of rotation along the X-axis of its own coordinate system (the coordinate system with the X-axis forward).
[0090] After determining the first face angle, determine the face contour area and the face facial feature area in combination with the first face angle, which will be described separately below.
[0091] A specific implementation manner of determining the face contour area according to the contour point set includes A1 and A2.
[0092] A1: Generate a first affine transformation matrix corresponding to the contour point set through the contour point set and the first face angle.
[0093] Among them, the affine transformation matrix is used to identify how to perform affine transformations (such as rotation, scaling, translation, etc.) so as to perform an affine transformation on the image according to the geometric structure of the face in the image, transform the face into a unified state, achieve face alignment, or eliminate the rotation that actually exists in the i-th image of the to-be-determined face, thereby improving the accuracy of subsequent face localization.
[0094] As a possible implementation, the circumscribed rectangle corresponding to the contour point set can be determined according to the contour point set, and the first affine transformation matrix corresponding to the contour point set can be determined according to the first face angle and the circumscribed rectangle corresponding to the contour point set.
[0095] A2: Extract the face contour region from the i-th image frame according to the first affine transformation matrix.
[0096] Continue to refer to Figure 5 , extract the image corresponding to the face contour region from the face box cutout according to the first affine transformation matrix, realize the correction of the face contour, thereby improving the accuracy of subsequent face localization.
[0097] A specific implementation of determining the facial feature regions according to the inlier set includes B1 and B2.
[0098] B1: Generate the second affine transformation matrix corresponding to the inlier set through the inlier set and the first face angle.
[0099] As a possible implementation, the circumscribed rectangle corresponding to the inlier set can be determined according to the inlier set, and the second affine transformation matrix corresponding to the inlier set can be determined according to the first face angle and the circumscribed rectangle corresponding to the inlier set.
[0100] B2: Extract the facial feature regions from the i-th image frame according to the second affine transformation matrix.
[0101] Continue to refer to Figure 5 , extract the image corresponding to the facial feature regions from the face box cutout according to the second affine transformation matrix, realize the correction of the facial features, thereby improving the accuracy of subsequent face localization.
[0102] In the embodiment of the present application, for the case where the to-be-localized face has rotation in the i-th image frame, after determining the first face angle through the inlier set, the first affine transformation matrix and the second affine transformation matrix are determined, and the rotation existing in the i-th image frame of the to-be-localized face is eliminated through affine transformation, thereby obtaining the corrected face contour region and facial feature regions, and thus improving the accuracy of subsequent face localization.
[0103] In a possible implementation manner, the embodiments of the present application provide two specific implementation manners for determining the face detection box described in S202 above based on whether the face to be located is located in the (i - 1)-th image frame, which will be described separately below. Among them, in N consecutive image frames, the (i - 1)-th image frame is the previous image frame of the i-th image frame.
[0104] First: The face to be located is not located in the (i - 1)-th image frame.
[0105] At this time, the i-th image frame is in a non-following state, that is, in the process of continuously locating the face in each image frame, the face to be located is not located in the (i - 1)-th image frame, or the face is lost. At this time, it is necessary to re-find the position of the face from the entire i-th image frame, that is, to determine the face detection box from the area range corresponding to the i-th image frame. The specific determination method can adopt the method for determining the face detection box described in S202 above, which will not be elaborated here.
[0106] Second: The face to be located is located in the (i - 1)-th image frame.
[0107] At this time, the i-th image frame is in a following state, that is, in the process of continuously locating the face in each image frame, the face to be located is located in the (i - 1)-th image frame, or the face is not lost. Since in the video content to be recognized, the movement change range of the same face to be located in adjacent image frames is small, the position of the face to be located in the (i - 1)-th image frame in this image frame, that is, the image position, can be used as the basis for the face detection box in the i-th image frame. That is, according to the image position of the face to be located in the (i - 1)-th image frame, a face detection area including the image position is determined from the i-th image frame, and the face detection box is determined in the face detection area.
[0108] Among them, the area range corresponding to the face detection area is smaller than the area range corresponding to the i-th image frame, so that when performing face positioning subsequently, the area to be recognized becomes smaller, and thus the calculation amount is reduced when the face is not lost in the video stream sequence. Moreover, using the image position of the face to be located in the (i - 1)-th image frame as the basis for the face detection box in the i-th image frame can ensure that the face is always followed when the face moves rapidly and can follow the face with high precision without loss.
[0109] As a possible implementation manner, according to the image position of the face to be located in the (i - 1)-th image frame (such as the face positioning result of the face to be located in the (i - 1)-th image frame, the circumscribed rectangle determined according to the face positioning result, etc.), the area where the image position is located can be found in the i-th image frame and expanded outward (such as expanding the area by a preset threshold multiple uniformly in all directions) to obtain a face detection area including the image position.
[0110] See Figure 6 , which is a schematic diagram of different regional ranges provided by an embodiment of the present application. Figure 6 The left figure shows the regional range corresponding to the i-th image frame, including the original image data of the entire i-th image frame. Figure 6 The right figure shows the face detection region obtained by expanding the region where the face localization result in the (i - 1)-th image frame is located by 3 times for the face to be localized. Two different input states correspond to different face detection models. Through experiments, it is found that the computational complexity of the face detection model in the full-image mode ( Figure 6 with the left figure as the input) is 4 times that of the face detection model in the local mode ( Figure 6 with the right figure as the input).
[0111] In an embodiment of the present application, when there is no following state, the full image of the current frame is used as the basis. Since the data volume of the full image is large, the face detection model corresponding to the full image (large model module) is called to determine the face detection frame. When there is a following state, the image position of the face to be localized in the previous frame is used as the basis to determine a face detection region with a smaller data volume, and the face detection model corresponding to the local image (small model module) is called to determine the face detection frame. Thus, according to whether there is a following state, the call logic of each module is automatically adjusted, which can reduce the computational complexity when the face is not lost in the video stream sequence, and ensure that the face can be accurately followed without loss when the face moves rapidly.
[0112] In a possible implementation manner, an embodiment of the present application provides a specific implementation manner for determining the face detection frame by means of detection frame mapping in the aforementioned second manner, that is, when a face to be localized is located in the (i - 1)-th image frame. The detection frame mapping method is specifically as follows: Based on the possibility of face existence, a plurality of pending detection frames are determined in the face detection region, and the face detection frame is determined from the plurality of pending detection frames according to the distances between the plurality of pending detection frames and the image position respectively.
[0113] In the i-th image frame, based on the possibility of face existence, a plurality of pending detection frames can be determined in the face detection region, and each pending detection frame identifies a region where a face may exist in the face detection region. In the video content to be recognized, the movement change amplitude of the same face to be localized in adjacent image frames is small, that is, for the same face to be localized, the change between the i-th image frame and the (i - 1)-th image frame is small. Therefore, the image position of the face to be localized in the (i - 1)-th image frame can be used as the basis to determine the face detection frame from the plurality of detection frames.
[0114] For example, according to the distances between multiple detection boxes to be determined and the image position respectively, the detection box to be determined with the shortest distance to the image position is used as the face detection box. The embodiments of the present application do not specifically limit the method for determining the distance between the detection box to be determined and the image position. Hereinafter, the Generalized Intersection over Union (GIOU) is taken as an example for illustration.
[0115] See Figure 7 , which is a schematic diagram for determining the distance between the detection box to be determined and the image position by GIOU provided in the embodiments of the present application. Figure 7 The left figure shows three rectangles, namely the circumscribed rectangle A (filled with left-slanted stripes) of the face localization result of the face to be determined in the (i - 1)-th image frame, a detection box B to be determined among multiple detection boxes to be determined in the i-th image frame (filled with right-slanted stripes), and the smallest rectangle C (represented by thick lines) that contains both A and B.
[0116] The union part of the circumscribed rectangle A and the detection box B to be determined is as shown in Figure 7 the middle figure, and the area of this union part is denoted as K. The overlapping part of the circumscribed rectangle A and the detection box B to be determined is as shown in Figure 7 the right figure, and the area of this overlapping part is denoted as J. Thus, the Intersection over Union (IOU) between the circumscribed rectangle A and the detection box B to be determined can be expressed as: IOU = J / K, and GIOU = IOU - (C - K) / C, where C represents the area of the smallest rectangle C. Thus, by calculating the distance between each detection box to be determined and the circumscribed rectangle A (image position), the detection box to be determined with the shortest distance is used as the face detection box.
[0117] In the embodiments of the present application, still based on the image position of the face to be located in the (i - 1)-th image frame, the face detection box is determined from the face detection area according to the distances between multiple detection boxes to be determined and the image position respectively. That is to say, by performing a mapping operation on the circumscribed rectangle corresponding to the image position of the previous frame, the face detection box in the current frame is determined according to the distance, so as to reduce the calculation amount without losing the face.
[0118] Furthermore, after the face detection box is determined by the above-mentioned method of detecting box mapping, since the face detection box is obtained in a rough way, the corresponding area may be larger or smaller than the area where the face to be located is, resulting in inaccurate face detection box results, and may affect the face localization result due to the unstable state of the face detection box results.
[0119] To solve this problem, the embodiment of the present application adjusts the face detection frame by means of target area adjustment, so as to improve the accuracy of the face detection frame result. The specific method of target area adjustment is as follows: obtain the face size parameter of the face to be located in the (i-1)-th image frame, and based on the face size parameter, adjust the size of the face detection frame according to the center position of the face detection frame. Wherein, the target area corresponds to the area identified by the face detection frame.
[0120] As can be seen from the foregoing, in the video content to be recognized, for the same face to be located, the change between the i-th image frame and the (i-1)-th image frame is small, or rather, the size change of the face to be located between adjacent image frames is not significant. Therefore, the face size parameter of the face to be located in the (i-1)-th image frame can be used as a basis, and based on the face size parameter, the size of the face detection frame is adjusted according to the center position of the face detection frame. Wherein, the face size of the (i-1)-th image frame can be determined according to the face positioning result corresponding to the image position of the face to be located in the (i-1)-th image frame, and the present application does not make specific limitations on this.
[0121] See Figure 8 , which is a schematic diagram of adjusting the face detection frame provided by the embodiment of the present application. In Figure 8 the left figure, rectangle E is the circumscribed rectangle of the face to be located determined according to the face size of the (i-1)-th image frame, and rectangle D is the face detection frame of the i-th image frame. It can be seen that the area corresponding to the face detection frame is larger than the area where the face to be located is located. Determine the center of rectangle D, which is generally the center of the face to be located. Based on this center, reduce rectangle D to the same size as rectangle E to obtain the adjusted face detection frame. At this time, the length and width of rectangle D are equal to the length and width of matrix E, as Figure 8 shown in the right figure.
[0122] In the embodiment of the present application, the size of the face detection frame in the current frame is adjusted based on the actual size of the face in the previous frame, so that the area corresponding to the face detection frame is closer to the area where the face to be located is located, improving the accuracy of the face detection frame result. When performing subsequent face key point positioning, it is more focused on the face detection frame with higher accurate results, thereby reducing the impact of the unstable state of the face detection frame result on the face positioning result.
[0123] Furthermore, after the face detection frame is determined by the foregoing method of detection frame mapping, if the rotation angle of the face to be located is relatively large, such as rotating 180 degrees, the difference between the face to be located and the face detection frame is large, and a face positioning result of no face will be obtained, resulting in missed recognition and reducing the accuracy of the face positioning result.
[0124] To solve this problem, in the embodiments of this application, the rotation angle of the (i-1)th image frame is used as the rotation basis, and the inclination degree of the face detection box in the ith image frame in the ith image frame is adjusted by means of target area correction. The specific method of target area correction is as follows: obtain the second face angle of the face to be located in the (i-1)th image frame, and adjust the inclination degree of the face detection box in the ith image frame according to the second face angle.
[0125] Among them, the second face angle is used to identify the inclination degree of the face to be located in the (i-1)th image frame, or the rotation angle of the face to be located in the (i-1)th image frame. As can be seen from the foregoing, for the same face to be located, the change between the ith image frame and the (i-1)th image is small. Therefore, the second face angle can be used as the adjustment basis to adjust the inclination degree of the face detection box in the ith image frame.
[0126] See Figure 9 , which is a schematic diagram of target area correction provided by the embodiments of this application. Figure 9 In the left figure, the face to be located is inclined, while the face detection box is not inclined, resulting in a low accuracy of the subsequent face positioning result. Through target area correction, the inclination degree of the face detection box in the ith image frame is adjusted according to the second face angle, so that the inclination degree of the face detection box is basically the same as that of the face to be located. For the terminal device, when performing subsequent face positioning, it will consider that the face to be located has been corrected, thus avoiding missed recognition caused by a large inclination degree and improving the accuracy of the face positioning result.
[0127] For example, the second face angle (used to identify the inclination degree of the face to be located in the (i-1)th image frame) is 150 degrees, and the first face angle (used to identify the inclination degree of the face to be located in the ith image frame) is 170 degrees. If face positioning is directly performed, a face positioning result of no face will be obtained due to the large first face angle, resulting in missed recognition and reducing the accuracy of the face positioning result. If the inclination degree of the face detection box in the ith image is first adjusted based on the second face angle, the inclination degree of the adjusted face detection box in the ith image frame is 20 degrees. Compared with 170 degrees, 20 degrees is smaller, and there will be no missed recognition caused by a large inclination degree, improving the accuracy of the face positioning result.
[0128] In the embodiment of the present application, based on the second face angle of the face to be located in the (i - 1)-th image frame, the inclination degree of the face detection frame in the i-th image frame is adjusted, thereby reducing the inclination difference between the face detection frame and the face to be located. This not only avoids obtaining a face localization result where there is no face, improving the accuracy of the face localization result, but also, through the method of target area correction, enables the angle between the face to be located in the (i - 1)-th image frame and the face to be located in the i-th image frame to not exceed 30 degrees, and the face recognition system supports 360-degree face rotation following.
[0129] As a possible implementation manner, after reducing the inclination difference between the face detection frame and the face to be located based on the adjustment of the second face angle, for the terminal device, the face to be located may still be inclined (due to the difference in the inclination degrees of the (i - 1)-th image frame and the i-th image frame). At this time, the face to be located can be corrected twice by the aforementioned methods A1 and A2, or B1 and B2 to determine the face contour area and the face facial feature area, thereby improving the accuracy of subsequent face localization.
[0130] As a possible implementation manner, in the aforementioned first case, that is, when the face to be located is not located in the (i - 1)-th image frame, if the face to be located rotates in the i-th image frame and the rotation angle of the (i - 1)-th image frame cannot be used as the rotation basis, the third face angle of the face to be located in the face detection frame can be determined at this time, and the inclination degree of the face detection frame in the i-th image frame is adjusted according to the third face angle. The third face angle is used to identify the inclination degree of the face to be located in the i-th image frame, and the third face angle can be determined through the position information of the face detection frame, etc. The present application does not make specific limitations on this. Thus, by adjusting the inclination degree of the face detection frame in the i-th image frame through the third face angle, the correction quality of the target area can be further improved, thereby improving the accuracy of face key point localization.
[0131] It should be noted that adjusting the inclination degree of the face detection frame in the i-th image frame through the third face angle is not only applicable to the aforementioned first case, that is, the case where the face to be located is not located in the (i - 1)-th image frame, but also applicable to the aforementioned second case, that is, the case where the face to be located is located in the (i - 1)-th image frame, as a specific implementation manner of the target area correction method.
[0132] Further, in the case where the face to be located is not located in the previous first, i.e., the (i - 1)-th image frame, after obtaining the face location result of the i-th image frame, cache the face location result of the i-th image frame in the memory, so as to use the image position of the face to be located in the i-th image frame as the basis for the face detection frame in the (i + 1)-th image frame. Obtain the (i + 1)-th image frame in the consecutive N image frames included in the video content to be recognized. If the face to be located is not located for the (i + 1)-th image frame, i.e., the (i + 1)-th image frame is in a non-following state, the face location result of the i-th image frame will not be used subsequently, and clear the face location result of the i-th image frame from the memory. Thus, the purpose of saving memory is achieved, so that the face location method provided by the embodiments of the present application is more applicable to terminal devices such as smart phones.
[0133] As a possible implementation manner, in the previous second case, i.e., when the face to be located is located in the (i - 1)-th image frame, after obtaining the face location result of the i-th image frame, cache the face location result of the i-th image frame in the memory, and update the timing data, only save the face location results of the first m image frames of the i-th image frame in the consecutive N image frames that will be used subsequently, and do not save the face location results of other image frames. Thus, the purpose of saving memory is achieved, so that the face location method provided by the embodiments of the present application is more applicable to terminal devices such as smart phones.
[0134] As a possible implementation manner, before caching the face location result of the i-th image frame in the memory, it is also necessary to determine the result confidence of the face location result of the i-th image frame. If the result confidence is greater than or equal to a preset threshold, it can be cached. If the result confidence is less than the preset threshold, clear all the currently stored timing data. Thus, the purpose of saving memory is achieved, so that the face location method provided by the embodiments of the present application is more applicable to terminal devices such as smart phones.
[0135] It should be noted that the timing data may include one or more of the face location results of the first m image frames of the i-th image frame in the consecutive N image frames, the second face angle value of the i-th image frame, or the circumscribed rectangle corresponding to the face location result of the i-th image frame.
[0136] In a possible implementation manner, an embodiment of the present application provides an S204, that is, a specific implementation manner for generating a face location result for the face to be located in the i-th image frame based on contour key points and facial feature key points, specifically including S2041 - S2045.
[0137] S2041: Generate an initial face location result for the face to be located in the i-th image frame based on contour key points and facial feature key points.
[0138] Among them, the initial face localization result includes multiple face key points for identifying face features. For example, the initial face localization result is a collection of contour key points and facial feature key points.
[0139] Under the video stream sequence, the initial face localization result may visually jitter, that is, there is a problem of jitter in the point positions. Based on this, the embodiments of the present application smooth the initial face localization result by using the face localization results of the first m image frames, which will be specifically described below.
[0140] S2042: Obtain the face localization results of the first m image frames among the consecutive N image frames of the i-th image frame.
[0141] Obtain the face localization result of the (i - 1)-th image frame, the face localization result of the (i - 2)-th image frame... the face localization result of the (i - m)-th image frame. Those skilled in the art can set according to the actual situation, and the present application does not make specific limitations. For example, m = 3.
[0142] It can be understood that if all or part of the face localization results of the first m image frames among the consecutive N image frames of the i-th image frame are not obtained, the smoothing process can be performed based on the obtained face localization results, or no smoothing process is performed.
[0143] S2043: Extract the target key points for the same face feature from the initial face localization result and the face localization results of the first m image frames.
[0144] For example, for the mouth of the face to be localized, extract the target key points for the mouth from the initial face localization result. It should be noted that if there are multiple key points for the mouth, one of the multiple key points can be further located.
[0145] S2044: Based on the position differences between the m actual positions corresponding to the target key points in the face localization results of the first m image frames and the undetermined position of the target key point in the initial face localization result, smooth and correct the undetermined position to obtain the actual position of the target key point in the i-th image frame.
[0146] Through research, it is found that the jitter of the point positions is a jitter without any rules and belongs to Gaussian jitter. Therefore, the smoothing correction can be performed based on the position differences. The following takes m = 3 as an example for illustration.
[0147] If the target key point is P0, the actual position corresponding to the face localization result of the (i - 1)-th image frame is P1, the actual position corresponding to the face localization result of the (i - 2)-th image frame is P2, and the actual position corresponding to the face localization result of the (i - 3)-th image frame is P3.
[0148] Assuming that the distance function between two points is Dis, the distance between P0 and P1 can be expressed as dis_1=Dis(P0,P1), the distance between P0 and P2 can be expressed as dis_2=Dis(P0,P2), and the distance between P0 and P3 can be expressed as dis_3=Dis(P0,P3).
[0149] Assuming that the width of the bounding rectangle obtained based on the initial face positioning result in the i-th image is W, the distance-based adjustment weight of the i-1-th image frame is fl=exp(-dis_1 / (2*W)), the distance-based adjustment weight of the i-2-th image frame is f2=exp(-dis_2 / (2*W)), and the distance-based adjustment weight of the i-3-th image frame is f3=exp(-dis_3 / (2*W)). It should be noted that the distance-based adjustment weight of the i-th image frame is 1.0.
[0150] The determined position is smoothed and corrected, and the actual position of the target key point in the i-th image frame is obtained as P = (fl*P1+f2*P2+f3*P3+P0) / sum_f. Among them, sum_f is the embodiment of normalization, sum_f = fl+f2+f3+1.0.
[0151] S2045: Generate a face positioning result for the face to be positioned in the i-th image frame according to the actual positions of the multiple facial key points in the i-th image frame.
[0152] Each of the multiple facial key points is used as a target key point, and the actual position of each facial key point in the i-th image frame is obtained, thereby generating a face positioning result for the face to be positioned in the i-th image frame.
[0153] In the embodiment of the present application, the timing information of the first m image frames is used for smoothing to ensure the smooth output of the face positioning result, reduce jitter, and improve the user experience. If the embodiment of the present application is used simultaneously with the embodiment corresponding to the aforementioned target area adjustment method, while adaptively adjusting the relevant smoothing coefficient (adaptively adjusting the smoothing strength) according to the face size parameter, the lag caused by excessive smoothing and the jitter caused by unclear smoothing strength are also taken into account.
[0154] Next, we will combine Figure 10 and Figure 11 , a face positioning system in which the face positioning method provided in an embodiment of the present application is installed and applied in a terminal device is used as an example for explanation.
[0155] The face localization system includes a following state judgment module, a large face detection model module, a small face detection model module, a key point hierarchical localization module, a detection box mapping module, a target area adjustment module, a target area correction module, a key point temporal smoothing module, and a temporal information update module, which will be described separately below.
[0156] See Figure 10 , which is a schematic diagram of the application scenario of a face localization method provided by an embodiment of the present application.
[0157] S1001: The following state judgment module judges whether the (i - 1)-th image frame is in the face following state.
[0158] If the face to be localized is not located, the i-th image frame is in the non-following state, corresponding to the first case described above, that is, the face to be localized is not located in the (i - 1)-th image frame, and S1002 is executed.
[0159] If the face to be localized is located, the i-th image frame is in the face following state, corresponding to the second case described above, that is, the face to be localized is located in the (i - 1)-th image frame, and S1005 is executed.
[0160] Among them, the face detection area used in S1005 is a local map relative to the full map of the i-th image frame used in S1002. Taking the face detection area as the input for subsequent face localization makes the area to be recognized smaller when performing face detection subsequently, thereby reducing the calculation amount when the face is not lost in the video stream sequence. Moreover, using the image position of the face to be localized in the (i - 1)-th image frame as the basis for the face detection box in the i-th image frame can ensure that the face is always followed when the face moves rapidly, and can follow the face with high precision without loss.
[0161] S1002: Face detection - full map.
[0162] Obtain the full map of the i-th image frame, use it as the input for the subsequent large face detection model module, and determine the face detection box from the area range corresponding to the i-th image frame.
[0163] S1003: Key point hierarchical localization.
[0164] Call the key point hierarchical localization module. Through the methods described in S202 and S203 above, that is, through the face detection box in the i-th image frame, determine the face contour area and the face facial feature areas of the face to be localized in the face detection box, identify the contour key points according to the face contour area, and identify the facial feature key points according to the face facial feature areas.
[0165] Next, in combination with Figure 5 and Figure 11The method of hierarchical key point localization within the hierarchical key point localization module will be described. Refer to Figure 11 , which is a flowchart of a hierarchical key point localization provided by an embodiment of the present application.
[0166] S1101: Face box cropping.
[0167] Refer to Figure 5 , and crop the face image of the corresponding area according to the position of the face detection box in the i-th image frame to obtain a face box cropping.
[0168] S1102: Coarse localization.
[0169] Continue to refer to Figure 5 , and input the face box cropping into a coarse localization model such as the CoarseNet model for coarse localization to obtain initial face key points.
[0170] S1103: Obtain a contour point set and an inlier point set.
[0171] Continue to refer to Figure 5 , and divide the initial face key points into a contour point set for identifying the initial contour of the face to be localized, and an inlier point set for identifying the initial facial features of the face to be localized.
[0172] S1104: Obtain the first face angle.
[0173] S1105: Affine transformation.
[0174] Generate a first affine transformation matrix corresponding to the contour point set through the contour point set and the first face angle, and generate a second affine transformation matrix corresponding to the inlier point set through the inlier point set and the first face angle.
[0175] S1106: Obtain the face contour area and the face facial feature area.
[0176] Continue to refer to Figure 5 , extract the face contour area from the i-th image frame according to the first affine transformation matrix, and extract the face facial feature area from the i-th image frame according to the second affine transformation matrix.
[0177] S1107: Fine localization.
[0178] Continue to refer to Figure 5 , input the face contour area into a fine localization model such as the ContourNet model for fine localization to obtain contour key points, and input the face facial feature area into a fine localization model such as the InlineNet model for fine localization to obtain facial feature key points.
[0179] S1108: Point position fusion.
[0180] Fuse the contour key points and facial feature key points to generate a face localization result for the face to be localized in the i-th image frame.
[0181] As a possible implementation, the matrix regression method can be used to replace the deep neural network model, map the sparse point results to the dense point results, further simplify the computational complexity of face key point localization, and thus reduce the computational time consumption.
[0182] It can be understood that Figure 4 and Figure 5 similar to the process of hierarchical key point localization used, since Figure 4 the first face angle of the face to be localized in is zero, so it is not necessary to execute S1104, that is, obtain the first face angle, and in S1105, that is, the affine transformation, generate the first affine transformation matrix corresponding to the contour point set through the contour point set, and generate the second affine transformation matrix corresponding to the inlier set through the inlier set.
[0183] Thus, through the hierarchical key point localization process, the method of separately localizing the face contour and facial features realizes the decoupling of the face change amplitude, which is equivalent to splitting the complex problems in the related art into multiple simple problems for processing, simplifies the difficulty of face key point localization, and significantly reduces the overall computational complexity of face localization. Moreover, it ensures that the contour key points and facial feature key points do not interfere with each other during the localization process and achieve accurate localization.
[0184] After the description of S1003, that is, hierarchical key point localization, the description of S1004, that is, temporal data update, is continued.
[0185] S1004: Temporal data update.
[0186] Call the temporal information update module to update the temporal data. Among them, the temporal data can include one or more of the face localization results of the first m image frames in the continuous N image frames of the i-th image frame, the second face angle value of the i-th image frame, or the circumscribed rectangle corresponding to the face localization result of the i-th image frame.
[0187] By updating the temporal data in a timely manner, only the temporal data that will be used later is saved, and the temporal data that will not be used is deleted in a timely manner, so as to achieve the purpose of saving memory, so that the face localization method provided by the embodiments of the present application is more applicable to terminal devices such as smart phones.
[0188] S1005: Face detection - local map.
[0189] According to the image position of the face to be localized in the (i - 1)-th image frame, determine a face detection region including the image position from the i-th image frame, and use it as the input of the subsequent face detection small model module, and determine a face detection box in the face detection region.
[0190] As a possible implementation, the small face detection model module can be replaced with a real following model to further improve the performance and effect of face following.
[0191] S1006: Detection box mapping.
[0192] Call the detection box mapping module. Based on the image position of the face to be located in the (i - 1)-th image frame, determine the face detection box from the face detection area by the distances between multiple pending detection boxes and the image position respectively. Thus, the amount of calculation can be reduced without losing the face.
[0193] S1007: Target area adjustment.
[0194] Call the target area adjustment module to adjust the size of the face detection box in the current frame based on the actual size of the face in the previous frame, so that the area corresponding to the face detection box is closer to the area where the face to be located is, improving the accuracy of the face detection box result. When performing subsequent face key point localization, it focuses more on the face detection box with a higher accurate result, thereby reducing the impact of the unstable state of the face detection box result on the face localization result.
[0195] S1008: Target area correction.
[0196] Call the target area correction module to adjust the inclination degree of the face detection box in the i-th image frame based on the second face angle of the face to be located in the (i - 1)-th image frame. Thus, the inclination difference between the face detection box and the face to be located is reduced. Not only can it avoid obtaining a face localization result of a non-existent face and improve the accuracy of the face localization result, but also through the way of target area correction, the face angle corresponding to the image input to the subsequent key point hierarchical localization module does not exceed 30 degrees, and the face recognition system supports 360-degree face rotation following.
[0197] S1009: Key point hierarchical localization.
[0198] Call the key point hierarchical localization module. Through the methods described in the foregoing S202 and S203, that is, through the face detection box in the i-th image frame, determine the face contour area and the face facial feature areas of the face to be located in the face detection box, identify the contour key points according to the face contour area, and identify the facial feature key points according to the face facial feature areas.
[0199] Through the collaborative cooperation of the three modules of the detection box mapping module, the target area adjustment module, and the target area correction module, realize stable, efficient, and 360-degree face rotation-supported face video stream following.
[0200] S1010: Key point temporal smoothing.
[0201] Call the key point timing smoothing module, and use the timing information of the previous m image frames for smoothing processing to ensure the smooth output of the face positioning result, reduce jitter, and improve the user experience. If the embodiment of the present application is used simultaneously with the corresponding embodiment of the aforementioned target area adjustment method, while adaptively adjusting the relevant smoothing coefficient (adaptive adjustment of the smoothing strength) according to the face size parameter, the hysteresis caused by excessive smoothing and the jitter caused by insufficient smoothing strength are also considered.
[0202] After executing S1010, that is, key point timing smoothing, execute S1004, that is, timing data update, so as to obtain the face positioning result for the (i + 1)-th image frame.
[0203] Therefore, the embodiment of the present application proposes a face positioning system, which optimizes the overall call logic, simplifies the face positioning difficulty on the premise of ensuring accurate face key point positioning, improves the processing speed of the algorithm in the video stream state, and enables it to be applied to multiple projects and product applications including face recognition systems, face beauty projects, 3D face reconstruction systems, virtual human driving systems, etc. It can locate the positions of hundreds of accurate face key points, and can achieve ultra-real-time processing speed in terminal devices such as smart phones, improving the smooth experience of users.
[0204] In daily application scenarios, the face positioning method provided by the embodiment of the present application has high robustness characteristics, is applicable to different environments (indoor, outdoor, normal light, underexposed images, overexposed images, etc.), different genders and age groups (male, female, child, elderly, middle-aged, etc.), and solves the problem of real-time face positioning following in the video stream state.
[0205] For the face positioning method provided in the above embodiment, the embodiment of the present application also provides a face positioning device.
[0206] See Figure 12 This figure is a schematic diagram of a face positioning device provided by the embodiment of the present application. As Figure 12 shown, the face positioning device 1200 includes: an acquisition unit 1201, a determination unit 1202, an identification unit 1203, and a generation unit 1204;
[0207] The acquisition unit 1201 is configured to acquire the i-th image frame of the video content to be recognized, where the i-th image frame is one of the continuous N image frames included in the video content to be recognized;
[0208] The determination unit 1202 is configured to determine the face contour area and the facial feature area of the face to be positioned in the face detection box through the face detection box in the i-th image frame;
[0209] The recognition unit 1203 is configured to recognize contour key points based on the face contour region and recognize facial feature key points based on the facial feature region of the face; wherein, the contour key points are used to identify the face contour of the to-be-localized face, and the facial feature key points are used to identify the facial features of the to-be-localized face.
[0210] The generation unit 1204 is configured to generate a face localization result for the to-be-localized face in the i-th image frame based on the contour key points and the facial feature key points.
[0211] As a possible implementation manner, the determination unit 1202 includes a first determination subunit, a recognition subunit, and a second determination subunit;
[0212] The first determination subunit is configured to determine initial face key points corresponding to the to-be-localized face in the face detection box through the face detection box in the i-th image frame;
[0213] The recognition subunit is configured to divide the initial face key points into a contour point set for identifying the initial contour of the to-be-localized face and an inlier set for identifying the initial facial features of the to-be-localized face;
[0214] The second determination subunit is configured to determine the face contour region according to the contour point set and determine the facial feature region of the face according to the inlier set.
[0215] As a possible implementation manner, the device further includes a first face angle determination unit, configured to determine a first face angle of the to-be-localized face according to the inlier set, where the first face angle is used to identify the inclination degree of the to-be-localized face in the i-th image frame;
[0216] The second determination subunit is configured to:
[0217] Generate a first affine transformation matrix corresponding to the contour point set through the contour point set and the first face angle;
[0218] Extract the face contour region from the i-th image frame according to the first affine transformation matrix;
[0219] Generate a second affine transformation matrix corresponding to the inlier set through the inlier set and the first face angle;
[0220] Extract the facial feature region of the face from the i-th image frame according to the second affine transformation matrix.
[0221] As a possible implementation manner, the device further includes a face detection box determination unit, configured to:
[0222] In response to the face to be located being located in the (i - 1)th image frame, it is determined that the ith image frame is in the face following state, where the (i - 1)th image frame is the previous image frame of the ith image frame in the consecutive N image frames;
[0223] According to the image position of the face to be located in the (i - 1)th image frame, a face detection region including the image position is determined from the ith image frame, and the region range corresponding to the face detection region is smaller than the region range corresponding to the ith image frame;
[0224] The face detection box is determined in the face detection region.
[0225] As a possible implementation manner, the face detection box determination unit is configured to:
[0226] Based on the face presence possibility, a plurality of pending detection boxes are determined in the face detection region;
[0227] According to the distances between the plurality of pending detection boxes and the image position respectively, the face detection box is determined from the plurality of pending detection boxes.
[0228] As a possible implementation manner, the device further includes a target region adjustment unit, configured to:
[0229] Obtain the face size parameter of the face to be located in the (i - 1)th image frame;
[0230] According to the face size parameter, the size of the face detection box is adjusted based on the center position of the face detection box.
[0231] As a possible implementation manner, the device further includes a target region correction unit, configured to:
[0232] Obtain the second face angle of the face to be located in the (i - 1)th image frame, where the second face angle is used to identify the inclination degree of the face to be located in the (i - 1)th image frame;
[0233] According to the second face angle, the inclination degree of the face detection box in the ith image frame is adjusted.
[0234] As a possible implementation manner, the device further includes a face detection box determination unit, configured to:
[0235] In response to the face to be located not being located in the (i - 1)th image frame, it is determined that the ith image frame is in the non - following state, where the (i - 1)th image frame is the previous image frame of the ith image frame in the consecutive N image frames;
[0236] Determine the face detection box from the area range corresponding to the i-th image frame.
[0237] As a possible implementation, the device further includes a target area correction unit, configured to:
[0238] Determine the third face angle of the face to be located in the face detection box, where the third face angle is used to identify the inclination degree of the face to be located in the i-th image frame;
[0239] Adjust the inclination degree of the face detection box in the i-th image frame according to the third face angle.
[0240] As a possible implementation, the generating unit 1204 is configured to:
[0241] Generate an initial face positioning result for the face to be located in the i-th image frame based on the contour key points and the facial feature key points, where the initial face positioning result includes multiple face key points for identifying face features;
[0242] Obtain the face positioning results of the first m image frames in the consecutive N image frames of the i-th image frame;
[0243] Extract target key points for the same face feature from the initial face positioning result and the face positioning results of the first m image frames;
[0244] Based on the position differences between the m actual positions corresponding to the target key points in the face positioning results of the first m image frames and the pending positions of the target key points in the initial face positioning result, smooth and correct the pending positions to obtain the actual positions of the target key points in the i-th image frame;
[0245] Generate a face positioning result for the face to be located in the i-th image frame according to the actual positions of the multiple face key points in the i-th image frame.
[0246] As a possible implementation, the device further includes a timing data update unit, configured to:
[0247] Cache the face positioning result of the i-th image frame in the memory;
[0248] Obtain the (i + 1)-th image frame in the consecutive N image frames included in the video content to be recognized;
[0249] If the face to be located is not located for the (i + 1)-th image frame, clear the face positioning result of the i-th image frame from the memory.
[0250] As can be seen from the above technical solutions, when performing face localization on the video content to be recognized including N consecutive image frames, the real-time requirement is relatively high. Since during the movement of the face, the change ranges of the face contour and facial features generally have obvious differences. For example, the change range of the face contour is larger than that of the facial features. Therefore, in order to reduce the processing delay in the related technologies, for the i-th image frame of the video content to be recognized, the face contour region and the facial feature region of the face to be localized are determined through the face detection box. Since the change ranges of the face parts within the face contour region or the facial feature region are relatively more consistent, separately localizing the face contour and facial features can actually achieve a more accurate effect. Compared with the related technologies that perform face localization by treating each face part with different change ranges as a whole, the method of separately localizing the face contour and facial features decouples the change range of the face, which is equivalent to splitting the complex problem in the related technologies into multiple simple problems for processing. This not only significantly reduces the overall computational amount of face localization, but also can more accurately identify the contour key points indicating the face contour based on the face contour region, and identify the feature key points indicating the facial features based on the facial feature region, and obtain the face localization result in the i-th image frame according to the contour key points and the feature key points. The reduced computational amount effectively reduces the time delay of face localization, enabling the face localization technology to be extended to image processing scenarios with relatively high real-time performance.
[0251] The embodiment of the present application further provides a computer device, which is the computer device described above. The computer device can be a server or a terminal device. The face localization device described above can be built into the server or the terminal device. Next, the computer device provided by the embodiment of the present application will be introduced from the perspective of hardware implementation. Among them, Figure 13 The structure diagram of the server is shown. Figure 14 The structure diagram of the terminal device is shown.
[0252] See Figure 13 , Figure 13FIG. 0 is a schematic structural diagram of a server provided by an embodiment of the present application. The server 1400 may vary greatly due to different configurations or performances, and may include one or more central processing units (CPUs) 1422 and a memory 1432, and a storage medium 1430 (such as one or more mass storage devices) for storing one or more application programs 1442 or data 1444. Among them, the memory 1432 and the storage medium 1430 may be transient storage or persistent storage. The program stored in the storage medium 1430 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Further, the CPU 1422 may be configured to communicate with the storage medium 1430 and execute a series of instruction operations in the storage medium 1430 on the server 1400.
[0253] The server 1400 may further include one or more power supplies 1426, one or more wired or wireless network interfaces 1450, one or more input / output interfaces 1458, and / or one or more operating systems 1441, such as Windows Server TM , Mac OS X TM , Unix TM , Linux TM , FreeBSD TM and so on.
[0254] The steps performed by the server in the above embodiments may be based on the Figure 13 server structure shown.
[0255] Among them, the CPU 1422 is used to perform the following steps:
[0256] Obtain the i-th image frame of the video content to be recognized, where the i-th image frame is one of the continuous N image frames included in the video content to be recognized;
[0257] Determine the face contour area and the facial feature areas of the face to be located in the face detection box through the face detection box in the i-th image frame;
[0258] Identify the contour key points according to the face contour area, and identify the facial feature key points according to the facial feature areas; wherein, the contour key points are used to identify the face contour of the face to be located, and the facial feature key points are used to identify the facial features of the face to be located;
[0259] Generate a face localization result for the face to be located in the i-th image frame based on the contour key points and the facial feature key points.
[0260] Optionally, the CPU 1422 may also execute the method steps of any specific implementation manner of the face localization method in the embodiments of the present application.
[0261] See Figure 14 , Figure 14 which is a schematic structural diagram of a terminal device provided by an embodiment of the present application. Figure 14 Shown is a block diagram of a part of the structure of a smart phone related to the terminal device provided by an embodiment of the present application. The smart phone includes components such as a Radio Frequency (RF) circuit 1510, a memory 1520, an input unit 1530, a display unit 1540, a sensor 1550, an audio circuit 1560, a Wireless Fidelity (WiFi) module 1570, a processor 1580, and a power supply 1590. Those skilled in the art can understand that Figure 14 the structure of the smart phone shown in
[0262] does not limit the smart phone, and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Figure 14 The following specifically introduces each component of the smart phone:
[0263] The RF circuit 1510 can be used for receiving and sending signals during information reception or call processes. Specifically, after receiving the downlink information from the base station, it is given to the processor 1580 for processing; in addition, the data designed for uplink is sent to the base station.
[0264] The memory 1520 can be used for storing software programs and modules. The processor 1580 realizes various functional applications and data processing of the smart phone by running the software programs and modules stored in the memory 1520.
[0265] The input unit 1530 can be used for receiving input digital or character information, and generating key signal inputs related to the user settings and function controls of the smart phone. Specifically, the input unit 1530 may include a touch panel 1531 and other input devices 1532. The touch panel 1531, also known as a touch screen, can collect touch operations of the user on or near it, and drive the corresponding connection device according to a pre-set program. In addition to the touch panel 1531, the input unit 1530 may further include other input devices 1532. Specifically, the other input devices 1532 may include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, a joystick, etc.
[0266] The display unit 1540 can be used to display information input by the user or information provided to the user, as well as various menus of the smart phone. The display unit 1540 may include a display panel 1541. Optionally, the display panel 1541 may be configured in the form of, for example, a liquid crystal display (LCD), an organic light-emitting diode (OLED), etc.
[0267] The smart phone may also include at least one sensor 1550, such as a light sensor, a motion sensor, and other sensors. As for other sensors that the smart phone may also be configured with, such as a gyroscope, a barometer, a hygrometer, a thermometer, an infrared sensor, etc., they will not be elaborated here.
[0268] The audio circuit 1560, the speaker 1561, and the microphone 1562 can provide an audio interface between the user and the smart phone. The audio circuit 1560 can transmit the electrical signal converted from the received audio data to the speaker 1561, and the speaker 1561 converts it into a sound signal for output. On the other hand, the microphone 1562 converts the collected sound signal into an electrical signal, which is received by the audio circuit 1560 and then converted into audio data. After the audio data is output to the processor 1580 for processing, it is sent through the RF circuit 1510 to, for example, another smart phone, or the audio data is output to the memory 1520 for further processing.
[0269] The processor 1580 is the control center of the smart phone. It connects various parts of the entire smart phone using various interfaces and lines. By running or executing software programs and / or modules stored in the memory 1520, and by calling data stored in the memory 1520, it executes various functions of the smart phone and processes data. Optionally, the processor 1580 may include one or more processing units.
[0270] The smart phone also includes a power supply 1590 (such as a battery) for supplying power to each component. Preferably, the power supply can be logically connected to the processor 1580 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system.
[0271] Although not shown, the smart phone may also include a camera, a Bluetooth module, etc., which will not be elaborated here.
[0272] In the embodiment of the present application, the memory 1520 included in the smart phone can store program codes and transmit the program codes to the processor.
[0273] The processor 1580 included in the smart phone can execute the face positioning method provided in the above embodiment according to the instructions in the program code.
[0274] An embodiment of the present application also provides a computer-readable storage medium for storing a computer program, and the computer program is used to execute the face positioning method provided in the above embodiment.
[0275] An embodiment of the present application also provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the face positioning method provided in various optional implementation manners in the above aspects.
[0276] Those of ordinary skill in the art can understand that all or part of the steps for implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium, and when the program is executed, it executes the steps including the above method embodiments; and the foregoing storage medium can be at least one of the following media: read-only memory (abbreviation: ROM), RAM, magnetic disk, or optical disk, etc., various media that can store program codes.
[0277] It should be noted that the various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device and system embodiments, since they are basically similar to the method embodiments, they are described relatively simply, and the relevant parts can be referred to the partial description of the method embodiments. The device and system embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or they may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement without creative efforts.
[0278] As described above, it is only a specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Moreover, on the basis of the implementation manners provided in the above aspects of the present application, further combinations can be made to provide more implementation manners. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A face localization method, characterized in that, The method includes: Obtaining the i-th image frame of the video content to be recognized, where the i-th image frame is one of the consecutive N image frames included in the video content to be recognized; Determining, through the face detection box in the i-th image frame, the face contour region and the facial feature regions of the face to be located in the face detection box; Identifying contour key points based on the face contour region, and identifying facial feature key points based on the facial feature regions; wherein, the contour key points are used to identify the face contour of the face to be located, and the facial feature key points are used to identify the facial features of the face to be located; Generating a face localization result for the face to be located in the i-th image frame based on the contour key points and the facial feature key points; The method further includes: In response to the face to be located being located in the (i - 1)-th image frame, determining that the i-th image frame is in a face tracking state, where the (i - 1)-th image frame is the previous image frame of the i-th image frame in the consecutive N image frames; Determining, according to the image position of the face to be located in the (i - 1)-th image frame, a face detection region including the image position from the i-th image frame, where the region range corresponding to the face detection region is smaller than the region range corresponding to the i-th image frame; Determining the face detection box in the face detection region; Obtaining a second face angle of the face to be located in the (i - 1)-th image frame, where the second face angle is used to identify the inclination degree of the face to be located in the (i - 1)-th image frame; Adjusting the inclination degree of the face detection box in the i-th image frame according to the second face angle.
2. The method according to claim 1, wherein The determining, through the face detection box in the i-th image frame, the face contour region and the facial feature regions of the face to be located in the face detection box includes: Determining, through the face detection box in the i-th image frame, initial face key points corresponding to the face to be located in the face detection box; Dividing the initial face key points into a set of contour points for identifying the initial contour of the face to be located and a set of inner points for identifying the initial facial features of the face to be located; Determining the face contour region according to the set of contour points, and determining the facial feature regions according to the set of inner points.
3. The method according to claim 2, characterized in that, The method further includes: Determining a first face angle of the face to be located according to the set of inner points, where the first face angle is used to identify the inclination degree of the face to be located in the i-th image frame; The determining the face contour region according to the set of contour points includes: Generating a first affine transformation matrix corresponding to the set of contour points through the set of contour points and the first face angle; Extracting the face contour region from the i-th image frame according to the first affine transformation matrix; The determining the facial feature regions according to the set of inner points includes: Generating a second affine transformation matrix corresponding to the set of inner points through the set of inner points and the first face angle; Extracting the facial feature regions from the i-th image frame according to the second affine transformation matrix.
4. The method according to claim 1, characterized in that, Determining the face detection box in the face detection area includes: Based on the likelihood of face presence, determining a plurality of candidate detection boxes in the face detection area; Determining the face detection box from the plurality of candidate detection boxes according to the distances between the plurality of candidate detection boxes and the image position respectively.
5. The method according to claim 1, wherein The method further includes: Obtaining the face size parameter of the face to be located in the (i - 1)-th image frame; Adjusting the size of the face detection box based on the center position of the face detection box according to the face size parameter.
6. The method according to any one of claims 1-5, characterized in that, The method further includes: In response to the face to be located not being located in the (i - 1)-th image frame, determining that the i-th image frame is in a non-following state, where the (i - 1)-th image frame is the previous image frame of the i-th image frame in the consecutive N image frames; Determining the face detection box from the area range corresponding to the i-th image frame.
7. The method according to claim 6, characterized in that, The method further includes: Determining the third face angle of the face to be located in the face detection box, where the third face angle is used to identify the inclination degree of the face to be located in the i-th image frame; Adjusting the inclination degree of the face detection box in the i-th image frame according to the third face angle.
8. The method according to any one of claims 1-5, characterized in that, Generating the face positioning result for the face to be located in the i-th image frame based on the contour key points and the facial feature key points includes: Generating an initial face positioning result for the face to be located in the i-th image frame based on the contour key points and the facial feature key points, where the initial face positioning result includes a plurality of face key points for identifying face features; Obtaining the face positioning results of the first m image frames among the consecutive N image frames corresponding to the i-th image frame; Extracting target key points for the same face feature from the initial face positioning result and the face positioning results of the first m image frames; Based on the position differences between the m actual positions corresponding to the target key points in the face positioning results of the first m image frames respectively and the pending positions of the target key points in the initial face positioning result, smoothing and correcting the pending positions to obtain the actual positions of the target key points in the i-th image frame; Generating the face positioning result for the face to be located in the i-th image frame according to the actual positions of the plurality of face key points in the i-th image frame.
9. The method according to claim 6, characterized in that The method further includes: Caching the face positioning result of the i-th image frame in memory; Obtaining the (i + 1)-th image frame among the consecutive N image frames included in the video content to be recognized; If the face to be located is not located for the (i + 1)-th image frame, clearing the face positioning result of the i-th image frame from memory.
10. A face localization device, characterized in that, The device includes: an acquisition unit, a determination unit, an identification unit, and a generation unit; The acquisition unit is configured to acquire the i-th image frame of the video content to be recognized, where the i-th image frame is one of the consecutive N image frames included in the video content to be recognized; The determining unit is configured to determine a face contour region and a face facial feature region of the face to be located in the face detection box through the face detection box in the i-th image frame; The recognition unit is configured to recognize contour key points according to the face contour region, and recognize facial feature key points according to the face facial feature region; wherein, the contour key points are used to identify the face contour of the face to be located, and the facial feature key points are used to identify the facial features of the face to be located; The generating unit is configured to generate a face positioning result for the face to be located in the i-th image frame based on the contour key points and the facial feature key points; The apparatus further includes a face detection box determining unit, configured to: In response to the face to be located being located in the (i-1)-th image frame, determine that the i-th image frame is in a face following state, where the (i-1)-th image frame is the previous image frame of the i-th image frame in the consecutive N image frames; According to the image position of the face to be located in the (i-1)-th image frame, determine a face detection region including the image position from the i-th image frame, where the region range corresponding to the face detection region is smaller than the region range corresponding to the i-th image frame; Determine the face detection box in the face detection region; The apparatus further includes a target region correction unit, configured to: Obtain a second face angle of the face to be located in the (i-1)-th image frame, where the second face angle is used to identify the inclination degree of the face to be located in the (i-1)-th image frame; Adjust the inclination degree of the face detection box in the i-th image frame according to the second face angle.
11. The device according to claim 10, characterized in that, The determining unit includes a first determining subunit, a recognition subunit, and a second determining subunit; The first determining subunit is configured to determine initial face key points corresponding to the face to be located in the face detection box through the face detection box in the i-th image frame; The recognition subunit is configured to divide the initial face key points into a set of contour points for identifying the initial contour of the face to be located, and a set of inner points for identifying the initial facial features of the face to be located; The second determining subunit is configured to determine the face contour region according to the set of contour points, and determine the face facial feature region according to the set of inner points.
12. The device according to claim 11, wherein The apparatus further includes a first face angle determining unit, configured to determine a first face angle of the face to be located according to the set of inner points, where the first face angle is used to identify the inclination degree of the face to be located in the i-th image frame; The second determining subunit is configured to: Generate a first affine transformation matrix corresponding to the set of contour points through the set of contour points and the first face angle; Extract the face contour region from the i-th image frame according to the first affine transformation matrix; Generate a second affine transformation matrix corresponding to the set of inner points through the set of inner points and the first face angle; Extract the face facial feature region from the i-th image frame according to the second affine transformation matrix.
13. The device according to claim 11, wherein, The face detection box determining unit is further configured to: Determine a plurality of pending detection frames in the face detection region based on the possibility of face presence; Determine the face detection frame from the plurality of pending detection frames according to the distances between the plurality of pending detection frames and the image position respectively.
14. The device according to claim 11, wherein, The apparatus further includes a target area adjustment unit, configured to: Obtain the face size parameter of the face to be located in the (i-1)th image frame; Adjust the size of the face detection frame based on the center position of the face detection frame according to the face size parameter.
15. The device according to any one of claims 11-14, characterized in that, The face detection frame determination unit is further configured to: In response to the face to be located not being located in the (i-1)th image frame, determine that the ith image frame is in a non-following state, where the (i-1)th image frame is the previous image frame of the ith image frame in the consecutive N image frames; Determine the face detection frame from the area range corresponding to the ith image frame.
16. The device according to claim 15, characterized in that, The apparatus further includes a target area correction unit, configured to: Determine the third face angle of the face to be located in the face detection frame, where the third face angle is used to identify the inclination degree of the face to be located in the ith image frame; Adjust the inclination degree of the face detection frame in the ith image frame according to the third face angle.
17. The device according to any one of claims 11-14, characterized in that, The generating unit is configured to: Generate an initial face positioning result for the face to be located in the ith image frame based on the contour key points and the facial feature key points, where the initial face positioning result includes a plurality of face key points for identifying face features; Obtain the face positioning results of the first m image frames among the consecutive N image frames of the video content to be recognized corresponding to the ith image frame; Extract target key points for the same face feature from the initial face positioning result and the face positioning results of the first m image frames; Based on the position differences between the m actual positions corresponding to the target key points in the face positioning results of the first m image frames and the pending positions of the target key points in the initial face positioning result respectively, perform smooth correction on the pending positions to obtain the actual positions of the target key points in the ith image frame; Generate a face positioning result for the face to be located in the ith image frame according to the actual positions of the plurality of face key points in the ith image frame.
18. The device according to claim 15, characterized in that, The apparatus further includes a timing data update unit, configured to: Cache the face positioning result of the ith image frame in the memory; Obtain the (i + 1)th image frame among the consecutive N image frames included in the video content to be recognized; If the face to be located is not located for the (i + 1)th image frame, clear the face positioning result of the ith image frame from the memory.
19. A computer device, characterized in that, The computer device includes a processor and a memory: The memory is used to store program codes and transmit the program codes to the processor; The processor is used to execute the face positioning method according to any one of claims 1-9 according to the instructions in the program codes.
20. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store a computer program, and the computer program is used to execute the face localization method described in any one of claims 1-9.
21. A computer program product including instructions, which, when running on a computer, causes the computer to execute the face localization method described in any one of claims 1-9.
Citation Information
Patent Citations
High-precision face key point positioning method and system based on deep learning
CN111209873A