Methods, devices, equipment, storage media, and software products for screening facial images
By detecting facial key points, calculating and correcting three-dimensional pose angles, and mapping the key points of the target area onto a virtual projection plane, and combining state parameters for dual-condition filtering, the problem of insufficient accuracy and efficiency in facial image filtering in existing technologies is solved, and efficient and accurate facial image filtering is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- XINHUA NEWS AGENCY
- Filing Date
- 2026-01-16
- Publication Date
- 2026-05-26
AI Technical Summary
Existing technologies lack integrated solutions and cannot simultaneously integrate face detection, pose estimation, eye and mouth state analysis, and hardware adaptation capabilities. This results in shortcomings in face pose estimation methods, making it difficult to balance accuracy and efficiency. Furthermore, existing methods are not effective in multi-dimensional screening requirements.
By acquiring multiple face images, detecting key facial points, calculating the first three-dimensional pose angle and performing error compensation correction, a second three-dimensional pose angle is obtained. Based on this pose angle, the coordinates of key points of the target area are mapped to the virtual projection plane of the face facing the camera. Combined with the state parameters of the target area, dual-condition filtering is performed to ensure the accuracy and efficiency of the filtering results.
It significantly improves the accuracy and efficiency of face image screening, avoids the misselection of invalid images that are correct in posture but not in the condition of the parts or correct in the condition of the parts but have incorrect posture, ensures the consistency and reliability of parameter calculation, and adapts to the batch processing needs of various application scenarios.
Smart Images

Figure CN121545205B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of facial image screening technology, and in particular relates to a method, apparatus, device, storage medium and program product for screening facial images. Background Technology
[0002] Rapid and accurate image screening technology has a wide and urgent application demand in news media, the photography industry, social media platforms, and security monitoring. With the popularization of digital photography equipment, the amount of image data is growing explosively. How to automatically and efficiently select high-quality images with natural expressions and compliant postures from massive amounts of photos has become a key technical link in improving the efficiency of workflows in related industries.
[0003] Currently, existing automatic image filtering schemes can be mainly divided into three categories: image similarity-based filtering, image content-based filtering, and image quality evaluation-based filtering. Similarity-based methods (such as those using SSIM, PSNR, or deep features) are mainly used for image deduplication and archiving; content-based methods (relying on technologies such as object detection and face recognition) focus on identifying and classifying people and scenes in images; while image quality evaluation-based methods assess images from two dimensions: technical quality (such as blurriness and noise) and aesthetic quality (such as composition and lighting). In addition, in terms of specific technologies, there are face pose estimation methods based on traditional geometric feature points or complex deep learning models, as well as eye state evaluation methods based on object detection or eye aspect ratio (EAR) geometric calculation.
[0004] However, existing technologies lack integrated solutions, failing to simultaneously integrate face detection, pose estimation, eye and mouth state analysis, and hardware adaptation capabilities, thus failing to efficiently meet multi-dimensional screening needs. Furthermore, face pose estimation methods have shortcomings: either they rely heavily on key point detection accuracy, are greatly affected by expression and pose, and have weak generalization ability; or they require a large amount of training data or complex calculations, which are time-consuming and resource-intensive, making it difficult to balance accuracy and efficiency. Summary of the Invention
[0005] This application provides a method, apparatus, device, storage medium, and program product for screening facial images, which can improve the accuracy and efficiency of the screening results.
[0006] On one hand, embodiments of this application provide a method for filtering face images, the method comprising: acquiring multiple face images; for each face image, performing steps A to E respectively: Step A: detecting facial key points and target facial regions from the face image; Step B: calculating a first three-dimensional pose angle of the face image based on the facial key points; Step C: performing error compensation correction on the first three-dimensional pose angle based on a preset mapping relationship to obtain a second three-dimensional pose angle; Step D: based on the second three-dimensional pose angle, mapping the coordinates of key points of the target facial region to a virtual projection plane when the face is facing the camera through geometric transformation to obtain the corrected coordinates of key points of the target region; Step E: calculating the state parameters of the target region based on the corrected coordinates of key points of the target region; filtering face images from multiple face images that satisfy the first preset condition for the state parameters of the target region and the second preset condition for the second three-dimensional pose angle to obtain a target image.
[0007] In some possible implementations, detecting facial key points from a face image includes: acquiring an image sequence containing multiple facial expressions of the same person in a fixed head posture; performing facial key point detection on the image sequence to obtain the coordinate sequence of each facial key point; calculating the coordinate offset of each facial key point relative to a reference facial expression image in the image sequence; calculating the total offset norm of each key point based on the coordinate offset; selecting the top N key points according to the total offset norm from smallest to largest to obtain expression-independent key points, and using the expression-independent key points as facial key points.
[0008] In some possible implementations, after correcting the first three-dimensional pose angle based on a preset mapping relationship to obtain the second three-dimensional pose angle through error compensation, the method further includes: performing a linear transformation on the second three-dimensional pose angle based on personalized parameters of the face to obtain a third three-dimensional pose angle; and mapping the key point coordinates of the target facial region to a virtual projection plane when the face is facing the camera through geometric transformation based on the second three-dimensional pose angle to obtain the corrected key point coordinates of the target facial region, including: mapping the key point coordinates of the target facial region to a virtual projection plane when the face is facing the camera through geometric transformation based on the third three-dimensional pose angle to obtain the corrected key point coordinates of the target facial region.
[0009] In some possible implementations, before performing a linear transformation on the second three-dimensional pose angle based on personalized parameters of the face to obtain the third three-dimensional pose angle, the method includes: acquiring multiple first images of the face corresponding to the face image; wherein the first image includes a manually annotated fourth three-dimensional pose angle; performing error compensation correction based on a preset mapping relationship on the first three-dimensional pose angle of the first image to obtain the second three-dimensional pose angle of the first image; and performing linear regression fitting on the second three-dimensional pose angle and the fourth three-dimensional pose angle of the first image to obtain personalized parameters of the face.
[0010] In some possible implementations, before performing a linear transformation on the second three-dimensional pose angle based on the personalized parameters of the face to obtain the third three-dimensional pose angle, the method further includes: obtaining the feature vector of the face corresponding to the face image; calculating the similarity between the feature vector and a preset face feature vector; and, if the similarity satisfies a third preset condition, performing a linear transformation on the second three-dimensional pose angle based on the personalized parameters of the face corresponding to the face image to obtain the third three-dimensional pose angle.
[0011] In some possible implementations, the method further includes: acquiring a face sample set; the sample set includes multiple face samples, each face sample containing a second image of a face and a corresponding true three-dimensional pose angle; based on the face sample set, calculating a first three-dimensional pose angle of the second image of each face sample to obtain a data pair composed of the first three-dimensional pose angle and a standard three-dimensional pose angle; based on the data pair, establishing a mapping function from the first three-dimensional pose angle to the standard three-dimensional pose angle using a curve fitting method, the mapping function serving as a preset mapping relationship.
[0012] In some possible implementations, the first preset condition is that the target part state parameter is greater than the first preset level, and the second preset condition is that the second three-dimensional pose angle is greater than the second preset level. The process involves selecting face images from multiple face images that satisfy both the first preset condition and the second preset condition, to obtain a target image. This includes: determining the target part state level based on a preset target part state parameter grading interval; determining the second three-dimensional pose level based on a preset second three-dimensional pose angle grading interval; and selecting face images from multiple face images that satisfy both the first preset level and the second preset level, to obtain the target image.
[0013] In some possible implementations, before selecting a face image from multiple face images that satisfies the first preset condition for the target part state parameters and the second three-dimensional pose angle satisfies the second preset condition to obtain the target image, the method further includes: calculating the numerical distribution of the target part state parameters and the second three-dimensional pose angle based on multiple third images containing the face; segmenting the numerical distribution of the target part state parameters to obtain a preset target part state parameter grading interval; and segmenting the numerical distribution of the second three-dimensional pose angle to obtain a preset target part state parameter grading interval.
[0014] In some possible implementations, the multiple third images are faces corresponding to preset facial feature vectors. Based on the multiple third images containing faces, the numerical distribution of target part state parameters and second three-dimensional pose angles is calculated, including: based on the multiple third images containing faces corresponding to preset facial feature vectors, the numerical distribution of target part state parameters and second three-dimensional pose angles is calculated.
[0015] In some possible implementations, the target region includes the eyes. Before calculating the target region state parameters based on the corrected target region key point coordinates, the method further includes: if the person in the face image is wearing glasses, classifying the eye region in the face image using a pre-trained eye state classification model to obtain eye state information; wherein the eye state information includes whether the eyes are open or closed; if the eye state information is open, calculating the target region state parameters based on the corrected target region key point coordinates.
[0016] In some possible implementations, the target region includes the mouth, and the target region state parameters include mouth shape parameters. Based on the corrected coordinates of key points of the target region, the target region state parameters are calculated, including: obtaining the coordinates of at least one pair of corner key points located on both sides of the midline of the face and symmetrical to each other; calculating the mouth width based on the coordinates of the pair of corner key points; obtaining the coordinates of at least one pair of lip key points located on the midline of the upper and lower lips; calculating the degree of mouth opening based on the coordinates of the pair of lip key points; and calculating the ratio of the degree of mouth opening to the mouth width as the mouth shape parameter.
[0017] On the other hand, embodiments of this application provide a face image filtering device, the device comprising: an image acquisition module for acquiring multiple face images; an image processing module for performing processing on each face image separately, including: a key point detection unit for detecting facial key points and facial target parts from the face images; a pose angle calculation unit for calculating a first three-dimensional pose angle of the face image based on the facial key points; an error compensation unit for performing error compensation correction on the first three-dimensional pose angle based on a preset mapping relationship to obtain a second three-dimensional pose angle; a geometric transformation unit for mapping the coordinates of key points of the facial target parts to a virtual projection plane when the face is facing the camera through geometric transformation based on the second three-dimensional pose angle to obtain corrected coordinates of key points of the target parts; a state parameter calculation unit for calculating state parameters of the target parts based on the corrected coordinates of key points of the target parts; and an image filtering module for filtering face images from multiple face images that satisfy the first preset condition for the state parameters of the target parts and the second preset condition for the second three-dimensional pose angle to obtain a target image.
[0018] In another aspect, embodiments of this application provide an apparatus, which includes: a processor and a memory storing computer program instructions; the processor executes the computer program instructions to implement a method for screening facial images.
[0019] In another aspect, embodiments of this application provide a computer storage medium on which computer program instructions are stored, and when the computer program instructions are executed by a processor, a method for screening face images is implemented.
[0020] In another aspect, embodiments of this application provide a computer program product in which instructions, when executed by the processor of an electronic device, cause the electronic device to perform a method for screening facial images.
[0021] The face image filtering method, apparatus, device, storage medium, and program product of this application embodiment, by combining a second three-dimensional pose angle and target part state parameters, effectively avoids the misselection of invalid images that are qualified in pose but not in part state or qualified in part state but have pose deviations through dual condition constraints, significantly improving the accuracy of the filtering results. Error compensation is performed on the first three-dimensional pose angle through a preset mapping relationship to obtain a precise second three-dimensional pose angle; then, based on this pose angle, the coordinates of key points of the target part are mapped to the virtual projection plane of the face facing the camera, eliminating perspective distortion interference caused by pose offset, ensuring that the target part state parameters under different poses have a unified calculation benchmark, avoiding parameter misjudgment due to pose differences, and improving the consistency and reliability of parameter calculation. Batch processing of multiple face images eliminates the need for manual parameter adjustment for individual images, and can quickly adapt to various application scenarios, meeting the efficiency requirements of batch image processing while ensuring filtering accuracy. Attached Figure Description
[0022] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0023] Figure 1 This is a flowchart illustrating a face image filtering method provided in one embodiment of this application;
[0024] Figure 2 This is a flowchart illustrating a face image filtering method provided in another embodiment of this application;
[0025] Figure 3 This is a flowchart illustrating a face image filtering method provided in another embodiment of this application;
[0026] Figure 4This is a flowchart illustrating a face image filtering method provided in another embodiment of this application;
[0027] Figure 5 This is a flowchart illustrating a face image filtering method provided in another embodiment of this application;
[0028] Figure 6 This is a flowchart illustrating a face image filtering method provided in another embodiment of this application;
[0029] Figure 7 This is a flowchart illustrating a face image filtering method provided in another embodiment of this application;
[0030] Figure 8 This is a schematic diagram of the structure of a face image screening device provided in another embodiment of this application;
[0031] Figure 9 This is a schematic diagram of the structure of a face image screening device provided in another embodiment of this application;
[0032] Figure 10 This is a schematic diagram of facial key points in a face image provided in another embodiment of this application. Detailed Implementation
[0033] The features and exemplary embodiments of various aspects of this application will be described in detail below. To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain this application and not to limit it. For those skilled in the art, this application can be implemented without some of these specific details. The following description of the embodiments is merely to provide a better understanding of this application by illustrating examples.
[0034] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes the element.
[0035] In the field of portrait photo screening, existing technologies (such as face detection, pose estimation, and eye and mouth state analysis) exist primarily as independent functions, lacking a collaborative and compatible overall architecture. The development of different technology modules is often led by different teams or fields. For example, face detection technology focuses on target localization accuracy, pose estimation technology focuses on angle calculation accuracy, and eye and mouth state analysis focuses on fine-grained feature extraction. The lack of unified technical standards, data interfaces, and computational logic for each module makes it difficult to directly integrate them into a coherent screening process.
[0036] The shortcomings of existing face pose estimation methods mainly stem from the challenge of balancing accuracy, efficiency, and generalization ability. Methods relying on traditional geometric feature points calculate pose angles using a small number of key points, reducing computational complexity. However, the selection of these key points doesn't adequately consider the impact of facial expressions and pose changes on point stability, and they overly depend on the output accuracy of the key point detection model. If the detection results deviate (e.g., facial occlusion or exaggerated expressions causing point shifts), the pose estimation results will be significantly distorted, naturally limiting generalization ability. While methods based on CNN neural networks or 3DMM improve accuracy and robustness through deep learning or 3D modeling, CNN models require large-scale labeled datasets for training to cover diverse scenarios, and 3DMM methods require complex 3D reconstruction calculations. Both consume significant computational resources and increase time overhead, failing to meet the real-time requirements of scenarios like news filtering. Ultimately, existing methods struggle to find a suitable balance between accuracy, efficiency, and generalization ability for practical applications.
[0037] This embodiment combines a second three-dimensional pose angle with target part state parameters, using dual-condition constraints to filter images. This effectively avoids the misselection of invalid images where the pose is acceptable but the part state is unacceptable, or vice versa, significantly improving the practicality and accuracy of the filtering results. Error compensation is applied to the first three-dimensional pose angle using a preset mapping relationship to obtain a precise second three-dimensional pose angle. Then, based on this pose angle, the coordinates of key points of the target part are mapped to a virtual projection plane of the face facing the camera, eliminating perspective distortion interference caused by pose shifts. This ensures that the target part state parameters have a unified calculation benchmark under different poses, avoiding parameter misjudgment due to pose differences and improving the consistency and reliability of parameter calculation. Batch processing of multiple face images eliminates the need for manual parameter adjustment for individual images, allowing for rapid adaptation to various application scenarios. While ensuring filtering accuracy, it also meets the efficiency requirements of batch image processing.
[0038] To address the problems of the prior art, embodiments of this application provide a method, apparatus, device, computer storage medium, and computer program product for filtering facial images. The method for filtering facial images provided in this application embodiment will be described first below.
[0039] Figure 1 A flowchart illustrating a face image filtering method according to an embodiment of this application is shown. Figure 1 As shown, the method includes steps S110-S130.
[0040] S110, acquire multiple face images.
[0041] As an example, acquiring multiple facial images can be achieved by collecting a set of images containing facial information to be filtered through a preset data acquisition or reading method. This step provides basic data support for subsequent feature analysis of individual images and global filtering, and is a prerequisite for starting the entire filtering process.
[0042] As an example, a face image can refer to image data containing at least one face region. This image data can be an original image with a face as the main subject, acquired directly by an image acquisition device (such as a camera or webcam); or it can be a sub-image of a face region extracted or identified from more complex photos of people or scenes using face detection technology.
[0043] Specifically, in practical applications, the sources of the facial images include, but are not limited to:
[0044] Automated face detection is performed on photos containing human subjects, and the face regions are located and cropped out.
[0045] Face regions are extracted in real time from video streams or image sequences using face tracking and localization technology.
[0046] Standard close-up images of faces from an image library, pre-cropped with facial regions.
[0047] Specifically, the method for acquiring multiple facial images can be flexibly selected according to the actual application scenario. For example: batch reading stored facial image files from local storage devices (such as computer hard drives, mobile storage media) (supporting common image formats such as JPG, PNG, TIFF, etc.); capturing and continuously acquiring image sequences containing faces in real time through image acquisition devices (such as cameras, webcams); batch acquiring facial image data from cloud storage systems, databases, or image management platforms through interface calls or data export functions; and acquiring a set of facial images that meet the filtering requirements from third-party image data sources (such as authorized image libraries, compliant public image resources).
[0048] As an example, the acquired multiple face images can meet the basic data validity requirements, that is, the images need to be clear and distinguishable, so as to ensure that the algorithms for facial key point detection and target part recognition in subsequent steps can extract features normally. If the application scenario has specific requirements for image quality (such as high-resolution images for news reporting scenarios and unobstructed face images for security scenarios), invalid images can be initially filtered out in this step by using preset quality screening rules (such as resolution threshold and blur detection) to ensure that the input image set meets the basic conditions for subsequent processing.
[0049] S120, for each face image, execute S121-S125 respectively.
[0050] As an example, step S120 can adopt the logic of batch input and single-image processing, performing independent feature extraction, calculation and correction operations on each of the multiple face images obtained in step S110.
[0051] As an example, different facial images may exhibit variations in facial pose (e.g., looking up, looking down, turning left or right), target feature state (e.g., eye opening / closing, mouth shape), and shooting environment (e.g., lighting, angle). Using a "batch uniform calculation" approach could easily lead to a decrease in feature extraction accuracy (e.g., some images cannot be accurately identified by a uniform algorithm due to their unique poses). Therefore, this step employs a step-by-step processing design, adapting corresponding algorithm parameters and processing logic to the individual differences of each image.
[0052] As an example, automated processing can be achieved by relying on computer hardware and software systems:
[0053] Specifically, as an example, at the software level, the set of face images obtained in step S110 can be distributed to different processing threads through multi-threading or batch processing scheduling algorithms, and steps S121-S125 can be executed in parallel for multiple images to improve overall processing efficiency. If the amount of image data is large, a task queue mechanism can be used to schedule images into the S121-S125 processing flow one by one in sequence to avoid system resource overload.
[0054] As an example, at the hardware level, CPU or GPU hardware resources can be adapted according to processing needs. For example, for computationally intensive steps such as facial landmark detection and 3D pose angle calculation, GPU resources can be called to accelerate the processing speed of a single image, ensuring that an efficient processing rhythm can still be maintained in batch processing scenarios (such as screening tens of thousands of facial images).
[0055] In this step, as an example, after each face image is processed by steps S121-S125, a set of standardized quantized feature data (such as the second three-dimensional pose angle and the corrected target part state parameters) can be output. This data can be temporarily stored in the system cache or database, and used as the basis for judgment in the subsequent global screening step S130 after all images have been processed by steps S121-S125.
[0056] S121, detect facial landmarks and target facial regions from a face image.
[0057] As an example, facial key points refer to characteristic points on the face surface that have stable geometric positions and can represent the facial contour and facial structure. The spatial distribution of these points can fully reflect the geometric shape and posture information of the face. For example, they include stable areas with less influence from facial expressions, such as the center of the forehead, the area under the nose, and the cheekbones on both sides, as well as functional points closely related to the state of the target area, such as the upper and lower eyelids, the corners of the eyes, the corners of the mouth, and the lips. They usually need to cover the core areas such as the facial contour, eyes, nose, and mouth to form a complete representation of the facial geometry.
[0058] Specifically, as an example, a mature and high-precision face landmark detection model (such as the deep learning-based FaceMesh framework) can be used. This type of model can automatically identify face regions in an image through a pre-trained neural network and output the three-dimensional coordinates (including x-axis, y-axis planar coordinates and z-axis depth coordinates) of a preset number (e.g., 468) of key points. The coordinate data must accurately correspond to the actual pixel positions of the key points in the image to ensure the accuracy of subsequent geometric calculations.
[0059] As an example, the detection process can meet the dual requirements of location accuracy and stability. On the one hand, the deviation between the key point coordinates and the actual facial feature positions must be controlled within a preset threshold (e.g., a pixel-level deviation of no more than 2 pixels) to avoid distortion in subsequent pose calculations or state assessments due to point offsets. On the other hand, it must have a certain degree of anti-interference capability, and be able to stably output key point coordinates even in scenarios such as slight changes in lighting or minor facial occlusions (e.g., slight hair occlusion) in the image, ensuring the robustness of the detection results.
[0060] As an example, the detected facial landmarks can be stored as a set of coordinates, with each landmark corresponding to a unique number and three-dimensional coordinate data (e.g., the landmark numbered 10 corresponds to the coordinates...). The coordinates of key point number 33 This facilitates subsequent steps by calling specific key points for calculation according to their numbers.
[0061] As an example, the target facial region can refer to a specific facial area that is directly related to the facial image screening criteria, and whose state (such as the degree of opening and closing, and shape) determines whether the image meets the screening standards. It mainly covers the two major areas of the eyes and mouth. Among them, the state of the eyes (open / closed, degree of closure) and the state of the mouth (degree of closure, mouth shape) are the key judgment criteria for screening portrait photos with expressions and postures that meet the requirements, and therefore become the target regions that this application focuses on detecting. If the screening scenario is expanded in the future, the nose, facial contours, and other regions can also be included in the scope of target regions as needed.
[0062] As an example, based on the facial key points detected above, target areas can be located by using the regions enclosed by these key points. For instance, the eye area can be determined by a polygonal region formed by key points around the eyes; the mouth area can be determined by a polygonal region formed by key points at the corners of the mouth and the center of the lips, ensuring accurate boundaries of the located target areas and complete coverage of the functional areas of the eyes and mouth.
[0063] As an example, the detection process must ensure the integrity of the target area to avoid missing areas due to positioning errors (e.g., the eye area not fully including the upper and lower eyelids, or the mouth area not including the corners of the mouth). Simultaneously, interference from non-target areas must be eliminated (e.g., the eye area excluding eyebrows, cheeks, and other irrelevant areas) to ensure that subsequent state analysis is performed only on the target area, reducing computational redundancy.
[0064] As an example, the detected facial target regions can be stored in the form of region coordinate ranges (such as the minimum x coordinate, maximum x coordinate, minimum y coordinate, and maximum y coordinate of the eye region), or the image sub-regions corresponding to the target regions can be directly output (i.e., local images of the eyes and mouth cropped from the original face image), providing independent analysis objects for subsequent calculation of the state parameters of the target regions.
[0065] As one implementation of S121, S121 may also include the following steps:
[0066] Collect image sequences containing multiple facial expressions of the same person in a fixed head posture;
[0067] Perform facial landmark detection on the image sequence to obtain the coordinate sequence of each facial landmark;
[0068] For each facial key point, calculate its coordinate offset relative to the reference expression image in the image sequence;
[0069] Calculate the total offset norm for each key point based on the coordinate offset;
[0070] Based on the total offset norm in ascending order, select the top N key points to obtain expression-independent key points, and use these expression-independent key points as facial key points.
[0071] As an example, an image sequence can be directed at a single target person, guiding them to make diverse facial expressions while keeping their head posture fixed, and continuously capturing facial images during the process to form an image set containing the same posture and multiple expressions.
[0072] Specifically, as an example, fixing the head posture can be achieved through physical fixation devices (such as head supports or positioning pillows) or visual positioning technology, ensuring that the person's head does not undergo any changes in posture such as pitch, lateral rotation, or tilt during the acquisition process, and that facial expressions are only altered by facial muscle activity. Based on this, the interference of posture changes on keypoint coordinates can be eliminated, retaining only the keypoint displacement caused by facial expressions. As another example, this can also cover common and extreme daily expressions, such as relaxed expression, open-mouthed laughter, pursed-lip smile, closed eyes, frowning, wide-eyed staring, pouting, and sticking out the tongue, ensuring that the acquired expression scenarios are sufficiently comprehensive to verify the stability of keypoints under different expressions. Furthermore, the image sequence must maintain uniform shooting parameters (such as resolution, focal length, light intensity, and shooting distance) to avoid keypoint detection errors due to differences in shooting conditions; the number of images must meet statistical validity requirements to ensure the reliability of subsequent offset calculations.
[0073] As an example, the coordinate sequence can be generated by using mature facial landmark detection algorithms (such as the FaceMesh model based on deep learning, the Dlib landmark detection framework, etc.) to process each frame of the above-mentioned image, identify and output the three-dimensional coordinates of all preset facial landmarks in the image, and integrate them in the order of the image frames to form the coordinate time sequence corresponding to each landmark.
[0074] It should be noted that, as an example, the coordinate sequence of all key points must correspond one-to-one with the image frame to ensure that the coordinate changes of each key point in different expression frames can be traced subsequently, providing data support for offset calculation.
[0075] As an example, the coordinate offset can be calculated by selecting a reference facial expression image from the acquired image sequence, using the coordinates of each key point in that frame as a reference, and calculating the difference between the coordinates of the same key point in other facial expression frames and the reference coordinates. This difference is the coordinate offset of the key point in the corresponding facial expression frame.
[0076] For example, let the coordinates of a key point P in the reference facial expression image be... The coordinates of the keypoint P in another expression frame in the image sequence are: (t represents the frame number of the non-reference frame), then the three-dimensional coordinate offset of the key point in that expression frame. It can be represented as: By performing the above calculation on all non-reference frames in the image sequence, multiple sets of coordinate offsets for each keypoint under different facial expressions can be obtained, thereby quantifying the degree of displacement of the keypoint affected by facial expressions.
[0077] As an example, the total offset norm can be used to comprehensively calculate the coordinate offset of each keypoint across all expression frames, quantifying the total displacement of the keypoint throughout the entire expression change process with a single value. The smaller the total offset norm, the less the keypoint is affected by the expression and the stronger its positional stability.
[0078] Specifically, the magnitude of the offset in a single frame can be calculated using the Euclidean norm, and then the offset norms of all expression frames can be summed to obtain the total offset norm S. Taking a keypoint P as an example, the formula for calculating its total offset norm S is as follows:
[0079] (1)
[0080] Where n is the number of non-reference frames in the image sequence. Represents the offset of frame t. The Euclidean norm.
[0081] By calculating the total offset norm, the multi-frame offset of each key point can be transformed into a single value that can be directly compared, providing a clear quantitative basis for subsequent key point selection.
[0082] As an example, the total offset norm of all facial key points is sorted, and the top N key points with the smallest total offset norm are selected (N is the preset number of key points, which can be determined according to the needs of subsequent three-dimensional pose angle calculation and target part analysis) as expression-irrelevant key points for subsequent processing in this embodiment of the application.
[0083] Specifically, as an example, keypoints can be sorted by their total offset norm from smallest to largest. Keypoints with smaller total offset norms are ranked higher, indicating more stable positions during facial expression changes and a lower likelihood of being affected by facial expressions. The value of N can balance computational accuracy and processing efficiency. If N is too small, it may lead to insufficient feature dimensions in subsequent attitude angle calculations, affecting accuracy; if N is too large, it may introduce some less stable keypoints, increasing computational redundancy. In practical applications, the value of N (e.g., N=4) can be determined based on the specific algorithm requirements (e.g., 4 keypoints are needed to calculate pitch, horizontal rotation, and tilt angles in 3D attitude angles).
[0084] As an example, the first N key points obtained through screening can be directly used in subsequent steps such as facial 3D pose angle calculation and target area key point correction. Their stability can minimize the interference of facial expression changes on subsequent feature calculations and ensure the accuracy of the entire screening process.
[0085] As an example, such as Figure 10 As shown, according to the above method, the selected facial key points are expression-independent key points. Expression-independent key points can correspond to stable areas in the human facial muscle anatomy, specifically including at least one of the following: the central tendon region of the forehead, the central region where the nasal columella connects to the skin of the upper lip, and the bilateral zygomatic arch regions.
[0086] The face image selection method in this application collects image sequences of various expressions under a fixed head posture, calculates the total offset norm of key points under different expressions, and selects the top N key points with the smallest offset. These key points are not affected by facial muscle activity, ensuring that the first three-dimensional pose angle calculated based on these key points will not generate additional errors due to changes in facial expression, thus improving the stability of pose angle calculation. Since the selected key points are expression-independent, subsequent pose angle calculation does not require separate adjustment of the model or parameters for different expressions. It can directly adapt to various common expression scenarios such as natural facial expressions, smiles, and slight frowns, without the need to collect a large amount of additional expression annotation data, reducing the cost of model training and application, while improving the generalization and adaptation capability of the solution to face images with different expressions.
[0087] S122, calculate the first three-dimensional pose angle of the face image based on facial key points.
[0088] As an example, the first three-dimensional pose angle can be three independent quantified angles used to fully describe the pose state of a face in three-dimensional space, respectively corresponding to the rotation state of the face around the three coordinate axes of the spatial rectangular coordinate system (constructed with the image plane as the reference, usually with the center of the face or a certain reference key point as the origin), specifically including: pitch angle, horizontal rotation angle and tilt angle.
[0089] As an example, the pitch angle can characterize the vertical posture change of a face, that is, the degree to which the head is tilted up or down. When the face is tilted up, the pitch angle value increases; when the face is tilted down, the pitch angle value decreases, thus quantifying the vertical posture shift of the face.
[0090] As an example, the horizontal rotation angle can characterize the change in facial posture in the horizontal direction, that is, the degree to which the head turns to the left or right. When the face turns to the left, the horizontal rotation angle increases; when it turns to the right, the angle decreases, which is used to quantify the facial posture shift in the left and right directions.
[0091] As an example, the tilt angle can characterize the change in facial posture in the tilt direction, that is, the degree to which the head tilts to the left or right, such as tilting the head to the left or right. When the face tilts to the left, the tilt angle value increases; when it tilts to the right, the value decreases, thus quantifying the facial posture shift in the tilt direction.
[0092] The three angles mentioned above can together constitute the first three-dimensional pose angle, which can completely and quantitatively describe the pose state of the face in three-dimensional space, avoiding the ambiguity of traditional qualitative descriptions (such as obvious head tilting and slight head turning), and providing a standardized quantitative basis for subsequent pose correction and screening.
[0093] As an example, before calculation, key points that are expression-independent and spatially evenly distributed can be selected from the detected facial key points as the calculation benchmark. The selected key points should meet two main conditions: first, they should have strong positional stability during expression changes (e.g., key points in the center of the forehead, the area below the nose, and the bilateral zygomatic arches) to avoid calculation errors caused by expression interference; second, the key points should be distributed in different areas of the face in three-dimensional space (e.g., vertically covering the forehead to the chin, and horizontally covering the left to right cheeks) to ensure that the overall facial posture can be accurately inferred from the spatial relationship between the key points. Typically, selecting 4-6 key points that meet the above conditions is sufficient to meet the calculation requirements for the three-dimensional pose angles.
[0094] Four reference key points were selected (denoted as ). Taking the following as an example (corresponding to the center of the forehead, below the nose, left cheekbone arch, and right cheekbone arch respectively), the calculation process is as follows:
[0095] As an example, we extract the 3D coordinates of key points. We use the 3D coordinate data of the four reference key points obtained earlier, denoted as follows: , where x and y represent the two-dimensional coordinates of the key point on the image plane, and z represents the depth coordinate of the key point (representing the distance between the key point and the camera lens).
[0096] As an example, pitch angle calculation can be based on key points distributed longitudinally (such as...). and The pitch angle is calculated by taking the difference between the z-axis (depth) and y-axis (vertical plane) coordinates and using the arctangent function. A specific formula example is as follows:
[0097] (2)
[0098] In the formula, reflect and Position difference in the longitudinal plane To reflect the depth difference between the two, the coordinate difference is converted into an angle using the arctangent function, then converted from radians to degrees, and finally the pitch angle is obtained. .
[0099] As an example, the horizontal rotation angle calculation can be based on key points distributed laterally (such as...). and The difference between the z-axis (depth) and x-axis (horizontal plane) coordinates is used to calculate the horizontal rotation angle using the arctangent function. A specific formula example is as follows:
[0100] (3)
[0101] In the formula, reflect and Positional difference in the horizontal plane Reflecting the depth difference between the two, the final horizontal rotation angle is obtained. .
[0102] As an example, tilt angle calculation can be based on key points distributed laterally (such as... and The difference between the y-axis (vertical plane) and x-axis (horizontal plane) coordinates is used to calculate the tilt angle using the arctangent function. A specific formula example is as follows:
[0103] (4)
[0104] In the formula, reflect and Position difference in the longitudinal plane Reflecting the difference in lateral position between the two, the tilt angle is ultimately obtained. .
[0105] Through the above calculations, the pitch angle can be obtained. Horizontal rotation angle Inclination angle The first three-dimensional attitude angle is formed, completing the transformation from key point coordinates to attitude quantization angle.
[0106] S123, based on the preset mapping relationship, the first three-dimensional attitude angle is corrected by error compensation to obtain the second three-dimensional attitude angle.
[0107] The first 3D pose angle calculated above, while providing a preliminary quantification of facial pose, may contain systematic errors due to multiple factors. Specifically, the depth coordinates (z-axis) output by facial landmark detection models (such as FaceMesh) may become distorted with changes in facial pose (such as large pitches or rotations), leading to a deviation between the initial pose angle calculated based on the coordinates and the actual pose. Uneven lighting during shooting, camera angle shifts, and image noise can reduce the accuracy of landmark coordinate extraction, indirectly causing errors in the initial pose angle calculation. The geometric formulas used in the initial pose angle calculation (such as the arctangent calculation based on the difference between two coordinates) are simplified models that do not fully cover the complex 3D geometry of the face, potentially leading to deviations in angle calculation. These errors typically exhibit statistically predictable systematic characteristics and cannot be eliminated through a single calculation; they require correction through a dedicated error compensation mechanism.
[0108] As an example, the preset mapping relationship can be a mapping function constructed using curve fitting algorithms (such as polynomial fitting or nonlinear least squares fitting) based on the error data pairs of the initial angle and the true angle.
[0109] As an example, each angle of the first three-dimensional attitude angle is substituted into the matching mapping function to calculate the corrected angle, which is the second three-dimensional attitude angle.
[0110] Specifically, targeting Preset mapping relationship as follows:
[0111] (5)
[0112] against Preset mapping relationship as follows:
[0113] (6)
[0114] against Preset mapping relationship as follows:
[0115] (7)
[0116] Where a, b, c, and d are the coefficients obtained from the fitting. When When the error is small, it can be simplified to The linear relationship, similarly, in , This can also be simplified to a linear relationship.
[0117] S124, based on the second three-dimensional pose angle, uses geometric transformation to map the key point coordinates of the target facial region to the virtual projection plane when the face is facing the camera, thus obtaining the corrected key point coordinates of the target region.
[0118] In real-world shooting scenarios, faces are often not directly facing the camera (e.g., there are tilting, horizontal rotation, or tilting postures). In such cases, key points of the target facial features (e.g., eyes, mouth) will exhibit perspective distortion in the 2D image. For example, when a face tilts its head back, the key points of the eyes will appear narrower at the top and wider at the bottom, and the key points of the mouth will shift downwards in the image. When a face turns to the left, the key points of the left eye will be partially obscured, and the key points of the right eye will be stretched and distorted. This distortion means that the original coordinates of the key points cannot accurately reflect the actual geometric shape of the target area. If these coordinates are directly used to calculate the state parameters of the target area, significant errors will occur, such as misinterpreting the narrowing of the eyes caused by tilting the head back as closed eyes. Therefore, geometric transformations are necessary to eliminate the distortion.
[0119] As an example, a virtual projection plane can refer to a virtual two-dimensional plane constructed in a computer system, parallel to the camera lens plane, assuming the face is perfectly facing the camera. This plane can have two main characteristics: first, the distribution of all facial key points within the plane corresponds to a frontal pose of the face without pitch, rotation, or tilt, serving as a standardized pose reference for key points of the target area; second, the plane has a coordinate scale (such as pixels) consistent with the original image, ensuring that the mapped key point coordinates can be directly used for subsequent parameter calculations. The purpose of constructing this plane is to provide a unified correction benchmark for key points of perspective distortion, enabling key points of facial target areas under different poses to be transformed into coordinates under the same pose standard, achieving cross-pose feature standardization.
[0120] As an example, for the perspective distortion of key points of facial target parts (such as eyes and mouth) caused by facial posture deviation (such as looking up, turning the head, or tilting) during actual shooting, the original coordinates of these key points can be projected onto the virtual plane of the face facing the camera through a preset spatial geometric transformation algorithm, based on the second three-dimensional posture angle with optimized accuracy. This eliminates the coordinate deviation caused by posture deviation and finally obtains the standardized key point coordinates consistent with the frontal facial posture, that is, the corrected key point coordinates of the target part.
[0121] As another implementation of S124, S124 may also include the following steps:
[0122] Construct a three-dimensional rotation matrix around the X-axis, Y-axis, and Z-axis based on the second three-dimensional attitude angle;
[0123] The coordinates of the key points of the target area are translated to the coordinate system of the target origin to obtain the coordinates of the translated key points;
[0124] Multiply the translated keypoint coordinates by the 3D rotation matrix to obtain the corrected coordinates on the virtual projection plane.
[0125] As an example, a three-dimensional rotation matrix can be a mathematical tool used to describe the rotation of a spatial point around a coordinate axis. Essentially, it transforms the coordinates of key points in three-dimensional space according to the rotation law corresponding to a second three-dimensional pose angle through linear algebraic operations, thereby compensating for coordinate deviations caused by facial pose shifts. Three independent rotation matrices around the X, Y, and Z axes respectively eliminate three pose deformations: facial pitch, tilt, and horizontal rotation. Their combined effect achieves complete correction of any pose shift.
[0126] As an example, the second three-dimensional attitude angle can include the pitch angle about the X-axis. Inclination angle around the Y-axis Horizontal rotation angle about the Z-axis For example, corresponding rotation matrices are constructed based on these three angles, and all angles can be converted from degrees to radians (conversion formula: radians = degrees × π / 180) to ensure the accuracy of trigonometric function calculations.
[0127] Specifically, as an example, a rotation matrix around the X-axis is used to counteract keypoint deformation caused by facial pitch (head up, head down). It is constructed with the X-axis as the rotation axis, causing the keypoints to rotate in the YZ plane. The mathematical expression is as follows:
[0128] (8)
[0129] in, The pitch angle is the second three-dimensional attitude angle. , These are the cosine and sine function values of the pitch angle in radians, respectively. The purpose of this matrix is to correct the coordinate offset of key points in the longitudinal (Y-axis) and depth (Z-axis) directions caused by pitch by using a linear combination of the Y-axis and Z-axis coordinates.
[0130] Specifically, as an example, a rotation matrix around the Y-axis is used to counteract keypoint distortion caused by facial tilting poses (tilting left or right). It is constructed with the Y-axis as the rotation axis, causing the keypoints to rotate within the XZ plane. The mathematical expression is as follows:
[0131] (9)
[0132] in, This refers to the tilt angle in the second three-dimensional attitude angle; , These are the cosine and sine function values of the tilt angle in radians, respectively. The purpose of this matrix is to correct the coordinate offset of key points in the horizontal (X-axis) and depth (Z-axis) directions caused by tilting by using a linear combination of X-axis and Z-axis coordinates.
[0133] Specifically, as an example, a rotation matrix around the Z-axis is used to counteract keypoint distortion caused by horizontal facial rotation (turning the head left or right). It is constructed with the Z-axis as the rotation axis, causing the keypoints to rotate in the XY plane. The mathematical expression is as follows:
[0134] (10)
[0135] in, This refers to the horizontal rotation angle in the second three-dimensional attitude angle. , These are the cosine and sine function values of the horizontal rotation angle in radians, respectively. The purpose of this matrix is to correct the coordinate offset of key points in the horizontal (X-axis) and vertical (Y-axis) directions caused by horizontal rotation through a linear combination of the X-axis and Y-axis coordinates.
[0136] As an example, the original coordinates of key points in the target area are in a pixel coordinate system with the top left corner of the image as the origin (default origin). However, the calculation of the 3D rotation matrix requires the geometric center of the face as the origin (target origin). If the original coordinates are used directly for rotation, the key points will rotate around the corners of the image, resulting in additional positional offsets. Therefore, the main purpose of coordinate translation is to transform the key point coordinates from the default origin coordinate system to the target origin coordinate system, providing a unified and reasonable coordinate reference for subsequent rotation matrix calculations, ensuring that the rotation effect only affects pose deformation and does not introduce additional positional errors.
[0137] As an example, the original coordinates of key points in the target area are in a pixel coordinate system with the top left corner of the image as the origin (default origin). However, the calculation of the 3D rotation matrix requires the geometric center of the face as the target origin. If the original coordinates are used directly for rotation, the key points will rotate around the corners of the image, resulting in additional positional offsets. Therefore, the core purpose of coordinate translation is to transform the key point coordinates from the default origin coordinate system to the target origin coordinate system, providing a unified and reasonable coordinate reference for subsequent rotation matrix calculations, ensuring that the rotation effect only affects pose deformation and does not introduce additional positional errors.
[0138] As an example, the origin of the target can be selected from stable keypoints corresponding to the geometric center of the face (such as the expression-irrelevant keypoints selected earlier), and its three-dimensional coordinates can be denoted as... This selection principle satisfies two conditions: first, the position is stable and unaffected by changes in facial expression; second, it is located in the central area of the face, ensuring that all key points of the target area are distributed around the center after translation, and the force is evenly distributed during rotation.
[0139] Let the original three-dimensional coordinates of a key point in a target area be... The coordinates of the key points after translation (denoted as...) ) is calculated using the following formula:
[0140] (11)
[0141] in, These represent the horizontal, vertical, and depth coordinates of the key points in the target origin coordinate system after translation. Through this calculation, the coordinates of all key points in the target area are redistributed with the center of the face as the origin. For example, a key point located on the right side of the face in the original coordinates will have a positive horizontal coordinate after translation, while a key point on the left side will have a negative horizontal coordinate, ensuring that subsequent rotation calculations can accurately offset posture deformation.
[0142] As an example, since the three attitude deformations corresponding to the second three-dimensional attitude angle have mutual influence, the rotation matrix can be combined in a fixed order around the Z-axis, Y-axis, and X-axis to translate the key point coordinates. Substitute the following formulas into the calculation:
[0143] (12)
[0144] in, , , These are the three-dimensional rotation matrices about the Z-axis, Y-axis, and X-axis, respectively. (Rotation matrix part) Responsible for eliminating pose distortion, transforming the coordinates of key points after translation into a spatial distribution consistent with the face facing the camera; at the end... The operation then restores the rotated coordinates from the target origin coordinate system to the pixel coordinate system of the original image, ensuring that the corrected coordinates... The coordinates fall within the valid pixel range of the original image. Its distribution is completely consistent with that of a person's face when facing the camera, and can be directly used for the calculation of the state parameters of the target part.
[0145] The face image screening method of this application constructs a three-dimensional rotation matrix around the X, Y, and Z axes based on a second three-dimensional pose angle. This fully simulates the three-dimensional spatial transformation of the face from its actual pose to its pose facing the camera, ensuring that the coordinates of the transformed key points accurately reflect the true position of the target area from the facing viewpoint and eliminating the interference of perspective distortion on subsequent parameter calculations. The standardized steps for constructing the rotation matrix, coordinate translation, and matrix multiplication are clearly defined, with each step having explicit mathematical logic. This avoids differences in correction results caused by unclear operation steps, ensuring that consistent corrected coordinates are obtained when different devices and operators perform these steps, thus improving the repeatability and reliability of the solution.
[0146] S125, calculate the state parameters of the target part based on the corrected coordinates of the key points of the target part.
[0147] As an example, the target part's state parameters can be numerical indicators used to quantitatively describe the core functional state or morphological characteristics of the target part, and their type is directly related to the selection requirements of the target part. These mainly include parameters of opening and closing degree and morphological characteristic parameters.
[0148] As an example, the degree of opening and closing parameter can be used to describe the opening and closing state of parts such as the eyes and mouth, such as the degree of eye closure and the degree of mouth closure. The larger the value, the greater the degree of opening of the part, which can accurately distinguish the state of open eyes, closed eyes, closed mouth, slightly open or wide open, etc.
[0149] As an example, morphological feature parameters can be used to describe the specific shape of a target part, such as mouth shape parameters. These parameters quantify mouth features by the distance ratio between key points, providing a basis for selecting images with specific mouth shape requirements.
[0150] The main purpose of these parameters is to transform traditional subjective visual judgment into objective numerical comparison, so that the state of target parts in different images has a unified comparison benchmark.
[0151] As another implementation of S125, the target part status parameters include the aspect ratio of the target part, and S125 may include the following steps:
[0152] Based on the corrected coordinates of the key points of the target area, the aspect ratio of the target area is calculated using the following formula:
[0153] (13)
[0154] in, and The coordinates of two corrected key points characterizing the width of the target area. , , , The coordinates of multiple pairs of corrected key points characterizing the height of the target area. , These can represent the Euclidean distance between corresponding key points on the upper and lower eyelids (representing eye height). It can represent the Euclidean distance between key points at the left and right corners of the eyes (characterizing the width of the eyes); the degree of eye closure is quantified by the ratio of the sum of the heights to the width, and the larger the ratio, the greater the degree of eye opening.
[0155] Similarly, mouth closure can be a parameter characterizing the degree of mouth opening, using the same geometric calculation logic as eye closure, as shown in the following formula:
[0156] (14)
[0157] in, The coordinates of the key points at the corners of the mouth after correction. The coordinates of the upper and lower lip key points are corrected; the numerator is the sum of the Euclidean distances between multiple pairs of upper and lower lip key points (representing mouth height), and the denominator is the Euclidean distance between key points at the corners of the mouth (representing mouth width). The degree of mouth closure is quantified by the ratio of the sum of heights to the width, and the larger the value, the greater the degree of mouth opening.
[0158] The face image screening method in this application calculates the aspect ratio using a standardized formula, transforming subjective states such as "openness" and "closedness" into objective values. Higher values indicate greater openness or closure of the facial features. The calculated parameters are quantifiable and comparable, avoiding biases caused by subjective judgment. The definition of key points in the formula is completely standardized. Regardless of changes in image capturing equipment, resolution, or facial features, the formula can be calculated based on the corrected key point coordinates, ensuring a unified calculation benchmark for the facial feature state parameters across different scenes and images. This allows for direct horizontal comparison and improves cross-scene consistency of parameters.
[0159] As another implementation of S125, the target part includes the mouth, and the target part state parameters include mouth shape parameters. S125 may also include the following steps:
[0160] Obtain the coordinates of at least one pair of key points at the corners of the mouth that are symmetrical to each other and located on both sides of the midline of the face;
[0161] Calculate the mouth width based on the coordinates of a pair of key points at the corners of the mouth;
[0162] Obtain the coordinates of at least one pair of key points located on the midline of the upper and lower lips;
[0163] The degree of mouth opening and closing is calculated based on the coordinates of a pair of key points in the lips;
[0164] Calculate the ratio of the degree of mouth opening to the mouth width, and use it as a mouth shape parameter.
[0165] As an example, the facial midline can refer to a virtual straight line that runs perpendicularly through the geometric center of the face (such as the subnasal keypoint or the central forehead keypoint). The corner keypoints of the mouth must be located on the left and right sides of this midline, respectively, and the distances from the two points to the midline must be equal (i.e., symmetrical). It should be noted that the coordinates of the selected keypoints can be the standardized coordinates after posture correction mentioned above, to ensure that the interference of posture deformation on the coordinates is eliminated and to ensure the accuracy of the mouth width calculation.
[0166] Specifically, as an example, the straight-line distance between a selected pair of key points at the corners of the mouth can be calculated using geometric methods. This distance represents the mouth width and can be used to characterize the lateral extent of the mouth. For instance, the Euclidean distance formula can be used to calculate the distance between two points. If multiple pairs of symmetrical key points at the corners of the mouth are selected, the distance between each pair can be calculated separately, and the arithmetic mean can be taken as the final mouth width to improve the stability of the result.
[0167] As an example, the midline of the upper and lower lips can be a virtual straight line coinciding with the midline of the face and extending longitudinally along the lips. Key points in the center of the lips should be located in the central areas of the upper and lower lips respectively, and the line connecting these two points should align with the direction of mouth opening (i.e., longitudinal distribution). The selected key points must precisely correspond to the center position of the lip thickness to avoid coordinate deviations caused by blurred lip edge contours. At least one pair of key points can be selected (one for the upper lip and one for the lower lip). To optimize accuracy, multiple pairs of key points in the center of the lips can be selected (e.g., one pair each for the upper, middle, and lower parts of the lips), and the average of the calculated results can be used to offset local contour interference. It should be noted that the coordinates of the selected key points and the coordinates of the key points at the corners of the mouth must be standardized coordinates after correction to ensure a consistent calculation basis for the horizontal and vertical dimensions, avoiding ratio distortion caused by scale differences. Specifically, the Euclidean distance formula can be used to calculate the distance between two points. If multiple pairs of key points in the center of the lips are selected, the distance between each pair can be calculated separately, and the arithmetic mean can be taken as the final degree of mouth opening.
[0168] As an example, the ratio of mouth opening depth to mouth width is calculated as a mouth shape parameter. This ratio accurately represents mouth shape features through the relative relationship between vertical opening depth and horizontal width. For example, a small ratio usually corresponds to a closed mouth with a large width (such as a naturally relaxed mouth shape); a medium ratio corresponds to a slightly open mouth in a natural expression; and a large ratio corresponds to a wide-open mouth or pouting. By using a ratio, the influence of differences in facial size (such as the difference in absolute mouth size between adults and children, but the ratio can reflect relative shape) is eliminated, providing a unified comparison benchmark for mouth shape parameters of different individuals and images, adapting to the screening needs across scenes and groups.
[0169] The face image filtering method of this application calculates the ratio of mouth width to mouth opening degree to quantify different mouth shape features, achieving multi-dimensional and accurate evaluation of mouth state. Based on mouth shape parameters, it can adapt to different mouth shape filtering needs. By simply setting the corresponding mouth shape parameter threshold, images that meet the mouth shape requirements can be accurately filtered out, expanding the application scope of the solution in mouth state-related filtering scenarios.
[0170] S130, select a face image from multiple face images whose target part state parameters meet the first preset condition and whose second three-dimensional pose angle meets the second preset condition, and obtain the target image.
[0171] As an example, the first preset condition can refer to the qualification criteria for the target part's state parameters, which are pre-set according to the actual screening scenario requirements. This condition can be presented in the form of parameter value ranges or parameter level thresholds. The first preset condition can be flexibly adjusted according to different scenarios to ensure that the screening results are adapted to personalized needs.
[0172] As an example, the second preset condition can refer to the qualified judgment standard of the second three-dimensional pose angle set in advance according to the requirements of the face pose in the screening scenario. It can also be presented in the form of angle value range or pose level threshold.
[0173] As an example, the target image can refer to a face image that simultaneously meets the first and second preset conditions, perfectly matching the filtering requirements of actual application scenarios.
[0174] As another implementation of S130, the first preset condition is that the target part's state parameter is greater than a first preset level, and the second preset condition is that the second three-dimensional attitude angle is greater than a second preset level. S130 may also include the following steps:
[0175] Based on the preset target part status parameter grading range, determine the target part status level of the target part status parameter.
[0176] Based on the preset second three-dimensional attitude angle grading range, the second three-dimensional attitude level of the second three-dimensional attitude angle is determined;
[0177] The target image is obtained by selecting face images from multiple face images whose target part state level is greater than the first preset level and whose second three-dimensional pose level is greater than the second preset level.
[0178] As an example, the target body part status parameter grading interval can be divided into multiple continuous and non-overlapping intervals based on the physical meaning and screening requirements of the target body part status parameters. Each interval corresponds to a unique level (the level value usually increases as the parameter compliance improves). For example, for the eye closure degree EAR, the preset grading intervals can be set as follows: Level 1: EAR < 0.2 (eye closed or nearly closed); Level 2: 0.2 ≤ EAR < 0.5 (eye half open); Level 3: 0.5 ≤ EAR ≤ 1 (eye fully open).
[0179] As an example, the state parameters of the target parts of a single face image can be compared one by one with preset classification intervals to determine the interval to which it belongs and the corresponding level. For example, if an image has an EAR=0.6, it falls into the level 3 interval, and its target part state level is 3. If there are multiple target part state parameters, the level of each parameter needs to be determined separately, and then they can be combined into a comprehensive target part state level through preset weighting rules (such as taking the average or the lowest level) to ensure the uniqueness of the classification result.
[0180] As an example, the second three-dimensional pose angle grading interval can be based on the compliance requirements of face pose, dividing the reasonable value range of each pose angle into multiple intervals, with each interval corresponding to a unique level (the level value increases as pose compliance improves). For example, for the horizontal rotation angle... The preset grading range can be set as follows: Grade 1: ≥60° (excessive head turn, posture unacceptable); Level 2: 30°≤ <60° (moderate head turn, posture relatively compliant); Level 3: <30° (slight head turn, proper posture).
[0181] As an example, after determining the levels of the three pose angles of a single face image, the lowest level is taken as the second three-dimensional pose level of the image, ensuring that the pose level reflects the least desirable pose dimension. For example, if... Corresponding level 3 Corresponding level 3 For a corresponding level of 2, the second 3D pose level of the image is 2. For directional pose angles (such as...) (Positive values indicate looking up, negative values indicate looking down) The hierarchical intervals can usually be set based on absolute values (e.g., ... <30°), to avoid interference from directional differences in grade determination.
[0182] As an example, the first preset level and the second preset level can be the qualification thresholds set according to the requirements of the filtering scenario, and a single image must meet both level conditions at the same time to be selected.
[0183] As an example, the target image can be a set of all qualified images that have been compared with the above-mentioned level of all images, which can be directly used in subsequent application scenarios (such as news release, portrait screening).
[0184] The face image filtering method in this application converts continuous values into discrete levels by using hierarchical intervals. During filtering, it is only necessary to determine whether the level meets the requirements, reducing the impact of numerical fluctuations on the filtering results, simplifying the filtering logic, and improving judgment efficiency. Different scenarios have different requirements for filtering standards. Since the level division is bound to the scenario requirements, the filtering standards can be quickly switched by simply adjusting the first preset level and the second preset level, without the need to readjust the parameter calculation logic, thus improving the flexibility of the solution to adapt to different filtering needs.
[0185] The face image filtering method of this application combines a second three-dimensional pose angle and target part state parameters, filtering images through dual constraints. This effectively avoids the misselection of invalid images where the pose is acceptable but the part state is unacceptable, or vice versa, significantly improving the practicality and accuracy of the filtering results. Error compensation is applied to the first three-dimensional pose angle using a preset mapping relationship to obtain a precise second three-dimensional pose angle. Then, based on this pose angle, the coordinates of key points of the target part are mapped to a virtual projection plane of the face facing the camera, eliminating perspective distortion interference caused by pose shifts. This ensures that the target part state parameters have a unified calculation benchmark under different poses, avoiding parameter misjudgment due to pose differences and improving the consistency and reliability of parameter calculation. Batch processing of multiple face images eliminates the need for manual parameter adjustment for individual images, quickly adapting to various application scenarios and meeting the efficiency requirements of batch image processing while ensuring filtering accuracy.
[0186] As another implementation of this application, in order to improve the accuracy of subsequent coordinate mapping and parameter calculation, such as Figure 2 As shown, after S123, the method may further include the following step S210.
[0187] S210, based on the personalized parameters of the face, performs a linear transformation on the second three-dimensional pose angle to obtain the third three-dimensional pose angle.
[0188] As an example, personalized facial parameters can refer to quantitative indicators that characterize the unique features of an individual's face, including but not limited to: facial proportion parameters (such as the ratio of eye distance to face width, and the ratio of nose length to face length); facial contour parameters (such as cheekbone width and jaw angle); and key point distribution density parameters (such as the density of key points around the eyes). These parameters are obtained through the initial facial feature extraction module and are linked to the individual, reflecting the personalized differences in that face.
[0189] As an example, the technical implementation of this step relies on the following three sets of independent linear transformation formulas, corresponding to attitude angle corrections in three dimensions:
[0190] As an example, the formula for the linear transformation of the pitch angle about the X-axis is:
[0191] (15)
[0192] As an example, the formula for the linear transformation of the tilt angle about the Y-axis is:
[0193] (16)
[0194] As an example, the linear transformation formula for the horizontal rotation angle about the Z-axis is:
[0195] (17)
[0196] in, The third three-dimensional attitude angle is obtained after linear transformation, corresponding to the personalized corrected attitude angles around the X-axis, Y-axis, and Z-axis, respectively; The scaling factor (slope) for the linear transformation is determined by the personalized parameters of the face and is used to adjust the correction range of the pose angle. The offset coefficient (intercept) of the linear transformation is determined by the individual parameters of the face and is used to compensate for systematic posture deviations caused by individual differences in facial structure.
[0197] Based on this, S124 may include mapping the key point coordinates of the target facial region to the virtual projection plane when the face is facing the camera through geometric transformation based on the third three-dimensional pose angle, so as to obtain the corrected key point coordinates of the target region.
[0198] As an example, using the third three-dimensional pose angle (tilt angle, tilt angle, and horizontal rotation angle that integrates personalized facial features) obtained after personalized optimization as a reference, the original coordinates of the key points of the target part (which are subject to perspective distortion due to facial pose) are mapped to the virtual projection plane when the face is facing the camera through geometric transformation operations such as three-dimensional rotation and coordinate translation, and finally the corrected coordinates of the key points of the target part are output.
[0199] Eliminating the positional offset of key points caused by facial postures such as looking up, looking down, turning, and tilting, the corrected coordinates are distributed on the virtual plane in a way that not only conforms to the individual's facial anatomy but also accurately reflects the actual shape of the target area (such as the eyes and mouth), providing high-precision basic data for subsequent calculation of the target area's state parameters.
[0200] The face image screening method of this application performs a linear transformation on the second three-dimensional pose angle using personalized parameters, binding the pose angle correction to the individual characteristics of the person. This makes the final third three-dimensional pose angle more closely match the actual pose of the person, further narrowing the gap between the calculated pose angle value and the true value, and improving the accuracy of subsequent coordinate mapping and parameter calculation. For scenarios requiring screening of specific individuals, the personalized correction of the third three-dimensional pose angle can ensure that the person's pose evaluation criteria match their individual characteristics, providing a reliable pose benchmark for subsequent accurate screening of that person.
[0201] As another implementation of this application, in order to make the third three-dimensional attitude angle infinitely close to the true value, such as Figure 3 As shown, before S210, the method may also include steps S310-S330.
[0202] S310, acquire multiple first images of the face corresponding to the face image; wherein, the first image includes a manually annotated fourth three-dimensional pose angle.
[0203] As an example, the first image may contain images of a face in a variety of typical poses, covering at least different angle ranges of pitch (head up, head down), tilt (left, right), and horizontal rotation (left, right) to ensure comprehensiveness of pose changes.
[0204] As an example, the fourth three-dimensional pose angle can be a manually annotated true pose angle, including the pitch angle around the X-axis, the tilt angle around the Y-axis, and the horizontal rotation angle around the Z-axis; and it can be completed by professional annotators using three-dimensional pose annotation tools. The annotations of different first images of the same face should also maintain a consistent benchmark (such as using the face midline as the benchmark for pose angle calculation) to avoid systematic deviations in the annotation system.
[0205] S320, perform error compensation correction based on a preset mapping relationship based on the first three-dimensional attitude angle of the first image to obtain the second three-dimensional attitude angle of the first image.
[0206] As an example, the second three-dimensional attitude angle of the first image can be calculated using the same method as the aforementioned second three-dimensional attitude angle. That is, for each first image, its first three-dimensional attitude angle is first calculated using a basic algorithm, and then corrected using a preset mapping relationship to output the second three-dimensional attitude angle.
[0207] S330: Perform linear regression fitting on the second and fourth three-dimensional pose angles of the first image to obtain personalized parameters of the face.
[0208] As an example, this step primarily uses linear regression analysis to extract personalized parameters that accurately reflect individual facial features. Specifically, for multiple first images of the same face, each image contains two sets of pose angle data: one set is the second three-dimensional pose angle obtained after general error compensation correction, and the other set is the manually labeled high-precision fourth three-dimensional pose angle (serving as the benchmark for the true pose). Fitting these two sets of pose angles using a linear regression algorithm essentially involves finding the linear correspondence between them, that is, determining how to adjust the value of the second three-dimensional pose angle to make it closer to the fourth three-dimensional pose angle. This adjustment rule is ultimately transformed into a series of quantitative parameters, which are the personalized parameters of the face. These parameters can accurately capture the pose angle deviation characteristics caused by the face's own facial structure (such as facial feature proportions, contour features, etc.), providing a specific basis for subsequent personalized correction of the face's pose angles.
[0209] The face image screening method in this application uses a fourth three-dimensional pose angle as a fitting benchmark, and then obtains personalized parameters through linear regression fitting of the second and fourth three-dimensional pose angles. This ensures that the parameters can offset individual biases to the greatest extent, making the linearly transformed third three-dimensional pose angle infinitely close to the true value. It provides a standardized personalized parameter generation process for data acquisition, pose angle calculation, and linear fitting, eliminating the need for manual parameter adjustment by professional technicians. Only the labeled image of the person needs to be input to automatically generate suitable personalized parameters, lowering the application threshold of personalized correction functions and facilitating rapid implementation in real-world scenarios.
[0210] As another implementation of this application, in order to improve the accuracy of the correction, such as Figure 4 As shown, before S210, the method may also include steps S410-S430.
[0211] S410, obtain the feature vector of the face corresponding to the face image.
[0212] As an example, feature vectors can be extracted from key regions of a face image (such as facial contours, facial textures, and key point distribution) using deep learning or traditional computer vision algorithms. For example, deep semantic features can be extracted using convolutional neural networks (CNNs), or texture features can be extracted based on local binary patterns (LBP). The extracted feature vectors are typically 128, 256, or 512 dimensional. The higher the dimension, the stronger the uniqueness of the facial features.
[0213] By extracting unique facial feature vectors, we can provide identity credentials for subsequent determination of whether the face to be processed has preset personalized parameters, ensuring that personalized corrections only apply to the target face.
[0214] S420 calculates the similarity between the feature vector and the preset face feature vector.
[0215] As an example, the preset face feature vector can be a feature vector stored in advance for the target face (such as a specific user who needs high-precision pose correction, or a face that has completed personalized parameter training), and associated with the corresponding personalized parameters to form a correspondence or mapping library of feature vector-personalized parameters.
[0216] As an example, similarity calculation methods can employ algorithms suitable for high-dimensional vector comparison, such as:
[0217] Euclidean distance: Calculates the straight-line distance between two feature vectors in a high-dimensional space. The smaller the distance, the higher the similarity.
[0218] Cosine similarity: Calculates the cosine of the angle between two feature vectors. The closer the value is to 1, the higher the similarity.
[0219] The specific method can be selected based on the type of feature vector (e.g., normalized vectors are suitable for cosine similarity) to ensure that the calculation results can accurately reflect the face matching degree.
[0220] S430, if the similarity meets the third preset condition, the second three-dimensional pose angle is linearly transformed based on the personalized parameters of the face corresponding to the face image to obtain the third three-dimensional pose angle.
[0221] As an example, the third preset condition can be a threshold standard for determining whether feature matching is effective. For example, setting "similarity ≥ 0.9" means that when the similarity between the face feature vector to be processed and the preset vector reaches or exceeds 0.9, it is determined to be a successful match. The threshold can be adjusted according to the accuracy requirements of the application scenario (e.g., the threshold is set to 0.95 for security scenarios and 0.85 for ordinary interactive scenarios).
[0222] As an example, after a successful match under the third preset condition, the system can retrieve the corresponding personalized parameters from the mapping library based on the index of the preset feature vector to ensure that the parameters completely correspond to the face to be processed and avoid parameter retrieval errors.
[0223] As an example, based on the personalized parameters invoked, the pitch angle, tilt angle, and horizontal rotation angle of the second three-dimensional attitude angle are linearly transformed respectively. For example, by using the logic of "corrected attitude angle = proportional coefficient × original attitude angle + offset coefficient", the attitude deviation caused by individual facial structure is eliminated, and the third three-dimensional attitude angle is finally obtained. The transformation process must maintain the same calculation benchmark (such as angle unit and coordinate system) as when the personalized parameters were trained in the early stage to ensure the accuracy of the correction.
[0224] The face image screening method of this application calculates the similarity between the face feature vector and a preset face feature vector. Personalized parameter correction is only initiated when the similarity meets the condition, avoiding unnecessary personalized calculations for non-target individuals. This ensures that personalized correction only applies to the target object, improving the accuracy of the correction. For non-target individuals, the second three-dimensional pose angle is directly used in subsequent processes, eliminating the need for linear transformation operations. This reduces redundant calculation steps in batch processing, lowers CPU and GPU resource consumption, and significantly improves the overall screening speed, especially when processing image sets containing a large number of non-target individuals.
[0225] As another implementation of this application, in order to improve the scenario adaptability and long-term availability of the solution, such as Figure 5 As shown, the method may further include steps S510-S530.
[0226] S510, Obtain a face sample set; the sample set includes multiple face samples, each containing a second image of the face and its corresponding real three-dimensional pose angle.
[0227] As an example, the composition of the face sample set can cover faces of different genders, ages (children, adults, the elderly), and facial features (round face, long face, high cheekbones, etc.) to ensure that the samples can represent a wide range of people; the second image can contain samples of faces in different three-dimensional poses, and the samples are evenly distributed in each angle range.
[0228] As an example, true 3D pose angles can be obtained directly from 3D scanning equipment or by a professional annotation team using high-precision pose annotation tools. The annotation results can be cross-validated (e.g., different annotators independently annotate the same image; if the error exceeds a threshold, the annotation is re-annotated) to ensure the reliability of the true values.
[0229] By constructing a large-scale, high-quality labeled sample set, we can provide data support for the subsequent learning of the general error rules of the pose angle, and ensure that the mapping function can adapt to the common characteristics of different groups of people.
[0230] S520, based on the face sample set, calculates the first three-dimensional pose angle of the second image of each face sample, and obtains a data pair composed of the first three-dimensional pose angle and the standard three-dimensional pose angle.
[0231] As an example, for the second image of the face of each sample, the same first three-dimensional pose angle calculation method as in the subsequent actual screening process is used (i.e., the geometric calculation method based on facial key points, such as calculating the initial pitch angle, horizontal rotation angle, and tilt angle by combining the preset expression-independent key points with trigonometric functions) to obtain the initial pose angle corresponding to the second image, which is denoted as the first three-dimensional pose angle.
[0232] Meanwhile, the real three-dimensional pose angles obtained in advance from the sample are used as standard three-dimensional pose angles. These standard values are the benchmark for measuring the accuracy of the first three-dimensional pose angles and represent the actual pose of the face in the second image.
[0233] Through the above operations, each face sample can generate a set of data pairs of first three-dimensional pose angle and standard three-dimensional pose angle; by traversing all samples in the sample set, multiple sets of such data pairs can be obtained, forming the basic dataset for subsequent fitting of the mapping function.
[0234] S530, based on data pairs, establishes a mapping function from the first three-dimensional attitude angle to the standard three-dimensional attitude angle through curve fitting, and this mapping function serves as a preset mapping relationship.
[0235] Based on multiple sets of data pairs generated by S520, a mapping function from the first three-dimensional attitude angle to the standard three-dimensional attitude angle is constructed using curve fitting. This function is the preset mapping relationship used for subsequent attitude angle correction.
[0236] In practice, the first step is to analyze the numerical relationship between the first 3D attitude angle and the standard 3D attitude angle in the data pair. Since the first 3D attitude angle is based on the geometric calculation of key points in the image, it may be affected by factors such as key point detection errors and perspective distortion during shooting, and there is usually a certain deviation between it and the standard 3D attitude angle. To address this deviation, curve fitting methods (such as polynomial fitting, nonlinear regression fitting, etc.) can be used to find a mathematical function that can optimally describe the transformation relationship from the first 3D attitude angle to the standard 3D attitude angle.
[0237] The final function obtained is the preset mapping relationship. In the subsequent actual face image screening process, the calculated first three-dimensional pose angle can be substituted into this mapping function to realize the error compensation and correction of the initial pose angle, and obtain the corrected pose angle that is closer to the true value.
[0238] The face image screening method in this application constructs data pairs using a face sample set, where the true three-dimensional pose angle serves as an objective benchmark, and the first three-dimensional pose angle is the actual calculated value. A mapping function between the two is established through curve fitting, which accurately captures the error pattern of the first three-dimensional pose angle, enabling the mapping function to have error compensation capabilities. Substituting the first three-dimensional pose angle, a second three-dimensional pose angle that approximates the true value can be output. If the application scenario changes subsequently, a new face sample set corresponding to the scenario can be added, and the mapping function can be refitted to achieve dynamic updates of the mapping relationship. This ensures that the preset mapping relationship always adapts to the current application scenario, improving the scenario adaptability and long-term availability of the solution.
[0239] As another implementation of this application, to avoid screening bias caused by empirical thresholds, such as Figure 6As shown, before S130, the method may also include steps S610-S630.
[0240] S610 calculates the numerical distribution of target region state parameters and second three-dimensional pose angles based on multiple third images containing a human face.
[0241] As an example, the third image can encompass people of different ages (e.g., children, youth, middle-aged, elderly), genders, and facial features (e.g., facial proportions, face shape, whether or not they are wearing glasses / masks), ensuring that the numerical distribution is not limited by specific human characteristics; it can include different shooting environments (e.g., indoor natural light, outdoor strong light, low light, night scene), different shooting devices (e.g., professional cameras, mobile phones, surveillance cameras), and different shooting angles (e.g., frontal, slight side shot, slight overhead / under-the-head shot), simulating the actual scene of acquiring human face images; it can cover different facial postures (e.g., slight head rotation, changes in pitch angle) and different states of the target parts (e.g., the gradual change of the eyes from completely closed to completely open, the different forms of the mouth from closed to open), ensuring that the numerical distribution can reflect the complete range of changes in the target parameters.
[0242] As an example, facial landmark detection models (such as deep learning models capable of recognizing hundreds of facial landmarks) can locate and extract the coordinates of key landmarks in target areas (such as the eyes and mouth), and calculate state parameters using corresponding mathematical formulas based on the type of the target area. For example, if the target area is the eyes, the eye aspect ratio formula is used to calculate the eye closure value by the ratio of the distance between the key landmarks at the upper and lower edges of the eyes and the key landmarks at the corners of the eyes; if the target area is the mouth, the mouth closure is calculated using the mouth aspect ratio, or the mouth shape parameters are calculated by the ratio of the distance between the key landmarks at the corners of the mouth and the distance between the key landmarks in the center of the lips.
[0243] As an example, the second three-dimensional pose angle can be obtained by first calculating the initial first three-dimensional pose angle through facial key points, then substituting it into a preset mapping function for error compensation, and finally obtaining the second three-dimensional pose angle corresponding to the image.
[0244] As an example, the numerical distribution can be the frequency pattern of the target part state parameters and the second three-dimensional pose angle in the actual application scenario. This pattern will serve as the basis for establishing the preset grade interval when screening face images, ensuring that the screening criteria conform to the parameter distribution characteristics of the actual image, and improving the rationality and practicality of the screening results.
[0245] As another implementation of S610, where multiple third images are faces corresponding to preset facial feature vectors, S610 may also include the following steps:
[0246] Based on multiple third images of faces containing preset facial feature vectors, the numerical distribution of target region state parameters and second three-dimensional pose angles is calculated.
[0247] As an example, a pre-stored facial feature vector can refer to pre-stored facial feature data that can uniquely represent a specific person. This vector can be extracted by a professional facial recognition model, contains key biometric information of the face, and is unique and stable, making it suitable for accurate matching of the corresponding person.
[0248] As an example, the face corresponding to the preset facial feature vector refers to all images containing the face of that specific person. When selecting multiple third images, the candidate images can be filtered through facial recognition technology. The facial feature vectors in the candidate images are compared with the preset facial feature vectors, and only images whose similarity meets the preset threshold are retained. This ensures that the multiple third images finally selected are all facial images of that specific person, rather than images of other people or mixed images containing that person and other people.
[0249] As an example, the method for obtaining the numerical distribution of the second three-dimensional pose angle for each third image is the same as described above. The final numerical distribution of the face corresponding to the preset facial feature vector can be a representation of the target part state and pose angle pattern that fits the individual characteristics of that person. This can provide a personalized grading standard for subsequent face image screening of that person, ensuring that the screening results are more in line with the actual characteristics and application needs of that person.
[0250] The face image screening method in this application calculates the numerical distribution based on multiple third images of a specific person. The divided hierarchical intervals perfectly match the individual parameter patterns of that person, ensuring that the screening criteria for that person accurately adapt to their individual characteristics and avoiding misscreening or omissions. For scenarios requiring focused screening of specific core individuals, personalized hierarchical intervals ensure that the screening results fully meet the optimal presentation requirements of that person, enhancing the application value of the screening results and meeting the needs of highly accurate personalized screening scenarios.
[0251] S620, the numerical distribution of the target part state parameters is divided to obtain the preset target part state parameter grading interval.
[0252] As an example, based on the obtained numerical distribution of state parameters for the target body part, the physical meaning of the parameters, actual screening requirements, and numerical distribution characteristics can be combined to determine preset grading intervals through numerical segmentation. For example, a larger eye closure EAR value indicates a higher degree of eye opening, and segmentation can ensure that the interval division corresponds one-to-one with the physical states of closed, half-open, and fully open eyes; a larger mouth shape parameter value indicates a greater degree of mouth opening or widening, and the interval can match the mouth shape changes logic of closed, slightly open, smiling, and laughing. Referring to the statistical indicators of the numerical distribution, segmentation should be prioritized at numerical inflection points or inflection points of central tendency. The number of levels can also be determined according to the accuracy requirements of the parameters in the application scenario. For example, in news photography scenarios, which require precise screening of fully open-eyed images, the eye EAR value can be divided into 5 levels (0-0.1: fully closed, 0.1-0.2: half closed, 0.2-0.3: slightly open, 0.3-0.4: mostly open, and above 0.4: fully open). In contrast, ordinary photo album screening scenarios can be simplified to 3 levels (0-0.2: closed, 0.2-0.35: half open, and above 0.35: open).
[0253] S630, the numerical distribution of the second three-dimensional attitude angle is divided to obtain the preset target part state parameter classification interval.
[0254] Regarding the numerical distribution of the second three-dimensional attitude angles (including pitch angle, horizontal rotation angle, and tilt angle), since the physical meaning and numerical range of pitch angle (representing the head's vertical rotation), horizontal rotation angle (representing the head's horizontal rotation), and tilt angle (representing the head's horizontal tilt) are completely independent, the numerical distribution of the second three-dimensional attitude angles can be first divided into three single-dimensional numerical distributions, and then further segmented.
[0255] The face image screening method of this application calculates the parameter and pose angle numerical distribution of multiple third images, divides intervals based on the distribution characteristics, and matches the parameter patterns of the actual images to ensure the objectivity of the grading standard and avoid screening bias caused by empirical thresholds. Since the third images contain different people, different shooting scenes, and different poses and body parts, their numerical distribution can reflect the overall pattern of parameters and pose angles in actual applications. The grading intervals divided based on this distribution can be adapted to most common scenarios, eliminating the need to divide intervals separately for a single scenario and improving the universality of the grading standard.
[0256] As another implementation of this application, in order to achieve a multi-dimensional and accurate assessment of the mouth's condition, such as Figure 7 As shown, before S125, the method may also include steps S710-S720.
[0257] During the face image screening process, if the person in the image is wearing glasses, the detection of key points of the eyes may be affected by factors such as lens reflection and frame obstruction, which in turn affects the accuracy of the calculation of eye state parameters.
[0258] S710: When the person in the face image is wearing glasses, a pre-trained eye state classification model is used to classify the eye region in the face image to obtain eye state information; wherein, the eye state information includes whether the eyes are open or closed.
[0259] As an example, when a person wearing glasses is detected in a face image (which can be determined by identifying features such as the outline of glasses frames and reflections in the lens area through a face feature detection model), a pre-trained eye state classification model can be invoked to classify the eye region in the image to obtain accurate eye state information.
[0260] As an example, the model training phase can utilize a large number of eye image samples containing images of people wearing glasses with their eyes open or closed, covering different types of glasses (such as framed glasses, rimless glasses, and sunglasses), different lens states (such as transparent lenses and reflective lenses), and different lighting conditions (such as lens reflection in strong light and eye shadows in low light), ensuring the model's robustness to scenes involving people wearing glasses. The model can use eye state as the classification target, and the output results only include two categories: open eyes or closed eyes, without needing to output intermediate states. Its main function is to quickly eliminate obvious closed-eye scenes, avoiding invalid processing of closed-eye states in subsequent parameter calculations. The model structure can be lightweight (such as using lightweight convolutional neural network architectures like MobileNet and SqueezeNet), ensuring that the classification inference time for a single eye image is extremely short, without affecting the speed of the overall screening process.
[0261] S720 calculates the target area state parameters based on the corrected key point coordinates of the target area when the eye state information is open.
[0262] As an example, if the eye status information is closed, the image is directly determined to not meet the screening requirements for eye status compliance. There is no need to perform subsequent calculations of the target part status parameters. The image is directly excluded and the process moves on to the next image.
[0263] As an example, if the eye status information is "eyes open", it means that the eyes are in an analyzable state. Based on the previously corrected coordinates of key points of the target eye area, the corresponding status parameters can be calculated to ensure that the parameters can accurately reflect the degree of eye opening and closing in the open eye state.
[0264] The face image filtering method of this application calculates the ratio of mouth width to mouth opening degree to quantify different mouth shape features, thereby achieving a multi-dimensional and accurate assessment of mouth state. Based on mouth shape parameters, it can adapt to different mouth shape filtering needs. By simply setting the corresponding mouth shape parameter threshold, images that meet the mouth shape requirements can be accurately filtered out, expanding the application scope of the solution in mouth state-related filtering scenarios.
[0265] Based on the face image filtering method provided in the above embodiments, this application also provides specific implementations of a face image filtering device. Please refer to the following embodiments.
[0266] First see Figure 8 The face image screening device 80 provided in this application embodiment includes the following modules:
[0267] Image acquisition module 810 is used to acquire multiple face images;
[0268] Image processing module 820 is used to perform processing on each face image separately, including:
[0269] The key point detection unit 821 is used to detect facial key points and facial target parts from a face image;
[0270] In some embodiments, the key point detection unit 821 may include:
[0271] The image sequence acquisition subunit is used to acquire image sequences containing multiple facial expressions of the same person in a fixed head posture.
[0272] The key point detection subunit is used to detect facial key points in the image sequence and obtain the coordinate sequence of each facial key point;
[0273] The offset calculation subunit is used to calculate the coordinate offset of each facial key point in the image sequence relative to the reference expression image;
[0274] The norm calculation subunit is used to calculate the total offset norm of each keypoint based on the coordinate offset.
[0275] The key point selection sub-unit is used to select the top N key points as facial key points in ascending order of total offset norm.
[0276] The pose angle calculation unit 822 is used to calculate the first three-dimensional pose angle of the face image based on the face key points;
[0277] The error compensation unit 823 is used to perform error compensation and correction on the first three-dimensional attitude angle based on a preset mapping relationship to obtain the second three-dimensional attitude angle;
[0278] The geometric transformation unit 824 is used to map the key point coordinates of the target part of the face to the virtual projection plane when the face is facing the camera based on the second three-dimensional pose angle, so as to obtain the corrected key point coordinates of the target part.
[0279] In some embodiments, the geometric transformation unit 824 may include:
[0280] The matrix construction sub-unit is used to construct a three-dimensional rotation matrix around the X-axis, Y-axis, and Z-axis based on the second three-dimensional attitude angle;
[0281] The coordinate translation sub-unit is used to translate the coordinates of key points of the target part to the coordinate system of the target origin, so as to obtain the coordinates of the key points after translation.
[0282] The coordinate rotation sub-unit is used to multiply the translated keypoint coordinates with the 3D rotation matrix to obtain the corrected coordinates on the virtual projection plane.
[0283] The state parameter calculation unit 825 is used to calculate the state parameters of the target part based on the corrected coordinates of the key points of the target part.
[0284] In some embodiments, the state parameter calculation unit 825 is specifically used to: calculate the aspect ratio of the target part based on the corrected coordinates of the key points of the target part using the following formula:
[0285] (13)
[0286] in, and The coordinates of two corrected key points characterizing the width of the target area. , , , The coordinates of multiple pairs of corrected key points characterizing the height of the target area.
[0287] In some embodiments, the state parameter calculation unit 825 is specifically used when the target part includes the mouth:
[0288] The corner of the mouth key point acquisition sub-unit is used to acquire the coordinates of at least one pair of corner of the mouth key points that are located on both sides of the midline of the face and are symmetrical to each other;
[0289] The mouth width calculation subunit is used to calculate the mouth width based on the coordinates of a pair of key points at the corners of the mouth;
[0290] The lip midpoint acquisition subunit is used to acquire the coordinates of at least one pair of lip midpoints located on the midline of the upper and lower lips;
[0291] The opening and closing calculation subunit is used to calculate the degree of mouth opening and closing based on the coordinates of a pair of key points in the lip.
[0292] The mouth shape parameter calculation subunit is used to calculate the ratio of the mouth opening degree to the mouth width, which is used as the mouth shape parameter.
[0293] The image filtering module 830 is used to filter out face images from multiple face images that meet the first preset condition for the state parameters of the target part and the second preset condition for the second three-dimensional pose angle, so as to obtain the target image.
[0294] In some embodiments, the image filtering module 830 may further include the following modules:
[0295] The first-level determination submodule is used to determine the target location status level of the target location status parameters based on the preset target location status parameter grading range.
[0296] The second level determination submodule is used to determine the second three-dimensional attitude level of the second three-dimensional attitude angle based on the preset second three-dimensional attitude angle classification range;
[0297] The image filtering submodule is used to filter out face images from multiple face images that have a target part state level greater than a first preset level and a second three-dimensional pose level greater than a second preset level, thereby obtaining the target image.
[0298] In some embodiments, the face image screening device 80 may further include the following modules:
[0299] The personalized correction module is used to perform a linear transformation on the second three-dimensional pose angle based on personalized parameters of the face to obtain the third three-dimensional pose angle.
[0300] The geometric transformation unit 824 is used to map the key point coordinates of the target facial region to the virtual projection plane when the face is facing the camera based on the third three-dimensional pose angle, thereby obtaining the corrected key point coordinates of the target region.
[0301] In some embodiments, the face image screening device 80 may further include the following modules:
[0302] The sample acquisition module is used to acquire multiple first images of the face corresponding to the face image; wherein, the first image includes a manually annotated fourth three-dimensional pose angle;
[0303] The parameter fitting module is used to perform error compensation correction based on a preset mapping relationship based on the first three-dimensional pose angle of the first image to obtain the second three-dimensional pose angle of the first image, and to perform linear regression fitting on the second three-dimensional pose angle and the fourth three-dimensional pose angle of the first image to obtain personalized parameters of the face.
[0304] In some embodiments, the face image screening device 80 may further include the following modules:
[0305] The feature extraction module is used to obtain the feature vector of the face corresponding to the face image;
[0306] The similarity calculation module is used to calculate the similarity between the feature vector and the preset face feature vector;
[0307] The personalized triggering module is used to trigger the personalized correction module to perform a linear transformation on the second three-dimensional pose angle based on the personalized parameters of the face corresponding to the face image when the similarity meets the third preset condition.
[0308] In some embodiments, the face image screening device 80 may further include the following modules:
[0309] The sample acquisition module is used to acquire a face sample set; the sample set includes multiple face samples, and each face sample contains a second image of the face and the corresponding real three-dimensional pose angle;
[0310] The data pair generation module is used to calculate the first three-dimensional pose angle of each face sample second image based on the face sample set, and obtain a data pair composed of the first three-dimensional pose angle and the standard three-dimensional pose angle.
[0311] The mapping relationship establishment module is used to establish a mapping function from the first three-dimensional attitude angle to the standard three-dimensional attitude angle based on data pairs and through curve fitting. This mapping function serves as a preset mapping relationship.
[0312] In some embodiments, the face image screening device 80 may further include the following modules:
[0313] The distribution calculation module is used to calculate the numerical distribution of the target part state parameters and the second three-dimensional pose angle based on multiple third images containing the face before the filtering module works.
[0314] In some embodiments, the distribution calculation module is specifically used to: calculate the numerical distribution of target part state parameters and second three-dimensional pose angles based on multiple third images of faces corresponding to preset facial feature vectors.
[0315] The first segmentation module is used to segment the numerical distribution of the target part's state parameters to obtain a preset target part state parameter grading interval.
[0316] The second segmentation module is used to segment the numerical distribution of the second three-dimensional attitude angle to obtain a preset second three-dimensional attitude angle grading interval.
[0317] In some embodiments, the face image screening device 80 may further include the following modules:
[0318] The eye condition inspection module performs the following operations before the second calculation module operates, when the target area includes the eyes:
[0319] The glasses detection submodule is used to classify the eye region in a face image when the person in the face image is wearing glasses, using a pre-trained eye state classification model to obtain eye state information; wherein, the eye state information includes whether the eyes are open or closed.
[0320] The state triggering submodule is used to trigger the second calculation module to calculate the state parameters of the target part based on the corrected coordinates of the key points of the target part when the eye state information is open.
[0321] Figure 9 A schematic diagram of the hardware structure of the face image screening device provided in an embodiment of this application is shown.
[0322] The face image screening device may include a processor 901 and a memory 902 storing computer program instructions.
[0323] Specifically, the processor 901 may include a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits that can be configured to implement the embodiments of this application.
[0324] Memory 902 may include mass storage for data or instructions. For example, and not limitingly, memory 902 may include a hard disk drive (HDD), floppy disk drive, flash memory, optical disk, magneto-optical disk, magnetic tape, or Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, memory 902 may include removable or non-removable (or fixed) media. Where appropriate, memory 902 may be internal or external to the integrated gateway disaster recovery device. In a particular embodiment, memory 902 is non-volatile solid-state memory.
[0325] Memory may include read-only memory (ROM), random access memory (RAM), disk storage media devices, optical storage media devices, flash memory devices, and electrical, optical, or other physical / tangible memory storage devices. Therefore, typically, memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the methods according to one aspect of this disclosure.
[0326] The processor 901 reads and executes computer program instructions stored in the memory 902 to implement any of the face image screening methods in the above embodiments.
[0327] In one example, the face image screening device may further include a communication interface 903 and a bus 910. Wherein, as Figure 9 As shown, the processor 901, memory 902, and communication interface 903 are connected through bus 910 and complete communication with each other.
[0328] The communication interface 903 is mainly used to realize communication between various modules, devices, units and / or equipment in the embodiments of this application.
[0329] Bus 910 includes hardware, software, or both, that couples components of a face image screening device together. For example, and not limitingly, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an Infinite Bandwidth Interconnect, a Low Pin Count (LPC) bus, a memory bus, a Microchannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses, or combinations of two or more of these. Where appropriate, bus 910 may include one or more buses. Although specific buses are described and illustrated in embodiments of this application, any suitable bus or interconnect is contemplated herein.
[0330] Furthermore, in conjunction with the face image filtering method in the above embodiments, this application embodiment can provide a computer storage medium for implementation. The computer storage medium stores computer program instructions; when these computer program instructions are executed by a processor, they implement any of the face image filtering methods in the above embodiments.
[0331] This application also provides a computer program product, including a computer program that, when executed, implements any of the methods for filtering facial images described in the above embodiments.
[0332] It should be clarified that this application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of this application is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of this application.
[0333] The functional blocks shown in the above-described block diagram can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, they can be, for example, electronic circuits, application-specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of this application are programs or code segments used to perform the required tasks. Programs or code segments can be stored on a machine-readable medium or transmitted over a transmission medium or communication link via data signals carried on a carrier wave. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, fiber optic media, radio frequency (RF) links, etc. Code segments can be downloaded via computer networks such as the Internet, intranets, etc.
[0334] It should also be noted that the exemplary embodiments mentioned in this application describe methods or systems based on a series of steps or apparatus. However, this application is not limited to the order of the above steps; that is, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.
[0335] The aspects of this disclosure have been described above with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block in the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that these instructions, executable via the processor of the computer or other programmable data processing apparatus, enable the implementation of the functions / actions specified in one or more blocks of the flowchart illustrations and / or block diagrams. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field-programmable logic circuit. It is also understood that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can also be implemented by special-purpose hardware performing the specified functions or actions, or can be implemented by a combination of special-purpose hardware and computer instructions.
[0336] The above description is merely a specific implementation of this application. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, modules, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. It should be understood that the protection scope of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the protection scope of this application.
Claims
1. A method for screening facial images, characterized in that, include: Acquire multiple facial images; For each face image, perform steps A through E respectively: Step A: Detect facial key points and facial target areas from the face image; wherein, the facial key points are expression-independent key points, which correspond to stable regions in the human facial muscle anatomy. Step B: Calculate the first three-dimensional pose angle of the face image based on the facial key points; Step C: Based on the preset mapping relationship, perform error compensation correction on the first three-dimensional attitude angle to obtain the second three-dimensional attitude angle; wherein, the preset mapping relationship is a mapping function from the first three-dimensional attitude angle to the standard three-dimensional attitude angle established in advance by the curve fitting method; Step D: Based on the second three-dimensional pose angle, the key point coordinates of the target facial region are mapped to the virtual projection plane when the face is facing the camera through geometric transformation, so as to obtain the corrected key point coordinates of the target region. Step E: Calculate the target location state parameters based on the corrected key point coordinates of the target location; The target image is obtained by selecting face images from the multiple face images whose target part state parameters satisfy the first preset condition and whose second three-dimensional pose angles satisfy the second preset condition; The step of detecting facial key points from the face image includes: Collect image sequences containing multiple facial expressions of the same person in a fixed head posture; Facial landmark detection is performed on the image sequence to obtain the coordinate sequence of each facial landmark; For each facial key point, calculate its coordinate offset relative to the reference expression image in the image sequence; Based on the coordinate offset, calculate the total offset norm of each key point; According to the total offset norm in ascending order, select the top N key points to obtain expression-independent key points, and use these expression-independent key points as the facial key points.
2. The method according to claim 1, characterized in that, After performing error compensation correction on the first three-dimensional attitude angle based on a preset mapping relationship to obtain the second three-dimensional attitude angle, the method further includes: Based on the personalized parameters of the face, the second three-dimensional pose angle is linearly transformed to obtain the third three-dimensional pose angle; The step of mapping the key point coordinates of the target facial region to the virtual projection plane when the face is facing the camera through geometric transformation based on the second three-dimensional pose angle, to obtain the corrected key point coordinates of the target region, includes: Based on the third three-dimensional pose angle, the key point coordinates of the target facial region are mapped to the virtual projection plane when the face is facing the camera through geometric transformation, thereby obtaining the corrected key point coordinates of the target region.
3. The method according to claim 2, characterized in that, Before performing a linear transformation on the second three-dimensional pose angle based on the personalized parameters of the face to obtain the third three-dimensional pose angle, the method includes: Obtain multiple first images of the face corresponding to the face image; wherein, the first image includes a manually annotated fourth three-dimensional pose angle; Based on the first three-dimensional attitude angle of the first image, error compensation correction based on a preset mapping relationship is performed to obtain the second three-dimensional attitude angle of the first image; Linear regression fitting is performed on the second three-dimensional pose angle and the fourth three-dimensional pose angle of the first image to obtain the personalized parameters of the face.
4. The method according to claim 2, characterized in that, Before performing a linear transformation on the second three-dimensional pose angle based on the personalized parameters of the face to obtain the third three-dimensional pose angle, the method further includes: Obtain the feature vector of the face corresponding to the face image; Calculate the similarity between the feature vector and the preset face feature vector; When the similarity meets the third preset condition, the second three-dimensional pose angle is linearly transformed based on the personalized parameters of the face corresponding to the face image to obtain the third three-dimensional pose angle.
5. The method according to claim 1, characterized in that, The method further includes: Obtain a face sample set; the sample set includes multiple face samples, each face sample containing a second image of a face and its corresponding true three-dimensional pose angle; Based on the face sample set, the first three-dimensional pose angle of each face sample second image is calculated to obtain a data pair composed of the first three-dimensional pose angle and the standard three-dimensional pose angle; Based on the data pair, a mapping function from the first three-dimensional attitude angle to the standard three-dimensional attitude angle is established by curve fitting method, and this mapping function serves as the preset mapping relationship.
6. The method according to claim 4, characterized in that, The first preset condition is that the target part's state parameter is greater than a first preset level, and the second preset condition is that the second three-dimensional pose angle is greater than a second preset level. The step of selecting face images from the plurality of face images whose target part's state parameter satisfies the first preset condition and whose second three-dimensional pose angle satisfies the second preset condition to obtain a target image includes: Based on the preset target part status parameter grading range, the target part status level of the target part status parameter is determined. Based on a preset second three-dimensional attitude angle grading range, the second three-dimensional attitude level of the second three-dimensional attitude angle is determined; The target image is obtained by selecting face images from the multiple face images whose target part state level is greater than a first preset level and whose second three-dimensional pose level is greater than a second preset level.
7. The method according to claim 6, characterized in that, Before selecting face images from the plurality of face images that satisfy the first preset condition for the state parameters of the target part and the second three-dimensional pose angle satisfies the second preset condition, and obtaining the target image, the method further includes: Based on multiple third images containing a human face, the numerical distribution of the target region state parameters and the second three-dimensional pose angle is calculated; The numerical distribution of the target part state parameters is divided to obtain the preset target part state parameter grading interval; The numerical distribution of the second three-dimensional attitude angle is divided to obtain the preset target part state parameter grading interval.
8. The method according to claim 7, characterized in that, The multiple third images are faces corresponding to the preset facial feature vectors. The step of calculating the numerical distribution of the target region state parameters and the second three-dimensional pose angle based on the multiple third images containing the faces includes: Based on multiple third images of the face corresponding to the preset facial feature vector, the numerical distribution of the target part state parameters and the second three-dimensional pose angle is calculated.
9. The method according to claim 1, characterized in that, The target region includes the eye. Before calculating the state parameters of the target region based on the corrected key point coordinates of the target region, the method further includes: When the person in the face image is wearing glasses, a pre-trained eye state classification model is used to classify the eye region in the face image to obtain eye state information; wherein, the eye state information includes whether the eyes are open or closed. When the eye state information indicates that the eyes are open, the target part state parameters are calculated based on the corrected target part key point coordinates.
10. The method according to claim 1, characterized in that, The target region includes the mouth, and the target region state parameters include mouth shape parameters. The calculation of the target region state parameters based on the corrected key point coordinates of the target region includes: Obtain the coordinates of at least one pair of key points at the corners of the mouth that are symmetrical to each other and located on both sides of the midline of the face; Calculate the mouth width based on the coordinates of the pair of key points at the corners of the mouth; Obtain the coordinates of at least one pair of key points located on the midline of the upper and lower lips; The degree of mouth opening and closing is calculated based on the coordinates of the key points in the pair of lips; The ratio of the degree of mouth opening to the width of the mouth is calculated and used as the mouth shape parameter.
11. A facial image screening device, characterized in that, The device includes: The image acquisition module is used to acquire multiple face images; The image processing module is used to perform processing on each face image separately, including: A key point detection unit is used to detect facial key points and facial target areas from the face image; wherein, the facial key points are expression-independent key points, and the expression-independent key points correspond to stable regions in the human facial muscle anatomy. A pose angle calculation unit is used to calculate the first three-dimensional pose angle of the face image based on the facial key points; An error compensation unit is used to perform error compensation and correction on the first three-dimensional attitude angle based on a preset mapping relationship to obtain a second three-dimensional attitude angle; wherein, the preset mapping relationship is a mapping function from the first three-dimensional attitude angle to the standard three-dimensional attitude angle established in advance by a curve fitting method; The geometric transformation unit is used to map the key point coordinates of the target part of the face to the virtual projection plane when the face is facing the camera based on the second three-dimensional pose angle, so as to obtain the corrected key point coordinates of the target part. The state parameter calculation unit is used to calculate the state parameters of the target part based on the corrected coordinates of the key points of the target part. The image filtering module is used to filter out face images from the multiple face images that satisfy the first preset condition for the state parameters of the target part and the second three-dimensional pose angle satisfies the second preset condition, so as to obtain the target image; The key point detection unit includes: The image sequence acquisition subunit is used to acquire image sequences containing multiple facial expressions of the same person in a fixed head posture. The key point detection subunit is used to perform facial key point detection on the image sequence and obtain the coordinate sequence of each facial key point. Offset calculation subunit is used to calculate the coordinate offset of each facial key point in the image sequence relative to the reference expression image; The norm calculation subunit is used to calculate the total offset norm of each key point based on the coordinate offset. The key point selection sub-unit is used to select the top N key points in order of increasing total offset norm to obtain expression-independent key points, and to use the expression-independent key points as the facial key points.
12. An electronic device, characterized in that, The electronic device includes: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, it implements the face image filtering method as described in any one of claims 1-10.
13. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer program instructions, which, when executed by a processor, implement the face image filtering method as described in any one of claims 1-10.
14. A computer program product, characterized in that, When the instructions in the computer program product are executed by the processor of the electronic device, the electronic device causes the electronic device to perform the face image screening method as described in any one of claims 1-10.