Face recognition task complexity analysis method and system

By performing detailed annotation and complexity calculation on the test data set of the face recognition algorithm, the problem of lack of quantitative evaluation methods in the existing technology is solved, and the precise quantification of the complexity of face recognition tasks and the objectivity and accuracy of algorithm performance evaluation are achieved.

CN119942609APending Publication Date: 2025-05-06BEIJING GUOKE FUNDAMENTAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411942083.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-26
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The prior art lacks objective and quantitative methods to evaluate the difficulty of the test task in facial recognition algorithm testing, resulting in deviations in algorithm performance evaluation.

Method used

By marking the bounding frames of the eyes, mouth and mouth corners of the test data set of the face recognition algorithm, combining the key points, width and height values ​​and intersection positions of the bounding frame, the posture, eye expression and mouth expression complexity of the sample is calculated, and the overall complexity is comprehensively calculated.

Benefits of technology

Accurate quantification of the complexity of face recognition tasks is achieved, evaluation deviations caused by subjective differences in labeling personnel are avoided, and objectivity and accuracy of evaluation are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942609A_ABST
    Figure CN119942609A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of AI algorithm development and evaluation, and provides a face recognition task complexity analysis method, which comprises the following steps: S1, labeling all samples in a test data set of a test task to be subjected to a face recognition algorithm, including labeling bounding boxes for eyes and mouth, and performing mouth corner connection; s2, acquiring a plurality of auxiliary lines based on key points in the left eye bounding box, the right eye bounding box and the mouth bounding box, and calculating the attitude complexity of the sample through included angles of the plurality of auxiliary lines; s3, calculating the eye expression complexity based on the width and height values of the left eye bounding box and the right eye bounding box, and calculating the mouth expression complexity based on the intersection point position of the mouth corner connecting line and the mouth bounding box; and S4, comprehensively considering the attitude complexity, the eye expression complexity and the mouth expression complexity, and calculating the overall complexity of the sample. The test task difficulty can be accurately quantified, algorithm performance evaluation deviation caused by subjective difference of labeling personnel is avoided, and evaluation objectivity and accuracy are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of AI algorithm development and evaluation, and in particular to a method and system for analyzing the complexity of face recognition tasks. Background Art

[0002] In today's era of rapid technological development, face recognition technology has shown extremely broad application prospects in many fields, such as security monitoring, access control systems, financial payment verification, and social entertainment. With the continuous expansion and diversification of application scenarios, higher and higher requirements are placed on the accuracy, reliability, and efficiency of face recognition algorithms.

[0003] In the development and evaluation of face recognition algorithms, the testing phase plays a pivotal role. As the key basis for evaluating algorithm performance, the quality and characteristics of the test set directly affect the effectiveness and credibility of the test results. A well-designed test set that comprehensively covers various situations can accurately reflect the performance of the algorithm in complex real-world environments. Otherwise, it may lead to a one-sided or even wrong evaluation of the algorithm's performance.

[0004] However, there is a significant problem in the field of face recognition algorithm testing. When evaluating the algorithm, people often only focus on the test results themselves, while seriously ignoring the difficulty of the test task itself. Currently, existing technical means mostly rely on the subjective judgment of the annotators to determine the difficulty of the test task. For example, in the difficulty assessment of the test task of face image data, the annotators only rely on their own experience and intuition to judge the impact of factors such as the posture changes of the face in the image, lighting conditions, and occlusion on the test difficulty. This subjective judgment method has many disadvantages.

[0005] On the one hand, due to factors such as personal experience, professional knowledge level, and subjective cognition, different annotators often have difficulty reaching a consensus on the difficulty of the same test task. For example, for a set of image test sets containing facial posture changes at a specific angle, an experienced annotator who is good at handling such situations may think that the difficulty is moderate, while an annotator with relatively less experience may overestimate the difficulty. This inconsistency makes the difficulty assessment of the test task lack objectivity and accuracy, which in turn leads to deviations in the evaluation of algorithm performance.

[0006] On the other hand, subjective judgment cannot provide a quantitative measurement method. This means that it is difficult to accurately measure the degree of difference in difficulty between different test tasks. For example, when comparing two groups of test sets that focus on different lighting conditions and facial occlusion, it is impossible to know exactly how much more difficult one set of test sets is than the other, and in which specific dimensions there are differences in difficulty based solely on subjective judgment. Without this quantitative analysis, it is impossible to gain an in-depth understanding of the specific performance of the algorithm when dealing with different difficulty factors, which is not conducive to algorithm developers to optimize and improve the algorithm in a targeted manner.

[0007] In summary, in the field of face recognition task complexity analysis, there is an urgent need for a method that can objectively and quantitatively evaluate the difficulty of the test task, so as to avoid misjudgment of the face recognition algorithm capabilities due to ignoring the difficulty of the test task, thereby promoting the further development and improvement of face recognition technology. Summary of the invention

[0008] In view of the above problems, the purpose of the present invention is to provide a method and system for analyzing the complexity of face recognition tasks, which can accurately quantify the difficulty of test tasks, avoid the deviation of algorithm performance evaluation caused by subjective differences of labelers, and improve the objectivity and accuracy of evaluation. It can effectively promote the development of face recognition technology and improve the reliability and efficiency of face recognition algorithms in complex application scenarios.

[0009] The above-mentioned object of the present invention is achieved through the following technical solutions:

[0010] A method for analyzing the complexity of a face recognition task comprises the following steps:

[0011] S1: All samples in the test data set for the face recognition algorithm test task are labeled, including the bounding boxes of the eyes and mouth and the lines connecting the mouth corners;

[0012] S2: obtaining a number of auxiliary lines based on key points in the left eye bounding box, the right eye bounding box, and the mouth bounding box, and calculating the posture complexity of the sample through the angles of the auxiliary lines;

[0013] S3: calculating the complexity of eye expression based on the width and height values ​​of the left eye bounding box and the right eye bounding box, and calculating the complexity of mouth expression based on the intersection position of the mouth corner line and the mouth bounding box;

[0014] S4: comprehensively considering the posture complexity, the eye expression complexity, and the mouth expression complexity, and calculating the overall complexity of the sample.

[0015] Furthermore, in step S1, the bounding boxes of the eyes and mouth are marked as follows:

[0016] For all samples in the test data set, a cascade classifier based on Haar features pre-trained on a large amount of face image data is used to detect the face area, and the image portion containing only the face is cropped;

[0017] Constructing a Haar feature template of the eye, sliding windows of different sizes in the face area and calculating Haar feature values, using a cascade classifier trained with the AdaBoost algorithm to classify and determine whether it is an eye area, and determining the left eye bounding box and the right eye bounding box according to the position and size of the detected eye area window;

[0018] A pixel-based classifier is constructed using the skin color and texture differences between the mouth area and the surrounding skin. The color and texture features are analyzed pixel by pixel in the face area to mark the possible mouth pixel area. The marked area is clustered to preliminarily determine the mouth position range. Based on the mouth being a horizontal strip, the preliminary mouth bounding box is determined according to the minimum circumscribed rectangle of the clustered mouth area. The mouth bounding box is further adjusted based on the geometric relationship of the mouth being located within a certain proportional position range below the nose in the face.

[0019] Furthermore, in step S1, the mouth corners are marked, including the line connecting the mouth corners, specifically:

[0020] Constructing a YCrCb skin color model, traversing pixels in the cropped face area according to the value range of the Cr and Cb channels in the YCrCb skin color model, judging and marking skin color pixels, preliminarily determining the skin color area including the corners of the mouth, performing edge detection on the preliminarily screened skin color area using the Canny edge detection algorithm, determining edge pixels by calculating the image gradient and connecting them into an edge contour, analyzing the shape features of the edge contour, calculating the curvature of each point on the contour, and selecting points with larger curvature and meeting the shape features of the corners of the mouth as candidate points of the corners of the mouth;

[0021] The left and right mouth corners are determined according to the left-right positional relationship of the mouth corner candidate points, and the left and right mouth corners are connected using a least squares straight line fitting algorithm to obtain a straight line equation for the mouth corner connection line.

[0022] Further, in step S2, a plurality of auxiliary lines are obtained based on the key points in the left eye bounding box, the right eye bounding box, and the mouth bounding box, and the posture complexity of the sample is calculated by the angles of the plurality of auxiliary lines, specifically:

[0023] The posture complexity is to analyze whether the face is facing the camera. The less facing the camera, the higher the posture complexity. The centers of the left eye bounding box, the right eye bounding box, and the mouth bounding box are taken as the key points to obtain a plurality of auxiliary lines.

[0024] If both the left eye bounding box and the right eye bounding box exist, the pose complexity is:

[0025] Connect the centers of the left eye bounding box and the right eye bounding box to obtain an auxiliary line L1;

[0026] Connect the centers of the right eye bounding box and the mouth bounding box to obtain an auxiliary line L2;

[0027] Connect the centers of the left eye bounding box and the mouth bounding box to obtain an auxiliary line L2;

[0028] Draw a perpendicular line to the auxiliary line L1 through the center of the mouth bounding box, the intersection angle between the perpendicular line and the auxiliary line L2 is α, and the intersection angle between the perpendicular line and the auxiliary line L3 is β;

[0029] Calculate the posture complexity λ p for:

[0030] λ p =2|α-β| / π

[0031] If there is only one of the left eye bounding box and the right eye bounding box, the posture complexity is:

[0032] λ p =1.

[0033] Further, in step S3, the eye expression complexity is calculated based on the width and height values ​​of the left eye bounding box and the right eye bounding box, specifically:

[0034] The eye expression complexity λ e for:

[0035] λ e =|the aspect ratio of the right eye bounding box−the aspect ratio of the left eye bounding box| / (the aspect ratio of the right eye bounding box+the aspect ratio of the left eye bounding box).

[0036] Further, in step S3, the complexity of mouth expression is calculated based on the intersection position of the mouth corner line and the mouth boundary box, specifically:

[0037] Assume that the distance from the intersection of the mouth corner line and the left border of the mouth bounding box to the upper edge of the mouth bounding box is b1, and the distance from the intersection of the mouth corner line and the left border of the mouth bounding box is b2;

[0038] Assume that the distance from the intersection of the mouth corner line and the right border of the mouth bounding box to the upper edge of the mouth bounding box is b3, and the distance from the intersection of the mouth corner line and the right border of the mouth bounding box is b4;

[0039] The mouth expression complexity λ m for:

[0040] λ m =(|b1-b2|+|b3-b4|) / (b1+b2+b3+b4).

[0041] Further, in step S4, the overall complexity of the sample is calculated by comprehensively considering the posture complexity, the eye expression complexity, and the mouth expression complexity, specifically:

[0042] The weights of the posture complexity, the eye expression complexity, and the mouth expression complexity are set to A, B, and C respectively, and the overall complexity is:

[0043] λ=A*λ p +B*λ e +C*λ m .

[0044] A face recognition task complexity analysis system for executing the above-mentioned face recognition task complexity analysis method comprises:

[0045] The sample initialization annotation module is used to annotate all samples in the test data set for the face recognition algorithm test task, including annotating the bounding boxes of the eyes and mouth and connecting the corners of the mouth;

[0046] A posture complexity calculation module is used to obtain a number of auxiliary lines based on key points in the left eye bounding box, the right eye bounding box, and the mouth bounding box, and calculate the posture complexity of the sample through the angles of the auxiliary lines;

[0047] An expression complexity calculation module, used to calculate the eye expression complexity based on the width and height values ​​of the left eye bounding box and the right eye bounding box, and calculate the mouth expression complexity based on the intersection position of the mouth corner line and the mouth bounding box;

[0048] The overall complexity calculation module is used to comprehensively consider the posture complexity, the eye expression complexity, and the mouth expression complexity to calculate the overall complexity of the sample.

[0049] A computer device comprises a memory and one or more processors, wherein the memory stores computer codes, and when the computer codes are executed by the one or more processors, the one or more processors execute the above method.

[0050] A computer-readable storage medium stores computer codes. When the computer codes are executed, the above method is executed.

[0051] Compared with the prior art, the present invention has at least one of the following beneficial effects:

[0052] (1) Simplified labeling method, saving labor costs: The present invention only needs to label the eyes, mouth boundary boxes and mouth corners of the samples in the test data set. Compared with the traditional complex facial feature labeling, this clear and focused labeling method greatly simplifies the labeling process and reduces the difficulty and workload of the labelers. Without the need for professional and tedious training, ordinary people can quickly get started, which significantly saves labor costs and improves the efficiency of labeling work.

[0053] (2) Providing quantitative standards to improve evaluation accuracy: The present invention provides quantitative standards for evaluating the complexity of face recognition tasks through a series of clear quantitative calculation methods. In terms of posture complexity calculation, auxiliary lines are obtained based on the key points of the bounding box and calculated through the angle, so that posture changes can be measured with specific values; the complexity of eye expression is calculated based on the width and height of the left and right eye bounding boxes, and the complexity of mouth expression is determined by the intersection of the line connecting the mouth corners and the mouth bounding box. These quantitative indicators avoid the ambiguity and inconsistency of subjective judgment. The overall complexity of the sample is calculated by combining these quantitative complexity indicators, which significantly improves the accuracy and scientificity of the evaluation of the complexity of face recognition tasks, and provides a reliable basis for subsequent algorithm performance evaluation and optimization.

[0054] (3) Low computational complexity, convenient for wide application: The calculation method adopted by the entire analysis method of the present invention is mainly based on simple information such as the key points, width and height values, and intersection positions of the bounding box, and the results are obtained through relatively simple geometric calculations and logical operations. Compared with some complex deep learning algorithms, this method has low computational complexity, does not require high performance of computing equipment, and can run quickly on ordinary computer devices. This makes the method more adaptable and scalable, and is convenient for wide use in different research and application scenarios, effectively promoting the application and development of complexity analysis of face recognition tasks in practical work. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] Figure 1 It is an overall flow chart of the complexity analysis method of face recognition task of the present invention;

[0056] Figure 2 A schematic diagram of face annotation of the present invention;

[0057] Figure 3 It is the overall structure diagram of the complexity analysis system of face recognition task of the present invention. DETAILED DESCRIPTION

[0058] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0059] Those skilled in the art will appreciate that, unless otherwise stated, the singular forms "a", "an", "said" and "the" used herein may also include plural forms. It should be further understood that the term "comprising" used in the specification of the present invention refers to the presence of the features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.

[0060] First embodiment

[0061] like Figure 1 As shown, this embodiment provides a method for analyzing the complexity of a face recognition task, comprising the following steps:

[0062] S1: All samples in the test data set for the test task of the face recognition algorithm are labeled, including marking the bounding boxes of the eyes and mouth and connecting the corners of the mouth.

[0063] In this embodiment, the marking includes marking the bounding boxes of the eyes and mouth and connecting the corners of the mouth. Figure 2 The following are specific examples of the technical solutions for marking the bounding boxes of the eyes and mouth and connecting the corners of the mouth:

[0064] (1) Label the bounding boxes for the eyes and mouth

[0065] For all samples in the test data set, a cascade classifier based on Haar features pre-trained on a large amount of face image data is used to detect the face area, and the image part containing only the face is cropped; a Haar feature template of the eye is constructed, windows of different sizes are slid in the face area and Haar feature values ​​are calculated, and a cascade classifier trained by the AdaBoost algorithm is used to classify and determine whether it is an eye area, and the left eye bounding box and the right eye bounding box are determined according to the position and size of the detected eye area window; a pixel-based classifier is constructed using the skin color and texture differences between the mouth area and the surrounding skin, color and texture features are analyzed pixel by pixel in the face area to mark possible mouth pixel areas, and the marked area is clustered to preliminarily determine the mouth position range, and according to the mouth being a horizontal strip, the preliminary mouth bounding box is determined according to the minimum circumscribed rectangle of the mouth area obtained by clustering, and then the mouth bounding box is further adjusted in combination with the geometric relationship that the mouth is located in a certain proportional position range below the nose in the face.

[0066] For the above technical solution, the process can be described as follows:

[0067] S111: Data preprocessing

[0068] Collect image datasets containing faces in various poses. Normalize the images to unify the image size, such as adjusting to 300×300 pixels, to facilitate subsequent operations. At the same time, convert the images to grayscale to reduce the amount of calculation and highlight the texture features of the image.

[0069] S112: Face Detection and Localization

[0070] Face detection is performed using a cascade classifier based on Haar features. The classifier is pre-trained on a large amount of face image data and can quickly locate the face area in the image.

[0071] The detected face area is cropped to extract the image part containing only the face, so as to subsequently detect the eyes and mouth within the face range.

[0072] S113: Eye Detection and Bounding Box Annotation

[0073] 1. Eye detection based on Haar features

[0074] Construct a Haar feature template for the eyes. The eye region usually has a specific grayscale change pattern, such as darker eyes and brighter surrounding skin. By sliding windows of different sizes in the face region and calculating the Haar feature values ​​in the window, the cascade classifier trained by the AdaBoost algorithm is used to classify the window to determine whether it is an eye region.

[0075] When possible eye regions are detected, a preliminary annotated bounding box of the eye is determined based on the position and size of the window.

[0076] 2. Fine-tuning based on geometric constraints

[0077] Considering the geometric structure of the face, under normal circumstances, the eyes have a certain degree of symmetry in the horizontal direction and the distance is relatively stable. According to the position of one detected eye and the approximate width ratio of the face, the annotated bounding box of the other eye is adjusted and verified to improve the detection accuracy.

[0078] For the annotated bounding box of a single eye, it is further corrected according to the common aspect ratio range of the eye (such as 1:2 to 1:3) to ensure that the complete eye area is framed.

[0079] S114: Mouth detection and bounding box annotation

[0080] 1. Detection based on skin color and texture features

[0081] The skin color of the mouth area is somewhat different from the surrounding skin and has a unique texture. These features are used to build a pixel-based classifier. Within the face area, the color and texture features are analyzed pixel by pixel, and the pixel area that may belong to the mouth is marked.

[0082] Cluster analysis is performed on the marked pixel areas, and adjacent pixels with similar characteristics are aggregated into regions to preliminarily determine the location range of the mouth.

[0083] 2. Determination of annotation bounding box based on shape constraints

[0084] The mouth is usually in the shape of a horizontal strip, and its preliminary annotation bounding box is determined based on the minimum enclosing rectangle of the mouth area obtained by clustering.

[0085] Taking into account the relative position of the mouth in the human face, which is generally located within a certain proportion below the nose, this geometric relationship is used to further adjust and optimize the annotated bounding box of the mouth to ensure that the mouth is accurately framed.

[0086] S115: Post-processing and verification

[0087] Check the overlap of the labeled eye and mouth bounding boxes. If there is severe overlap or unreasonable box position relationship, make corrections or re-detections based on prior knowledge and face geometry.

[0088] Randomly select some of the annotation results for manual verification to check the accuracy of the annotation. If errors or inaccurate annotations are found, analyze the reasons and adjust the corresponding detection parameters or algorithms, and then reprocess the entire data set until the annotation accuracy reaches an acceptable level.

[0089] (2) Annotate the mouth corners, including the line connecting them. Specifically:

[0090] A YCrCb skin color model is constructed, pixels are traversed in the cropped face area according to the value range of Cr and Cb channels in the YCrCb skin color model, skin color pixels are judged and marked, and the skin color area including the corners of the mouth is preliminarily determined. The Canny edge detection algorithm is used to perform edge detection on the preliminarily screened skin color area, edge pixels are determined by calculating image gradients and connected into edge contours, the shape features of the edge contours are analyzed, the curvatures of each point on the contours are calculated, and points with larger curvatures and meeting the shape features of the corners of the mouth are taken as candidate corners of the mouth; the left corner of the mouth and the right corner of the mouth are determined according to the left-right positional relationship of the candidate corners of the mouth, and the left corner of the mouth and the right corner of the mouth are connected using the least squares straight line fitting algorithm to obtain the straight line equation of the mouth corner connection line.

[0091] For the above technical solution, the process can be described as follows:

[0092] S121: Mouth corner detection

[0093] 1. Preliminary screening based on skin color model

[0094] Construct a skin color model, such as a YCrCb skin color model, and determine skin color pixels based on the value range of the Cr and Cb channels in the model. In the cropped face area, traverse each pixel to determine whether it is a skin color pixel, mark the points that belong to skin color pixels, and preliminarily obtain the skin color area that may contain the corners of the mouth.

[0095] 2. Mouth corner positioning based on edge detection

[0096] The Canny edge detection algorithm can be used to detect the edge of the initially selected skin color area. The algorithm calculates the image gradient, determines the edge pixels, and connects the edge pixels to form the edge contour.

[0097] Analyze the shape characteristics of the edge contour. The corners of the mouth are usually located in the concave part of the edge contour and have a certain curvature change. By calculating the curvature of each point on the edge contour, find the point with larger curvature and conforming to the shape characteristics of the corner of the mouth (such as a relatively sharp concave) as the candidate point of the corner of the mouth.

[0098] S122: Mouth corner connection

[0099] From the mouth corner candidate points, the left and right mouth corners are determined according to the left and right position relationship. Generally, the mouth corner candidate point on the left side of the image is the left mouth corner, and the one on the right side is the right mouth corner.

[0100] The determined left and right mouth corners are connected using a straight line fitting algorithm, such as the least squares method. The goal of the least squares method is to find a straight line that minimizes the sum of the distances from all data points (here, the mouth corner points) to the straight line, thereby obtaining the straight line equation of the mouth corner connection line.

[0101] S123: Post-processing and verification

[0102] Check the rationality of the mouth corner connection line and the face contour. If the connection line intersects with the detected face contour or exceeds the face area, it means that there may be a misdetection, and the mouth corner detection and connection operation need to be repeated.

[0103] Manually annotate some images for verification, and compare the mouth corner lines generated by the algorithm with the manually annotated results. Calculate the error indicators between the two, such as the average distance error or the intersection over union (IoU). If the error indicator exceeds the set threshold, analyze the cause, which may be due to false detection caused by factors such as lighting changes and extreme postures. Adjust the algorithm parameters or improve the algorithm steps for these problems, and then reprocess the entire data set until satisfactory accuracy is achieved.

[0104] S2: Obtain several auxiliary lines based on key points in the left eye bounding box, the right eye bounding box, and the mouth bounding box, and calculate the posture complexity of the sample through the angles of the auxiliary lines.

[0105] In this embodiment, step S2 is specifically as follows:

[0106] The posture complexity is to analyze whether the face is facing the camera. The less facing the camera, the higher the posture complexity. The centers of the left eye bounding box, the right eye bounding box, and the mouth bounding box are taken as the key points to obtain a plurality of auxiliary lines.

[0107] If both the left eye bounding box and the right eye bounding box exist, the pose complexity is:

[0108] Connect the centers of the left eye bounding box and the right eye bounding box to obtain an auxiliary line L1;

[0109] Connect the centers of the right eye bounding box and the mouth bounding box to obtain an auxiliary line L2;

[0110] Connect the centers of the left eye bounding box and the mouth bounding box to obtain an auxiliary line L2;

[0111] Draw a perpendicular line to the auxiliary line L1 through the center of the mouth bounding box, the intersection angle between the perpendicular line and the auxiliary line L2 is α, and the intersection angle between the perpendicular line and the auxiliary line L3 is β;

[0112] Calculate the posture complexity λ p for:

[0113] λ p =2|α-β| / π

[0114] If there is only one of the left eye bounding box and the right eye bounding box, the posture complexity is:

[0115] λ p =1.

[0116] S3: Calculate the complexity of eye expression based on the width and height values ​​of the left eye bounding box and the right eye bounding box, and calculate the complexity of mouth expression based on the intersection position of the mouth corner line and the mouth bounding box.

[0117] In this embodiment, the eye expression complexity is calculated based on the width and height values ​​of the left eye bounding box and the right eye bounding box, specifically:

[0118] The eye expression complexity λ e for:

[0119] λ e =|the aspect ratio of the right eye bounding box−the aspect ratio of the left eye bounding box| / (the aspect ratio of the right eye bounding box+the aspect ratio of the left eye bounding box).

[0120] In this embodiment, the complexity of mouth expression is calculated based on the intersection position of the mouth corner line and the mouth boundary box, specifically:

[0121] Assume that the distance from the intersection of the mouth corner line and the left border of the mouth bounding box to the upper edge of the mouth bounding box is b1, and the distance from the intersection of the mouth corner line and the left border of the mouth bounding box is b2;

[0122] Assume that the distance from the intersection of the mouth corner line and the right border of the mouth bounding box to the upper edge of the mouth bounding box is b3, and the distance from the intersection of the mouth corner line and the right border of the mouth bounding box is b4;

[0123] The mouth expression complexity λ m for:

[0124] λ m =(|b1-b2|+|b3-b4|) / (b1+b2+b3+b4).

[0125] S4: comprehensively considering the posture complexity, the eye expression complexity, and the mouth expression complexity, and calculating the overall complexity of the sample.

[0126] The weights of the posture complexity, the eye expression complexity, and the mouth expression complexity are set to A, B, and C respectively, and the overall complexity is:

[0127] λ=A*λ p+B*λ e +C*λ m .

[0128] For example, we can take A = 0.5, B = 0.25, C = 0.25, then the overall complexity is:

[0129] λ=0.5*λ p +0.25*λ e +0.25*λ m

[0130] The larger the λ value, the higher the overall complexity.

[0131] The present invention is applicable to scenarios where, when testing a face recognition intelligent algorithm, not only the performance of the intelligent algorithm needs to be considered, but also a comprehensive evaluation needs to be performed in combination with the complexity of the task. The present invention provides a method for measuring the complexity of the task.

[0132] It should be noted that, in addition to the above labeling schemes, the present invention can also label multiple points on the face, such as the left and right eye corners, eyeball points, left and right side points of the nose, nose vertices, left and right corners of the mouth, upper and lower vertices of the mouth, etc., and then analyze each point. The scheme is exemplified as follows:

[0133] (1) Left and right eye corners

[0134] Analyzing posture complexity:

[0135] By connecting the left and right eye corners, a line segment can be obtained. The angle between this line segment and the horizontal direction can reflect the degree of inclination of the face in the horizontal direction. For example, if the angle between this line segment and the horizontal direction is large, it means that the face may have a large left-right rotation, thereby increasing the complexity of the posture.

[0136] When forming a triangle with other key points (such as the left and right corners of the mouth), the shape and angle changes of the triangle can also reflect the complexity of the posture. For example, when a face turns from the front to the side, the quadrilateral formed by the left and right corners of the eye and the left and right corners of the mouth will be significantly deformed. By analyzing this deformation, the complexity of the posture can be quantified.

[0137] Analyzing expression complexity:

[0138] Observe the change in the distance between the left and right eye corners. When a person's expression changes, such as smiling or frowning, the muscles around the eyes will contract or relax, causing the distance between the left and right eye corners to change. The greater the change in distance, the richer the eye expression and the more complex the expression.

[0139] (2) Eyeball Point

[0140] Analyzing posture complexity:

[0141] The position of the eyeball point can reflect the orientation of the face. For example, by observing the position of the eyeball point relative to the eye socket, it can be determined whether the face is looking straight ahead, looking askance, or turning the head. If the eyeball point is more inclined to one side of the eye socket, it may indicate that the face has turned at a larger angle, increasing the complexity of the posture.

[0142] Compare the relative positions of the left and right eyeballs. When the face changes posture, the relative positions of the left and right eyeballs may change. By calculating this position change, the posture complexity can be measured.

[0143] Analyzing expression complexity:

[0144] The direction and amplitude of eyeball rotation are closely related to facial expressions. For example, when a person is surprised, their eyes may be wide open and rotate to the sides; when a person is focused, their eyes may focus on a certain point. By analyzing the direction and amplitude of eyeball rotation, the complexity of eye expression can be assessed.

[0145] (3) The left and right side points and the vertex of the nose

[0146] Analyzing posture complexity:

[0147] By connecting the left and right points of the nose, we can get a line segment in the width direction of the nose. The change in the angle between this line segment and the horizontal direction can reflect the change in the posture of the face in the horizontal direction.

[0148] Taking the nose vertex as the reference point, the geometric shape and angle changes formed by other key points (such as eye or mouth key points) can be used to judge posture complexity. For example, the triangle formed by the nose vertex and the left and right eye corners will change in shape and angle when the face posture changes. The posture complexity is evaluated by quantifying this change.

[0149] Analyzing expression complexity:

[0150] When a person's expression changes, the muscles around the nose will move to a certain extent, resulting in slight changes in the relative positions of the left and right points and the vertex of the nose. For example, when laughing, the nose may expand, and measuring the change in the distance between the left and right points of the nose can reflect the complexity of this expression change.

[0151] (4) The left and right corners of the mouth and the upper and lower vertices of the mouth

[0152] Analyzing posture complexity:

[0153] By connecting the left and right points of the mouth corners, we can get a line segment in the width direction of the mouth. The change in the angle between this line segment and the horizontal direction can be used to determine the change in the horizontal posture of the face.

[0154] The shape changes of the quadrilateral formed by the left and right points of the mouth corners and the upper and lower vertices of the mouth in different postures are compared to quantify the posture complexity. For example, when the face turns from the front to the side, the quadrilateral will be obviously deformed.

[0155] Analyzing expression complexity:

[0156] The change in the distance between the left and right points of the mouth corners directly reflects the changes in the expression of the mouth. For example, when smiling, the corners of the mouth will rise, and the distance between the left and right points of the mouth corners will increase; when frowning, the corners of the mouth may fall down, and the distance will decrease. By analyzing this distance change, the complexity of the mouth expression can be evaluated.

[0157] The change in distance between the upper and lower vertices of the mouth can also reflect the expression. For example, when the mouth is surprised, it may open and the distance between the upper and lower vertices will increase. By measuring this distance change, the complexity of the expression can be quantified.

[0158] Second embodiment

[0159] like Figure 3 As shown, this embodiment provides a face recognition task complexity analysis system for executing the face recognition task complexity analysis method of the first embodiment, including:

[0160] The sample initialization labeling module 1 is used to label all samples in the test data set for the test task of the face recognition algorithm, including labeling the bounding boxes of the eyes and mouth and connecting the corners of the mouth.

[0161] A posture complexity calculation module 2 is used to obtain a plurality of auxiliary lines based on key points in the left eye bounding box, the right eye bounding box, and the mouth bounding box, and calculate the posture complexity of the sample through the angles of the plurality of auxiliary lines;

[0162] Expression complexity calculation module 3, used to calculate the complexity of eye expression based on the width and height values ​​of the left eye bounding box and the right eye bounding box, and calculate the complexity of mouth expression based on the intersection position of the mouth corner line and the mouth bounding box;

[0163] The overall complexity calculation module 4 is used to comprehensively consider the posture complexity, the eye expression complexity, and the mouth expression complexity to calculate the overall complexity of the sample.

[0164] A computer-readable storage medium stores a computer code. When the computer code is executed, the above method is executed. A person skilled in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing related hardware through a program. The program can be stored in a computer-readable storage medium. The storage medium can include: a read-only memory (ROM), a random access memory (RAM), a disk or an optical disk, etc.

[0165] The above is only a preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions under the concept of the present invention belong to the protection scope of the present invention. It should be pointed out that for ordinary technicians in this technical field, some improvements and modifications without departing from the principle of the present invention should also be regarded as the protection scope of the present invention.

[0166] The technical features of the above-described embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0167] It should be noted that the above embodiments can be freely combined as needed. The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered as the protection scope of the present invention.

Claims

1. A method for analyzing the complexity of face recognition tasks, characterized in that: The following steps are involved: S1: All samples in the test data set for the face recognition algorithm test task are labeled, including the bounding boxes of the eyes and mouth and the lines connecting the mouth corners; S2: obtaining a number of auxiliary lines based on key points in the left eye bounding box, the right eye bounding box, and the mouth bounding box, and calculating the posture complexity of the sample through the angles of the auxiliary lines; S3: calculating the complexity of eye expression based on the width and height values ​​of the left eye bounding box and the right eye bounding box, and calculating the complexity of mouth expression based on the intersection position of the mouth corner line and the mouth bounding box; S4: comprehensively considering the posture complexity, the eye expression complexity, and the mouth expression complexity, and calculating the overall complexity of the sample.

2. The method for analyzing the complexity of face recognition tasks according to claim 1, characterized in that: In step S1, the bounding boxes of the eyes and mouth are marked as follows: For all samples in the test data set, a cascade classifier based on Haar features pre-trained on a large amount of face image data is used to detect the face area, and the image portion containing only the face is cropped; Constructing a Haar feature template of the eye, sliding windows of different sizes in the face area and calculating Haar feature values, using a cascade classifier trained with the AdaBoost algorithm to classify and determine whether it is an eye area, and determining the left eye bounding box and the right eye bounding box according to the position and size of the detected eye area window; A pixel-based classifier is constructed using the skin color and texture differences between the mouth area and the surrounding skin. The color and texture features are analyzed pixel by pixel in the face area to mark the possible mouth pixel area. The marked area is clustered to preliminarily determine the mouth position range. Based on the mouth being a horizontal strip, the preliminary mouth bounding box is determined according to the minimum circumscribed rectangle of the clustered mouth area. The mouth bounding box is further adjusted based on the geometric relationship of the mouth being located within a certain proportional position range below the nose in the face.

3. The method for analyzing the complexity of face recognition tasks according to claim 1, characterized in that: In step S1, the mouth corners are marked, including the line connecting the mouth corners, as follows: Constructing a YCrCb skin color model, traversing pixels in the cropped face area according to the value range of the Cr and Cb channels in the YCrCb skin color model, judging and marking skin color pixels, preliminarily determining the skin color area including the corners of the mouth, performing edge detection on the preliminarily screened skin color area using the Canny edge detection algorithm, determining edge pixels by calculating the image gradient and connecting them into an edge contour, analyzing the shape features of the edge contour, calculating the curvature of each point on the contour, and selecting points with larger curvature and meeting the shape features of the corners of the mouth as candidate points of the corners of the mouth; The left and right mouth corners are determined according to the left-right positional relationship of the mouth corner candidate points, and the left and right mouth corners are connected using a least squares straight line fitting algorithm to obtain a straight line equation for the mouth corner connection line.

4. The method for analyzing the complexity of face recognition tasks according to claim 1, characterized in that: In step S2, a plurality of auxiliary lines are obtained based on the key points in the left eye bounding box, the right eye bounding box, and the mouth bounding box, and the posture complexity of the sample is calculated by the angles of the auxiliary lines, specifically: The posture complexity is to analyze whether the face is facing the camera. The less facing the camera, the higher the posture complexity. The centers of the left eye bounding box, the right eye bounding box, and the mouth bounding box are taken as the key points to obtain a plurality of auxiliary lines. If both the left eye bounding box and the right eye bounding box exist, the pose complexity is: Connect the centers of the left eye bounding box and the right eye bounding box to obtain an auxiliary line L1; Connect the centers of the right eye bounding box and the mouth bounding box to obtain an auxiliary line L2; Connect the centers of the left eye bounding box and the mouth bounding box to obtain an auxiliary line L2; Draw a perpendicular line to the auxiliary line L1 through the center of the mouth bounding box, the intersection angle between the perpendicular line and the auxiliary line L2 is α, and the intersection angle between the perpendicular line and the auxiliary line L3 is β; Calculate the posture complexity λ p for: l p =2|α-β| / π If there is only one of the left eye bounding box and the right eye bounding box, the posture complexity is: l p =1.

5. The method for analyzing the complexity of face recognition tasks according to claim 4, characterized in that: In step S3, the eye expression complexity is calculated based on the width and height values ​​of the left eye bounding box and the right eye bounding box, specifically: The eye expression complexity λ e for: λ e =|the aspect ratio of the right eye bounding box−the aspect ratio of the left eye bounding box| / (the aspect ratio of the right eye bounding box+the aspect ratio of the left eye bounding box).

6. The method for analyzing the complexity of face recognition tasks according to claim 5, characterized in that: In step S3, the complexity of mouth expression is calculated based on the intersection position of the mouth corner line and the mouth boundary box, specifically: Assume that the distance from the intersection of the mouth corner line and the left border of the mouth bounding box to the upper edge of the mouth bounding box is b1, and the distance from the intersection of the mouth corner line and the left border of the mouth bounding box is b2; Assume that the distance from the intersection of the mouth corner line and the right border of the mouth bounding box to the upper edge of the mouth bounding box is b3, and the distance from the intersection of the mouth corner line and the right border of the mouth bounding box is b4; The mouth expression complexity λ m for: <h2 style=";text-align:left;direction:ltr">λ<h2 style=";text-align:left;direction:ltr"> m <h2 style=";text-align:left;direction:ltr"> (|b1-b2|+|b3-b4|) / (b1+b2+b3+b4) 7. The method for analyzing the complexity of face recognition tasks according to claim 6, characterized in that: In step S4, the overall complexity of the sample is calculated by comprehensively considering the posture complexity, the eye expression complexity, and the mouth expression complexity, specifically: The weights of the posture complexity, the eye expression complexity, and the mouth expression complexity are set to A, B, and C respectively, and the overall complexity is: λ=A*λ p +B*λ e +C*λ m 。 8. A face recognition task complexity analysis system for executing the face recognition task complexity analysis method according to any one of claims 1 to 7, characterized in that: include: The sample initialization annotation module is used to annotate all samples in the test data set for the face recognition algorithm test task, including annotating the bounding boxes of the eyes and mouth and connecting the corners of the mouth; A posture complexity calculation module is used to obtain a number of auxiliary lines based on key points in the left eye bounding box, the right eye bounding box, and the mouth bounding box, and calculate the posture complexity of the sample through the angles of the auxiliary lines; An expression complexity calculation module, used to calculate the eye expression complexity based on the width and height values ​​of the left eye bounding box and the right eye bounding box, and calculate the mouth expression complexity based on the intersection position of the mouth corner line and the mouth bounding box; The overall complexity calculation module is used to comprehensively consider the posture complexity, the eye expression complexity, and the mouth expression complexity to calculate the overall complexity of the sample.

9. A computer device comprising a memory and one or more processors, wherein the memory stores computer codes, and when the computer codes are executed by the one or more processors, the one or more processors execute the method according to any one of claims 1 to 7. 10 . A computer-readable storage medium storing a computer code. When the computer code is executed, the method according to claim 1 is executed.