Face image classification method and device, electronic equipment, medium and program product

By employing a multi-layer feature fusion and dynamic classification threshold adjustment mechanism, the problem of insufficient generalization ability of face image classification systems under long-tailed distributions is solved, achieving high recognition accuracy and robustness in complex environments and improving user experience.

CN121330747APending Publication Date: 2026-01-13INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511738257.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-25
Publication Date
2026-01-13

AI Technical Summary

Technical Problem

Existing face image classification systems struggle to learn the feature distribution of rare categories under long-tailed distributions, resulting in insufficient generalization ability in complex scenarios, manifested as classification result bias, confidence fluctuations, and misclassification.

Method used

By introducing a multi-layer feature fusion and dynamic classification threshold adjustment mechanism, a pre-trained deep feature extraction model is used to extract multi-layer feature information, and the classification boundary is dynamically adjusted by calculating attention weights and confidence distribution, thereby realizing multi-layer feature expression and adaptive decision-making.

Benefits of technology

It improves the recognition accuracy and robustness of face image classification systems in complex environments, reduces false recognition and rejection, and provides a faster, smoother and more accurate face classification experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121330747A_ABST
    Figure CN121330747A_ABST
Patent Text Reader

Abstract

The invention provides a face image classification method and device, electronic equipment, a medium and a program product, and can be applied to the technical field of artificial intelligence and the field of financial science and technology. The method comprises the steps that a to-be-classified face image is acquired, multi-layer feature information of the face image is extracted based on a pre-trained depth feature extraction model, and the multi-layer feature information comprises feature mapping results output by different layers of the depth feature extraction model; performing fusion processing on the multi-layer feature information to obtain fused feature representation; the fusion feature representation is input into a classification model, multiple face categories and corresponding confidence distribution are output, a dynamic classification threshold value is calculated according to the confidence distribution, and the dynamic classification threshold value is used for adjusting classification boundaries of the multiple face categories; and obtaining the category of the face image based on the plurality of face categories and the dynamic classification threshold.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence and the technical field of financial technology, and more particularly to a face image classification method, device, equipment, medium and program product. BACKGROUND

[0002] As an important branch of computer vision, face image classification technology has been widely used in identity recognition, image retrieval, financial risk control and other scenarios. However, in a real data environment, face image samples generally exhibit long-tail distribution characteristics, that is, a small number of common categories and standard pose image samples are sufficient in quantity, while a large number of rare categories, special poses, extreme illumination, partially occluded or low definition samples are scarce, resulting in problems such as sample sparsity, class imbalance, high noise and distribution tailing. Due to the limited number and unstable quality of tail samples, the classification model is difficult to fully learn the feature distribution of the tail categories in the training process, resulting in insufficient generalization ability for rare categories or complex scenarios in the classification stage, which manifests as classification result deviation, confidence fluctuation and error classification. Especially under multi-source data collection or different device imaging conditions, the feature distribution difference further amplifies the long-tail effect, making the features of the same category of face images inconsistent under different scenarios, thereby affecting the stability and generalization performance of the face image classification system. SUMMARY

[0003] In view of the above problems, the present application provides a face image classification method, device, equipment, medium and program product.

[0004] According to a first aspect of the present application, a face image classification method is provided, the method comprising: obtaining a face image to be classified, extracting multi-layer feature information of the face image based on a pre-trained deep feature extraction model, wherein the multi-layer feature information comprises feature mapping results output by different layers of the deep feature extraction model; performing fusion processing on the multi-layer feature information to obtain a fused feature representation; inputting the fused feature representation into a classification model to output multiple face categories and corresponding confidence distributions, calculating a dynamic classification threshold according to the confidence distributions, wherein the dynamic classification threshold is used to adjust the classification boundaries of the multiple face categories; and obtaining the category of the face image based on the multiple face categories and the dynamic classification threshold.

[0005] According to an embodiment of the present application, the deep feature extraction model comprises a shallow convolutional layer, a middle pooling layer and a deep attention layer; the multi-layer feature information comprises local texture features extracted by the shallow convolutional layer, structure features extracted by the middle pooling layer and semantic features extracted by the deep attention layer.

[0006] According to an embodiment of the present application, the method further comprises: performing fusion processing on the multi-layer feature information by using a dynamic weighting mechanism based on attention weights to obtain a fused feature representation, wherein the fusion weights corresponding to the multi-layer feature information are adjusted according to the clarity and local saliency information of the face image.

[0007] According to an embodiment of the present application, the deep feature extraction model comprises a spatial feature extraction branch and a frequency domain feature extraction branch, wherein the spatial feature extraction branch is used to extract spatial texture and structure features of the face image; and the frequency domain feature extraction branch is used to obtain spectral features of the face image through fast Fourier transform.

[0008] According to an embodiment of the present application, the method further comprises: calculating a noise intensity map based on the pixel gradient distribution of the face image; and generating a feature mask according to the noise intensity map, and adjusting the fusion weights corresponding to the multi-layer feature information according to the feature mask.

[0009] According to an embodiment of the present application, the method further comprises: performing quantile calculation on the confidence distribution corresponding to each face class to obtain a target quantile point; determining the confidence value corresponding to the target quantile point as an initial dynamic classification threshold; and adjusting the initial dynamic classification threshold according to the sample number of each face class in the training data to obtain the dynamic classification threshold.

[0010] According to an embodiment of the present application, the method further comprises: sorting the confidence distribution in descending order to obtain a confidence sequence; selecting the first k target confidence classes in the confidence sequence, calculating the mean of the confidence of the target confidence classes as a main threshold, and k is a positive integer; calculating the confidence dispersion based on the remaining confidence classes in the confidence sequence except the target confidence classes to obtain an auxiliary threshold; and combining the main threshold and the auxiliary threshold to generate the dynamic classification threshold.

[0011] According to an embodiment of the present application, the method further comprises: performing distribution statistics on multiple face classes and corresponding confidence distributions to determine a target sparse class and a corresponding confidence interval; and calculating a class compensation factor based on the confidence interval, and adjusting the classification boundaries of multiple face classes based on the class compensation factor and the dynamic classification threshold.

[0012] According to an embodiment of the present application, the method further comprises: recording a confidence distribution sequence corresponding to adjacent multi-frame face images in a continuous classification task; calculating a time fluctuation rate of the confidence distribution sequence, and comparing the time fluctuation rate with a preset stability threshold; and in response to the time fluctuation rate exceeding the preset stability threshold, performing timing correction on the dynamic classification threshold.

[0013] The second aspect of the present application provides a face image classification device, the device comprising: a feature extraction module, configured to: obtain a face image to be classified, and extract multi-layer feature information of the face image based on a pre-trained deep feature extraction model, wherein the multi-layer feature information comprises feature mapping results output by different layers of the deep feature extraction model; a fusion processing module, configured to: perform fusion processing on the multi-layer feature information to obtain a fused feature representation; a dynamic threshold optimization module, configured to: input the fused feature representation into a classification model to output multiple face categories and corresponding confidence distributions, and calculate a dynamic classification threshold according to the confidence distributions, wherein the dynamic classification threshold is used to adjust the classification boundaries of the multiple face categories; and a classification module, configured to: obtain the category of the face image based on the multiple face categories and the dynamic classification threshold.

[0014] The third aspect of the present application provides an electronic device, comprising: one or more processors; a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the above method.

[0015] The fourth aspect of the present application further provides a computer-readable storage medium having a computer program or instructions stored thereon, wherein the computer program or instructions are executed by a processor to implement the steps of the above method.

[0016] The fifth aspect of the present application further provides a computer program product comprising a computer program or instructions, wherein the computer program or instructions are executed by a processor to implement the steps of the above method.

[0017] According to the embodiments of the present application, by introducing the multi-layer feature fusion and dynamic classification threshold adjustment mechanism in the face image classification process, multi-level expression of features is realized, and the strategy of self-adaptive adjustment of classification boundary combining confidence distribution is combined, so that the model can dynamically optimize the judgment standard for different class distribution and sample features. The embodiments of the present application not only improve the integrity and discriminability of feature representation, but also effectively alleviate the problems of long-tail distribution, sample imbalance and confidence bias existing in face data, and can maintain high recognition accuracy and robustness in various practical scenes such as complex light, posture change, expression difference and sample sparsity, thereby improving the generalization ability and stability of the face image classification system in diversified environment. At the same time, the misrecognition and rejection phenomenon in the classification process is reduced, so that the user can obtain a faster, smoother and more accurate face classification experience in the use process. BRIEF DESCRIPTION OF DRAWINGS

[0018] The above and other objects, features and advantages of the present application will become more apparent from the following description of the embodiments of the present application taken in conjunction with the accompanying drawings, in which:

[0019] Figure 1 An application scenario diagram of a face image classification method, apparatus, device, medium and program product according to embodiments of the present application is schematically shown;

[0020] Figure 2 A flowchart of a face image classification method according to embodiments of the present application is schematically shown;

[0021] Figure 3 A flowchart of a method of calculating a dynamic classification threshold according to some example embodiments of the present application is schematically shown;

[0022] Figure 4 A flowchart of a method of calculating a dynamic classification threshold according to some example embodiments of the present application is schematically shown;

[0023] Figure 5 A structural block diagram of a face image classification apparatus according to embodiments of the present application is schematically shown; and

[0024] Figure 6 A block diagram of an electronic device suitable for implementing a face image classification method according to embodiments of the present application is schematically shown. DETAILED DESCRIPTION

[0025] Embodiments of the present application will be described below with reference to the accompanying drawings. However, it should be understood that the description is merely exemplary and is not intended to limit the scope of the present application. In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the embodiments of the present application. However, it will be apparent to one skilled in the art that one or more embodiments can be practiced without these specific details. In addition, in the following description, descriptions of well-known structures and techniques have been omitted to avoid unnecessarily obscuring the concepts of the present application.

[0026] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the present application. As used herein, the term "includes" and tautological expressions thereof, means the inclusion of the stated features, steps, operations, and / or elements but not to the exclusion of one or more other features, steps, operations, and / or elements.

[0027] All terms used herein, including technical and scientific terms, have the same meaning as commonly understood by one of ordinary skill in the art unless otherwise defined herein. It should be noted that the terms used herein should be interpreted as having a meaning that is consistent with the context of the specification, and should not be interpreted in an idealized or overly formal manner.

[0028] In the case of using expressions similar to "at least one of A, B, and C, etc.", it should generally be interpreted to include any of one, all, or a combination thereof unless otherwise specified. For example, "a system having at least one of A, B, and C" should be interpreted to include a system having A alone, a system having B alone, a system having C alone, a system having A and B together, a system having A and C together, a system having B and C together, and / or a system having A, B, and C together, etc.

[0029] Face image classification technology, as an important research direction in the field of computer vision, has been widely applied in security monitoring, behavior analysis, emotion recognition, user portrait, and human-computer interaction scenarios. Existing face image classification systems are mostly based on deep learning models, which extract discriminative deep feature vectors by training a large number of labeled samples, and calculate similarity or perform classification discrimination in the feature space to distinguish different types or attributes of face images. However, in real application environments, face image data generally presents obvious long-tail distribution characteristics, i.e., a small number of common categories (such as main viewing angle, positive illumination, and standard posture) have sufficient sample quantities, while a large number of rare categories (such as deflection angle, extreme illumination, partial occlusion, low definition, or special expression) have extremely small sample quantities, forming a typical long-tail dataset.

[0030] Long-tailed face image data has the characteristics of sample sparsity, class imbalance, high noise, and distribution tailing. Due to the limited number of tail class samples and unstable data quality, the model is difficult to fully learn the feature distribution of the tail class during training, resulting in a significant lack of discrimination ability for rare types or extreme scenarios during the classification stage. At the same time, tail samples are often accompanied by high noise, such as blur, polarization or occlusion, which makes manual labeling difficult and increases the error rate, further reducing the effectiveness of feature learning. In addition, the deep model tends to learn the features of the head class during the learning process, making the representation ability of the tail features insufficient, and the decision boundary of the model biased towards the high-frequency class, resulting in an imbalance in the classification results. Especially in complex environments, such as backlight, low illumination or expression change conditions, the stability of the features extracted by the model decreases, the confidence of class discrimination fluctuates, and the classification accuracy decreases.

[0031] Therefore, the characteristics of long-tailed data become a key bottleneck restricting the robustness and generalization performance of the face image classification system. Existing models perform well on head classes, but the classification accuracy decreases in rare classes, low-quality samples and few-sample scenarios, resulting in an imbalance in overall performance.

[0032] Based on this, the embodiments of the present application provide a face image classification method, which comprises: obtaining a face image to be classified, extracting multi-layer feature information of the face image based on a pre-trained deep feature extraction model, wherein the multi-layer feature information includes feature mapping results output by different layers of the deep feature extraction model; performing fusion processing on the multi-layer feature information to obtain a fused feature representation; inputting the fused feature representation into a classification model to output multiple face classes and corresponding confidence distributions, and calculating a dynamic classification threshold based on the confidence distributions, wherein the dynamic classification threshold is used to adjust the classification boundary of the multiple face classes; and obtaining the class of the face image based on the multiple face classes and the dynamic classification threshold. According to the embodiments of the present application, by introducing a multi-layer feature fusion and dynamic classification threshold adjustment mechanism in the face image classification process, multi-level feature expression is realized, and the strategy of adaptively adjusting the classification boundary combined with the confidence distribution is used, so that the model can dynamically optimize the decision criteria for different class distributions and sample features. The embodiments of the present application not only improve the completeness and discriminability of the feature representation, but also effectively alleviate the problems of long-tailed distribution, sample imbalance and confidence bias in face data, and can maintain high recognition accuracy and robustness in complex lighting, posture change, expression difference and sample sparsity, etc. multiple practical scenarios, thereby improving the generalization ability and stability of the face image classification system in a diversified environment. At the same time, it also reduces the misrecognition and rejection phenomenon in the classification process, so that users can obtain a faster, smoother and more accurate face classification experience in the use process.

[0033] It should be noted that the face image classification method, device, equipment, medium and program product provided by the present application can be used in the field of artificial intelligence technology and the field of financial technology, and can also be used in various fields other than the field of artificial intelligence technology and the field of financial technology. The application field of the face image classification method, device, equipment, medium and program product provided by the embodiments of the present application is not limited.

[0034] In the technical solutions of the present application, the user information (including but not limited to user personal information, user image information, user device information such as location information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved are information and data authorized by the user or authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of related data comply with relevant laws, regulations and standards, necessary security measures are taken, do not violate public order and good customs, and provide corresponding operation portal for user to choose authorization or refusal.

[0035] In the scenario of using personal information for automated decision-making, the method, device and system provided by the embodiments of the present application all provide corresponding operation portal for the user to choose to agree or refuse the automated decision-making result; if the user chooses to refuse, the expert decision-making process is entered. The expression "expert decision-making" here refers to the decision-making activities of personnel who are engaged in a certain field of work, have special experience, knowledge and skills, and reach a certain professional level.

[0036] Figure 1 The application scenario diagram of the face image classification method, device, equipment, medium and program product according to the embodiments of the present application is schematically shown.

[0037] As shown in Figure 1 The application scenario 100 according to the embodiments can include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104 and a server 105. The network 104 is used as a medium to provide a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103 and the server 105. The network 104 can include various connection types, such as wired, wireless communication links or optical fiber cables, etc.

[0038] The user can use the first terminal device 101, the second terminal device 102, the third terminal device 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications can be installed on the first terminal device 101, the second terminal device 102, the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only as examples).

[0039] In the embodiments of the present application, the first terminal device 101 can be taken as an example of the first device, and the second terminal device 102 and / or the third terminal device 103 can be taken as an example of at least one second device. The first device and the second device can communicate with each other through a client internal mechanism, to implement the data distribution and rendering logic in the face image classification method.

[0040] In some embodiments, the first device and the at least one second device can be different display modules, windows or screens on the same computing terminal (such as a host computer), or can be multiple physical devices working cooperatively through a network, for example, different client instances deployed on a desktop computer, a tablet terminal or a mobile device respectively.

[0041] The first terminal device 101, the second terminal device 102 and the third terminal device 103 can be various electronic devices with display screens and supporting web browsing, including but not limited to smart mobile terminals, tablet computers, laptop computers and desktop computers, etc.

[0042] The server 105 can be a server providing various services, for example, a background management server supporting a website browsed by a user using the first terminal device 101, the second terminal device 102 and the third terminal device 103 (only as an example). The background management server can analyze and process received user requests and other data, and feed back the processing results (such as a webpage, information or data generated or obtained according to a user request) to the terminal device.

[0043] It should be noted that the face image classification method provided in the embodiments of the present application can generally be executed by the server 105. Correspondingly, the face image classification apparatus provided in the embodiments of the present application can generally be arranged in the server 105. The face image classification method provided in the embodiments of the present application can also be executed by a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105. Correspondingly, the face image classification apparatus provided in the embodiments of the present application can also be arranged in a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105.

[0044] It should be understood that Figure 1 The number of terminal devices, networks and servers in the above scenario is only illustrative. According to the implementation needs, there can be any number of terminal devices, networks and servers.

[0045] The following will be described based on Figure 1 the scenario described above, through Figures 2 to 4The face image classification method of the disclosed embodiment is described in detail.

[0046] Figure 2 A flowchart of the face image classification method according to the embodiment of the application is schematically shown.

[0047] As shown in Figure 2 The face image classification method 200 of the embodiment includes operations S210-S240.

[0048] In operation S210, a face image to be classified is acquired, and multi-layer feature information of the face image is extracted based on a pre-trained deep feature extraction model, wherein the multi-layer feature information includes feature mapping results output by different layers of the deep feature extraction model.

[0049] In the embodiment of the application, the face image to be classified can be received from a camera, a terminal or a database. The face image can be a single frame of static image or a continuous frame image in a video stream. In order to ensure the quality and consistency of the input image, an image preprocessing operation can be performed before feature extraction, such as face detection, geometric correction, illumination normalization, size adjustment and color standardization on the input face image, to eliminate differences caused by different acquisition devices and environmental factors. The preprocessed face image can be input into the pre-trained deep feature extraction model to obtain multi-level feature information. The deep feature extraction model can be a convolutional neural network, a visual Transformer or a hybrid structure network, which has learned general face feature expression ability based on a large-scale face dataset in the pre-training stage.

[0050] In some embodiments, the multi-layer structure of the deep feature extraction model can include a shallow convolutional layer, a middle pooling layer and a deep attention layer. The shallow convolutional layer is mainly responsible for capturing local texture information of the face, such as eye corners, lip edges, skin details, etc.; the middle pooling layer is used to extract geometric structure features of the face, such as face shape contour, relative position relationship of facial features; the deep attention layer can model semantic features and identity-related patterns of the face from a global perspective, thereby forming a more discriminative high-level expression in the feature space.

[0051] In another embodiment, the deep feature extraction model can be adjusted according to different application scenarios. For example, in a high-resolution monitoring scenario, a feature extraction model based on a residual network can be preferred to fully retain high-frequency details; while in a mobile terminal or real-time recognition scenario, a lightweight network structure can be used to reduce computational complexity while ensuring recognition accuracy. In addition, the pre-training of the model can use a public face recognition dataset or an enterprise self-built dataset to obtain a feature distribution more consistent with the target application domain.

[0052] Further, to enhance the adaptability of the model to complex environments, the system can normalize or standardize the multi-layer feature output. For example, the feature mapping results of different layers can be standardized using L2 normalization, batch normalization, or feature rescaling mechanisms to ensure that the features of each layer remain consistent in numerical scale, thereby avoiding feature dominance effects in subsequent fusion processes. The system can also introduce a feature pyramid or multi-scale receptive field in the feature extraction stage to adapt to different resolutions of face regions, making the extracted features more robust in scale.

[0053] In yet another embodiment, the deep feature extraction model can also incorporate a frequency domain analysis mechanism. Specifically, a fast Fourier transform or discrete wavelet transform can be performed on the image after the shallow or middle layer feature output to extract frequency spectrum features representing high-frequency detail information and low-frequency structure distribution in the image. The frequency domain features can be combined with the spatial domain features to form part of the multi-layer feature information, thereby improving the feature expression ability of the model under conditions of varying illumination, noise interference, or partial occlusion.

[0054] In some embodiments, the deep feature extraction model can be configured as a dynamic structure, i.e., automatically selecting different levels of feature output as part of the multi-layer feature information according to the clarity, resolution, or scene conditions of the input face image. For example, in low-light or high-noise scenarios, the system can preferentially extract middle and deep layer features to obtain stable structure and semantic information; in high-quality image scenarios, the system can retain shallow layer detail features to improve discrimination and accuracy.

[0055] In operation S220, fusion processing is performed on the multi-layer feature information to obtain a fused feature representation.

[0056] In embodiments of the present application, after completing multi-layer feature extraction, fusion processing can be performed on the multi-layer feature information to obtain a globally consistent and discriminative fused feature representation. Illustratively, the feature mapping results from different levels of the deep feature extraction model typically have semantic level differences: shallow layer features contain local information such as texture and edges, middle layer features reflect structure and shape patterns, and deep layer features embody high-dimensional semantic information related to identity. To fully utilize the complementary relationship between these multi-layer features, the system can align, normalize, and weight combine the features of each layer to achieve information integration through layer-by-layer fusion. The fused feature representation maintains global semantic consistency while retaining detailed features, enabling the subsequent classification model to more accurately distinguish different face categories in complex scenarios.

[0057] In some embodiments, the fusion process can be implemented by combining feature concatenation and dimension reduction. The system first concatenates the feature mapping results from different layers in the channel dimension to form a high-dimensional composite feature tensor. Since direct concatenation can lead to excessively high feature dimensions and increase the computational cost, the system further adopts methods such as principal component analysis, linear discriminant analysis, or autoencoder dimension reduction to compress the concatenated features into a unified feature space, so as to retain the main information and remove redundant features. Through the fusion method of concatenation and dimension reduction, the system can reduce the computational complexity while ensuring the integrity of the information, thereby achieving efficient inference.

[0058] In other embodiments, to improve the adaptability of fusion, the system can introduce a dynamic weight allocation mechanism. Under this mechanism, the fusion weights of the features of different layers can be automatically adjusted according to the quality indicators of the input face images.

[0059] In some embodiments, the fusion process can be based on feature correlation modeling. The system can calculate the correlation matrix between the features of different layers to measure the complementarity and similarity of the outputs of different layers in the feature space, and allocate fusion weights based on this. For example, when the shallow and middle layer features have high correlation but differ greatly from the deep layer features, the system can preferentially strengthen the fusion between the shallow and deep layers, thereby introducing more semantic difference information.

[0060] In another embodiment, the fusion process can be extended to the frequency domain to enhance the detail resolution capability of the features. The system can perform fast Fourier transform or discrete wavelet transform on the feature mappings of different layers to obtain frequency domain feature representations, and calculate the fusion weights according to the spectral energy distribution in the frequency domain space.

[0061] To further improve the stability of the fusion results, normalization and regularization strategies can also be introduced in the feature fusion stage. Specifically, batch normalization or layer normalization can be performed on the features of different layers in the channel dimension or spatial dimension to reduce the scale difference between the outputs of different layers. The system can also introduce L2 regularization or feature sparsification constraints to suppress the influence of noise features on the fusion results.

[0062] In another embodiment, the fusion process can also enable cross-layer information interaction based on graph neural networks or attention mechanisms. The system can regard the feature mappings of different layers as nodes to construct a feature correlation graph, propagate context information through graph convolution operations, and thereby realize deep fusion between the feature layers. Alternatively, self-attention mechanisms can be used to globally weight the multi-layer features, dynamically highlighting key feature channels by calculating the similarity matrix between the features. Such structures can capture complex cross-layer dependencies, further enhancing the semantic consistency and global perception of the fused features.

[0063] In operation S230, the fusion feature representation is input into a classification model, and multiple face categories and corresponding confidence distributions are output. A dynamic classification threshold is calculated according to the confidence distributions, and the dynamic classification threshold is used to adjust the classification boundaries of the multiple face categories.

[0064] In an embodiment of the present application, after obtaining the fusion feature representation, the fusion feature can be input into a pre-constructed classification model to output multiple face categories and corresponding confidence distributions. The classification model can be a multi-classification structure of a deep neural network or a support vector machine. Specifically, the fusion feature is first projected to a fixed-dimensional classification space through a fully connected layer or a feature mapping layer, and then the output probability distribution of each category is calculated to obtain the confidence value corresponding to each face category. Through statistical analysis of the confidence value, the recognition confidence degree of the classification model for different categories can be represented.

[0065] In a traditional method, a fixed threshold is often used to determine the classification boundary, which is difficult to cope with the case where the confidence difference between categories is significant under long-tail distribution. In an embodiment of the present application, the system realizes adaptive decision by calculating a dynamic classification threshold, so that the model can automatically adjust the recognition boundary according to the feature distribution of different categories, thereby improving the discrimination accuracy under the condition of sample imbalance.

[0066] In some embodiments, the calculation of the confidence distribution can be optimized in combination with the hierarchical contribution of multi-layer features. The system can introduce a hierarchical weighting mechanism in the classification model to dynamically divide the confidence weight according to the discrimination ability of different hierarchical features. For example, in the scene where shallow texture information contributes more to the classification result, the confidence proportion of the corresponding channel can be improved; and in the complex recognition task dominated by deep semantic features, the confidence output of the high semantic layer can be preferentially referred to.

[0067] In other embodiments, the calculation process of the dynamic classification threshold can be based on the statistical features of the confidence distribution. The system can perform quantile analysis, mean-variance calculation or probability density estimation on the output confidence distribution to obtain the confidence interval of each face category. According to the distribution features of different category samples, the system can adaptively determine the dynamic threshold corresponding to each category. For example, for high-frequency categories with concentrated confidence, the system can set a higher classification threshold to reduce false recognition; and for long-tail categories with sparse samples or fuzzy features, the threshold can be reduced to improve the recall rate.

[0068] To improve the accuracy and stability of threshold adjustment, dynamic correction can also be made based on the time series of confidence distribution or task context. For example, in continuous video recognition or real-time monitoring scenarios, the system can record the confidence change curve corresponding to adjacent multiple frames of face images, and calculate the confidence fluctuation rate. When the confidence fluctuation exceeds the preset stable interval, the system automatically corrects the dynamic classification threshold, so that it tends to the current confidence trend, thereby maintaining stable classification judgment results.

[0069] In another embodiment, the system can also calibrate the dynamic classification threshold in combination with the category statistical characteristics. For the sparse categories with extremely small sample size or highly skewed confidence distribution, the system can calculate a category compensation factor based on the category frequency information in the training set, and introduce the factor into the threshold calculation formula to correct the confidence offset. For example, if a certain category appears extremely low frequency in the training set, it may present systematic low confidence in the inference stage. The system can increase the recognition sensitivity of this category by increasing the corresponding compensation factor.

[0070] In some embodiments, the determination of the dynamic classification threshold can also be optimized in combination with the uncertainty estimation mechanism. The system can evaluate the uncertainty of the current recognition result by calculating the entropy value of the classification model output or the confidence interval based on Monte Carlo random inactivation. When the model presents high uncertainty in the classification confidence distribution of a certain input sample, the system can correspondingly increase the classification threshold to avoid false recognition; while when the model confidence is concentrated and the output is stable, the threshold is reduced to improve the response speed.

[0071] In another embodiment, to achieve efficient real-time response, the system can perform the calculation of the dynamic classification threshold and the classification inference process in parallel. Specifically, while the classification model generates the confidence distribution, another processing module can update the threshold estimate in real time based on the statistical information of the previous input, so as to directly output the final classification result after the inference is completed without additional delay.

[0072] In operation S240, a category of the face image is obtained based on the plurality of face categories and the dynamic classification threshold.

[0073] In the embodiments of the present application, after the classification model outputs the plurality of face categories and the corresponding confidence distribution, and determines the dynamic classification threshold, the final decision stage can be entered, i.e., determining the target category to which the face image belongs based on the plurality of face categories and the dynamic classification threshold. Specifically, the confidence value corresponding to each category can be compared with the dynamic classification threshold of the category, and when the confidence value of a certain category exceeds the corresponding threshold, it is determined that the category is the candidate recognition result. If there are multiple categories that meet the condition at the same time, the system can further compare the confidence difference or the confidence growth rate to determine the optimal category label.

[0074] In another embodiment, the system can employ a multi-stage decision mechanism to enhance the reliability of the discrimination. Specifically, a preliminary candidate class set is first screened according to the comparison result of the confidence and the dynamic threshold, and then a secondary decision is performed in the set based on the feature similarity between classes, spatial distribution consistency or historical recognition records. For example, in a video stream or multi-frame image scene, the system can combine the recognition results of the previous and subsequent frames to calculate the time continuity score of the class confidence, so as to improve the stability of the final decision. When it is detected that the confidence is continuously higher than the threshold in multiple adjacent frames, the system confirms that the class is the final recognition result; when the confidence fluctuates greatly, the decision is delayed and waits for subsequent frame information to supplement.

[0075] In some embodiments, to further optimize the user experience of recognition, the system can introduce a confidence fusion and explainability analysis mechanism before the final class output. Specifically, the system calculates the confidence interval or confidence level of the decision result at the same time, so as to quantify the reliability of the classification result. For example, when the target class confidence is far from the threshold, the system can mark the result as a high confidence level and directly output the recognition result; when the confidence is close to the threshold boundary, the system can mark it as a medium or low confidence result, and can prompt the user for secondary verification.

[0076] Further, the system can also dynamically update or self-correct the threshold strategy according to the stability of the recognition result. When it is detected that the continuous multiple recognition results fluctuate on the same class but the confidence is long-term in the threshold critical region, the system can automatically fine-tune the dynamic threshold corresponding to the class, so that the subsequent decision is more consistent with the actual confidence distribution. For example, for the tail class that is in a low confidence state for a long time but appears frequently, the system can appropriately lower the threshold of the class to improve the recall rate; and for the high-frequency class that is frequently misrecognized, the threshold can be appropriately increased to reduce the false matching.

[0077] In some embodiments, to support diversified face recognition scenarios, the system can introduce a multi-task decision mechanism in the final output stage. In addition to outputting the face class label, auxiliary indicators such as identity similarity, expression state or posture stability can also be generated simultaneously. These indicators can be calculated based on the intermediate layer output or feature fusion result of the classification model, and are used to evaluate the comprehensive reliability of the recognition result. For example, when a dramatic change in expression or a posture deflection is detected, the system can automatically reduce the weight of the final confidence output, so as to maintain the consistency of recognition in dynamic scenes.

[0078] According to the embodiments of the present application, the face image classification method can be widely applied in financial risk control and intelligent banking business, and is used for realizing identity verification, customer behavior monitoring and credit risk assessment based on face recognition. For example, in the bank counter and online credit approval scene, first, the face image of the customer to be classified is obtained, the multi-level texture, structure and semantic features are extracted through the pre-trained deep feature extraction model, and then the robust feature representation is obtained through fusion processing to cope with the feature differences under different light, posture and camera conditions. Subsequently, the system inputs the fused features into the classification model, outputs multiple face categories and corresponding confidence distribution, and calculates a dynamic classification threshold according to the uncertainty of customer identity recognition, so as to adaptively adjust the recognition boundary of high-risk and low-risk customers. In actual application, the dynamic threshold mechanism enables the system to automatically increase the verification standard for abnormal faces (such as abnormal posture or suspicious customers with multiple failed recognitions), while moderately relaxing the judgment boundary for frequently visited known customers, balancing safety and convenience. Finally, the system determines the customer identity category based on the dynamic classification result, and realizes business operations such as safe login, account unlocking or credit limit determination. This method not only improves the recognition accuracy of the financial system for long-tail face samples (such as a small number of access customers and edge lighting environment), but also enhances the sensitivity and response speed of the system to abnormal behavior through the dynamic threshold adjustment mechanism, thereby improving the overall operation safety and customer experience.

[0079] The face image classification method of the embodiments of the present application will be specifically described in the preferred embodiments.

[0080] In the embodiments of the present application, the deep feature extraction model includes a shallow convolutional layer, a middle pooling layer and a deep attention layer, each layer undertakes different semantic abstraction functions in the feature extraction process, and collectively constructs multi-level face feature representation from local to global and from low-dimensional to high-dimensional.

[0081] Specifically, the shallow convolutional layer is used to extract the local texture features of the input face image, and through local perception operation of multiple groups of small-scale convolution kernels on the image, it can capture facial detail information such as skin texture, eye corner lines, lip contour and other low-level features, thereby retaining the subtle differences of the appearance of the face; the middle pooling layer extracts the structural features of the face by performing down-sampling and spatial aggregation operations on the convolution feature map, reflecting the spatial relationship and geometric layout of the key parts of the face, such as the relative position between the features, the face shape contour and the proportion feature, thereby forming a more stable middle-level semantic representation; the deep attention layer further models the feature map in a global range, calculates the correlation between different regions through a self-attention mechanism, adaptively highlights the feature regions related to identity recognition, and suppresses the interference of non-key information such as light, posture and occlusion, thereby extracting high-level semantic features.

[0082] The above multi-layer features complement each other, the shallow layer emphasizes detail sensitivity, the middle layer provides structural stability, and the deep layer enhances semantic discrimination ability. Through the hierarchical feature extraction structure, the model can obtain a robust face feature representation in a complex environment, providing a high-quality feature basis for subsequent feature fusion and classification decision.

[0083] In the embodiments of the present application, a dynamic weighting mechanism based on attention weight can be used for adaptive fusion of multi-layer features. Specifically, the dynamic weighting mechanism based on attention weight can dynamically determine the weight distribution of shallow convolutional features, middle structural features and deep semantic features in the fusion process by calculating and analyzing the clarity and local saliency information of the input face image. The clarity information can be measured by image gradient, spectral energy or edge strength, etc., which is used to judge the degree of detail fidelity of the image; the local saliency information can be obtained by saliency detection or heat map calculation, which is used to identify the region containing key identity features in the face region.

[0084] For example, when the system detects that the image clarity is high, the weight of the shallow layer feature can be appropriately increased to fully retain the local texture details; when the image has blur, uneven illumination or partial occlusion, the fusion proportion of the middle and deep layer features can be increased to strengthen the stability of the global structure and semantic features. Through this dynamic weighting fusion method based on attention weight, the system can adaptively balance the importance of different layer features, so that the fused features achieve optimal coordination between detail preservation and semantic abstraction, thereby still outputting face feature representation with high discriminability and high robustness in complex environments.

[0085] In some embodiments, the deep feature extraction model can be set to include a spatial feature extraction branch and a frequency domain feature extraction branch. The spatial feature extraction branch mainly extracts features in the spatial domain of the image, and perceives and encodes the pixel distribution of the input face image through a multi-layer convolutional neural network to extract its spatial texture and structure features. This branch can capture local details of the face, such as wrinkles, skin texture, facial feature edges, etc., while modeling global geometric relationships, including eye distance, nose shape, jaw curve, etc. structural information, providing significant geometric stability and appearance consistency for identity recognition.

[0086] The frequency domain feature extraction branch is used to mine image information from the frequency space, and converts the input face image from the spatial domain to the frequency domain through Fast Fourier Transform (FFT). After transformation, the high-frequency component reflects the edge details and noise characteristics of the image, and the low-frequency component contains the overall illumination distribution and structural outline. The frequency domain feature extraction branch extracts stable face spectral features by designing convolution operations or attention modules in the frequency domain to analyze different frequency components.

[0087] In this embodiment, the spatial feature extraction branch and the frequency domain feature extraction branch can be trained collaboratively using a parallel structure. The system inputs the same face image into two separate feature extraction pathways. The spatial branch captures local and global texture patterns in the spatial domain, while the frequency domain branch learns periodic features and energy distribution in the frequency space. Subsequently, the output features of the two branches are aligned and synthesized in the fusion module to obtain a unified feature representation containing multi-scale, multi-frequency, and multi-semantic features.

[0088] To further enhance feature representation capabilities, the frequency domain feature extraction branch can incorporate a multi-channel filtering mechanism to extract features from signals in different frequency bands (such as high-frequency, mid-frequency, and low-frequency bands), and then obtain a complete spectral description through weighted fusion. For example, the system can use a bandpass filter to extract mid-frequency detail information to identify facial features with local variations; simultaneously, it can utilize low-frequency components to maintain the overall stability of the facial contour. Through this multi-band analysis structure, the frequency domain branch can more accurately distinguish non-identity factors caused by changes in expression, lighting, or resolution, thereby further enhancing the model's ability to capture true identity features.

[0089] In another embodiment, an information exchange mechanism can also exist between the spatial feature extraction branch and the frequency domain feature extraction branch. Specifically, the system can achieve bidirectional information transfer between spatial domain features and frequency domain features during the training phase through a mutual attention mechanism or a feature mapping projection module.

[0090] Through the above embodiments, the deep feature extraction model achieves feature modeling of face images from both spatial and frequency domains through the synergistic effect of spatial feature extraction branches and frequency domain feature extraction branches. This improves the generalization performance of the model under different shooting conditions, device environments, and long-tailed samples, thereby enhancing the richness and stability of feature representation.

[0091] In some embodiments, to improve the robustness and stability of face image feature fusion, a noise intensity map can be calculated based on the pixel gradient distribution of the face image, and a feature mask can be generated based on the noise intensity map to dynamically adjust the fusion weights of multi-layer feature information. Specifically, the system first performs pixel gradient calculation on the input face image to obtain the gradient magnitude distribution of each pixel in the image. The system establishes a noise intensity model based on these gradient magnitude features to evaluate the signal-to-noise ratio of different regions. For example, in regions with blurriness, compression artifacts, or illumination reflection, the gradient direction changes randomly and is unevenly distributed, resulting in higher noise intensity values; while in clear, uniform image regions, the gradient changes smoothly, resulting in lower noise intensity.

[0092] Subsequently, a feature mask can be generated based on the noise intensity map to dynamically adjust the weight allocation of multi-layer feature information during the feature fusion stage. This feature mask can be represented as a weighted matrix of the same size as the input feature map, with its element values ​​reflecting the reliability of the corresponding region. For regions with high noise intensity, the system reduces the weight of shallow features to weaken the interference of local texture noise on the overall fusion result; while for regions with low noise intensity and clear details, the weight of shallow features is increased accordingly to preserve effective texture information. Simultaneously, mid-layer and deep features achieve higher semantic stability under mask weighting, thereby ensuring the overall consistency of feature fusion.

[0093] In this embodiment, the generation and application of the feature mask can be performed synchronously with the deep feature extraction process. The system can introduce an explicit mask generation branch into the network structure, which automatically outputs the corresponding mask matrix based on the gradient statistics or noise-aware features of the feature map, and performs end-to-end training with the feature fusion module.

[0094] Through the above embodiments, the noise intensity estimation and feature mask generation mechanism based on pixel gradient distribution enables the system to perform differentiated feature processing on different image regions, thereby effectively suppressing local noise interference in the feature fusion stage and improving the effective feature ratio of face images.

[0095] In the embodiments of this application, the method of calculating the dynamic classification threshold based on the confidence distribution may include the following two schemes.

[0096] Figure 3 A flowchart illustrating a method for calculating a dynamic classification threshold according to some exemplary embodiments of this application is shown.

[0097] like Figure 3 As shown, the method for calculating the dynamic classification threshold includes operations S310 to S330.

[0098] In operation S310, quantile calculations are performed on the confidence distribution corresponding to each face category to obtain the corresponding target quantile. The confidence distribution is output by the classification model during the inference phase, representing the set of predicted probabilities for different face categories. Since the confidence distributions vary significantly across categories, especially in categories with sparse samples or ambiguous features, confidence values ​​often concentrate in lower intervals. Therefore, quantile calculations can effectively reflect the statistical characteristics of confidence within a category. For example, the 80th or 90th percentile can be chosen as the target quantile to represent the typical high-confidence interval for that category, thus avoiding bias in threshold estimation caused by extreme values ​​or outliers.

[0099] In operation S320, the confidence value corresponding to the target quantile is determined as the initial dynamic classification threshold.

[0100] The initial dynamic classification threshold reflects the natural confidence boundary of the classification model for that category, i.e., the lower limit of the model's confidence in classifying samples of that category as "positive samples". Unlike the global decision criterion of a fixed threshold, the initial dynamic threshold is defined differently at the category level, which allows high-confidence categories to maintain a high decision standard, while automatically lowering the decision threshold for tail categories with generally low confidence, thereby achieving a balance between precision and recall.

[0101] In operation S330, the initial dynamic classification threshold is adjusted according to the number of samples of each face category in the training data to obtain the dynamic classification threshold.

[0102] This adjustment process can incorporate sample size information to correct for confidence statistical bias. Specifically, for head categories with a large number of training samples, the system can appropriately increase their dynamic threshold to reduce the risk of false identification; while for sparse or long-tail categories, the threshold is lowered accordingly to improve the recall probability of that category. Through the sample size weighted adjustment strategy, the system can maintain overall decision fairness and stability despite differences in sample distribution across different categories.

[0103] According to embodiments of this application, a dynamic classification threshold calculation method combining quantile statistics and sample size weighting enables the face classification model to adaptively address issues such as uneven class distribution, confidence drift, and sparse tail samples. Embodiments of this application do not require additional model complexity; they can update the threshold decision boundary in real time during the inference stage, thereby improving recognition accuracy and stability while maintaining classification efficiency.

[0104] Figure 4 A flowchart illustrating a method for calculating a dynamic classification threshold according to some other exemplary embodiments of this application is shown.

[0105] like Figure 4 As shown, the method for calculating the dynamic classification threshold may include operations S410 to S440.

[0106] In operation S410, the confidence distribution is sorted in descending order to obtain a confidence sequence.

[0107] In operation S420, the first k target confidence categories are selected from the confidence sequence, and the mean confidence level of the target confidence categories is calculated and used as the main threshold, where k is a positive integer.

[0108] In the embodiments of this application, the principal threshold reflects the central tendency of the main identification results and can be regarded as a stable judgment boundary under high confidence conditions. For example, when k=5, the principal threshold represents the average confidence level of the top five most likely categories, thereby effectively smoothing the influence of local abnormally high values. This calculation method enables the system to dynamically determine a reliable benchmark decision standard based on the distribution of confidence peaks, thereby improving the core stability of identification.

[0109] In operation S430, based on the remaining confidence categories in the confidence sequence other than the target confidence category, the confidence dispersion is calculated to obtain an auxiliary threshold.

[0110] In the embodiments of this application, the auxiliary threshold is used to reflect the confidence fluctuation in non-primary categories. By calculating indicators such as standard deviation, variance, or coefficient of variation for the remaining category confidence scores, the system can measure the uncertainty of the model in secondary categories. When the dispersion is large, it indicates that the model's confidence in distinguishing different categories differs significantly, and the auxiliary threshold can be appropriately increased to strengthen the decision boundary; when the dispersion is small, it indicates that the confidence scores between categories are similar, and the auxiliary threshold can be correspondingly decreased to avoid overly conservative rejection.

[0111] In operation S440, the main threshold is combined with the auxiliary threshold to generate the dynamic classification threshold.

[0112] Specifically, the combination method can employ weighted linear combination, nonlinear function mapping, or confidence interval-based fusion strategies. For example, the system can control the weight ratio between the primary and secondary thresholds through parameters to achieve flexible adaptation to different recognition scenarios. When the model is designed for high-precision scenarios (such as identity verification), the weight of the primary threshold can be increased to ensure strict decision-making; while in scenarios requiring improved recall (such as video surveillance), the influence of the secondary threshold can be increased to enhance sensitivity.

[0113] According to the embodiments of this application, by introducing a primary threshold, a stable core decision criterion can be formed for high-confidence recognition results, thereby improving the recognition accuracy of the main categories; by introducing an auxiliary threshold, the confidence fluctuation of low-confidence categories can be adjusted, reducing misjudgments caused by edge samples and blurred samples; through the synergistic fusion of primary and auxiliary thresholds, robust adaptation to long-tail categories, low-quality images, and complex lighting environments is achieved.

[0114] Preferably, to further improve the recognition balance of the classification model in long-tail data scenarios, the embodiments of this application, based on dynamic classification threshold calculation, also include statistical analysis of multiple face categories and their corresponding confidence distributions to identify target sparse categories and perform boundary compensation adjustments. Specifically, the system first performs global statistical analysis on the confidence distribution of multiple face categories output by the classification model. By analyzing the concentration intervals of confidence and the distribution of sample numbers for different categories, it identifies sparse categories belonging to the long tail. The target sparse category can be a category whose sample size in the training data is significantly less than that of the mainstream category, and whose confidence distribution is relatively dispersed or low. Due to insufficient sample representativeness, such categories often exhibit low recognition confidence and ambiguous boundaries during inference, making them prone to misclassification by high-frequency categories.

[0115] After identifying the target sparse category, its confidence interval can be further calculated. The confidence interval can be obtained by calculating the mean, standard deviation, or quantile range of the confidence distribution for that category, reflecting the predictive stability of the classification model for samples of that category. For categories with wide intervals or low confidence, they can be identified as weak categories requiring threshold compensation adjustments. Therefore, the system calculates a category compensation factor based on the confidence interval. The category compensation factor can be defined as a function of the confidence interval width, sample sparsity, and the degree of prediction confidence shift, used to measure the adjustment required for that category in decision-making. For example, when a category has a small sample size and its average confidence is significantly lower than the global mean, the system can automatically assign a larger positive compensation factor to appropriately lower the decision threshold for that category, thereby increasing its probability of being correctly identified.

[0116] Subsequently, the classification boundaries of multiple face categories can be adjusted based on the category compensation factor and the aforementioned dynamic classification threshold. The adjustment process can be implemented using linear offset, non-linear mapping, or adaptive weight updates. For example, for the head category, the system maintains a high threshold to ensure accuracy; for the tail category, the system lowers the corresponding threshold based on the compensation factor, thereby achieving a balance between overall recognition accuracy and recall.

[0117] According to embodiments of this application, by performing statistical modeling on the confidence distribution and introducing a category compensation factor, long-tail category decision optimization based on dynamic threshold adjustment is achieved. This embodiment not only improves the model's recognition recall rate under low-sample, high-uncertainty categories, but also improves the recognition consistency between different categories globally, enabling the system to have higher stability and fairness in complex, imbalanced face data environments.

[0118] In some embodiments, to improve the stability and consistency of face image classification in continuous video or real-time monitoring scenarios, the system may further include temporal correction of the dynamic classification threshold based on time-series information. Specifically, in continuous classification tasks, the system can record the confidence distribution sequence corresponding to multiple adjacent frames of face images and determine the stability of the model output by analyzing the temporal fluctuation characteristics of this sequence.

[0119] In video recognition or multi-frame acquisition scenarios, input images may be affected by factors such as changes in lighting, pose shifts, and occlusion interference. Even for the same identity, the confidence value output by the model may fluctuate significantly over time, leading to discontinuous recognition results or momentary misjudgments. To address this issue, after obtaining the confidence distribution sequence, its temporal volatility can be calculated, which is the amplitude or variance of the confidence change across consecutive frames, used to measure the stability of the recognition results over time. A higher temporal volatility indicates instability in the classification model's recognition of the current target, potentially due to environmental interference or input anomalies. The system compares this temporal volatility with a preset stability threshold. If the volatility is below the threshold, it indicates that the recognition results are relatively consistent over time, and a decision can be made directly based on the current dynamic classification threshold. If the volatility exceeds the preset threshold, it indicates significant fluctuations in the recognition results, and the system needs to perform temporal correction to the dynamic classification threshold.

[0120] When performing temporal correction, the dynamic classification threshold can be adjusted through methods such as smoothing filtering, time-weighted averaging, or recursive updates. For example, the system can recalculate the threshold correction term based on the average confidence level and trend information of several past frames, making the new dynamic classification threshold closer to the historical stable decision range. In this way, even if the confidence level temporarily decreases due to interference in the recognition result of a single frame, the system can maintain the overall recognition consistency by relying on the temporal smoothing mechanism, thereby avoiding frequent category switching or misidentification.

[0121] Through the aforementioned temporal correction mechanism, the embodiments of this application introduce temporal stability constraints on top of the original dynamic classification threshold, enabling the model to have adaptive decision-making capabilities under multi-frame input. This mechanism not only enhances the system's robustness to environmental noise, device jitter, and dynamic lighting changes, but also ensures the consistency and reliability of output results in continuous sampling or real-time video recognition tasks.

[0122] According to embodiments of this application, high-precision and robust recognition in complex environments is achieved through multi-layer feature extraction, dynamic fusion, and adaptive decision mechanisms. By constructing a deep feature extraction model including shallow convolutional layers, mid-level pooling layers, and deep attention layers, facial features can be modeled from multiple dimensions, including local texture, structural morphology, and semantics. Dynamic weighted fusion based on attention weights and frequency domain feature extraction mechanisms effectively enhance the global consistency and noise resistance of feature representation. In the classification stage, the system calculates a dynamic classification threshold based on confidence distribution. Combining quantile statistics, confidence ranking, and category compensation factors, it optimizes and balances the recognition of long-tail categories, thereby improving classification accuracy and fairness under imbalanced sample conditions. Furthermore, invalid features are suppressed by using a pixel gradient-based noise-aware mask, and confidence time fluctuation analysis is introduced in continuous frame tasks to temporally correct the dynamic threshold, ensuring the smoothness and stability of the recognition results. Overall, the embodiments of this application can maintain high recognition reliability under low-quality images, complex lighting and dynamic scenes, and are suitable for intelligent face recognition systems with high security requirements such as financial risk control, identity verification and security monitoring.

[0123] Corresponding to the above-described method for classifying facial images, embodiments of this application also provide a device for classifying facial images.

[0124] Figure 5 A schematic block diagram of a face image classification apparatus according to an embodiment of this application is shown.

[0125] like Figure 5 As shown, the face image classification device 500 of this embodiment includes a feature extraction module 510, a fusion processing module 520, a dynamic threshold optimization module 530, and a classification module 540.

[0126] The feature extraction module 510 can be used to acquire a face image to be classified, and extract multi-layer feature information of the face image based on a pre-trained deep feature extraction model. The multi-layer feature information includes feature mapping results output by different layers of the deep feature extraction model. In one embodiment, the feature extraction module 510 can be used to perform the operation S210 described above, which will not be repeated here.

[0127] The fusion processing module 520 can be used to perform fusion processing on the multi-layer feature information to obtain a fused feature representation. In one embodiment, the fusion processing module 520 can be used to perform the operation S220 described above, which will not be repeated here.

[0128] The dynamic threshold optimization module 530 can be used to input the fused feature representation into a classification model, output multiple face categories and their corresponding confidence distributions, and calculate a dynamic classification threshold based on the confidence distributions. The dynamic classification threshold is used to adjust the classification boundaries of the multiple face categories. In one embodiment, the dynamic threshold optimization module 530 can be used to perform the operation S230 described above, which will not be repeated here.

[0129] The classification module 540 can be used to obtain the category of the face image based on the multiple face categories and the dynamic classification threshold. In one embodiment, the classification module 540 can be used to perform the operation S240 described above, which will not be repeated here.

[0130] According to an embodiment of this application, for the feature extraction module 510, the deep feature extraction model includes a shallow convolutional layer, a middle pooling layer, and a deep attention layer; the multi-layer feature information includes local texture features extracted by the shallow convolutional layer, structural features extracted by the middle pooling layer, and semantic features extracted by the deep attention layer.

[0131] According to an embodiment of this application, for the feature extraction module 510, the deep feature extraction model includes a spatial feature extraction branch and a frequency domain feature extraction branch, wherein the spatial feature extraction branch is used to extract the spatial texture and structural features of the face image; the frequency domain feature extraction branch obtains the spectral features of the face image through fast Fourier transform.

[0132] According to an embodiment of this application, the fusion processing module 520 can also be used to perform fusion processing on the multi-layer feature information using a dynamic weighting mechanism based on attention weights to obtain a fused feature representation, wherein the fusion weights corresponding to the multi-layer feature information are adjusted according to the clarity and local saliency information of the face image.

[0133] According to an embodiment of this application, the fusion processing module 520 can also be used to calculate a noise intensity map based on the pixel gradient distribution of the face image; and generate a feature mask based on the noise intensity map, and adjust the fusion weights corresponding to the multi-layer feature information based on the feature mask.

[0134] According to an embodiment of this application, the dynamic threshold optimization module 530 can also be used to perform quantile calculation on the confidence distribution corresponding to each face category to obtain the corresponding target quantile; determine the confidence value corresponding to the target quantile as the initial dynamic classification threshold; and adjust the initial dynamic classification threshold according to the number of samples of each face category in the training data to obtain the dynamic classification threshold.

[0135] According to an embodiment of this application, the dynamic threshold optimization module 530 can also be used to sort the confidence distribution in descending order to obtain a confidence sequence; select the first k target confidence categories in the confidence sequence, calculate the mean confidence of the target confidence categories and use it as the main threshold, where k is a positive integer; calculate the confidence dispersion based on the remaining confidence categories in the confidence sequence other than the target confidence categories to obtain an auxiliary threshold; and combine the main threshold and the auxiliary threshold to generate the dynamic classification threshold.

[0136] According to an embodiment of this application, the dynamic threshold optimization module 530 can also be used to perform distribution statistics on multiple face categories and their corresponding confidence distributions to determine the target sparse category and its corresponding confidence interval; and to calculate a category compensation factor based on the confidence interval, and to adjust the classification boundaries of multiple face categories based on the category compensation factor and the dynamic classification threshold.

[0137] According to an embodiment of this application, the anomaly recognition device 500 may further include a correction module. The correction module can be used to, in a continuous classification task, record a confidence distribution sequence corresponding to multiple adjacent frames of face images; calculate the temporal volatility of the confidence distribution sequence and compare the temporal volatility with a preset stability threshold; and, in response to the temporal volatility exceeding the preset stability threshold, perform temporal correction on the dynamic classification threshold.

[0138] According to embodiments of this application, any multiple modules among the feature extraction module 510, fusion processing module 520, dynamic threshold optimization module 530, and classification module 540 can be combined into one module, or any one of these modules can be split into multiple modules. Alternatively, at least some of the functions of one or more of these modules can be combined with at least some of the functions of other modules and implemented in one module. According to embodiments of this application, at least one of the feature extraction module 510, fusion processing module 520, dynamic threshold optimization module 530, and classification module 540 can be at least partially implemented as hardware circuitry, such as a field-programmable gate array (FPGA), a programmable logic array (PLA), a system-on-a-chip, a system-on-a-substrate, a system-on-package, an application-specific integrated circuit (ASIC), or implemented in hardware or firmware by any other reasonable means of integrating or packaging the circuitry, or implemented in software, hardware, or firmware, or in any suitable combination of any of these three implementation methods. Alternatively, at least one of the feature extraction module 510, the fusion processing module 520, the dynamic threshold optimization module 530, and the classification module 540 may be implemented at least partially as a computer program module, which can perform corresponding functions when the computer program module is run.

[0139] Figure 6 A block diagram schematically illustrates an electronic device suitable for implementing a method for classifying face images according to an embodiment of this application.

[0140] like Figure 6 As shown, an electronic device 600 according to an embodiment of this application includes a processor 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage portion 608 into a random access memory (RAM) 603. The processor 601 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or an associated chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 601 may also include onboard memory for caching purposes. The processor 601 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of this application.

[0141] RAM 603 stores various programs and data required for the operation of electronic device 600. Processor 601, ROM 602, and RAM 603 are interconnected via bus 604. Processor 601 executes various operations of the method flow according to embodiments of this application by executing programs in ROM 602 and / or RAM 603. It should be noted that the programs may also be stored in one or more memories other than ROM 602 and RAM 603. Processor 601 may also execute various operations of the method flow according to embodiments of this application by executing programs stored in said one or more memories.

[0142] According to embodiments of this application, the electronic device 600 may further include an input / output (I / O) interface 605, which is also connected to a bus 604. The electronic device 600 may also include one or more of the following components connected to the input / output (I / O) interface 605: an input section 606 including a keyboard, mouse, etc.; an output section 607 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 608 including a hard disk, etc.; and a communication section 609 including a network interface card such as a LAN card, modem, etc. The communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the input / output (I / O) interface 605 as needed. A removable medium 611, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 610 as needed so that computer programs read from it can be installed into the storage section 608 as needed.

[0143] This application also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or it may exist independently and not assembled into the device / apparatus / system. The computer-readable storage medium carries one or more programs, which, when executed, implement the method according to the embodiments of this application.

[0144] According to embodiments of this application, the computer-readable storage medium can be a non-volatile computer-readable storage medium, such as including but not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this application, the computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to embodiments of this application, the computer-readable storage medium may include ROM 602 and / or RAM 603 and / or one or more memories other than ROM 602 and RAM 603 described above.

[0145] Embodiments of this application also include a computer program product comprising a computer program containing program code for performing the methods shown in the flowchart. When the computer program product is run on a computer system, the program code enables the computer system to implement the face image classification method provided in the embodiments of this application.

[0146] When the computer program is executed by the processor 601, it performs the functions defined in the system / apparatus of this application embodiment. According to the embodiments of this application, the systems, apparatuses, modules, units, etc., described above can be implemented by computer program modules.

[0147] In one embodiment, the computer program may rely on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may also be transmitted and distributed in the form of signals over a network medium, and downloaded and installed via the communication section 609, and / or installed from the removable medium 611. The program code contained in the computer program can be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination thereof.

[0148] In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 609, and / or installed from the removable medium 611. When the computer program is executed by the processor 601, it performs the functions defined in the system of this application embodiment. According to the embodiments of this application, the systems, devices, apparatuses, modules, units, etc., described above can be implemented by computer program modules.

[0149] According to embodiments of this application, program code for executing the computer programs provided in the embodiments of this application can be written in any combination of one or more programming languages. Specifically, these computational programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages ​​include, but are not limited to, languages ​​such as Java, C++, Python, "C", or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0150] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0151] Those skilled in the art will understand that the features described in the various embodiments of this application can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in this application. In particular, the features described in the various embodiments of this application can be combined and / or combined in various ways without departing from the spirit and teachings of this application. All such combinations and / or combinations fall within the scope of this application.

[0152] The embodiments of this application have been described above. However, these embodiments are merely illustrative and not intended to limit the scope of this application. Although various embodiments have been described above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. Without departing from the scope of this application, those skilled in the art can make various substitutions and modifications, all of which should fall within the scope of this application.

Claims

1. A method for classifying human face images, characterized in that, The method includes: A face image to be classified is obtained, and multi-layer feature information of the face image is extracted based on a pre-trained deep feature extraction model. The multi-layer feature information includes feature mapping results output by different layers of the deep feature extraction model. The multi-layer feature information is fused to obtain a fused feature representation; The fused feature representation is input into a classification model, which outputs multiple face categories and their corresponding confidence distributions. A dynamic classification threshold is calculated based on the confidence distributions, wherein the dynamic classification threshold is used to adjust the classification boundaries of the multiple face categories; and The category of the face image is obtained based on the multiple face categories and the dynamic classification threshold.

2. The method according to claim 1, characterized in that, The deep feature extraction model includes shallow convolutional layers, mid-level pooling layers, and deep attention layers; the multi-layer feature information includes local texture features extracted by the shallow convolutional layers, structural features extracted by the mid-level pooling layers, and semantic features extracted by the deep attention layers.

3. The method according to claim 1 or 2, characterized in that, The step of performing fusion processing on the multi-layer feature information to obtain a fused feature representation includes: The multi-layer feature information is fused using a dynamic weighting mechanism based on attention weights to obtain a fused feature representation, wherein the fusion weights corresponding to the multi-layer feature information are adjusted according to the sharpness and local saliency information of the face image.

4. The method according to claim 1, characterized in that, The deep feature extraction model includes a spatial feature extraction branch and a frequency domain feature extraction branch. The spatial feature extraction branch is used to extract the spatial texture and structural features of the face image. The frequency domain feature extraction branch obtains the spectral features of the face image through a fast Fourier transform.

5. The method according to claim 2 or 4, characterized in that, The method further includes: Calculate the noise intensity map based on the pixel gradient distribution of the face image; and A feature mask is generated based on the noise intensity map, and the fusion weights corresponding to the multi-layer feature information are adjusted based on the feature mask.

6. The method according to claim 1, characterized in that, The calculation of the dynamic classification threshold based on the confidence distribution includes: Quantile calculations are performed on the confidence distribution corresponding to each face category to obtain the corresponding target quantiles; The confidence value corresponding to the target quantile is determined as the initial dynamic classification threshold; and The initial dynamic classification threshold is adjusted based on the number of samples for each face category in the training data to obtain the dynamic classification threshold.

7. The method according to claim 1, characterized in that, The calculation of the dynamic classification threshold based on the confidence distribution includes: The confidence distribution is sorted in descending order to obtain a confidence sequence; Select the first k target confidence categories from the confidence sequence, calculate the mean confidence score of the target confidence categories and use it as the main threshold, where k is a positive integer; Based on the remaining confidence categories in the confidence sequence excluding the target confidence category, the confidence dispersion is calculated to obtain an auxiliary threshold; and The main threshold and the auxiliary threshold are combined to generate the dynamic classification threshold.

8. The method according to claim 6 or 7, characterized in that, The method further includes: Perform distribution statistics on multiple face categories and their corresponding confidence distributions to determine the target sparse category and its corresponding confidence interval; and The category compensation factor is calculated based on the confidence interval, and the classification boundaries of multiple face categories are adjusted based on the category compensation factor and the dynamic classification threshold.

9. The method according to claim 1, characterized in that, The method further includes: In a continuous classification task, the confidence distribution sequence corresponding to multiple adjacent frames of face images is recorded. Calculate the time volatility of the confidence distribution sequence and compare the time volatility with a preset stability threshold; as well as In response to the time volatility exceeding the preset stability threshold, the dynamic classification threshold is time-series corrected.

10. A face image classification device, characterized in that, The device includes: The feature extraction module is used to: acquire a face image to be classified, and extract multi-layer feature information of the face image based on a pre-trained deep feature extraction model, wherein the multi-layer feature information includes feature mapping results output by different layers of the deep feature extraction model; The fusion processing module is used to: perform fusion processing on the multi-layer feature information to obtain a fused feature representation; A dynamic threshold optimization module is used to: input the fused feature representation into a classification model, output multiple face categories and their corresponding confidence distributions, and calculate a dynamic classification threshold based on the confidence distributions, wherein the dynamic classification threshold is used to adjust the classification boundaries of the multiple face categories; and The classification module is used to: obtain the category of the face image based on the multiple face categories and the dynamic classification threshold.

11. An electronic device, comprising: One or more processors; Memory, used to store one or more computer programs. The characteristic feature is that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 9.

12. A computer-readable storage medium having a computer program or instructions stored thereon, characterized in that, When the computer program or instructions are executed by a processor, they implement the steps of the method according to any one of claims 1 to 9.

13. A computer program product, comprising a computer program or instructions, characterized in that, When the computer program or instructions are executed by a processor, they implement the steps of the method according to any one of claims 1 to 9.