A Face Part Localization Method, Device and Medium Based on Multi-Region Perception

Through the multi-region-aware face positioning method, combined with HRNet and recursive feature pyramid network, the problem that traditional diagnostic methods cannot provide real-time health monitoring is solved, and the precise positioning and division of face areas is achieved, which improves detection accuracy and robustness, and is suitable for applications in resource-constrained environments.

CN119851330BActive Publication Date: 2025-07-22NINGBO FIRST HOSPITAL
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510330537.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-07-22
Estimated Expiration
2045-03-20

AI Technical Summary

Technical Problem

Traditional methods for diagnosing coronary heart disease cannot provide real-time, non-invasive health status monitoring, and the equipment is complex, limiting its application in daily health management.

Method used

The face positioning method based on multi-region perception is adopted. By extracting the rectangular boundary image of the face region, sub-region division is performed, geometric relationship and feature correlation model is constructed, face positioning is combined with the recursive feature pyramid network and the dual detection head architecture, key point detection is used to optimize the calculation of complexity through lightweight pruning and knowledge distillation.

Benefits of technology

It realizes accurate positioning and division of face areas, provides real-time and non-invasive health status monitoring, improves detection accuracy and robustness, reduces calculation complexity, and is suitable for applications in resource-constrained environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119851330B_ABST
    Figure CN119851330B_ABST
Patent Text Reader

Abstract

The present invention discloses a face part localization method, device and medium based on multi-region perception, which relates to the technical field of image processing. The method includes the steps of: extracting a rectangular boundary image of a face region in an input image, and dividing the face region in the rectangular boundary image into sub-regions based on face key points; constructing a geometric relationship and feature correlation model between sub-regions by extracting multi-scale features of each sub-region and performing interactive optimization; using the constructed model as prior knowledge, further fusing and optimizing multi-scale features through a recursive feature pyramid; using the fused features as the input of the main detection head and the features before fusion as the input of the auxiliary detection head, and performing face part localization through a dual-detection head architecture. The present invention ensures the accurate localization and division of the face region, improves the detection accuracy, and also enhances the robustness of the system in complex scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly relates to a face part localization method, device and medium based on multi-region perception. Background Art

[0002] As a common and serious cardiovascular disease, the early symptoms of coronary heart disease are often very hidden and difficult to detect. However, during the development of the disease, some subtle physiological changes may occur on the patient's face. Although these changes are subtle, they may become important clues for early diagnosis. Traditional diagnostic methods such as blood tests, electrocardiograms, echocardiograms, and CT scans, although they play an important role in diagnosing coronary heart disease, usually cannot provide real-time and non-invasive health status monitoring. In addition, these detection methods often require complex equipment and professional operations, which limit their wide application in daily health management.

[0003] In view of this, the face part localization technology based on facial images has gradually emerged and become an emerging tool for assisting in the diagnosis of coronary heart disease. This technology can accurately capture the subtle facial changes related to coronary heart disease through an efficient face part localization model combined with multi-scale feature extraction technology. Summary of the Invention

[0004] In order to achieve accurate localization of each part of the face, the present invention proposes a face part localization method based on multi-region perception, including the steps of:

[0005] S1: Extract the rectangular boundary image of the face region in the input image, and divide the face region in the rectangular boundary image into sub-regions based on face key points;

[0006] S2: Construct a geometric relationship and feature correlation model between sub-regions by extracting multi-scale features of each sub-region and performing interactive optimization;

[0007] S3: Using the constructed model as prior knowledge, further fuse and optimize multi-scale features through a recursive feature pyramid;

[0008] S4: Using the fused features as the input of the main detection head and the features before fusion as the input of the auxiliary detection head, perform face part localization through a dual detection head architecture.

[0009] Further, in the step S1, the following steps are further included:

[0010] By calculating the proportion of the face region in the rectangular boundary image and the image acquisition quality, when the proportion is lower than the preset proportion range and / or the image acquisition quality is lower than the preset threshold, prompt to adjust the shooting distance and / or angle, and return to step S1.

[0011] Further, the proportion of the face region is calculated by the following formula:

[0012]

[0013] In the formula, is the proportion of the face region, and are the upper-left coordinate and the lower-right coordinate of the rectangular boundary image respectively, and W and H are the width and height of the input image respectively.

[0014] Further, the image acquisition quality is obtained by the following formula:

[0015]

[0016] In the formula, Q is the image acquisition quality, is the preset weight coefficient, C is the image sharpness, B is the image contrast, and L is the image brightness.

[0017] Further, in the S1 step, the HRNet with depthwise separable convolution replacing the depth feature extraction module is used to identify the facial key points.

[0018] Further, in the S2 step, in each sub-region, first, the depthwise separable convolution is used to extract the detailed features within the current sub-region, and then the multi-head self-attention mechanism is used for feature interaction between sub-regions, modeling the geometric relationship and feature correlation between sub-regions, and integrating the global information and local details through the shared feature representation layer.

[0019] Further, in the S3 step, the ordinary convolution is replaced with the depthwise separable convolution, and the recursive feature pyramid network combined with the channel attention and spatial attention mechanisms is used to achieve the recursive fusion of multi-layer features through the residual connection.

[0020] Further, in the S2 to S4 steps, the computational complexity is optimized by the lightweight pruning and knowledge distillation methods.

[0021] The present invention also includes a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the method for face part localization based on multi-region perception are implemented.

[0022] It also includes a device for processing data, including:

[0023] A memory, on which a computer program is stored;

[0024] A processor, configured to execute the computer program in the memory to implement the steps of the method for face part localization based on multi-region perception.

[0025] Compared with the prior art, the present invention has at least the following beneficial effects:

[0026] (1) For the method, device and medium for facial part positioning based on multi-region perception of the present invention, a pre-trained face detection model and a high-resolution network (HRNet) are used for key point detection, ensuring the accurate positioning and division of the face region. This high-precision positioning lays a solid foundation for subsequent feature extraction and analysis;

[0027] (2) Compared with traditional blood tests and imaging methods, facial image analysis technology provides a real-time and non-invasive way to monitor health status. Patients do not need to undergo complex detection processes and avoid the risk of radiation exposure, making health management more convenient and safe;

[0028] (3) By using a recursive feature pyramid network (RFPN), the system can fuse multi-scale features, enhancing the ability to capture different sizes and detailed features. This method not only improves the detection accuracy but also enhances the robustness of the system in complex scenarios;

[0029] (4) A dual-detection head architecture is adopted. The main detection head is responsible for high-precision classification and boundary regression, while the auxiliary detection head uses a loose label strategy to improve the recall rate. This design balances the accuracy and coverage of detection and reduces the possibility of missed detection;

[0030] (5) Through lightweight pruning and knowledge distillation methods, the system significantly reduces the computational complexity while maintaining high detection performance, achieving efficient real-time processing. This is of great significance for large-scale applications and deployments, especially in resource-constrained environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 It is a step diagram of a method for facial part positioning based on multi-region perception;

[0032] Figure 2 It is a schematic diagram of a dual-detection head architecture. DETAILED DESCRIPTION OF THE INVENTION

[0033] The following are specific embodiments of the present invention in combination with the drawings to further describe the technical solutions of the present invention, but the present invention is not limited to these embodiments.

[0034] Coronary heart disease is a heart disease caused by the stenosis or obstruction of the blood vessels that supply blood to the heart (coronary arteries). Traditionally, the diagnosis of coronary heart disease mainly relies on invasive examinations such as coronary angiography, non-invasive stress tests, or imaging examinations. However, these methods are usually costly and carry certain risks. In recent years, with the increasing widespread application of artificial intelligence in the medical field, health status assessment based on facial features has become an emerging research direction.

[0035] The present invention focuses on developing an efficient face part localization model, which can accurately identify and extract multi-scale features of different parts of the face. Specifically, through deep learning algorithms, the model can capture subtle changes and landmarks on the face at different resolutions, and these changes may be associated with the presence of coronary heart disease. As Figure 1 shown, the present invention proposes a face part localization method based on multi-region perception, including the steps of:

[0036] S1: Extract the rectangular boundary image of the face region in the input image, and divide the face region in the rectangular boundary image into sub-regions based on face key points;

[0037] S2: By extracting multi-scale features of each sub-region and performing interactive optimization, construct a geometric relationship and feature correlation model between sub-regions;

[0038] S3: Using the constructed model as prior knowledge, further fuse and optimize multi-scale features through a recursive feature pyramid;

[0039] S4: Using the fused features as the input of the main detection head and the features before fusion as the input of the auxiliary detection head, perform face part localization through a dual detection head architecture.

[0040] According to the solution adopted by the present invention, first, the input image is processed by a pre-trained face detection model. Here, the face detection model we adopt can be any one of mature face recognition technologies such as the Viola-Jones framework, HOG+SVM, DPM deformable part model, CNN convolutional neural network, and YOLO11. By training on a large-scale dataset, learn the general face feature representation, and then transfer the learned knowledge to the small-scale dataset of the face recognition task to improve the model performance and generalization ability.

[0041] Assume the input image is L, and the size of this image is W×H. Detect the complete face region in the image through the pre-trained face detection model. Assume the detected face region is a rectangular box , where and are the upper left and lower right coordinates of the face rectangular box. Then, calculate the proportion P of the face region occupying the image:

[0042]

[0043] Among them, the image quality P needs to be within a certain range to ensure the quality of the captured images. If it is detected that P is not within this range, the system will adjust the angle or distance of image capture in real time according to the current situation, guiding the collector to adjust the shooting position so that the proportion of the face reaches the set range.

[0044] Furthermore, to improve the accuracy and quality of the captured images, the system will analyze and judge in real time whether the image quality of the captured images is qualified. The image quality is obtained through the following formula:

[0045]

[0046] In the formula, Q is the image capture quality, is the preset weight coefficient, C is the image clarity, B is the image contrast, and L is the image brightness.

[0047] Based on the analysis of the image quality value calculated based on factors such as image clarity, contrast, and lighting conditions, if it is then prompted that the image meets the capture standard, otherwise it is prompted to readjust. Finally, the number of unqualified images is reduced to ensure that each captured image has good details and clarity. Among them, According to the ROC curve analysis of sensitivity and specificity, the best balance point is selected. In this embodiment, the Laplacian algorithm is used for clarity, and it is set that C > 10 indicates clarity, the contrast is selected as B > 0.5, and the lighting condition is selected as L ∈ [80, 200]. And the preliminary weight coefficients are 0.5, 0.3, and 0.2 respectively, and are appropriately adjusted according to the image quality feedback by the user in different environments.

[0048] Here, the present invention marks the images that meet the capture standard as positive classes by setting a threshold T, and vice versa as negative classes. The ROC curve is generated by adjusting the decision threshold T of the binary classification model and calculating the sensitivity (True Positive Rate, TPR) and false positive rate (False Positive Rate, FPR) at different thresholds. First, the samples are sorted from high to low according to the predicted scores output by the model, and then the decision threshold is gradually adjusted. The corresponding TPR and FPR are calculated at each threshold, and these points are plotted on a two-dimensional coordinate axis, with the horizontal axis being FPR and the vertical axis being TPR, finally forming a curve. The point closest to the upper left corner is selected as the best point balance on the ROC curve.

[0049] After obtaining the face region image that meets the conditions, it enters the face image division stage. The present invention constructs an efficient region division process by combining geometric prior knowledge and deep learning technology. This method first designs a facial geometric structure template, which is defined based on a series of key ratio parameters, including but not limited to the ratio of the eye width to the nose bridge length, the ratio of the mouth width to the face jaw width, etc. Through the statistical analysis of a large-scale facial dataset, these ratio parameters are carefully selected to ensure that they can adapt to various face morphologies, thus providing a general and flexible basic model.

[0050] To construct this facial geometric structure template, feature points are first extracted from a large number of facial images, and the distances and angles between these feature points are measured and analyzed in detail. By comprehensively considering data from different races, genders, and age groups, the most representative ratio parameters are determined. These parameters not only consider static facial features (such as the relative positions of facial features), but also incorporate dynamic features (such as the impact of facial expressions on facial structure).

[0051] Considering the various poses, sizes, or facial expression changes that a face may present in practical applications, an Affine Transformation is introduced in the template design. This mathematical transformation allows the template to be dynamically adjusted in scale and angle, enabling it to precisely fit the face region under different conditions. The Affine Transformation can not only maintain the parallelism of straight lines and the relative distance relationship in the original image, but also effectively handle problems such as rotation, scaling, and translation, thereby improving the applicability and robustness of the model in complex environments. The mathematical expression of this process is:

[0052]

[0053] Among them, represents the template point set, is the Affine Transformation matrix, is the transformed point set.

[0054] To accurately obtain the positions of facial key points, a High-Resolution Network (HRNet) is used for key point detection. HRNet achieves high-precision key point localization by gradually fusing feature maps of different resolutions. This method can not only capture local details but also maintain the consistency of the global structure, thereby improving the robustness and accuracy of key point detection.

[0055] The design concept of HRNet lies in its unique multi-resolution parallel processing mechanism. Traditional convolutional neural networks usually process images by gradually downsampling and upsampling, which may lead to the loss of detailed information. In contrast, HRNet maintains a high-resolution representation throughout the network process and exchanges information through multiple cross-resolution connections, ensuring the effective fusion of features from low-level to high-level. This design gives HRNet a significant advantage in dealing with complex scenarios, especially in face keypoint detection.

[0056] In the specific implementation, the system uses HRNet to extract 68 keypoints from the face region. These keypoints are widely distributed, covering important parts such as eyes, eyebrows, nose, mouth, and facial contours. The position of each keypoint is precisely calculated to ensure that they can accurately reflect the geometric features of each part of the face. For example, the eye keypoints can be used to analyze the degree of eyelid opening and the position of the eyeballs; the mouth keypoints help to evaluate the size of the mouth opening and the change of lip shape.

[0057] Although HRNet provides powerful feature extraction capabilities, in practical applications, due to the influence of pose changes, occlusion, or lighting conditions, keypoint localization may still deviate. Therefore, a geometric consistency check module is introduced to verify the keypoints output by HRNet. This module calculates the Euclidean distance and angular difference between adjacent keypoints to filter out abnormal points that do not conform to the expected distribution.

[0058] This method combining HRNet with geometric consistency check has shown significant advantages in various application scenarios. For example, in a face recognition system, accurate keypoint localization helps to improve the recognition rate, especially in the face of complex backgrounds or non-ideal shooting conditions.

[0059] To further enhance the robustness of keypoint localization, the present invention also introduces the Random Sample Consensus (RANSAC) algorithm as a post-processing step. RANSAC filters out the incorrect keypoints affected by noise through iterative sampling and model fitting, and ensures that the final result has higher global consistency. In this process, the system divides the 68 keypoints into multiple functional groups, such as the eyebrow group, the eye group, and the mouth group, etc., and then applies local geometric fitting models to each group of keypoints respectively. Through this grouping strategy, it can better adapt to the deformation of local features and reduce the problem of error accumulation that may occur in global fitting.

[0060] After optimizing the key point positions, the next step is to define the boundaries of the facial regions based on these key points. To achieve this goal, the system adopts the Delaunay triangulation algorithm to divide the facial regions. This method can not only maximize the minimum angle of each triangle, thus avoiding the generation of overly flat triangles, but also ensure the rationality and uniformity of the region distribution. This characteristic is crucial for accurately capturing and representing complex facial structures.

[0061] Taking the eye region as an example, a polygonal region is formed by connecting the key points of the upper and lower eyelids. Specifically, the system will first identify the key points related to the eyelids, and then use the Delaunay triangulation algorithm to generate a series of triangles, which together constitute a closed eye region. This method can not only accurately describe the shape of the eyes, but also adapt to the eye structures under different poses and expressions.

[0062] Similarly, the mouth region generates a closed region based on the contour points of the upper and lower lips. By identifying the key points related to the lips and applying the same triangulation method, the system can create an accurate and flexible representation of the lip region. This helps in the subsequent analysis of lip movements, mouth opening sizes and other features, providing a solid foundation for fields such as emotion recognition and speech synthesis.

[0063] Although the Delaunay triangulation algorithm can effectively generate reasonable region divisions, in practical applications, the boundaries of adjacent regions may overlap or conflict. To avoid this situation, the system further combines the Non-Maximum Suppression (NMS) algorithm to optimize the region boundaries.

[0064] By using the Delaunay triangulation algorithm for facial region division and combining the non-maximum suppression algorithm to optimize the region boundaries, this method not only improves the accuracy of region division, but also enhances the robustness of the system. It can provide reliable facial region information under different environments and conditions, providing strong support for various facial feature-based applications. Future research can further explore how to optimize the parameter settings of the triangulation algorithm and how to more efficiently integrate other types of geometric prior knowledge to address more complex practical problems. This comprehensive method opens up new possibilities for facial feature analysis and lays the foundation for further improving the practical application effects of computer vision technology.

[0065] After the region division is completed, the system refines the image information of each region by combining the lightweight object detection network YOLO11n and the local feature extraction module. Taking the eye region as an example, a specific network branch focuses on extracting the pupil position, eyelid curve, and texture details; the nose region focuses on analyzing the curvature of the nose bridge and the shape of the nasal wings; the mouth region focuses on the texture features and dynamic changes of the upper and lower lips. To further improve the independence of feature extraction, a decoupled feature extraction strategy is adopted. Through the decoupled head structure, the region features are independently encoded while avoiding interference between different regions.

[0066] After feature extraction is completed, a shared feature representation layer is used to achieve information interaction and optimization between regions. The shared feature representation layer models the geometric and feature correlation relationships between different regions through the multi-head self-attention mechanism, thereby capturing complementary information of each region of the face. For example, the curvature information of the nose region can be used as a geometric constraint to optimize the spatial distribution of key points in the eye region; while the symmetry information of the mouth region can help correct the features of the cheek region. In addition, to further improve the accuracy and efficiency during the sharing of features, depthwise separable convolution and channel attention mechanism are adopted, and cross-region feature interaction is strengthened through recursive iteration. The finally output feature representation contains both global context information and retains the refined details of local features.

[0067] After that, to further fuse and optimize multi-scale features, the system introduces a Recursive Feature Pyramid Network (RFPN). The core of this network lies in its unique recursive fusion strategy, which generates a richer multi-scale feature representation by gradually exchanging and superimposing information between multiple layers of feature maps.

[0068] At the initial stage of the network, the system extracts basic features from feature maps of different resolutions. These feature maps cover different levels from low resolution to high resolution, and each layer captures information at a specific scale. Specifically:

[0069] Low-resolution top-level feature map: mainly retains global context information, which is crucial for understanding the overall layout of the entire facial structure;

[0070] High-resolution bottom-level feature map: focuses on capturing fine-grained local details, such as the specific shapes and texture features of parts like eyes and lips.

[0071] The Recursive Feature Pyramid Network performs information flow and feature enhancement between different levels through the recursive fusion strategy, and the feature transformation operations during the recursive process are uniformly managed through shared parameters.

[0072] Through the recursive fusion strategy, the system can capture rich feature information at different scales, thereby improving the accuracy and robustness of feature extraction. This is particularly important for tasks such as facial recognition and sentiment analysis. By adopting an efficient parameter sharing mechanism, the entire system can achieve real-time processing while maintaining high accuracy. This is crucial for application scenarios that require quick responses (such as virtual reality, augmented reality, real-time monitoring, etc.).

[0073] On this basis, in order to balance the computational cost and performance, the present invention further adopts a lightweight design strategy for the Recursive Feature Pyramid Network (RFPN). Specifically, by introducing depthwise separable convolution, channel attention mechanism and spatial attention mechanism, the system not only reduces the number of parameters and computational complexity, but also significantly improves the quality of feature extraction, especially the performance in capturing key local information.

[0074] Due to its fully connected nature, traditional convolution operations often require a large number of parameters and computational resources when processing high-resolution images. To alleviate this burden, the present invention uses depthwise separable convolution instead of traditional convolution operations. Depthwise separable convolution decomposes the standard convolution into two steps:

[0075] 1. Depthwise convolution: Convolution operations are performed on each input channel separately to independently process the features of each channel.

[0076] 2. Pointwise convolution: Use a 1×1 convolutional kernel to linearly combine the results of the depthwise convolution to generate the final output feature map.

[0077] This method can significantly reduce the number of parameters and computational complexity while maintaining a high quality of feature extraction. For example, in the facial recognition task, depthwise separable convolution can greatly reduce the computational overhead of the model without sacrificing accuracy, making the system more suitable for deployment in resource-constrained environments.

[0078] To further improve the quality of feature fusion, the present invention introduces a channel attention mechanism. The channel attention mechanism dynamically weights the importance of different channels, enabling the model to adaptively focus on the features that are most helpful for the task. In addition to the channel attention mechanism, the present invention also introduces a spatial attention mechanism to enhance the ability to capture features in key spatial regions.

[0079] By adopting depthwise separable convolutions, channel attention mechanism, and spatial attention mechanism, the present invention significantly reduces the computational complexity and the number of parameters while ensuring the quality of feature extraction. This enables the system to operate efficiently even in resource-constrained environments (such as mobile devices or embedded systems), while maintaining high detection accuracy and generalization ability.

[0080] The lightweight design strategy not only reduces the complexity of the model but also improves the quality of feature fusion by introducing attention mechanisms. Especially when capturing key local information, the channel attention mechanism and the spatial attention mechanism can dynamically adjust the weights of each channel and spatial region, enabling the model to more accurately locate and analyze facial features. This is particularly important for tasks such as face recognition and sentiment analysis.

[0081] By comprehensively applying a variety of optimization techniques, the present invention enhances the robustness of the model, enabling it to better cope with challenges such as different lighting conditions, pose variations, and expression changes. For example, in practical application scenarios, even when faced with low-light environments or face images at extreme angles, the system can still maintain high recognition accuracy and stability.

[0082] In addition, as shown in Figure 2 an auxiliary detection head is added in the middle layer of the detection network and a looser label standard is used for object detection to improve the recall rate. Loose labels have a larger tolerance range, which can cover blurred or smaller objects, thereby capturing more candidate objects and reducing missed detections. The main detection head uses precise labels and focuses on improving the accuracy of object localization. The loss function still focuses on optimizing object classification and bounding box regression to achieve the detection goal of "few but precise". During the training process, the auxiliary detection head adopts a "coarse-to-fine" label strategy. The loose labels first generate larger candidate object regions to provide training samples, and then the precise labels further screen within these regions to optimize the ability of the main detection head. In addition, the intermediate features of the auxiliary detection head are used as the input of the main detection head. Through feature fusion, the performance of the main detection head is enhanced, and context information is provided through cross-layer feature flow, further improving the detection accuracy and robustness.

[0083] After training is completed, calculate the L1 norm (the sum of the absolute values of each weight) of the weights of each convolutional layer or fully connected layer. For the weight matrix W, the L1 norm calculation formula is:

[0084]

[0085] where is the i-th element in the weight matrix, and is the absolute value of the i-th element.

[0086] After calculating the L1 norm of all weights, sort and select the smallest 30% of the weights, and then select the element with the smallest L1 norm for pruning. Shrink its weight to near zero, but not exactly zero, to maintain the smoothness of the network.

[0087] In the final stage of model optimization, namely the pruning and fine-tuning stage, a technique called knowledge distillation is adopted to recover the accuracy that may be lost due to pruning. This process not only focuses on reducing the complexity and computational cost of the model, but also aims to maintain or even improve the performance of the model.

[0088] First, a powerful teacher model needs to be constructed - RT-DETR-R50 in this example. This model is known for its large number of parameters and rich feature representation ability, and can provide high-quality prediction results. At the same time, the pruned YOLO11n model is selected as the student model. Although pruning effectively reduces the size and computational burden of the model, it may also lead to a decline in the model's expressiveness. Therefore, the method of knowledge distillation is introduced to make up for this defect, enabling the student model to learn more knowledge from the teacher model.

[0089] The core of knowledge distillation is to let the student model enhance its feature representation ability by mimicking the behavior of the teacher model. Specifically, the student model not only needs to learn the final output of the teacher model, but also needs to mimic the activation values of the intermediate layers of the teacher model. This multi-level learning method helps the student model capture richer feature information and improve its understanding ability of the input data.

[0090] To achieve the above goals, a composite distillation loss function is designed, which consists of two parts:

[0091] 1. Output layer loss: This part of the loss is used to measure the difference between the final outputs of the student model and the teacher model. Usually, cross-entropy loss is used to calculate it to ensure that the student model can simulate the prediction results of the teacher model as accurately as possible.

[0092] 2. Intermediate layer loss: In addition to the output layer, the student model also needs to match the activation values of the intermediate layers of the teacher model. This part of the loss helps the student model learn the advanced feature representations used by the teacher model when processing the input. By minimizing the difference between the intermediate layers, the feature extraction ability of the student model can be further strengthened.

[0093] Through this method, even after pruning, the student model can still achieve good performance. On the one hand, due to the high-quality guidance provided by the teacher model, the student model can learn complex feature representations without adding too many parameters; on the other hand, this cross-layer knowledge transfer mechanism also helps to solve the overfitting problem and makes the student model more robust.

[0094] In addition, this method is of great significance for model deployment in resource-constrained environments. For example, when running deep learning models on mobile devices or embedded systems, it is often necessary to achieve efficient and accurate predictions with limited computing resources. By using knowledge distillation techniques, the computing cost can be significantly reduced while ensuring the model accuracy, providing strong support for practical applications.

[0095] In summary, as an effective model compression and optimization method, knowledge distillation can not only help reduce the model size and improve the inference speed, but also significantly enhance the model's generalization ability and feature learning ability by the way of the teacher model teaching knowledge to the student model. This has important value for promoting the wide application of deep learning technology in more fields.

[0096] By matching the intermediate layer features of the student model and the teacher model, the student model not only learns the final class predictions, but also can imitate the feature representations learned by the teacher model at different network levels. This method greatly improves the performance of the student model, especially after pruning or compression, helps to recover some detection performance that may be lost due to model simplification, and enhances its generalization ability.

[0097] This is because traditional knowledge distillation methods mainly focus on the loss of the output layer, that is, the student model tries to imitate the final prediction results of the teacher model. However, this method ignores the important information contained in the intermediate layer features. By introducing the intermediate layer feature matching mechanism, the student model can learn the knowledge of the teacher model at each layer, thus significantly enhancing its feature expression ability.

[0098] The rich semantic information contained in the intermediate layer features is crucial for understanding the input data. For example, in a face recognition task, low-level features may capture pixel-level details (such as edges and textures), while high-level features contain more abstract concepts (such as the positional relationship of facial features and the overall facial structure). By allowing the student model to imitate the intermediate layer features of the teacher model, it can ensure that the student model can not only make accurate classification decisions, but also better understand and process complex patterns in the input data.

[0099] Finally, the loss function can be expressed as:

[0100]

[0101] In the formula, is the cross-entropy loss, is the mean squared error loss of the intermediate layer, controls the weights of the two, represents the predicted output of the student model, represents the actual label, and They respectively represent the intermediate layer features in the knowledge distillation of the student model and the teacher model.

[0102] By matching the intermediate layer features of the student model and the teacher model, the student model can not only learn the final class prediction but also imitate the feature representations learned by the teacher model at different network levels. This method effectively restores some detection performance lost due to pruning or compression and significantly improves the generalization ability of the student. Whether deploying an efficient model in a resource-constrained environment or achieving high-precision recognition in complex and changing real-world application scenarios, this strategy demonstrates its unique advantages and potential.

[0103] The present invention also includes a computer-readable storage medium having a computer program stored thereon, and when the program is executed by a processor, the steps of the method for face part localization based on multi-region perception are implemented.

[0104] It also includes a device for processing data, including:

[0105] A memory having a computer program stored thereon;

[0106] A processor for executing the computer program in the memory to implement the steps of the method for face part localization based on multi-region perception.

[0107] In summary, for the method, device, and medium for face part localization based on multi-region perception of the present invention, a pre-trained face detection model and a high-resolution network (HRNet) are used for key point detection, ensuring the accurate positioning and division of the face region. This high-precision positioning lays a solid foundation for subsequent feature extraction and analysis.

[0108] Compared with traditional blood tests and imaging methods, facial image analysis technology provides a real-time and non-invasive way to monitor health status. Patients do not need to undergo complex detection processes and avoid the risk of radiation exposure, making health management more convenient and safe;

[0109] Using a recursive feature pyramid network (RFPN), the system can fuse multi-scale features, enhancing the ability to capture features of different sizes and details. This method not only improves the detection accuracy but also enhances the robustness of the system in complex scenarios.

[0110] Adopting a dual-detection head architecture, the main detection head is responsible for high-precision classification and boundary regression, while the auxiliary detection head uses a loose label strategy to improve the recall rate. This design balances the accuracy and coverage of detection and reduces the possibility of missed detections.

[0111] Through lightweight pruning and knowledge distillation methods, the system significantly reduces the computational complexity while maintaining high detection performance, achieving efficient real-time processing, which is of great significance for large-scale applications and deployments, especially in resource-constrained environments.

[0112] It should be noted that all directional indications (such as up, down, left, right, front, back...) in the embodiments of the present invention are only used to explain the relative positional relationship and movement conditions between components in a specific posture (as shown in the attached drawings). If the specific posture changes, the directional indications will also change accordingly.

[0113] In addition, in the present invention, descriptions such as "first", "second", "one", etc. are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of the present invention, the meaning of "a plurality" is at least two, such as two, three, etc., unless otherwise specifically defined.

[0114] In the present invention, unless otherwise clearly specified and limited, the terms "connection", "fixation", etc. shall be understood in a broad sense. For example, "fixation" may be a fixed connection, a detachable connection, or integrated; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the internal connection of two components or the interaction relationship between two components, unless otherwise clearly limited. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0115] In addition, the technical solutions between the various embodiments of the present invention can be combined with each other, but it must be based on the ability of those of ordinary skill in the art to implement. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the protection scope required by the present invention.

Claims

1. A face part localization method based on multi-region perception, characterized in that, Including the steps: S1: Process the input image with a pre-trained face detection model, extract the rectangular boundary image of the face region in the input image, and divide the face region in the rectangular boundary image into sub-regions based on face key points; S2: Construct a face part localization model; the face part localization model includes: the target detection network YOLO11n, a local feature extraction module, a decoupled head structure, a shared feature representation layer, a recursive feature pyramid network, an auxiliary detection head, and a main detection head; refine the image information of each sub-region through the target detection network YOLO11n and the local feature extraction module; adopt a decoupled feature extraction strategy, and independently encode the regional features through the decoupled head structure; Subsequently, use the shared feature representation layer to construct the geometric and feature correlation relationships between sub-regions by using the multi-head self-attention mechanism, and adopt depthwise separable convolution and channel attention mechanism on the basis of this model to strengthen cross-region feature interaction in a recursive iteration manner; S4: Gradually exchange and stack information between multi-layer feature maps through the recursive feature pyramid network to obtain and fuse multi-scale features; S5: Use the features before fusion as the input of the auxiliary detection head, and the features after fusion as the input of the main detection head to locate the face parts through a dual detection head architecture.

2. The face part localization method based on multi-region perception according to claim 1, wherein In the step S1, it further includes the steps: By calculating the proportion of the face region in the rectangular boundary image and the image acquisition quality, when the proportion is lower than the preset proportion range and / or the image acquisition quality is lower than the preset threshold, prompt to adjust the shooting distance and / or angle.

3. The face part positioning method based on multi-region perception according to claim 2, wherein Calculate the proportion of the face region through the following formula: ; In the formula, represents the proportion of the face area, and are respectively the upper left coordinate and the lower right coordinate of the rectangular boundary image, and W and H are respectively the width and height of the input image.

4. The face part positioning method based on multi-region perception according to claim 2, characterized in that The image acquisition quality is obtained through the following formula: ; Where Q is the image acquisition quality, is the preset weight coefficient, C is the image sharpness, B is the image contrast, and L is the image brightness.

5. The face part localization method based on multi-region perception according to claim 1, characterized in that In the step S1, replace the deep feature extraction module with HRNet with depthwise separable convolution for face key point recognition.

6. The face part localization method based on multi-region perception according to claim 1, characterized in that In the step S2, in each sub-region, first extract the detailed features in the current sub-region by using depthwise separable convolution, then perform feature interaction between cross-sub-regions through the multi-head self-attention mechanism, model the geometric relationship and feature correlation between sub-regions, and integrate global information and local details through the shared feature representation layer.

7. The face part localization method based on multi-region perception according to claim 1, characterized in that In the step S3, replace the ordinary convolution with depthwise separable convolution, and use a recursive feature pyramid network combined with channel attention and spatial attention mechanisms to achieve recursive fusion of multi-layer features through residual connections.

8. The face part localization method based on multi-region perception according to claim 1, characterized in that In the steps S2 to S4, optimize the computational complexity through lightweight pruning and knowledge distillation methods.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of the localization method described in any one of claims 1 to 8.

10. A device for processing data, characterized in that, Including: A memory, on which a computer program is stored; A processor, configured to execute the computer program in the memory to implement the steps of the localization method described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Coronary heart disease preliminary screening method and device based on facial feature analysis and medium

    CN119560141A

  • Face mask detection method for resisting shielding counterfeit attack based on YOLOv8

    CN119600670A