Face analysis method and system based on self-supervised representation learning
By constructing a local-to-global mask distillation network model and an attention-guided multi-level distillation module, the problem of lack of spatial sensitivity in face analysis by self-supervised learning methods is solved, achieving high-precision face analysis tasks that are suitable for various application scenarios.
Patent Information
- Application Number
- CN202510980514.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-16
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2045-07-16
AI Technical Summary
Existing self-supervised learning methods in face analysis lack spatial sensitivity and ignore the complex relationship between local and global perspectives, which limits the performance of fine-grained facial tasks.
A face analysis method employing self-supervised representation learning is proposed. By constructing a local-to-global mask distillation network model and combining it with an attention-guided multi-level distillation module, the mask loss and distillation loss are optimized to capture the deep relationship between local and global data, thereby enhancing spatial sensitivity and semantic consistency.
It significantly improves the accuracy and reliability of various face analysis tasks without requiring large-scale labeled data, and is suitable for a variety of application scenarios such as face recognition and emotion analysis, demonstrating high accuracy.
Smart Images

Figure CN120472521B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision technology, and in particular relates to a face analysis method and system based on self-supervised representation learning. Background Art
[0002] In recent years, face analysis has become an important research direction in computer vision, encompassing tasks such as facial expression recognition, facial attribute recognition, face parsing, and face alignment. The key to these tasks lies in learning high-quality facial representations that accurately capture the fine-grained features of the face and its global context. However, traditional supervised learning methods typically rely on large amounts of annotated data and tend to learn task-specific features, limiting their generalizability across tasks.
[0003] Self-supervised learning, a technique for extracting visual features from unlabeled data, has been widely used in recent years for tasks such as image classification, object detection, and semantic segmentation. A. Bulat et al. used unsupervised clustering to learn task-independent features from large-scale unlabeled data, while Y. Liu et al. used "pose decoupling" to achieve pose-invariant features. X. Di used prototype-based self-distillation, combined with data augmentation and prototype matching loss, and Z. Gao et al. used heat maps to learn global-local consistency in facial representations. In the latest method, Y. Li et al. enhanced the correlation between facial images and text representations by optimizing the cosine similarity of image-text pairs.
[0004] However, Y. Li et al. found that such methods focus on global representations but often lack spatial sensitivity. To address this shortcoming, Y. Zheng et al. combined contrastive learning with mask image modeling to capture both high-level semantics and low-level details. However, this approach ignores the complex relationship between local and global perspectives, limiting performance on fine-grained facial tasks. Summary of the Invention
[0005] The purpose of the present invention is to provide a face analysis method and system based on self-supervised representation learning, which is conducive to obtaining more robust, accurate and generalized results of face analysis tasks.
[0006] To achieve the above objectives, the present invention adopts a technical solution: a face analysis method based on self-supervised representation learning, comprising the following steps:
[0007] 1) Obtain face images from the dataset and preprocess them to generate local face image token sequences and global face image token sequences;
[0008] 2) Construct a self-supervised local-to-global mask distillation network model to extract universal face representation; for the input local face image Token sequence and global face image Token sequence, the local-to-global mask distillation network model removes the mask part of the local face image Token sequence and inputs it into the online local encoder, and the online local encoder extracts the visible local feature Token sequence; at the same time, the local-to-global mask distillation network model inputs the global face image Token sequence into the target global encoder, and the target global encoder extracts the target global feature Token sequence; then, the local-to-global decoder combines the relative position relationship between the local and global, and predicts the global feature Token sequence through the visible local feature Token sequence; during the model training process, the weighted sum of the mask loss and the distillation loss is calculated, and the parameters of the online local encoder are optimized through gradient backpropagation to achieve iterative training of the model; wherein, the mask loss measures the gap between the predicted global features and the target global features and learns the global context information through the dense mask loss and the global mask loss, and the distillation loss is used to measure the difference between the local features and the global features;
[0009] 3) Apply the trained online local encoder to face parsing, facial key point detection, facial attribute recognition or facial expression recognition tasks through transfer learning.
[0010] Furthermore, the implementation method of step 1) is:
[0011] A1) Obtain face images from the VGG-Face2 face dataset;
[0012] A2) Obtain the facial image Cropping to generate partial views , taking the uncropped image as the global view ; Divide the local view and the global view into a series of fixed-size image blocks , and apply spatial and color enhancements Generate enhanced local face image Token sequence and global face image Token sequence ; In addition, mask enhancement is applied to the local view Generate a masked partial face image token sequence .
[0013] Furthermore, in step 2), the local-to-global mask distillation network model is implemented as follows:
[0014] The local-to-global mask distillation network model adopts a local-to-global mask image modeling framework, including an online local encoder, a target global encoder, a local-to-global decoder, and an attention-guided multi-stage distillation module;
[0015] The local-to-global mask image modeling framework has two branches, consisting of an online local encoder and target global encoder Implementation: The online local encoder and target global encoder Each consists of multiple Transformer encoder blocks and a projection head; the global face image Token sequence is input to the target global encoder , through the target global encoder Extract target global feature Token sequence At the same time, the local face image Token sequence removes the mask part and inputs it into the online local encoder , through the online local encoder Extract visible local feature Token sequence ; Then the local feature Token sequence will be visible Input the local to global decoder, and obtain the predicted global feature Token sequence through the local to global decoder to combine the target global feature Token sequence in the iterative training process of the model Calculate mask loss;
[0016] In the local-to-global mask distillation network model, the online local encoder As a distilled student network, the target global encoder Serving as both the distilled teacher network and the source of global features, the target global encoder With online local encoder Shared architecture, through previous iterations Update the parameters of
[0017] During the iterative training of the model, the attention-guided multi-level distillation module establishes a connection between the local features extracted by the online local encoder and the global features extracted by the target global encoder. The attention weights are calculated using the query, key and value mechanism, and the most relevant global features are selected for weighted summation. By minimizing the difference between local and global features, the detail capture capability of the local encoder is improved.
[0018] Furthermore, the local-to-global decoder is implemented as follows:
[0019] B1) Constructing the relative position relationship between the local view and the global view: Local view In the global view The position in the upper left corner is determined by the coordinates of ,high and width OK, and the global view The size of ; The number of tokens in , The number of tokens in ; Calculate each position in the global view Relative position relationship with respect to the local view , which is expressed as follows:
[0020]
[0021] in, Indicates the scaling ratio of the global view to the local view in the height dimension, Represents the scaling ratio of the global view to the local view in the width dimension; the position in the local view is mapped to the corresponding position in the global view by scaling; in addition, Indicates the relative position height offset of the local view in the global view. Indicates the relative position width offset of the local view in the global view; in order to Embedded into the model’s feature representation, the relative scale changes between views are connected, and the relative scale changes of height and width between views are , and encode the sine-cosine position Applied to relative position relationship and relative proportion changes Then, the relative position relationship and relative scale change after sine-cosine position encoding are connected and passed to the linear layer to adjust the embedding dimension to obtain the final relative position relationship , which is expressed as follows:
[0022]
[0023] in, is the sine-cosine position encoding function, which is used to convert the position coordinates into a high-dimensional vector representation; represents the connection function, represents a linear transformation;
[0024] B2) Local to Global Decoder Combined position relationship and mask markers , guiding the online local encoder Output visible local feature Token sequence Reconstructing the global view; by integrating local Mark, corresponding position embed , visible local feature Token sequence , position embedding of local features and mask markers , predicting dense feature representations of the global view , which is expressed as follows:
[0025]
[0026] Predicted dense feature representation of the global view That is to predict the global feature Token sequence.
[0027] Furthermore, the implementation method of the attention-guided multi-stage distillation module is as follows:
[0028] C1) is respectively passed through the online local encoder and the target global encoder The encoder blocks extract features, , Set a set of encoder block numbers for the online local encoder and the target global encoder, where the numbers correspond to the encoder block numbers of the local information processing layer, the middle grammatical semantic layer, the high-level semantic abstraction layer, and the global context understanding layer in the online local encoder and the target global encoder respectively; then, the encoder block numbers of the online local encoder and the target global encoder are set to the number of the encoder block numbers of the online local encoder and the target global encoder respectively; The multi-level distillation module between the encoder blocks performs linear mapping on the extracted local and global features to generate query ,key Sum ,in 、 and is the linear transformation matrix, is the first The features extracted by the encoder blocks, is the target global encoder The features extracted by the encoder block; then, the scaled dot product attention mechanism is used to calculate the The attention weight between the local features and global features extracted by the encoder block is defined as ,in is the dimension of each attention head;
[0029] C2) Select for each local feature The most relevant global features; ,in Before screening The maximum attention weight, For the front The index of the largest attention weight; through The index of the target global encoder is The corresponding features are selected from the encoder blocks as the most relevant local features. Global features , a weighted sum is performed to generate a weighted global feature, defined as ,in Represents the first The weighted global features aligned with the local features of the encoder blocks, Before The largest attention weight matrix The attention weights, represents the kth global feature that is most correlated with the local feature A global feature.
[0030] Furthermore, in step C2), average pooling is performed on the local features and the weighted global features to obtain and ,in and Represent the pooled local features and global features respectively, Indicates the number of visible local feature tokens, Indicates the number of global feature tokens;
[0031] By aligning the online local encoder and the target global encoder The mean square error between the local features and global features extracted by the set encoder blocks is summed to construct the distillation loss between the local features and the global features:
[0032]
[0033] in, Represents distillation loss.
[0034] Furthermore, in step 2), the mask loss is constructed as follows:
[0035] D1) Construct the feature matrix of all tokens in the global view ,in Indicates the number of global feature tokens, Represents the target global feature Token sequence, and then calculates the predicted global feature Token sequence Each predicted global feature Token With the matrix The inner product of is used as the negative sample loss term; dense mask loss is used For each predicted global feature Token Each target global feature Token in the target global feature Token sequence Comparison to reconstruct global information from local input; dense mask loss The definition is as follows:
[0036]
[0037] in, Indicates that for each prediction Token and the corresponding target Token expectations, The weight for controlling the contribution of negative sample loss;
[0038] D2) Constructing global mask loss To fully learn the global context information; predict the global feature Token sequence through the local to global decoder Perform average pooling to obtain a compact global feature representation of the prediction ; Through the target global encoder to the target global feature Token sequence Perform average pooling to obtain a compact target global feature representation ; Using cosine similarity To maximize the similarity of positive samples and control the similarity of negative samples; global mask loss The definition is as follows:
[0039]
[0040] in, Represents the weight of negative sample loss;
[0041] Finally, the total mask loss is obtained, which is defined as follows:
[0042]
[0043] in, and are the weights of dense mask loss and global mask loss, respectively.
[0044] Furthermore, in step 2), the overall loss of the local to global mask distillation network model is constructed by calculating the weighted sum of the mask loss and the distillation loss. as follows:
[0045]
[0046] in, and The mask loss is and distillation losses The weight of .
[0047] The present invention also provides a face multi-task system based on self-supervised representation learning, including a memory, a processor, and computer program instructions stored in the memory and capable of being executed by the processor. When the processor executes the computer program instructions, the above method can be implemented.
[0048] Compared with the existing technology, the present invention has the following advantages: The present invention provides a face analysis method and system based on self-supervised representation learning. This method constructs a self-supervised local-to-global mask distillation network model. The local-to-global mask image modeling (L2G-MIM) framework reconstructs global features from local views, capturing the deep relationship between the local and global. The attention-guided multi-level distillation (AGMD) module uses a cross-scale attention mechanism and layer-by-layer feature distillation to ensure the alignment of local features with global contextual information and enhance the ability to capture fine-grained features. At the same time, the method jointly optimizes the mask loss and distillation loss to generate facial representations that balance spatial sensitivity and semantic consistency for a variety of downstream tasks. This method and system can significantly improve the accuracy and reliability of various face analysis tasks without the need for large-scale labeled data. It not only effectively solves the problem of a small number of labeled samples, but also maintains high accuracy when migrated to various face analysis tasks. It is suitable for a variety of application scenarios such as face recognition and emotion analysis. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 This is a flowchart of an implementation of a face analysis method based on self-supervised representation learning provided by an embodiment of the present invention;
[0050] Figure 2 This is an architectural diagram of the self-supervised local-to-global mask distillation network model in an embodiment of the present invention. DETAILED DESCRIPTION
[0051] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0052] It should be noted that the following detailed descriptions are exemplary and are intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present application belongs.
[0053] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or combinations thereof.
[0054] like Figure 1 As shown, this embodiment provides a face analysis method based on self-supervised representation learning, including the following steps:
[0055] 1) Obtain face images from the dataset and preprocess them to generate local face image token sequences and global face image token sequences;
[0056] 2) Construct a self-supervised local-to-global mask distillation network model to extract universal face representation; for the input local face image Token sequence and global face image Token sequence, the local-to-global mask distillation network model removes the mask part of the local face image Token sequence and inputs it into the online local encoder, and the online local encoder extracts the visible local feature Token sequence; at the same time, the local-to-global mask distillation network model inputs the global face image Token sequence into the target global encoder, and the target global encoder extracts the target global feature Token sequence; then, the local-to-global decoder combines the relative position relationship between the local and global, and predicts the global feature Token sequence through the visible local feature Token sequence; during the model training process, the weighted sum of the mask loss and the distillation loss is calculated, and the parameters of the online local encoder are optimized through gradient backpropagation to achieve iterative training of the model; among them, the mask loss measures the gap between the predicted global features and the target global features through dense mask loss and global mask loss, and fully learns the global context information, and the distillation loss is used to measure the difference between local features and global features;
[0057] 3) Apply the trained online local encoder to face parsing, facial key point detection, facial attribute recognition or facial expression recognition tasks through transfer learning.
[0058] In this embodiment, the specific implementation method of step 1) is as follows.
[0059] A1) Obtain face images from the VGG-Face2 face dataset.
[0060] A2) Obtain the facial image Cropping to generate partial views , taking the uncropped image as the global view ; Divide the local view and the global view into a series of fixed-size image blocks ,in represents the number of image blocks, represents the size of each image block, Indicates the number of channels of the image block, and Represent the height and width of the face image respectively; then, apply spatial and color enhancement Generate enhanced local face image Token sequence and global face image Token sequence ; In addition, mask enhancement is applied to the local view Generate a masked partial face image token sequence .
[0061] Figure 2 This is the architecture diagram of the self-supervised local to global mask distillation network model in this embodiment. Figure 2 As shown, in step 2), the specific implementation method of the local to global mask distillation network model is as follows.
[0062] The local-to-global mask distillation network model adopts a local-to-global mask image modeling (L2G-MIM) framework, including an online local encoder, a target global encoder, a local-to-global decoder, and an attention-guided multi-stage distillation (AGMD) module.
[0063] The L2G-MIM framework has two branches, consisting of an online local encoder and target global encoder Implementation: The online local encoder and target global encoder They have the same structure, consisting of multiple Transformer encoder blocks and a projection head. In this embodiment, the online local encoder and target global encoder Both consist of 12 Transformer encoder blocks and a projection head.
[0064] Global face image token sequence input target global encoder , through the target global encoder Extract target global feature Token sequence At the same time, the local face image Token sequence removes the mask part and inputs it into the online local encoder , through the online local encoder Extract visible local feature Token sequence ; Then the local feature Token sequence will be visible Input the local to global decoder, and obtain the predicted global feature Token sequence through the local to global decoder to combine the target global feature Token sequence in the iterative training process of the model Calculate the mask loss.
[0065] In the local-to-global mask distillation network model, the online local encoder As a distilled student network, the target global encoder Serving as both the distilled teacher network and the source of global features, the target global encoder With online local encoder Shared architecture, through previous iterations Update the parameters.
[0066] During the iterative training of the model, the attention-guided multi-level distillation (AGMD) module establishes a connection between the local features extracted by the online local encoder and the global features extracted by the target global encoder. The attention weights are calculated using the query, key, and value mechanism, and the most relevant global features are selected for weighted summation. By minimizing the difference between local and global features, the detail capture capability of the local encoder is improved.
[0067] In this embodiment, the specific implementation method of the local-to-global decoder is as follows.
[0068] B1) Constructing the relative position relationship between the local view and the global view: Local view In the global view The position in the upper left corner is determined by the coordinates of ,high and width OK, and the global view The size of ; The number of tokens in , The number of tokens in ; Calculate each position in the global view Relative position relationship with respect to the local view , which is expressed as follows:
[0069]
[0070] in, Indicates the scaling ratio of the global view to the local view in the height dimension, Represents the scaling ratio of the global view to the local view in the width dimension; the position in the local view is mapped to the corresponding position in the global view by scaling; in addition, Indicates the relative position height offset of the local view in the global view. Indicates the relative position width offset of the local view in the global view, which is used to ensure that the starting position of the local view is correctly encoded; in order to Embedded into the model’s feature representation, the relative scale changes between views are connected, and the relative scale changes of height and width between views are , and encode the sine-cosine position Applied to relative position relationship and relative proportion changes Then, the relative position relationship and relative scale change after sine-cosine position encoding are connected and passed to the linear layer to adjust the embedding dimension to obtain the final relative position relationship , which is expressed as follows:
[0071]
[0072] in, is the sine-cosine position encoding function, which is used to convert the position coordinates into a high-dimensional vector representation; represents the connection function, Represents a linear transformation.
[0073] B2) Local to Global Decoder Combined position relationship and mask markers , guiding the online local encoder Output visible local feature Token sequence Reconstructing the global view; by integrating local Mark, corresponding position embed , visible local feature Token sequence , position embedding of local features and mask markers , predicting dense feature representations of the global view , which is expressed as follows:
[0074]
[0075] Predicted dense feature representation of the global view That is to predict the global feature Token sequence.
[0076] In this embodiment, the specific implementation method of the attention-guided multi-stage distillation (AGMD) module is as follows.
[0077] C1) is respectively passed through the online local encoder and the target global encoder The encoder blocks extract features, , Set the numbered set of encoder blocks in the online local encoder and the target global encoder. In this embodiment, The numbers 3, 6, 9, and 12 correspond to the encoder block numbers of the local information processing layer, the middle syntax and semantic layer, the high-level semantic abstraction layer, and the global context understanding layer in the online local encoder and the target global encoder, respectively.
[0078] Then, the online local encoder and the target global encoder are The multi-level distillation module between the encoder blocks performs linear mapping on the extracted local and global features to generate query ,key Sum ,in 、 and is the linear transformation matrix, is the first The features extracted by the encoder blocks, is the target global encoder The features extracted by the encoder block; then, the scaled dot product attention mechanism is used to calculate the The attention weight between the local features and global features extracted by the encoder block is defined as ,in is the dimension of each attention head.
[0079] C2) In order to reduce computational complexity and focus on the most relevant global features, this method introduces a Top-K strategy to select The most relevant global features. ,in Before screening The maximum attention weight, For the front The index of the largest attention weight; through The index of the target global encoder is The corresponding features are selected from the encoder blocks as the most relevant local features. Global features , a weighted sum is performed to generate a weighted global feature, defined as ,in Represents the first The weighted global features aligned with the local features of the encoder blocks, Before The largest attention weight matrix The attention weights, represents the kth global feature that is most correlated with the local feature A global feature.
[0080] In order to align local and global features with different numbers of markers, average pooling is performed on the local features and the weighted global features to obtain and ,in and Represent the pooled local features and global features respectively, Indicates the number of visible local feature tokens, Indicates the number of global feature tokens. Here, Summarizes the visible local details, while We focus on the global context that is most relevant to the local features.
[0081] By aligning the online local encoder and the target global encoder The mean square error (MSE) between the local features extracted by the set encoder blocks and the global features is summed to construct the distillation loss between the local features and the global features to ensure that the local features can capture key information from the global features during the learning process:
[0082]
[0083] in, Represents distillation loss.
[0084] In this embodiment, the mask loss is constructed as follows.
[0085] D1) Construct the feature matrix of all tokens in the global view ,in Indicates the number of global feature tokens, Represents the target global feature Token sequence, and then calculates the predicted global feature Token sequence Each predicted global feature Token With the matrix The inner product of is used as the negative sample loss term; dense mask loss is used For each predicted global feature Token Each target global feature Token in the target global feature Token sequence Comparison to reconstruct global information from local input; dense mask loss The definition is as follows:
[0086]
[0087] in, Indicates that for each prediction Token and the corresponding target Token expectations, The weight for controlling the contribution of negative sample loss.
[0088] D2) To prevent the model from focusing only on local information, a global mask loss is constructed To fully learn the global context information; predict the global feature Token sequence through the local to global decoder Perform average pooling to obtain a compact global feature representation of the prediction ; Through the target global encoder to the target global feature Token sequence Perform average pooling to obtain a compact target global feature representation ; Using cosine similarity To maximize the similarity of positive samples and control the similarity of negative samples; global mask loss The definition is as follows:
[0089]
[0090] in, Represents the weight of the negative sample loss.
[0091] Finally, the total mask loss is obtained, which is defined as follows:
[0092]
[0093] in, and are the weights of dense mask loss and global mask loss, respectively.
[0094] Finally, the overall loss of the local-to-global mask distillation network model is constructed by calculating the weighted sum of the mask loss and the distillation loss. as follows:
[0095]
[0096] in, and The mask loss is and distillation losses The weight of .
[0097] In this embodiment, the results of the method proposed in the present invention on a variety of face analysis tasks are compared with those of other face representation learning methods and specific task supervision methods. Table 1 compares the results of the method proposed in the present invention with those of other face representation learning methods and specific task supervision methods on two mainstream datasets for face parsing tasks: the LaPa dataset and the CelebAMask-HQ dataset. Table 2 compares the results of the method proposed in the present invention with those of other face representation learning methods and specific task supervision methods on two mainstream datasets for face key point positioning tasks: the WFLW dataset and the 300W dataset. Table 3 compares the results of the method proposed in the present invention with those of other face representation learning methods and specific task supervision methods on two mainstream datasets for face attribute recognition tasks: the LFWA dataset and the CelebA dataset. Table 4 compares the results of the method proposed in the present invention with those of other face representation learning methods and specific task supervision methods on two mainstream datasets for face expression recognition tasks: the FERPlus dataset and the RAF-DB dataset. It can be seen from the comparison that the performance of the method proposed in the present invention is better than other existing methods.
[0098] Table 1
[0099]
[0100] Table 2
[0101]
[0102] Table 3
[0103]
[0104] Table 4
[0105]
[0106] This embodiment also provides a face multi-task system based on self-supervised representation learning, including a memory, a processor, and computer program instructions stored in the memory and capable of being executed by the processor. When the processor executes the computer program instructions, the above method can be implemented.
[0107] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0108] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0109] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0110] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0111] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any other manner. Any person skilled in the art may utilize the above-disclosed technical content to modify or modify the present invention into equivalent embodiments. However, any simple modifications, equivalent variations, and modifications to the above embodiments that do not depart from the technical content of the present invention and are based on the technical essence of the present invention remain within the scope of protection of the present invention.
Claims
1. A face analysis method based on self-supervised representation learning, characterized in that: The following steps are involved: 1) Obtain face images from the dataset and preprocess them to generate local face image token sequences and global face image token sequences; 2) Construct a self-supervised local-to-global mask distillation network model to extract universal face representation; for the input local face image Token sequence and global face image Token sequence, the local-to-global mask distillation network model removes the mask part of the local face image Token sequence and inputs it into the online local encoder, and the online local encoder extracts the visible local feature Token sequence; at the same time, the local-to-global mask distillation network model inputs the global face image Token sequence into the target global encoder, and the target global encoder extracts the target global feature Token sequence; then, the local-to-global decoder combines the relative position relationship between the local and global, and predicts the global feature Token sequence through the visible local feature Token sequence; during the model training process, the weighted sum of the mask loss and the distillation loss is calculated, and the parameters of the online local encoder are optimized through gradient backpropagation to achieve iterative training of the model; wherein, the mask loss measures the gap between the predicted global features and the target global features and learns the global context information through the dense mask loss and the global mask loss, and the distillation loss is used to measure the difference between the local features and the global features; 3) Apply the trained online local encoder to face parsing, facial key point detection, facial attribute recognition or facial expression recognition tasks through transfer learning.
2. The face analysis method of self-supervised representation learning according to claim 1, characterized in that The implementation method of step 1) is: A1) Obtain face images from the VGG-Face2 face dataset; A2) Obtain the facial image Cropping to generate partial views , taking the uncropped image as the global view ; Divide the local view and the global view into a series of fixed-size image blocks , and apply spatial and color enhancements Generate enhanced local face image Token sequence and global face image Token sequence ; In addition, mask enhancement is applied to the local view Generate a masked partial face image token sequence .
3. The face analysis method of self-supervised representation learning according to claim 1, characterized in that In step 2), the implementation method of the local to global mask distillation network model is: The local-to-global mask distillation network model adopts a local-to-global mask image modeling framework, including an online local encoder, a target global encoder, a local-to-global decoder, and an attention-guided multi-stage distillation module; The local-to-global mask image modeling framework has two branches, consisting of an online local encoder and target global encoder Implementation: The online local encoder and target global encoder Each consists of multiple Transformer encoder blocks and a projection head; the global face image Token sequence is input to the target global encoder , through the target global encoder Extract target global feature Token sequence At the same time, the local face image Token sequence removes the mask part and inputs it into the online local encoder , through the online local encoder Extract visible local feature Token sequence ; Then the local feature Token sequence will be visible Input the local to global decoder, and obtain the predicted global feature Token sequence through the local to global decoder to combine the target global feature Token sequence in the iterative training process of the model Calculate mask loss; In the local-to-global mask distillation network model, the online local encoder As a distilled student network, the target global encoder Serving as both the distilled teacher network and the source of global features, the target global encoder With online local encoder Shared architecture, through previous iterations Update the parameters of During the iterative training of the model, the attention-guided multi-level distillation module establishes a connection between the local features extracted by the online local encoder and the global features extracted by the target global encoder. The attention weights are calculated using the query, key and value mechanism, and the most relevant global features are selected for weighted summation. By minimizing the difference between local and global features, the detail capture capability of the local encoder is improved.
4. The face analysis method of self-supervised representation learning according to claim 3, characterized in that The local-to-global decoder is implemented as follows: B1) Constructing the relative position relationship between the local view and the global view: Local view In the global view The position in the upper left corner is determined by the coordinates of ,high and width OK, and the global view The size of ; The number of tokens in , The number of tokens in ; Calculate the global view for each position Relative position relationship with respect to the local view , which is expressed as follows: in, Indicates the scaling ratio of the global view to the local view in the height dimension, Represents the scaling ratio of the global view to the local view in the width dimension; the position in the local view is mapped to the corresponding position in the global view by scaling; in addition, Indicates the relative position height offset of the local view in the global view. Indicates the relative position width offset of the local view in the global view; in order to Embedded into the model’s feature representation, the relative scale changes between views are connected, and the relative scale changes of height and width between views are , and encode the sine-cosine position Applied to relative position relationship and relative proportion changes Then, the relative position relationship and relative scale change after sine-cosine position encoding are connected and passed to the linear layer to adjust the embedding dimension to obtain the final relative position relationship , which is expressed as follows: in, is the sine-cosine position encoding function, which is used to convert the position coordinates into a high-dimensional vector representation; represents the connection function, represents a linear transformation; B2) Local to Global Decoder Combined position relationship and mask markers , guiding the online local encoder Output visible local feature Token sequence Reconstructing the global view; by integrating local Mark, corresponding position embed , visible local feature Token sequence , position embedding of local features and mask markers , predicting dense feature representations of the global view , which is expressed as follows: Predicted dense feature representation of the global view That is to predict the global feature Token sequence.
5. The face analysis method of self-supervised representation learning according to claim 3, characterized in that The implementation method of the attention-guided multi-stage distillation module is as follows: C1) is respectively passed through the online local encoder and the target global encoder The encoder blocks extract features, , Set a set of encoder block numbers for the online local encoder and the target global encoder, where the numbers correspond to the encoder block numbers of the local information processing layer, the middle grammatical semantic layer, the high-level semantic abstraction layer, and the global context understanding layer in the online local encoder and the target global encoder respectively; then, the encoder block numbers of the online local encoder and the target global encoder are set to the number of the encoder block numbers of the online local encoder and the target global encoder respectively; The multi-level distillation module between the encoder blocks performs linear mapping on the extracted local and global features to generate query ,key Sum ,in 、 and is the linear transformation matrix, is the first The features extracted by the encoder blocks, is the target global encoder The features extracted by the encoder block; then, the scaled dot product attention mechanism is used to calculate the The attention weight between the local features and global features extracted by the encoder block is defined as ,in is the dimension of each attention head; C2) Select for each local feature The most relevant global features; ,in Before screening The maximum attention weight, For the front The index of the largest attention weight; through The index of the target global encoder is The corresponding features are selected from the encoder blocks as the most relevant local features. Global features , a weighted sum is performed to generate a weighted global feature, defined as ,in Represents the first The weighted global features aligned with the local features of the encoder blocks, Before The largest attention weight matrix The attention weights, represents the kth global feature that is most correlated with the local feature A global feature.
6. The face analysis method of self-supervised representation learning according to claim 5, characterized in that In step C2), average pooling is performed on the local features and the weighted global features to obtain and ,in and Represent the pooled local features and global features respectively, Indicates the number of visible local feature tokens, Indicates the number of global feature tokens; By aligning the online local encoder and the target global encoder The mean square error between the local features and global features extracted by the set encoder blocks is summed to construct the distillation loss between the local features and the global features: in, Represents distillation loss.
7. The face analysis method of self-supervised representation learning according to claim 1, characterized in that In step 2), the mask loss is constructed as follows: D1) Construct the feature matrix of all tokens in the global view ,in Indicates the number of global feature tokens, Represents the target global feature Token sequence, and then calculates the predicted global feature Token sequence Each predicted global feature Token With the matrix The inner product of is used as the negative sample loss term; dense mask loss is used For each predicted global feature Token Each target global feature Token in the target global feature Token sequence Comparison to reconstruct global information from local input; dense mask loss The definition is as follows: in, Indicates that for each prediction Token and the corresponding target Token expectations, The weight for controlling the contribution of negative sample loss; D2) Constructing global mask loss To fully learn the global context information; predict the global feature Token sequence through the local to global decoder Perform average pooling to obtain a compact global feature representation of the prediction ; Through the target global encoder to the target global feature Token sequence Perform average pooling to obtain a compact target global feature representation ; Using cosine similarity To maximize the similarity of positive samples and control the similarity of negative samples; global mask loss The definition is as follows: in, Represents the weight of negative sample loss; Finally, the total mask loss is obtained, which is defined as follows: in, and are the weights of dense mask loss and global mask loss, respectively.
8. The face analysis method of self-supervised representation learning according to claim 1, characterized in that In step 2), the overall loss of the local to global mask distillation network model is constructed by calculating the weighted sum of the mask loss and the distillation loss. as follows: in, and The mask loss is and distillation losses The weight of .
9. A face multi-task system based on self-supervised representation learning, characterized by: The method comprises a memory, a processor, and computer program instructions stored in the memory and capable of being executed by the processor. When the processor executes the computer program instructions, the method according to any one of claims 1 to 8 can be implemented.
Citation Information
Patent Citations
Method from image pre-training model to video facial expression recognition
CN117456581A
Self-supervision face AU detection method without label guidance, equipment and medium
CN118470774A