Optimization Method of Cosine Optimal Loss Function Based on Global Information

Through the cosine optimal loss function optimized by global information, the problem that existing loss functions fail to effectively apply weight and feature normalization is solved, and more efficient face recognition performance is achieved, especially significant improvements on specific data sets.

CN115761851BActive Publication Date: 2025-07-11HANGZHOU SHANHAIGUAN TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211442334.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-16
Publication Date
2025-07-11
Estimated Expiration
2042-11-16

AI Technical Summary

Technical Problem

The existing loss functions fail to effectively apply weight and feature normalization in convolutional neural networks, and fail to clearly follow the goal of minimizing intra-class changes and maximizing inter-class changes. At the same time, only small batches of feedback information are considered and the distribution information of the entire training set is ignored.

Method used

A cosine optimal loss function based on global information is proposed. Through L2 weight normalization and feature normalization, combined with AM-Softmax loss, the lightweight version of the cosine optimal loss function is used to update the cosine similarity between the class center and the class edge, and the distance between adjacent class centers is optimized to integrate into the standard version of the cosine optimal loss function.

Benefits of technology

More efficient performance improvements were achieved, especially the face recognition effect on LFW, SLLFW and YTF datasets were significantly improved, proving the effectiveness of the cosine optimal loss function.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0003946439500000021
    Figure BDA0003946439500000021
  • Figure BDA0003946439500000031
    Figure BDA0003946439500000031
  • Figure BDA0003946439500000032
    Figure BDA0003946439500000032
Patent Text Reader

Abstract

The present invention proposes an optimization method for a cosine optimal loss function based on global information, including: S1. Combining the advantages of existing loss functions with some important new attributes and applying L2 weight normalization; S2. Clearly following the two objectives of minimizing intra-class variation and maximizing inter-class variation, relying on a new algorithm to learn the cosine similarity between class centers and class margins, and respectively proposing two lightweight versions of the cosine optimal loss function; S3. Integrating the above two lightweight versions to create a standard version of the cosine optimal loss function. The present invention mainly aims at the problems that existing loss functions do not apply weight and feature normalization or do not clearly follow the objectives of minimizing intra-class variation and maximizing inter-class variation. By using global information as the feedback for face recognition, a cosine optimal loss function based on global information is proposed. Compared with existing loss functions, this loss function is more effective and achieves more advanced performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of artificial intelligence, machine learning, and face recognition, and particularly relates to an optimization method of a cosine optimal loss function that can be applied to face recognition and is based on global information. Background Art

[0002] Convolutional neural networks (CNNs) have shown impressive performance in face recognition, and the loss function plays an important role in this process. In recent years, many different loss functions have been proposed to learn highly discriminative features. Currently, the best-performing loss functions in face recognition can be divided into two categories - loss functions based on Euclidean distance and loss functions based on cosine similarity.

[0003] The Softmax loss can be expressed as: where N represents the batch size, P represents the number of classes in the entire training set, f i ∈R d is the feature vector of the i-th sample belonging to the y i -th class, W i ∈R d is the last fully connected layer of the j-th column of the weight matrix W, and b j is the bias term of the j-th class. Typical Euclidean distance-based losses include center loss, margin loss, and range loss. They all add additional penalties to achieve joint supervision with the Softmax loss and are designed based on the following two objectives: minimizing intra-class variation and maximizing inter-class variation. Both of these objectives contribute to performance improvement. Cosine similarity-based loss functions include L-Softmax loss, A-Softmax loss, and AM-Softmax loss. They are derived from the Softmax loss by adding additional margin constraints. L2 weight normalization improves performance, although the improvement is very limited. The advantages brought by feature normalization include better performance and better geometric interpretation.

[0004] Among the loss functions that have been proposed so far, some do not apply weight and feature normalization, such as contrast loss, triplet loss, center loss, range loss, and margin loss; some do not clearly follow the two objectives of improving discriminative ability, such as L-Softmax loss, A-Softmax loss, AM-Softmax loss, and ArcFace.

[0005] Currently, deep neural networks are trained by iteratively updating network parameters based on the feedback information of each mini - batch. This is a viable solution because there are two limitations: the computing power and memory size of GPUs, TPUs, or other similar processing units. In the absence of computing power limitations, deep neural networks can be trained using the entire training set as the source of feedback information to directly optimize the sample distribution of the entire training set. In the absence of memory size limitations, deep neural networks would input the entire training set into memory instead of processing data mini - batch by mini - batch. Perhaps due to these two constraints above, none of the losses use the entire data set as the source of feedback information to optimize CNNs in face recognition.

[0006] We propose a new loss function, namely the cosine optimal loss function based on global information. The cosine optimal loss function has all four properties of optimizing intra - class and inter - class variations, as well as weight and feature normalization. And, the cosine optimal loss function is guided by the distribution information of the entire training set. Compared with the previously proposed loss functions, the cosine optimal loss function is more effective and exhibits more advanced performance. Summary of the Invention

[0007] (1) Technical problems to be solved

[0008] Loss functions play an important role in CNNs (Convolutional Neural Networks). However, existing loss functions either do not apply weight and feature normalization or do not explicitly follow the two goals of improving discrimination ability: minimizing intra - class variation and maximizing inter - class variation. Moreover, all these functions only consider the feedback information of mini - batches and do not consider the distribution information of the entire training set.

[0009] (2) Technical solutions

[0010] A cosine optimal loss function applicable to face recognition and based on global information, including the following steps:

[0011] a) Softmax loss is the most commonly used loss function in deep learning and can be expressed as:

[0012]

[0013] where N represents the batch size, P represents the number of classes in the entire training set, f i ∈R d is the feature vector of the i - th sample belonging to the y i - th class, W j ∈R d is the j - th column of the weight matrix W of the last fully - connected layer, b j is the bias term of the j - th class;

[0014] Fix b in Softmax loss j = 0 and ||W j || = 1 to apply L2 weight normalization. At the same time, apply L2 normalization to the feature vector f i and rescale ||f i || to S, and then combine it with the AM-Softmax loss. The resulting total loss is L = L AM + λL G . Where S is a specified constant, L G is the proposed cosine optimal loss function, λ is a hyperparameter used to adjust the influence of these two losses, and L AM is the functional expression of the AM-Softmax loss.

[0015] b) To minimize the within-class variation, first propose a lightweight version of the cosine optimal loss function whose formula is as follows: R(j) = cos(c j , e j ), where P is the number of classes in the entire training set, c j is the center of class j, and e j represents the margin of class j (i.e., the farthest sample in class j). R(j) represents the cosine range of class j, that is, the cosine similarity between the class center and the margin of class j. We use W j as an approximate substitute for c j and propose an algorithm to recursively update the range of each class.

[0016] c) According to the algorithm mentioned in step b), initially, R(j) is initialized to 1. Then we use the following iterative method to update R(j):

[0017]

[0018] where j = 1, 2,..., P,

[0019]

[0020] where, when y i = j, φ(y i , j) = 1, otherwise φ(y i , j) = 0. β is the shrinkage rate, which is used to adjust the shrinkage speed of learning the class range.

[0021] According to the learning algorithm proposed in step b), its basic idea involves two scenarios: ① If the cosine similarity between the input sample and its corresponding class center is less than the recorded class range, directly replace the class range with their cosine similarity; ② On the contrary, if the cosine similarity between the input sample and its corresponding class center is not less than the recorded class range, shrink the class range by scaling their cosine similarity with β. Scenario ① keeps the learned class range up-to-date. As the training progresses, the true class range will become smaller and smaller. Scenario ② is used to help the learned class range shrink to the true value.

[0022] d) To maximize the between-class variation, another lightweight version of the cosine optimal loss function is proposed

[0023]

[0024] where ∑ Top (A, k) represents the sum of the K largest elements in set A, and W a and W b are approximate substitution values of the class centers of any two different classes. The purpose of the cosine optimal loss function is to find K pairs of the nearest class centers in the entire training set and calculate the sum of their distances. Compared with non-adjacent class centers, the corresponding classes of adjacent centers are likely to have smaller intervals or overlaps. If all adjacent classes have appropriate intervals, non-adjacent classes will have larger intervals. Therefore, it is not necessary to consider all center pairs. The most effective method is to optimize the distances of all adjacent centers. Here, the value of K is set to P, where P is the number of classes. Because when all class centers are arranged in a circle on the hypersphere, the minimum number of adjacent center pairs is P.

[0025] e) Integrate the two lightweight versions proposed in step b) and step d) to create a standard version of the cosine optimal loss function

[0026]

[0027] (3) Beneficial effects

[0028] The cosine optimal loss function of the present invention combines the advantages of the optimal loss function proposed in face recognition in recent years. And it is the first attempt to use global information as feedback for face recognition. The cosine optimal loss function uses a new algorithm to learn the cosine similarity between class centers and class margins. The cosine optimal loss function proposed in this patent has been extensively experimented on the LFW, SLLFW, and YTF datasets. The results prove its effectiveness and show that the cosine optimal loss function achieves state-of-the-art performance. Specific implementation manners

[0029] The following further describes the present invention.

[0030] A design method of a cosine optimal loss function applicable to face recognition and based on global information includes the following steps:

[0031] a) Fix b in the Softmax loss j = 0 and ||W j || = 1 to apply L2 weight normalization. At the same time, apply L2 normalization to the feature vector f i and rescale ||f i || to S, and then combine it with the AM-Softmax loss. The resulting total loss is L = L AM + λL G . Where S is a specified constant, L G is the proposed cosine optimal loss function, λ is a hyperparameter used to adjust the influence of these two losses, and L AM is the functional expression of the AM-Softmax loss.

[0032] The Softmax loss is the most commonly used loss function in deep learning and can be expressed as:

[0033]

[0034] where N represents the batch size, P represents the number of classes in the entire training set, f i ∈R d is the feature vector of the i-th sample belonging to the y i -th class, W j ∈R d is the last fully connected layer of the j-th column of the weight matrix W, and b j is the bias term of the j-th class;

[0035] b) To minimize the intra-class variation, first propose a lightweight version of the cosine optimal loss function whose formula is as follows:

[0036] R(j) = cos(c j , e j )

[0037] where P is the number of classes in the entire training set, c j is the center of class j, and e j represents the margin of class j (i.e., the farthest sample in class j). R(j) represents the cosine range of class j, that is, the cosine similarity between the class center and the margin of class j. We use W j as an approximate substitute for c j and propose an algorithm to recursively update the range of each class.

[0038] c) According to the algorithm mentioned in step b), initially, R(j) is initialized to 1. Then we use the following iterative method to update R(j):

[0039]

[0040] where j = 1, 2,..., P,

[0041]

[0042] where, when y i = j, φ(y i , j) = 1, otherwise φ(y i , j) = 0. β is the shrinkage rate, which is used to adjust the shrinkage speed of the learning class range.

[0043] According to the learning algorithm proposed in step b), its basic idea involves two cases: ① If the cosine similarity between the input sample and its corresponding class center is less than the recorded class range, then directly replace the class range with their cosine similarity; ② On the contrary, if the cosine similarity between the input sample and its corresponding class center is not less than the recorded class range, then shrink the class range by scaling their cosine similarity with β. Case ① keeps the learned class range up-to-date. As the training progresses, the true class range will become smaller and smaller. Case ② is used to help the learned class range shrink to the true value.

[0044] d) To maximize the between-class variation, another lightweight version of the cosine optimal loss function is proposed

[0045]

[0046] where ∑ Top (A, k) represents the sum of the k largest elements in set A, and W a and W b are approximate substitutes for the class centers of any two different classes. The purpose of the cosine optimal loss function is to find K pairs of the closest class centers in the entire training set and calculate the sum of their distances. Compared with non-adjacent class centers, the corresponding classes of adjacent centers are likely to have a smaller gap or overlap. If all adjacent classes have appropriate gaps, then non-adjacent classes will have a larger gap. Therefore, it is not necessary to consider all center pairs. The most effective method is to optimize the distances of all adjacent centers. Here, the value of K is set to P, where P is the number of classes. Because when all class centers are arranged in a circle on the hypersphere, the minimum number of adjacent center pairs is P.

[0047] e) Integrate the two lightweight versions proposed in step b) and step d) to create a standard version of the cosine optimal loss function

[0048]

[0049] The cosine optimal loss function combines the advantages of the optimal loss functions proposed in face recognition in recent years. And it attempts to use global information as the feedback for face recognition for the first time. The cosine optimal loss function uses a new algorithm to learn the cosine similarity between the class center and the class margin. A large number of experiments have been carried out on the LFW, SLLFW, and YTF datasets for the cosine optimal loss function proposed in this patent. The results prove its effectiveness and show that the cosine optimal loss function achieves state-of-the-art performance.

Claims

1. An optimization method for the cosine optimal loss function based on global information, where the cosine optimal loss function is used for face recognition, and is characterized in that, It includes the following steps: S1. Combine the advantages of the existing loss function with several new attributes, apply L2 weight normalization, and obtain the total loss function; S2. Clearly follow the two objectives of minimizing intra-class variation and maximizing inter-class variation, rely on a new algorithm to learn the cosine similarity between the class center and the class margin, and respectively propose two lightweight versions of the cosine optimal loss function; S3. Integrate the above two lightweight versions to create a standard version of the cosine optimal loss function; The specific steps of S2 include: S21. To minimize the intra-class variation, the first lightweight version of the cosine optimal loss function is proposed. Its formula is as follows: R(j) = cos(c j , e j ) where P is the number of classes in the entire training set, and c j is the center of class j, and e j represents the margin of class j, and R(j) represents the cosine range of class j, that is, the cosine similarity between the class center and the margin of class j; W j is used as an approximate substitute for c j , and a learning algorithm is used to recursively update the range of each class; S22. To maximize the between-class variation, a second lightweight version of the cosine optimal loss function is proposed. wherein, ∑ Top (A, k) represents the sum of the K largest elements in set A, and W a and W b are approximate substitution values of any two different class centroids; Cosine Optimal Loss Function The purpose is to find K pairs of the nearest class centers in the entire training set and calculate the sum of their distances; optimize the distances between all adjacent centers and set the value of K to P , where P is the number of classes, because when all class centers are arranged in a circle on the hypersphere, the minimum number of adjacent center pairs is P; The specific steps of S3 include: integrating the two lightweight versions proposed in S2 to create a standard version of the cosine optimal loss function 2. The optimization method of the cosine optimal loss function based on global information according to claim 1, characterized in that, The specific steps of S1 include: The Softmax loss function is expressed as: where N represents the batch size, P represents the number of classes in the entire training set, f i ∈R d is the feature vector of the i-th sample belonging to the y i -th class, W j ∈R d is the last fully connected layer of the j-th column of the weight matrix W, b j is the bias term for the j-th class; Fix b in the Softmax loss j = 0 and ||W j || = 1 to apply L2 weight normalization; at the same time, apply L2 normalization to the feature vector f i and rescale ||f i || to s, where s is a specified constant, and then combine it with the AM-Softmax loss. The resulting total loss is L = L AM + λL G ; where L G is the proposed cosine optimal loss function, λ is a hyperparameter used to adjust the influence of these two losses, and L AM is the functional representation of the AM-Softmax loss.

3. The optimization method of the cosine optimal loss function based on global information according to claim 1, characterized in that The learning algorithm of S21 is: R(j) is initialized to 1; then the following iterative method is used to update R(j): where j = 1, 2,..., P; When y i = j, φ(y i, j) = 1, otherwise φ(y i, j) = 0; β is the shrinkage rate, which is used to adjust the shrinkage speed of the learning category range.

4. The optimization method of the cosine optimal loss function based on global information according to claim 1, characterized in that: The learning algorithm of S21 involves two cases: ① If the cosine similarity between the input sample and its corresponding class center is less than the recorded class range, directly replace the class range with their cosine similarity; ② On the contrary, if the cosine similarity between the input sample and its corresponding class center is not less than the recorded class range, shrink the class range by scaling their cosine similarity with β; Case ① keeps the learned class range up-to-date, and as training progresses, the true class range will become smaller and smaller; Case ② is used to help the learned class range shrink to the true value.

Citation Information

Patent Citations

  • Face recognition model acquisition method and device, equipment and medium

    CN110598603A

  • Face recognition method, device and equipment and computer readable storage medium

    CN114627533A