Structural information alignment training method based on feature comparison task
By adopting the structural information alignment training method in the feature comparison task, the structural information alignment problem in complex scenarios is solved, the generalization ability and recognition accuracy of the model are improved, which is significantly better than the traditional training method.
Patent Information
- Application Number
- CN202411962475.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-27
AI Technical Summary
In complex scenarios, it is difficult for the prior art to effectively align structural information between the input space and the latent space, resulting in structural collapse in the latent space, affecting the generalization ability and recognition accuracy of the model.
The structural information alignment training method based on feature comparison tasks is adopted to optimize the model to maintain the structural consistency of the input space and the latent space through data enhancement, topological structure alignment, identification of samples with large structural differences and weighted classification losses.
It improves the generalization ability and robustness of the model, enhances the processing ability of difficult samples, significantly improves the recognition accuracy in complex environments, and accelerates training efficiency.
Smart Images

Figure BDA0005217525070000021 
Figure BDA0005217525070000023 
Figure BDA0005217525070000026
Abstract
Description
Technical Field
[0001] The present invention relates to a training method, and more specifically to a structural information alignment training method based on a feature comparison task. Background Art
[0002] With the continuous development of face recognition technology and pedestrian re-identification (ReID) technology in surveillance scenarios, it is particularly important to optimize the model's ability to remove duplicate faces and pedestrians in complex scenes. The recent success of unsupervised learning and graph neural networks has demonstrated the effectiveness of data structure information. Considering that feature matching tasks can utilize large-scale training data, these data inherently contain important structural information. As the amount of data increases, directly aligning the structural information between the input space and the latent space will inevitably encounter overfitting problems, leading to structural collapse in the latent space. Summary of the invention
[0003] In view of the shortcomings of the prior art, the purpose of the present invention is to provide a structural information alignment training method based on feature comparison tasks, which aligns the topological structure between the input image and the feature hidden layer through a continuous pairing method to increase the structural diversity of the feature layer space and enhance the effect of the model.
[0004] To achieve the above object, the present invention provides the following technical solution: a structural information alignment training method based on feature comparison task, characterized in that it comprises the following steps:
[0005] Step 1: Perform data enhancement, which perturbs the input image to increase the structural diversity of the latent space.
[0006] Step 2: align the topological structure. During the forward propagation process, the Vietoris-Rips complex is constructed based on the point cloud, and the topological features of the complex are calculated through persistent homology, and the topological structure alignment loss is calculated at the same time;
[0007] Step 3: Identify samples with large structural differences. For each sample, estimate the prediction uncertainty through the Gaussian-Uniform Mixture Model to identify difficult samples with large structural destructiveness.
[0008] Step 4: Calculate the weighted classification loss;
[0009] Step 5: Combine the topological alignment loss and weighted classification loss calculated in step 2 to optimize the model. The optimization goal is
[0010]
[0011] Among them, α is the balance coefficient, which is used to adjust the impact of the two losses on model training to complete the training.
[0012] As a further improvement of the present invention, the data enhancement operation in step 1 is as follows: applying the enhancement operation A to the input sample x i , generate perturbation samples The specific formula is:
[0013]
[0014] As a further improvement of the present invention, the specific method of calculating the topological structure alignment loss in step 2 is: comparing the topological difference between the input point cloud x and the perturbed latent space feature z, and defining the topological structure alignment loss L sa :
[0015]
[0016] Among them, M X and are the distance matrices of the original input space and the perturbed latent space, γ X and Pair with the corresponding persistence.
[0017] As a further improvement of the present invention, the damage score of the difficult sample with large structural difference in step 3 is calculated as follows:
[0018] ω(x i )=(1+h φ (x i )) λ ·(1-g gt )
[0019] Among them, h φ (x i ) represents the prediction uncertainty, g gt is the predicted probability of the correct label, and λ is the control coefficient. As a further improvement of the present invention, the weighted classification loss in step 4 is calculated by the following formula: Classification loss function L cls for
[0020] L cls =ω(x i )·L arc (x i ,y i )
[0021] Among them, L arc is the basic classification loss such as Arcface.
[0022] Beneficial effects of the present invention:
[0023] Improve generalization ability: Through the topological structure alignment strategy, the structural consistency of the input space and the latent space is effectively maintained, the risk of structural collapse is reduced, and the model is more robust in complex environments.
[0024] Enhanced processing capabilities for difficult samples: Prioritize optimization of more difficult samples based on their damage scores, significantly improving the recognition accuracy of the model under complex conditions such as occlusion and uneven lighting.
[0025] Accelerate training efficiency: The topological metric of the present invention is fast in calculation speed, avoiding the computational burden of traditional metrics such as bottleneck distance and Wasserstein distance, and is suitable for application on large-scale face and pedestrian re-identification datasets. DETAILED DESCRIPTION
[0026] The present invention will be further described in detail with reference to the given embodiments below.
[0027] In this embodiment, a structural information alignment training method based on a feature comparison task is used, and the most commonly used ResNet50 model is selected, including the following steps:
[0028] Data augmentation: First, the input image is perturbed through a series of data augmentation operations (such as grayscale, color perturbation, etc.) to increase the structural diversity of the latent space. This process is done by applying an augmentation operation A to the input sample x. i , generate perturbation samples Data enhancement uses only four common data enhancement operations (random erasing, Gaussian blur, grayscale, and color enhancement):
[0029]
[0030] Align topology: During the forward propagation, a Vietoris-Rips complex is constructed based on the point cloud, and the topological features of the complex are calculated through persistent homology (PH). The topological difference between the input point cloud x and the perturbed latent space features z is compared, and the topological alignment loss L is defined sa :
[0031]
[0032] Among them, M X and are the distance matrices of the original input space and the perturbed latent space, γ X and Pair with the corresponding persistence.
[0033] Identify samples with large structural differences: For each sample, the prediction uncertainty is estimated through the Gaussian-Uniform Mixture Model (GUM) to identify difficult samples with large structural damage. The damage score is calculated as:
[0034] ω(x i )=(1+h φ (x i )) λ ·(1-g gt )
[0035] Among them, h φ (x i ) represents the prediction uncertainty, g gt is the predicted probability of the correct label, and λ is the control coefficient. Different samples are weighted according to their difficulty, and difficult samples with high structural differences are optimized first.
[0036] Weighted classification loss: The weighted loss function guides the model to pay more attention to difficult samples. The classification loss function L cls for
[0037] L cls =ω(x i )·L arc (x i ,y i )
[0038] Among them, L arc is the basic classification loss such as Arcface.
[0039] Model optimization: The total loss function combines the topology alignment loss and the weighted classification loss, and the optimization goal is
[0040]
[0041] Among them, α is the balance coefficient, which is used to adjust the impact of the two losses on model training.
[0042] Through the above strategies, the model can maintain a high sensitivity to the original image input structure during training, allowing the feature layer to retain more input image structure information.
[0043] To summarize, the structural information alignment training method based on the feature comparison task in this embodiment aligns the topological structure between the input image and the feature hidden layer through a continuous pairing method to increase the structural diversity of the feature layer space and enhance the effect of the model. This method is significantly superior to traditional training methods in monitoring face recognition and pedestrian deduplication tasks, and has higher accuracy, generalization and practicality.
[0044] The above is only a preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions under the concept of the present invention belong to the protection scope of the present invention. It should be pointed out that for ordinary technicians in this technical field, some improvements and modifications without departing from the principle of the present invention should also be regarded as the protection scope of the present invention.
Claims
1. A structural information alignment training method based on feature comparison task, characterized by: The steps include: Step 1: Perform data enhancement, which perturbs the input image to increase the structural diversity of the latent space. Step 2: align the topological structure. During the forward propagation process, the Vietoris-Rips complex is constructed based on the point cloud, and the topological features of the complex are calculated through persistent homology, and the topological structure alignment loss is calculated at the same time; Step 3: Identify samples with large structural differences. For each sample, estimate the prediction uncertainty through the Gaussian-Uniform Mixture Model to identify difficult samples with large structural destructiveness. Step 4: Calculate the weighted classification loss; Step 5: Combine the topological alignment loss and weighted classification loss calculated in step 2 to optimize the model. The optimization goal is Among them, α is the balance coefficient, which is used to adjust the impact of the two losses on model training to complete the training.
2. The structural information alignment training method based on feature comparison task according to claim 1 is characterized in that: The data enhancement operation in step 1 is as follows: Apply the enhancement operation A to the input sample x i , generate perturbation samples The specific formula is:
3. The structural information alignment training method based on feature comparison task according to claim 2 is characterized in that: The specific method of calculating the topological structure alignment loss in step 2 is: compare the topological difference between the input point cloud x and the perturbed latent space feature z, and define the topological structure alignment loss L sa : Among them, M X and are the distance matrices of the original input space and the perturbed latent space, γ X and Pair with the corresponding persistence.
4. The structural information alignment training method based on feature comparison task according to claim 3 is characterized in that: The damage score of the difficult sample with large structural difference in step 3 is calculated as follows: ω(x i )=(1+h φ (x i )) λ ·(1-g gt ) Among them, h φ (x i ) represents the prediction uncertainty, g gt is the predicted probability of the correct label, and λ is the control coefficient.
5. The structural information alignment training method based on feature comparison task according to claim 4 is characterized in that: The weighted classification loss in step 4 is calculated by the following formula: Classification loss function L cls for L cls =ω(x i )·L arc (x i ,y i ) Among them, L arc is the basic classification loss such as Arcface.