Negative Sample Sampling Method Based on Multi-Model Collaborative Contrastive Learning
By building multiple contrast learning models, using diversity constraints and data enhancement, combining potential positive sample identification and difficult negative sample mining, the problem of negative sample sampling deviation in the existing technology is solved, and the training effect and generalization ability of the model are improved.
Patent Information
- Application Number
- CN202211515939.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-30
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-11-30
AI Technical Summary
The random selection of negative samples in existing comparison learning is easy to introduce model deviation, resulting in the impact of the model convergence speed and performance. The existing methods rely on the similarity calculation in the model's own feature space, and there is sampling deviation.
Multiple comparison learning models are constructed, and the model learns different feature subspaces through diversity constraints and data augmentation methods are used to ensure that the model learns different feature subspaces, combines potential positive samples and eliminates potential positive samples and selects high-quality difficult-to-negative samples.
It improves the sampling quality and accuracy of negative samples, improves the generalization ability of the model, reduces the generalization error and the possibility of local optimal trapping, and enhances the training effect of the model.
Smart Images

Figure CN115759205B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of contrast learning and negative sample sampling, and in particular to a negative sample sampling method based on multi-model collaborative contrast learning. Background Art
[0002] As a self-supervised learning, the advantage of contrast learning is that in the scenario without data labels, by means of contrast loss, the model can learn the feature information of data by pulling the distance of positive samples closer and pushing the distance of negative samples farther in the feature embedding space. In the common contrast learning framework, each sample instance is regarded as a category, that is, except for the anchor sample, negative samples are randomly selected from the remaining all samples in the dataset. The problem caused by this approach is that there are likely to be some samples in the selected negative sample set that are very similar to the anchor sample, thus affecting the convergence speed and final performance of the model.
[0003] To solve the problem caused by randomly selecting negative samples, the existing technologies for negative sample sampling in contrast learning can be divided into the following two aspects: 1. Detect and eliminate potential positive samples in the negative sample set through clustering results or similarity-based methods. 2. Inspired by hard negative sample mining in metric learning, improve the efficiency and performance of model training by selecting difficult negative samples. Such methods mainly select high-quality hard negative samples from existing samples based on similarity, or generate hard negative samples through data mixing and adversarial generative networks. However, the above existing methods all need to select or synthesize by means of the similarity between samples, and the existing methods only calculate the similarity within the feature space of the model itself when calculating the similarity, resulting in the reliability of the similarity depending on the current feature expression ability of the model; at the same time, due to the randomness of deep learning training, a single model often only learns part of the feature space of the data, making the selected negative sample set have its own model bias, and there is a sampling bias problem. Summary of the Invention
[0004] The purpose of the present invention is to overcome the deficiencies of the prior art, and propose a negative sample sampling method based on multi-model collaborative contrast learning, which can eliminate the sampling bias introduced by the bias of the model's own feature space, more comprehensively eliminate potential positive samples in the negative sample set, improve the quality of hard negative samples, thereby improving the generalization ability of the contrast learning model and better applying it to downstream tasks.
[0005] To achieve the above purpose, the technical solution provided by the present invention is: a negative sample sampling method based on multi-model collaborative contrast learning, including the following steps:
[0006] 1) Construct two or more contrast learning models, and then use the diversity constraint method to ensure that different models can learn different feature sub-spaces of the dataset;
[0007] 2) Each contrastive learning model calculates the similarity between the anchor sample and the candidate negative sample set within its respective feature space, and then uses a potential positive sample identification algorithm to select a potential positive sample set from the candidate negative sample set;
[0008] 3) Combine the potential positive sample sets selected by different models to obtain the final positive sample set, and remove the final positive sample set from the candidate negative sample set to obtain a preliminary negative sample set;
[0009] 4) Use a hard negative sample mining algorithm to select hard negative samples from the preliminary negative sample set as the final negative sample set for participating in contrastive learning training.
[0010] Furthermore, in step 1), m contrastive learning models are constructed. Each contrastive learning model consists of a feature encoder and a mapper. After the input data x is data-augmented, it is sent into the i-th contrastive learning model to obtain the corresponding metric embedding representation That is:
[0011]
[0012] In the formula, x is the input data, is the metric embedding representation corresponding to the input data x, and h i (·) is the representation space function of the mapper of the i-th model, and f i (·) is the representation space function of the feature encoder of the i-th model, t(·) represents the data augmentation function, and m is the total number of contrastive learning models, m≥2;
[0013] At the same time, in order to ensure that different contrastive learning models can finally learn different feature subspaces of the dataset, the following aspects are used to ensure the diversity among different contrastive learning models:
[0014] Different contrastive learning models use different data augmentation transformations. The data augmentation methods include randomly cropping the size, randomly flipping horizontally, randomly changing the image attributes, and randomly converting to grayscale. Since the data augmentation methods are random, for the same input data x, applying a data augmentation method once before inputting it into different contrastive learning models can obtain different data augmentation transformation results;
[0015] The network layer parameters and initialization methods used in the construction of different contrastive learning models are not exactly the same. Each layer parameter of the contrastive learning model includes the convolution kernel size and number, the stride, and the zero-padding mode. The network layer weight initialization methods include uniform distribution, normal distribution, Xavier initialization, and kaiming initialization;
[0016] To further ensure the differences in the features extracted by different contrastive learning models, a feature diversity constraint between models is imposed during training, that is, the similarity between the features extracted by the feature encoders of different contrastive learning models is calculated, and this similarity is minimized:
[0017]
[0018] In the formula, min means minimizing the value of the right - hand side equation, Loss similarity is the sum of the similarities of the features between all pairs of models, cos(·) is used to calculate the cosine similarity between two data, λ represents the adjustment amplitude for the value of the entire equation, and f j (·) is the representation space function of the feature encoder of the j - th model.
[0019] Furthermore, step 2) includes the following steps:
[0020] 2.1) Given an anchor sample and a set of candidate negative samples, each model calculates the similarity between the anchor sample and each candidate negative sample in its own feature space:
[0021] similarity(x a ,x nc ; θ i ) = cos(h i (f i (t(x a ))),h i (f i (t(x nc )))),x nc ∈NC
[0022] In the formula, x a represents the anchor sample, NC represents the set of candidate negative samples, x nc represents the candidate negative sample, similarity represents the similarity measure between two samples, θ i is the entire representation space of the i - th model, including the feature encoder and the mapper, cos is used to calculate the cosine similarity between two data, h i (·) is the representation space function of the mapper of the i - th model, f i (·) is the representation space function of the feature encoder of the i - th model, and t(·) represents the data augmentation function;
[0023] 2.2) By means of the potential positive sample recognition algorithm, each model respectively selects a set of potential positive samples from the candidate negative sample set according to the obtained similarity. Here, the potential positive sample recognition algorithm means that when the similarity of a sample to the anchor sample is higher than the specified threshold, the sample is defined as a potential positive sample. The set of potential positive samples selected by each model is denoted as:
[0024] Pos i ={x p |x p ∈NC∧similarity(x a ,x p ;θ i )≥α}
[0025] In the formula, Pos i refers to the set of potential positive samples selected by model i, x p represents the potential positive sample, α is the specified threshold, and the determination method of this threshold is:
[0026]
[0027] In the formula, is used to calculate the k-th largest value in the input sequence , represents the similarity sequence of the anchor sample and all candidate negative samples obtained by the i-th model:
[0028]
[0029] In the formula, s nc represents the similarity between the anchor sample and a candidate negative sample.
[0030] Furthermore, in step 3), in order to eliminate all potential positive samples from the candidate negative sample set to the greatest extent, if a sample is considered a potential positive sample by one of the models, then the sample is confirmed as a positive sample. Therefore, the final positive sample set is:
[0031]
[0032] In the formula, Pos final refers to the final positive sample set, Pos i refers to the set of potential positive samples selected by model i, m is the total number of models. By removing the final positive sample set from the candidate negative sample set, the preliminary negative sample set can be obtained:
[0033]
[0034] In the formula, Neg represents the preliminary negative sample set, x nIt represents the preliminary negative samples, and NC is the set of candidate negative samples.
[0035] Furthermore, in step 4), the hard negative sample mining algorithm refers to selecting samples that are very similar to the anchor sample from the preliminary negative sample set, which can effectively improve the convergence speed and final performance of model training; the said step 4) includes the following steps:
[0036] 4.1) To avoid the bias introduced by a single model, the average similarity between the anchor sample and the preliminary negative sample set is calculated using all models as the final similarity score between the anchor sample and each negative sample:
[0037]
[0038] In the formula, score represents the final similarity score between the anchor sample and the negative sample, similarity represents the measure of similarity between two samples, specifically the cosine similarity between the anchor sample and the negative sample calculated by model i here, m is the total number of models, x a represents the anchor sample, x n represents the negative sample, and θ i is the entire representation space of the i-th model, including the feature encoder and the mapper;
[0039] 4.2) Sort the final similarity scores of all negative samples in descending order, and regard the negative samples with the top β in the sorting ratio as hard negative samples, which are used as the final negative sample set participating in the contrastive learning training. Here, β is a percentage.
[0040] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0041] 1. By means of multi-model collaborative sampling, the present invention can comprehensively consider the characteristics of different feature subspaces, reduce the sampling bias and error accumulation caused by the bias of its own model, and improve the quality and accuracy of negative sample sampling.
[0042] 2. By means of the multi-model collaborative method, the present invention identifies the final potential positive sample set, and can identify more potential positive samples compared with the existing methods, making the final negative sample set cleaner.
[0043] 3. In the sampling process, each model of the present invention learns the feature subspace information of other models through cooperation with other models, so that the generalization ability of the final contrastive learning is improved.
[0044] 4. In the present invention, multi-model training can obtain a relatively stable feature space and reduce the generalization error.
[0045] 5. By means of multi-model collaborative sampling, the present invention can reduce the possibility of a single model falling into a local optimum. Brief Description of the Drawings
[0046] Figure 1 This is a schematic diagram of the logic flow of the present invention.
[0047] Figure 2 This is a schematic diagram of multiple contrast model instances constructed by the present invention. Detailed Description of the Preferred Embodiments
[0048] The present invention will be further described in detail below in conjunction with the embodiments and the drawings, but the embodiments of the present invention are not limited thereto.
[0049] As Figure 1 and Figure 2 shown, this embodiment discloses a negative sample sampling method based on multi-model collaborative contrast learning, and the specific situation is as follows:
[0050] 1) Construct m contrast learning models, each contrast learning model is composed of a feature encoder and a mapper. After the input data x is data-augmented, it is sent to the i-th contrast learning model to obtain the corresponding metric embedding representation That is:
[0051]
[0052] In the formula, x is the input data, is the metric embedding representation corresponding to the input data x, h i (·) is the representation space function of the mapper of the i-th model, f i (·) is the representation space function of the feature encoder of the i-th model, t(·) represents the data augmentation function, m is the total number of contrast learning models, and m≥2. In this embodiment, m = 3 is taken as an example for illustration.
[0053] At the same time, in order to ensure that different contrast learning models can finally learn different feature subspaces of the dataset, the following aspects are used to ensure the diversity between different contrast learning models:
[0054] Different contrast learning models use different data augmentation transformations. The data augmentation methods include randomly cropping the size, randomly flipping horizontally, randomly changing the image attributes, and randomly converting to grayscale. Since the data augmentation methods are random, for the same input data x, applying a data augmentation method once before inputting it into different contrast learning models can obtain different data augmentation transformation results;
[0055] The network layer parameters and initialization methods used in the construction of different models are not exactly the same. The parameters of each layer of the contrastive learning model include the size and number of convolutional kernels, the stride, and the zero-padding mode. The network layer weight initialization methods include uniform distribution, normal distribution, Xavier initialization, and kaiming initialization. As Figure 2 shown, the parameters of the feature encoders of the three contrastive learning models are the same, while the number of network layers and the output channels of each layer of the mapper are not the same. In addition, the first model initializes the parameters in the way of normal distribution, the second model initializes the parameters in the way of Xavier initialization, and the third model initializes the parameters in the way of kaiming initialization.
[0056] To further ensure that the features extracted by different models are different, a feature diversity constraint between models is imposed during the training process, that is, calculating the similarity between the features extracted by the feature encoders of different models and minimizing this similarity:
[0057]
[0058] In the formula, min means minimizing the value of the right-hand side equation, Loss similarity is the sum of the similarities of the features between all pairs of models, cos(·) is to calculate the cosine similarity between two data, f j (·) is the representation space function of the feature encoder of the j-th model, and λ represents the adjustment amplitude of the value size of the entire equation. Here, according to experience, λ is taken as 0.05.
[0059] 2) According to the given anchor points and the set of candidate negative samples, each contrastive learning model calculates the similarity between samples in its own feature space, and then uses the potential positive sample recognition algorithm to select the set of potential positive samples from the set of candidate negative samples, including the following steps:
[0060] 2.1) Given the anchor point samples and the set of candidate negative samples, each model calculates the similarity between the anchor point samples and each candidate negative sample in its own feature space:
[0061] similarity(x a ,x nc ;θ i )=cos(h i (f i (t(x a ))),h i (f i (t(x nc )))),x nc ∈NC
[0062] In the formula, x aDenote the anchor sample as \(x\), \(NC\) represents the set of candidate negative samples. nc Denote the candidate negative sample as \(x\), and \(similarity\) represents the similarity measure between two samples, \(\theta\). i \(\mathcal{H}_i\) is the entire representation space of the \(i\)-th model, including the feature encoder and the mapper. \(cos\) is used to calculate the cosine similarity between two data, \(h\). i \(f_i(\cdot)\) is the representation space function of the mapper of the \(i\)-th model. i \(t_i(\cdot)\) is the representation space function of the feature encoder of the \(i\)-th model; \(t(\cdot)\) represents the data augmentation function.
[0063] 2.2) By means of the potential positive sample identification algorithm, each model separately selects a set of potential positive samples from the set of candidate negative samples according to the obtained similarity. Here, the potential positive sample identification algorithm means that when the similarity of a sample to the anchor sample is higher than a specified threshold, then this sample is defined as a potential positive sample. The set of potential positive samples selected by each model can be denoted as:
[0064] \(Pos_i\) i \(=\{x\) p \(|x\) p \(\in NC\land similarity(x\) a , x\) p ; \(\theta\) i )\geq\alpha\}\)
[0065] In the formula, \(Pos_i\) i refers to the set of potential positive samples selected by model \(i\), \(x\) p represents a potential positive sample, and \(\alpha\) is the specified threshold. The determination method of this threshold is:
[0066]
[0067] In the formula, is used to calculate the \(k\)-th largest value in the input sequence . represents the similarity sequence of the anchor sample and all candidate negative samples obtained by the \(i\)-th model:
[0068]
[0069] In the formula, \(s\) nc represents the similarity between the anchor sample and a candidate negative sample. Here, according to experience, \(k\) is set to the floor of 1% of the set of candidate negative samples. For example, if the size of the set of candidate negative samples here is 4096, then the value of \(k\) is the floor of \(4096\times1\% = 40.96\), that is, 40.
[0070] 3) To eliminate all potential positive samples from the candidate negative sample set to the greatest extent, if a sample is considered a potential positive sample by one of the models, then the sample is confirmed as a positive sample. Therefore, the final positive sample set is:
[0071]
[0072] In the formula, Pos final refers to the final positive sample set, Pos i refers to the set of potential positive samples selected by model i, and m is the total number of models. In this embodiment, the finally determined positive sample set is the union of the three sets of potential positive samples selected by the three models. Removing the final positive sample set from the candidate negative sample set can obtain the preliminary negative sample set:
[0073]
[0074] In the formula, Neg represents the preliminary negative sample set, x n represents a preliminary negative sample, and NC is the candidate negative sample set.
[0075] 4) Use the hard negative sample mining algorithm to select hard negative samples from the preliminary negative sample set as the final negative sample set participating in the contrastive learning training. Among them, the hard negative sample mining algorithm refers to selecting samples that are very similar to the anchor sample from the preliminary negative sample set, which can effectively improve the convergence speed and final performance of model training; it includes the following steps:
[0076] 4.1) To avoid the bias introduced by a single model, use all models to calculate the average similarity between the anchor sample and the preliminary negative sample set as the final similarity score between the anchor sample and each negative sample:
[0077]
[0078] In the formula, score represents the final similarity score between the anchor sample and the negative sample, similarity represents the measure of the similarity between two samples, and here specifically represents the cosine similarity between the anchor sample and the negative sample calculated by model i, m is the total number of models, x a represents the anchor sample, and x n represents the negative sample. That is, for each negative sample, the three models calculate the similarity between the negative sample and the anchor respectively, and then the average of the three similarities is used as the final similarity between the negative sample and the anchor.
[0079] 4.2) Sort the final similarity scores of all negative samples in descending order, and regard the negative samples with the top β proportion in the ranking as hard negative samples, which are used as the final negative sample set participating in the contrastive learning training. Here, β is a percentage. According to experience, β is set to 50% here.
[0080] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, or simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.
Claims
1. A negative sample sampling method based on multi-model collaborative contrastive learning, characterized in that It includes the following steps: 1) Construct two or more contrastive learning models, and then use the diversity constraint method to ensure that different models can learn different feature subspaces of the dataset; 2) Each contrastive learning model calculates the similarity between the anchor sample and the candidate negative sample set in its respective feature space, and then uses the potential positive sample recognition algorithm to select the potential positive sample set from the candidate negative sample set; 3) Combine the potential positive sample sets selected by different models to obtain the final positive sample set, and remove the final positive sample set from the candidate negative sample set to obtain the preliminary negative sample set; 4) Use the hard negative sample mining algorithm to select hard negative samples from the preliminary negative sample set as the final negative sample set for participating in the contrastive learning training.
2. The negative sample sampling method based on multi-model collaborative contrastive learning according to claim 1, wherein In step 1), m contrastive learning models are constructed. Each contrastive learning model consists of a feature encoder and a mapper. After the input data x is augmented, it is fed into the i-th contrastive learning model to obtain the corresponding metric embedding representation That is: Where x is the input data, is the metric embedding representation corresponding to the input data x, h i (·) is the representation space function of the mapper of the i-th model, f i (·) is the representation space function of the feature encoder of the i-th model, t(·) represents the data augmentation function, and m is the total number of contrastive learning models, where m ≥ 2; To ensure that different contrastive learning models can ultimately learn different feature subspaces of the dataset, the diversity between different contrastive learning models is ensured from the following aspects: Different contrastive learning models use different data augmentation transformations. The data augmentation methods include randomly cropping the size, randomly flipping horizontally, randomly changing the image attributes, and randomly converting to grayscale. Since the data augmentation methods are random, for the same input data x, applying a data augmentation method once before inputting it into different contrastive learning models can obtain different data augmentation transformation results; The network layer parameters and initialization methods used in the construction of different contrastive learning models are not exactly the same. Each layer parameter of the contrastive learning model includes the convolution kernel size and number, stride, and padding mode. The network layer weight initialization methods include uniform distribution, normal distribution, Xavier initialization, and kaiming initialization; To further ensure the difference in the features extracted by different contrastive learning models, a feature diversity constraint between models is imposed during the training process, that is, calculate the similarity between the features extracted by the feature encoders of different contrastive learning models and minimize this similarity: In the formula, min refers to minimizing the value of the right - hand side equation, and Loss similarity is the sum of the feature similarities between all pairs of models. cos(·) is used to calculate the cosine similarity between two data, λ represents the adjustment amplitude for the value of the entire equation, and f j (·) is the representation space function of the feature encoder of the j - th model.
3. The negative sample sampling method based on multi-model collaborative contrastive learning according to claim 1, characterized in that The step 2) includes the following steps: 2.1) Given the anchor sample and the candidate negative sample set, each model calculates the similarity between the anchor sample and each candidate negative sample in its own feature space: similarity(x a ,x nc ; θ i ) = cos(h i (f i (t(x a ))), h i (f i (t(x nc )))), x nc ∈ NC where x a represents the anchor sample, NC represents the set of candidate negative samples, and x nc represents the candidate negative sample, similarity represents the similarity measure between two samples, and θ i is the entire representation space of the i-th model, including the feature encoder and the mapper, cos is the cosine similarity for calculating between two data, and h i (·) is the representation space function of the mapper of the i-th model, and f i (·) is the representation space function of the feature encoder of the i-th model, and t(·) represents the data augmentation function; 2.2) By means of the potential positive sample recognition algorithm, each model separately selects the potential positive sample set from the candidate negative sample set according to the obtained similarity. Here, the potential positive sample recognition algorithm means that when the similarity between a sample and the anchor sample is higher than the specified threshold, the sample is defined as a potential positive sample. The potential positive sample set selected by each model is denoted as: Pos i = {x p | x p ∈ NC ∧ similarity(x a , x p ; θ i ) ≥ α} where Pos i refers to the set of potential positive samples selected by model i, and x p represents a potential positive sample, and α is a specified threshold, and the determination method of this threshold is as follows: In the formula, is used to calculate the k-th largest value in the input sequence , represents the similarity sequence of the anchor sample obtained by the i-th model and all candidate negative samples: where s nc represents the similarity between an anchor sample and a candidate negative sample.
4. The negative sample sampling method based on multi-model collaborative contrastive learning according to claim 1, characterized in that: In step 3), in order to remove all potential positive samples from the candidate negative sample set to the greatest extent, if a sample is considered a potential positive sample by one of the models, then the sample is confirmed as a positive sample. Therefore, the final positive sample set is: where Pos final refers to the final positive sample set, Pos i refers to the potential positive sample set selected by model i, m is the total number of models. The preliminary negative sample set can be obtained by removing the final positive sample set from the candidate negative sample set: In the formula, Neg represents the initial negative sample set, and x n represents the initial negative sample, and NC is the candidate negative sample set.
5. The negative sample sampling method based on multi-model collaborative contrastive learning according to claim 1, characterized in that: In step 4), the hard negative sample mining algorithm means selecting samples that are very similar to the anchor sample from the preliminary negative sample set, which can effectively improve the convergence speed and final performance of model training; the step 4) includes the following steps: 4.1) To avoid the bias introduced by a single model, the mean similarity between the anchor samples and the initial set of negative samples is calculated using all models as the final similarity score between the anchor samples and each negative sample: Where score represents the final similarity score between the anchor sample and the negative sample, similarity represents the similarity measurement between two samples, specifically the cosine similarity between the anchor sample and the negative sample calculated by model i here, m is the total number of models, x a represents the anchor sample, x n represents the negative sample, θ i is the entire representation space of the i-th model, including the feature encoder and the mapper; 4.2) The final similarity scores of all negative samples are sorted in descending order from largest to smallest, and the negative samples with the top β in the sorting ratio are regarded as hard negative samples and used as the final set of negative samples participating in the contrastive learning training, where β is a percentage.
Citation Information
Patent Citations
Expression recognition method based on positive and negative sample comparative learning
CN114998960A
Determining user authenticity with face liveness detection
US20180357501A1