Sample screening method and device, storage medium and electronic equipment

By calculating the distribution distance between the sample sets, the candidate sample set that is consistent with the distribution of the original sample set is selected as the training samples to be marked, which solves the problem of the model overfitting a few types of samples and improves the performance of the model in the main categories.

CN120578960APending Publication Date: 2025-09-02TSINGHUA UNIVERSITY
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510724957.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

In the prior art, the method of selecting the samples to be marked based on uncertainty indicators causes the model to overfit a few types of samples, ignore the distribution of data subjects, and lead to a degradation of the model's performance in the main categories.

Method used

By dividing the original sample set into candidate sample sets and calculating the distribution distance between the candidate sample set and the original sample set, selecting the candidate sample set with a distribution distance smaller than the critical distance as the training samples to be marked, avoiding excessive attention to a few types of samples and ensuring that the selected training samples can accurately characterize the subject distribution of the original sample data.

Benefits of technology

This improves the performance of the model in main categories, avoids the model's excessive attention to a few class samples, and ensures that the training samples can accurately characterize the subject distribution characteristics of the original sample data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120578960A_ABST
    Figure CN120578960A_ABST
Patent Text Reader

Abstract

The invention provides a sample screening method and device, a storage medium and electronic equipment, and the method comprises the steps: obtaining an original sample set which comprises a plurality of original sample data; after the multiple pieces of original sample data are divided into at least two candidate sample sets, the distribution distance between each candidate sample set and the original sample set is calculated, and the distribution distance is used for representing the distribution difference between the candidate sample set and the sample data in the original sample set. And when it is determined that the distribution distance is smaller than or equal to a first critical distance, taking the original sample data in the corresponding candidate sample set as a to-be-labeled training sample. According to the screening mechanism based on the distribution consistency, on one hand, excessive attention of the model on minority class samples can be effectively avoided, and on the other hand, it can be ensured that the selected training samples can accurately represent the main body distribution characteristics of original sample data, so that the performance of the model on main classes is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technical solution disclosed herein relates to the field of machine learning technology, and in particular to a sample screening method and device, a storage medium, and an electronic device. Background Art

[0002] The core idea of ​​active learning is to maximize model performance while minimizing the labeling cost by selecting the samples that are most valuable for model training.

[0003] In related technical solutions, samples to be labeled are mostly selected based on uncertainty indicators, such as maximum entropy or minimum confidence. This approach has obvious limitations: since the model often shows high uncertainty about outliers in the data distribution, that is, minority class samples, these samples will be repeatedly selected for labeling, causing the model to overfit the minority class samples and ignore the learning of the main data distribution, which ultimately leads to a decline in the performance of the model on the main categories. Summary of the Invention

[0004] In view of this, embodiments of the present disclosure provide a sample screening method and apparatus, a storage medium, and an electronic device.

[0005] According to a first aspect of the present disclosure, a sample screening method is proposed, the method comprising:

[0006] Acquire an original sample set, where the original sample set contains a plurality of original sample data;

[0007] After dividing the plurality of original sample data into at least two candidate sample sets, respectively calculating a distribution distance between each candidate sample set and the original sample set, the distribution distance being used to characterize a distribution difference between sample data in the candidate sample set and the original sample set;

[0008] When it is determined that the distribution distance is less than or equal to the first critical distance, the original sample data in the corresponding candidate sample set is used as a training sample to be labeled.

[0009] In combination with any embodiment provided in the present disclosure, when the feature dimension of the original sample data is greater than or equal to a preset dimension threshold;

[0010] After dividing the plurality of original sample data into at least two candidate sample sets, respectively calculating the distribution distance between each candidate sample set and the original sample set includes:

[0011] Mapping each original sample data in the original sample set to a latent space to obtain a reduced-dimensionality original sample set, wherein the reduced-dimensionality original sample set includes a plurality of reduced-dimensionality sample data;

[0012] After dividing the plurality of dimension-reduced sample data into at least two dimension-reduced candidate sample sets, respectively calculating a distribution distance between each of the dimension-reduced candidate sample sets and the dimension-reduced original sample set;

[0013] The method of using the original sample data in the corresponding candidate sample set as training samples to be labeled includes:

[0014] The original sample data corresponding to the dimension reduction sample data in the corresponding dimension reduction candidate sample set is used as the training sample to be labeled.

[0015] In combination with any embodiment provided in the present disclosure, when the original sample data has at least two characteristic dimensions;

[0016] The respectively calculating the distribution distance between each candidate sample set and the original sample set includes:

[0017] For each candidate sample set, respectively calculating the single-dimensional distribution distance between the candidate sample set and the original sample set in each feature dimension;

[0018] Based on a preset fusion algorithm, the single-dimensional distribution distances under each feature dimension are fused to obtain the distribution distance between the candidate sample set and the original sample set.

[0019] In combination with any embodiment provided in the present disclosure, the method further includes:

[0020] When a preset sample update condition is met, the original sample set is updated based on the newly added sample data to obtain an updated sample set;

[0021] Dividing the newly added sample data to obtain at least one newly added candidate sample set;

[0022] Calculating the distribution distance between each of the newly added candidate sample sets and the updated sample set respectively;

[0023] When it is determined that the distribution distance is less than or equal to the second critical distance, the newly added sample data in the corresponding newly added candidate sample set is used as a training sample to be labeled.

[0024] In combination with any embodiment provided in the present disclosure, the preset sample update condition includes any one of the following:

[0025] The amount of the newly added sample data reaches a preset threshold;

[0026] The time interval from the last sample set update reaches the preset time threshold.

[0027] According to a second aspect of the present disclosure, a sample screening device is provided, comprising:

[0028] A sample set acquisition module is used to acquire an original sample set, wherein the original sample set includes a plurality of original sample data;

[0029] a distribution distance calculation module, configured to, after dividing the plurality of original sample data into at least two candidate sample sets, respectively calculate a distribution distance between each candidate sample set and the original sample set, wherein the distribution distance is used to characterize the distribution difference between the sample data in the candidate sample set and the sample data in the original sample set;

[0030] The sample screening module is configured to use the original sample data in the corresponding candidate sample set as training samples to be labeled when it is determined that the distribution distance is less than or equal to the first critical distance.

[0031] According to a third aspect of the present disclosure, a computer-readable storage medium is provided, wherein the machine-readable storage medium stores machine-readable instructions, which, when called and executed by a processor, prompt the processor to implement the sample screening method of any embodiment of the present disclosure.

[0032] According to a fourth aspect of the present disclosure, there is provided an electronic device comprising

[0033] processor;

[0034] a memory for storing processor-executable instructions;

[0035] The processor is configured to execute the sample screening method of any embodiment of the present disclosure.

[0036] The technical solutions provided by the embodiments of the present disclosure may have the following beneficial effects:

[0037] In the sample screening method and apparatus, storage medium, and electronic device provided by the embodiments of the present disclosure, first, the multiple original sample data contained in the original sample set are divided into several candidate sample sets. Then, by calculating the distribution distance between each candidate sample set and the original sample set, only the sample data in the candidate sample set whose distribution distance is less than a first critical distance is selected as the training samples to be labeled. This screening mechanism based on distribution consistency can, on the one hand, effectively avoid the model's excessive focus on minority class samples, and on the other hand, ensure that the selected training samples can accurately represent the main distribution characteristics of the original sample data, thereby improving the performance of the model on the main categories.

[0038] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0040] Figure 1 is a flow chart of a sample screening method according to an exemplary embodiment of the present disclosure;

[0041] Figure 2 is a flow chart of another sample screening method according to an exemplary embodiment of the present disclosure;

[0042] Figure 3 1 is a schematic diagram of a training loss curve of a variational autoencoder according to an exemplary embodiment of the present disclosure;

[0043] Figure 4 is a flow chart of another sample screening method according to an exemplary embodiment of the present disclosure;

[0044] Figure 5 is a flow chart of another sample screening method according to an exemplary embodiment of the present disclosure;

[0045] Figure 6 is a flow chart of another sample screening method according to an exemplary embodiment of the present disclosure;

[0046] Figure 7 is a schematic diagram showing the relationship between errors within a model distribution and outside a model distribution according to an exemplary embodiment of the present disclosure;

[0047] Figure 8 is a structural schematic diagram of a sample screening device according to an exemplary embodiment of the present disclosure;

[0048] Figure 9 It is a structural diagram of an electronic device according to an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION

[0049] Here, exemplary embodiments will be described in detail, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, unless otherwise indicated, the same numerals in different drawings represent the same or similar elements.

[0050] The core idea of ​​active learning is to maximize model performance while minimizing the labeling cost by selecting the samples that are most valuable for model training.

[0051] In related technical solutions, samples to be labeled are mostly selected based on uncertainty indicators, such as maximum entropy or minimum confidence. This approach has obvious limitations: since the model often shows high uncertainty about outliers in the data distribution, that is, minority class samples, these samples will be repeatedly selected for labeling, causing the model to overfit the minority class samples and ignore the learning of the main data distribution, which ultimately leads to a decline in the performance of the model on the main categories.

[0052] In view of this, the present disclosure provides a sample screening method. Based on this method, samples with a distribution that is relatively consistent with the original sample data can be screened as training samples to be labeled. The model is trained based on these training samples to effectively avoid the model's excessive focus on minority samples.

[0053] The sample screening method according to the embodiment of the present disclosure is described in detail below with reference to the accompanying drawings.

[0054] Figure 1 This is a flow chart of a sample screening method according to an exemplary embodiment of the present disclosure, which can be executed by various computing devices, including but not limited to computer devices. Figure 1 As shown, the exemplary embodiment method may include the following steps:

[0055] In step 101, an original sample set is obtained, where the original sample set includes a plurality of original sample data.

[0056] Each original sample data has a corresponding feature representation.

[0057] The original sample set is an unlabeled sample data set, which can be recorded as D = {M1, M2, ..., M N}, where M N Represents the Nth sample data, where N is the total amount of sample data.

[0058] For ease of understanding, the following embodiments are described using N being 50,000 as an example.

[0059] In step 102, after the plurality of original sample data are divided into at least two candidate sample sets, the distribution distance between each candidate sample set and the original sample set is calculated respectively.

[0060] In practical applications, the multiple original sample data contained in the original sample set D can be reasonably divided according to actual needs to obtain at least two candidate sample sets. For example, 100 candidate sample sets {S1, S2, ..., S 100}, at this time, each candidate sample set contains 500 original sample data.

[0061] Then, for each candidate sample set S i , the candidate sample set S can be calculated i The distribution distance L from the original sample set D i , the distribution distance L i Used to characterize the candidate sample set S i The distribution difference between the sample data in the original sample set D.

[0062] Optionally, the candidate sample set S can be calculated in a variety of ways. i Distribution distances from the original sample set D include, but are not limited to, Wasserstein distance, maximum mean difference, Kullback-Leibler divergence, and Jensen-Shannon divergence. The Wasserstein distance quantifies the global distribution difference by calculating the "minimum transmission cost" of the sample data distributions in the candidate sample set S and the original sample set D. It accurately describes the overall deviation between distributions and is applicable to distribution matching problems in high-dimensional spaces. In this example, when using the Wasserstein distance method, the entropy-constrained Sinkhorn algorithm can be introduced. By adding an entropy regularization term to the optimization process, complex problems can be transformed into more easily solvable optimization problems, thereby improving computational efficiency while reducing computational complexity. The maximum mean difference measures distribution distance by using the difference in mean embeddings in the reproducing kernel Hilbert space and is applicable to distribution matching of high-dimensional or structured data. The Kullback-Leibler divergence is an information-theoretic metric that measures the overall deviation by calculating the relative entropy difference between two probability distributions. It is suitable for capturing scenarios with large differences in probability densities. The aforementioned Jensen-Shannon divergence is a symmetric form of the Kullback-Leibler divergence. It is more stable when dealing with overlapping distributions and can more balancedly characterize the similarities and differences between distributions.

[0063] In practical applications, a suitable distribution distance calculation method can be selected according to the data characteristics of different sample data, and this disclosure does not limit this.

[0064] In step 103, when it is determined that the distribution distance is less than or equal to the first critical distance, the original sample data in the corresponding candidate sample set is used as a training sample to be labeled.

[0065] In this example, the aforementioned first critical distance can be determined through Monte Carlo algorithm simulation.

[0066] The distribution distance {L1, L2, ..., L 100}After that, each distribution distance Li Compare it with the aforementioned first critical distance, and when it is determined to be less than or equal to the aforementioned first critical distance, determine that the distribution distance L i The corresponding candidate sample set has a high consistency in data distribution with the original sample set. At this time, the original sample data contained in the candidate sample set can be used as training samples to be labeled to train the model based on the training samples to be labeled.

[0067] In the sample screening method provided by the embodiments of the present disclosure, the multiple original sample data contained in the original sample set are first divided into several candidate sample sets. Then, by calculating the distribution distance between each candidate sample set and the original sample set, only the sample data in the candidate sample set whose distribution distance is less than a first critical distance is selected as the training samples to be labeled. This screening mechanism based on distribution consistency can effectively avoid the model's excessive focus on minority class samples on the one hand, and on the other hand, it can ensure that the selected training samples can accurately represent the main distribution characteristics of the original sample data, thereby improving the model's performance on the main categories.

[0068] In an optional embodiment, when the feature dimension of the original sample data is greater than or equal to the preset dimension threshold, that is, the original sample data is high-dimensional sample data, such as Figure 2 As shown, the aforementioned step 102 may specifically include:

[0069] In step 201, each original sample data in the original sample set is mapped to a latent space to obtain a reduced-dimensionality original sample set, where the reduced-dimensionality original sample set contains a plurality of reduced-dimensionality sample data.

[0070] Due to the curse of dimensionality problem, it is impossible to calculate the distribution distance of high-dimensional sample data well. Therefore, in this example, the high-dimensional sample data in the original sample set D can be first reduced in dimension to obtain the reduced-dimensional original sample set D'.

[0071] Continuing with the previous example, when the original sample set D contains 50,000 high-dimensional sample data, these 50,000 high-dimensional sample data can be mapped into the latent space through the trained variational autoencoder to obtain the corresponding 50,000 reduced-dimensional sample data. These 50,000 reduced-dimensional sample data constitute the aforementioned reduced-dimensional original sample set D'.

[0072] Optionally, after mapping the original sample data to the latent space, standardization processing (normalization processing) can be performed on the sample data in the latent space to eliminate the scale differences between feature dimensions and provide high-quality data input for subsequent calculation of distribution distance.

[0073] Figure 3 An example of a variational autoencoder training loss curve diagram, from Figure 3 The variational autoencoder is significantly effective in reducing reconstruction error. Mean Squared Error (MSE_loss) quickly decreases to a low value and then remains stable. KL Divergence (KLD_loss) initially decreases and then recovers, indicating that under the current training strategy and hyperparameter settings, the variational autoencoder tends to further reduce reconstruction error. This is consistent with the primary goal of reducing MSE_loss and maintaining a certain regularization constraint on the latent space.

[0074] In step 202, after the plurality of dimension-reduced sample data are divided into at least two dimension-reduced candidate sample sets, a distribution distance between each of the dimension-reduced candidate sample sets and the dimension-reduced original sample set is calculated respectively.

[0075] Similar to the above, the multiple dimensionality reduction original sample data contained in the dimensionality reduction original sample set D' can be reasonably divided according to actual needs to obtain at least two dimensionality reduction candidate sample sets. For example, 100 dimensionality reduction candidate sample sets {S1', S2', ..., S 100 '}, at this time, each dimensionality reduction candidate sample set contains 500 dimensionality reduction original sample data.

[0076] Then, for each dimensionality reduction candidate sample set S i ', the dimensionality reduction candidate sample set S can be calculated i The distribution distance L between ' and the original sample set D' i ', the distribution distance L i 'Used to characterize the candidate sample set S for dimensionality reduction i The distribution difference between the sample data in the original sample set D'.

[0077] At this time, the aforementioned step 103 may specifically include: using the original sample data corresponding to the dimension reduction sample data in the corresponding dimension reduction candidate sample set as the training samples to be labeled.

[0078] The distribution distance {L1', L2', ..., L between each of the dimensionality reduction candidate sample sets and the dimensionality reduction original sample set is calculated respectively. 100 '}After that, each distribution distance L i 'Compare with the preset first critical distance, and when it is determined to be less than or equal to the first critical distance, determine the distribution distance L i The data distribution of the corresponding dimensionality reduction candidate sample set is highly consistent with that of the dimensionality reduction original sample set. At this time, the original sample data corresponding to the dimensionality reduction original sample data contained in the dimensionality reduction candidate sample set can be used as the training samples to be labeled.

[0079] The sample screening method provided by the disclosed embodiments uses a variational autoencoder to reduce the dimensionality of high-dimensional sample data to obtain low-dimensional sample data, and then calculates the distribution distance. This effectively solves the curse of dimensionality problem and is highly practical. Furthermore, this solution uses a variational autoencoder to reduce the dimensionality of sample data, preserving the global structure and distribution characteristics of the sample data, thereby improving the accuracy of sample data screening.

[0080] In an optional embodiment, when the aforementioned original sample data has at least two feature dimensions, such as Figure 4 As shown, the aforementioned step 102 may specifically include:

[0081] In step 401 , for each candidate sample set, the single-dimensional distribution distance between the candidate sample set and the original sample set in each feature dimension is calculated respectively.

[0082] Taking the original sample data having two feature dimensions as an example, for each candidate sample set, the single-dimensional distribution distance 1 between the candidate sample set and the original sample set in feature dimension 1, and the single-dimensional distribution distance 2 between the candidate sample set and the original sample set in feature dimension 2 can be calculated.

[0083] It is understandable that different distribution distance calculation methods can be used for different feature dimensions to improve calculation efficiency.

[0084] In step 402, based on a preset fusion algorithm, the single-dimensional distribution distances under each feature dimension are fused to obtain the distribution distance between the candidate sample set and the original sample set.

[0085] The preset fusion algorithm may include but is not limited to arithmetic mean, geometric mean, Lp norm, etc.

[0086] In this example, the calculated single-dimensional distribution distance 1 and the single-dimensional distribution distance 2 may be fused based on a preset fusion algorithm to obtain the distribution distance between the candidate sample set and the original sample set.

[0087] It should be noted that the aforementioned description of the raw sample data having two characteristic dimensions is merely illustrative, intended to facilitate a better understanding of the technical solutions of the embodiments of the present disclosure by those skilled in the art. In practical applications, the characteristic dimensions of the raw sample data can reach hundreds or even thousands of dimensions, and this disclosure does not limit this.

[0088] The sample screening method provided by the disclosed embodiments significantly reduces computational complexity through a strategy of dimensional calculation and fusion, avoiding the computational burden of performing complex operations directly in high-dimensional space. Furthermore, the use of a configurable fusion algorithm allows for flexible adaptation to the metric characteristics and practical application requirements of different feature dimensions, preserving the differential characteristics of each dimension while achieving scientific quantification of overall distribution differences.

[0089] In an optional embodiment, if Figure 5 As shown, in Figure 1 Based on the process shown, the sample screening method may further include the following steps:

[0090] In step 501, when a preset sample update condition is met, the original sample set is updated based on the newly added sample data to obtain an updated sample set.

[0091] The preset sample update condition may include any one of the following: the number of newly added sample data reaches a preset number threshold, and the time interval from the last sample set update reaches a preset time threshold.

[0092] When the preset sample update condition includes: the number of new sample data reaches a preset number threshold, and the preset number threshold is 50,000, for example, whenever the number of new sample data collected reaches 50,000, the 50,000 collected sample data will be added to the original sample set to obtain an updated sample set.

[0093] When the preset sample update conditions include: the time interval from the last sample set update reaches a preset time threshold, and the preset time threshold is, for example, 3 hours, the collected new sample data can be added to the original sample set every 3 hours to obtain an updated sample set.

[0094] In step 502, the newly added sample data is divided to obtain at least one newly added candidate sample set.

[0095] In this example, the newly added sample data may be divided into a new candidate sample set with every 500 newly added sample data, to obtain multiple new candidate sample sets.

[0096] In step 503, the distribution distance between each of the newly added candidate sample sets and the updated sample set is calculated respectively.

[0097] For each newly added candidate sample set, a distribution distance between the newly added candidate sample set and the aforementioned updated sample set may be calculated. The distribution distance is used to characterize the distribution difference of sample data in the newly added candidate sample set and the updated sample set.

[0098] In step 504, when it is determined that the distribution distance is less than or equal to the second critical distance, the newly added sample data in the corresponding newly added candidate sample set is used as a training sample to be labeled.

[0099] After calculating the distribution distance between each newly added candidate sample set and the updated sample set respectively, each distribution distance can be compared with the second critical distance. When a distribution distance is less than or equal to the second critical distance, it is determined that the data distribution of the newly added candidate sample set corresponding to the distribution distance and the updated sample set has a high consistency. At this time, the newly added sample data contained in the newly added candidate sample set can be used as training samples to be labeled, and the model can be trained based on the training samples to be labeled.

[0100] In this example, after the original sample set is updated based on the newly added sample data to obtain an updated sample set, the aforementioned second critical distance can be determined using a Monte Carlo simulation algorithm based on the updated sample set. In other words, in this example, the training sample screening strategy can be adaptively adjusted based on the dynamic changes in the sample set, allowing the model to quickly respond to dynamic changes in the data distribution, ensuring real-time and robustness.

[0101] Optionally, when calculating the distribution distance based on the Wasserstein distance method, a progressive optimization strategy can be adopted. Specifically, in the initial stage, the first-order Wasserstein distance with lower computational complexity can be used first to quickly screen out representative samples to be labeled. In subsequent stages, for example, when the number of sample data in the updated sample set reaches a preset number, the high-precision second-order Wasserstein distance can be used to capture the fine-grained characteristics of the distribution, thereby improving the accuracy of sample screening.

[0102] The sample screening method provided by the disclosed embodiments utilizes a periodic or quantitatively triggered update mechanism, which can quickly adapt to changes in the distribution of dynamic data streams, ensuring the timeliness of the sample set while also avoiding the computational overhead associated with frequent updates. Furthermore, the use of distribution distance comparison ensures the consistency of the distribution of newly added sample data with that of existing data, effectively preventing the negative impact of data distribution shifts on model training.

[0103] Figure 6 This is a flow chart of another sample screening method according to an exemplary embodiment of the present disclosure. In this embodiment, the same steps as those in the previous embodiment will be briefly described and will not be described in detail. For details, please refer to any of the previous embodiments. Figure 6 As shown, the exemplary embodiment method may include the following steps:

[0104] In step 601, an original sample set is obtained.

[0105] The original sample set contains a plurality of original sample data, and the characteristic dimension of the original sample data is greater than or equal to a preset dimension threshold, that is, the original sample data is high-dimensional sample data.

[0106] In step 602, each original sample data in the original sample set is mapped to a latent space to obtain a reduced-dimensional original sample set.

[0107] The dimensionality reduction original sample set includes a plurality of dimensionality reduction sample data.

[0108] In step 603, the plurality of dimension-reduced sample data in the dimension-reduced original sample set is divided into at least two dimension-reduced candidate sample sets.

[0109] In step 604, for each dimensionality reduction candidate sample set, the single-dimensional distribution distance between the dimensionality reduction candidate sample set and the dimensionality reduction original sample set in each feature dimension is calculated respectively.

[0110] In step 605, based on a preset fusion algorithm, the single-dimensional distribution distances under each feature dimension are fused to obtain the distribution distance between the dimensionality reduction candidate sample set and the dimensionality reduction original sample set.

[0111] In step 606, when it is determined that the distribution distance is less than or equal to the first critical distance, the original sample data corresponding to the dimension reduction sample data in the dimension reduction candidate sample set is used as a training sample to be labeled.

[0112] In the sample screening method provided by the embodiment of the present disclosure, the distribution distance comparison method is used to ensure that the screened training samples to be labeled have a high distribution consistency with the sample data in the original sample set, that is, it can ensure that the screened training samples to be labeled have a high representativeness, thereby minimizing the labeling cost while maximizing the model performance.

[0113] Figure 7 This example shows a schematic diagram of the error relationship between the model distribution and the error outside the model distribution. Figure 7 In the example, Set A is a set where all sample data are within the critical distance, Set B is a set where all sample data deviate from the critical distance, and Set C is a set where all sample data have a large gap from the critical distance.

[0114] from Figure 7 From the above, we can see that the errors within the model distribution and outside the model distribution in Set A approximately satisfy a linear relationship. Therefore, by selecting sample data in Set A that meets the critical distance for training, we can obtain excellent generalization ability of the model, that is, lower out-of-distribution error.

[0115] Specifically, we can find a model with a larger In-Distribution MAE (mean absolute error within the distribution) in Set A to obtain a smaller Out-of-Distribution MAE (mean absolute error outside the distribution).

[0116] For the sake of simplicity, the aforementioned method embodiments are all expressed as a series of action combinations. However, those skilled in the art should know that the present disclosure is not limited to the order of the actions described, because according to the present disclosure, certain steps can be performed in other orders or simultaneously.

[0117] Corresponding to the aforementioned method embodiments, the present disclosure also provides apparatus embodiments.

[0118] Figure 8 FIG. 1 is a schematic structural diagram of a sample screening device according to an exemplary embodiment of the present disclosure. Figure 8 The sample screening device may include:

[0119] The sample set acquisition module 81 is configured to acquire an original sample set, where the original sample set includes a plurality of original sample data.

[0120] The distribution distance calculation module 82 is used to calculate the distribution distance between each candidate sample set and the original sample set after dividing the multiple original sample data into at least two candidate sample sets. The distribution distance is used to characterize the distribution difference between the sample data in the candidate sample set and the original sample set.

[0121] The sample screening module 83 is configured to use the original sample data in the corresponding candidate sample set as training samples to be labeled when it is determined that the distribution distance is less than or equal to the first critical distance.

[0122] Optionally, when the feature dimension of the original sample data is greater than or equal to a preset dimension threshold;

[0123] The distribution distance calculation module 82, when used to calculate the distribution distance between each candidate sample set and the original sample set after dividing the plurality of original sample data into at least two candidate sample sets, includes:

[0124] Each original sample data in the original sample set is mapped to a latent space to obtain a reduced-dimensional original sample set, wherein the reduced-dimensional original sample set contains a plurality of reduced-dimensional sample data.

[0125] After the plurality of dimension-reduced sample data are divided into at least two dimension-reduced candidate sample sets, a distribution distance between each of the dimension-reduced candidate sample sets and the dimension-reduced original sample set is calculated respectively.

[0126] The sample screening module 83, when used to use the original sample data in the corresponding candidate sample set as training samples to be labeled, includes:

[0127] The original sample data corresponding to the dimension reduction sample data in the corresponding dimension reduction candidate sample set is used as the training sample to be labeled.

[0128] Optionally, when the original sample data has at least two feature dimensions;

[0129] The distribution distance calculation module 82, when used to respectively calculate the distribution distance between each candidate sample set and the original sample set, includes:

[0130] For each candidate sample set, the single-dimensional distribution distance between the candidate sample set and the original sample set in each feature dimension is calculated respectively.

[0131] Based on a preset fusion algorithm, the single-dimensional distribution distances under each feature dimension are fused to obtain the distribution distance between the candidate sample set and the original sample set.

[0132] Optional, in Figure 8 Based on the modules shown, the sample screening device may further include:

[0133] The sample set updating module is configured to update the original sample set based on the newly added sample data to obtain an updated sample set when a preset sample updating condition is met.

[0134] The sample data partitioning module is configured to partition the newly added sample data to obtain at least one newly added candidate sample set.

[0135] The distribution distance calculation module is further configured to respectively calculate the distribution distance between each of the newly added candidate sample sets and the updated sample set.

[0136] The sample screening module is further configured to, when it is determined that the distribution distance is less than or equal to a second critical distance, use the newly added sample data in the corresponding newly added candidate sample set as training samples to be labeled.

[0137] Optionally, the preset sample update condition includes any of the following:

[0138] The amount of the newly added sample data reaches a preset threshold.

[0139] The time interval from the last sample set update reaches the preset time threshold.

[0140] As for the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can be referred to the partial description of the method embodiment.

[0141] Figure 9 FIG1 is a schematic diagram showing the structure of an electronic device 900 according to an exemplary embodiment of the present disclosure. The electronic device may be any type of computing device, including but not limited to a computer device.

[0142] Reference Figure 9 , the electronic device 900 may include one or more of the following components: a processing component 902 , a memory 904 , a power component 906 , a multimedia component 908 , an audio component 910 , an input / output (I / O) interface 912 , a sensor component 914 , and a communication component 916 .

[0143] The processing component 902 generally controls the overall operation of the electronic device 900, such as operations associated with display, phone calls, data communications, camera operation, and recording operations. The processing component 902 may include one or more processors 920 to execute instructions to perform all or part of the steps of the above-described method. In addition, the processing component 902 may include one or more modules to facilitate interaction between the processing component 902 and other components. For example, the processing component 902 may include a multimedia module to facilitate interaction between the multimedia component 908 and the processing component 902.

[0144] The memory 904 is configured to store various types of data to support operations on the device 900. Examples of such data include instructions for any application or method operating on the electronic device 900, contact data, phone book data, messages, pictures, videos, etc. The memory 904 can be implemented by any type of volatile or non-volatile storage device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.

[0145] The power supply component 906 provides power to the various components of the electronic device 900. The power supply component 906 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the electronic device 900.

[0146] The multimedia component 908 includes a screen that provides an output interface between the electronic device 900 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, slides, and gestures on the touch panel. The touch sensors may not only sense the boundaries of a touch or slide action, but also detect the duration and pressure associated with the touch or slide action. In some embodiments, the multimedia component 908 includes a front camera and / or a rear camera. When the electronic device 900 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera may receive external multimedia data. Each front camera and rear camera may be a fixed optical lens system or have focal length and optical zoom capabilities.

[0147] The audio component 910 is configured to output and / or input audio signals. For example, the audio component 910 includes a microphone (MIC), and when the electronic device 900 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode, the microphone is configured to receive an external audio signal. The received audio signal can be further stored in the memory 904 or transmitted via the communication component 916. In some embodiments, the audio component 910 also includes a speaker for outputting audio signals.

[0148] I / O interface 912 provides an interface between processing component 902 and peripheral interface modules, such as a keyboard, click wheel, buttons, etc. These buttons may include but are not limited to: a home button, volume buttons, a start button, and a lock button.

[0149] The sensor assembly 914 includes one or more sensors for providing various aspects of status assessment for the electronic device 900. For example, the sensor assembly 914 can detect the open / closed state of the electronic device 900, the relative positioning of components, such as the display and keypad of the electronic device 900. The sensor assembly 914 can also detect changes in the position of the electronic device 900 or a component of the electronic device 900, the presence or absence of user contact with the electronic device 900, the orientation or acceleration / deceleration of the electronic device 900, and the temperature change of the electronic device 900. The sensor assembly 914 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 914 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 914 may also include an accelerometer, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0150] The communication component 916 is configured to facilitate wired or wireless communication between the electronic device 900 and other devices. The electronic device 900 can access a wireless network based on a communication standard, such as WiFi, 4G or 5G, 4G LTE, 5G NR or a combination thereof. In an exemplary embodiment, the communication component 916 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 916 further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.

[0151] In an exemplary embodiment, the electronic device 900 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above-described methods.

[0152] In an exemplary embodiment, a non-temporary computer-readable storage medium is also provided, such as a memory 904 including instructions, which, when the instructions in the storage medium are executed by the processor 920 of the electronic device 900, enables the electronic device 900 to perform the sample screening method of any embodiment of the present disclosure.

[0153] The non-transitory computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, and the like.

[0154] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

Claims

1. A sample screening method, characterized in that: The method comprises: Acquire an original sample set, where the original sample set contains a plurality of original sample data; After dividing the plurality of original sample data into at least two candidate sample sets, respectively calculating a distribution distance between each candidate sample set and the original sample set, the distribution distance being used to characterize a distribution difference between sample data in the candidate sample set and the original sample set; When it is determined that the distribution distance is less than or equal to the first critical distance, the original sample data in the corresponding candidate sample set is used as a training sample to be labeled.

2. The method according to claim 1, characterized in that When the feature dimension of the original sample data is greater than or equal to a preset dimension threshold; After dividing the plurality of original sample data into at least two candidate sample sets, respectively calculating the distribution distance between each candidate sample set and the original sample set includes: Mapping each original sample data in the original sample set to a latent space to obtain a reduced-dimensionality original sample set, wherein the reduced-dimensionality original sample set includes a plurality of reduced-dimensionality sample data; After dividing the plurality of dimension-reduced sample data into at least two dimension-reduced candidate sample sets, respectively calculating a distribution distance between each of the dimension-reduced candidate sample sets and the dimension-reduced original sample set; The method of using the original sample data in the corresponding candidate sample set as training samples to be labeled includes: The original sample data corresponding to the dimension reduction sample data in the corresponding dimension reduction candidate sample set is used as the training sample to be labeled.

3. The method according to claim 1, characterized in that When the original sample data has at least two feature dimensions; The respectively calculating the distribution distance between each candidate sample set and the original sample set includes: For each candidate sample set, respectively calculating the single-dimensional distribution distance between the candidate sample set and the original sample set in each feature dimension; Based on a preset fusion algorithm, the single-dimensional distribution distances under each feature dimension are fused to obtain the distribution distance between the candidate sample set and the original sample set.

4. The method according to claim 1, wherein The method further comprises: When a preset sample update condition is met, the original sample set is updated based on the newly added sample data to obtain an updated sample set; Dividing the newly added sample data to obtain at least one newly added candidate sample set; Calculating the distribution distance between each of the newly added candidate sample sets and the updated sample set respectively; When it is determined that the distribution distance is less than or equal to the second critical distance, the newly added sample data in the corresponding newly added candidate sample set is used as a training sample to be labeled.

5. The method according to claim 4, characterized in that The preset sample update condition includes any of the following: The amount of the newly added sample data reaches a preset threshold; The time interval from the last sample set update reaches the preset time threshold.

6. A sample screening device, characterized in that: The device comprises: A sample set acquisition module is used to acquire an original sample set, wherein the original sample set includes a plurality of original sample data; a distribution distance calculation module, configured to, after dividing the plurality of original sample data into at least two candidate sample sets, respectively calculate a distribution distance between each candidate sample set and the original sample set, wherein the distribution distance is used to characterize the distribution difference between the sample data in the candidate sample set and the sample data in the original sample set; The sample screening module is configured to use the original sample data in the corresponding candidate sample set as training samples to be labeled when it is determined that the distribution distance is less than or equal to the first critical distance.

7. The device according to claim 6, characterized in that When the feature dimension of the original sample data is greater than or equal to a preset dimension threshold; The distribution distance calculation module, when used to calculate the distribution distance between each candidate sample set and the original sample set after dividing the plurality of original sample data into at least two candidate sample sets, includes: Mapping each original sample data in the original sample set to a latent space to obtain a reduced-dimensionality original sample set, wherein the reduced-dimensionality original sample set includes a plurality of reduced-dimensionality sample data; After dividing the plurality of dimension-reduced sample data into at least two dimension-reduced candidate sample sets, respectively calculating a distribution distance between each of the dimension-reduced candidate sample sets and the dimension-reduced original sample set; The sample screening module, when used to use the original sample data in the corresponding candidate sample set as training samples to be labeled, includes: The original sample data corresponding to the dimension reduction sample data in the corresponding dimension reduction candidate sample set is used as the training sample to be labeled.

8. The device according to claim 6, characterized in that When the original sample data has at least two feature dimensions; The distribution distance calculation module, when used to respectively calculate the distribution distance between each candidate sample set and the original sample set, includes: For each candidate sample set, respectively calculating the single-dimensional distribution distance between the candidate sample set and the original sample set in each feature dimension; Based on a preset fusion algorithm, the single-dimensional distribution distances under each feature dimension are fused to obtain the distribution distance between the candidate sample set and the original sample set.

9. A computer-readable storage medium having a computer program stored thereon, wherein when the program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.

10. An electronic device comprising: processor; a memory for storing processor-executable instructions; The processor is configured to execute the steps of any one of the methods of claims 1-5.

Citation Information

Cited By

  • Target detection method and device based on registration sample, storage medium and electronic equipment

    CN121388632A

  • Model quantification method, device, equipment, medium and product

    CN122222049A